SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

The Reality Posed by the Small Giant HRM: AI Evolution Shifts from 'Size' to 'Structural Intelligence'

Introduction: A Counterpunch to the AI Scaling Race

In recent years, AI development has pushed forward on a path of massive scaling driven by the 'scaling law.' Parameter counts have grown from billions to trillions, training data has expanded to a scale comparable to the entire internet, and computational resources have ballooned by orders of magnitude. However, what awaited at the end of this path were astronomical costs and inference capabilities that were hitting a ceiling.

Challenging this conventional wisdom is the 'Hierarchical Reasoning Model (HRM)' developed by Sapient Intelligence, a startup based in Singapore. Surprisingly, HRM is a small model with only 27 million parameters, yet it has successively broken through reasoning tasks that even massive LLMs, including ChatGPT, struggle with.



The Limits of Scaling Laws and the Fragility of 'Chain-of-Thought'

Conventionally, a technique called 'Chain-of-Thought' has been used to improve reasoning accuracy. While it improves accuracy by having the LLM write out intermediate reasoning steps, it has suffered from fatal flaws:
• Vulnerability: One wrong step can cause the entire conclusion to collapse.
• Inefficiency: Reasoning becomes slow.
These were critical issues.

Sapient researchers dismissed this as 'nothing more than a crutch' and sought a more robust reasoning mechanism. They found the answer in the human brain.



HRM Structure: A Two-Layer Structure of Strategists and Execution Units

The core of HRM lies in mimicking the 'hierarchy of the brain.'
• High-level H-module: Operates slowly, planning and strategizing (equivalent to the prefrontal cortex).
• Low-level L-module: Operates at high speed, processing specific tasks assigned by H (the execution unit).

'Hierarchical convergence' is repeated between the two, testing local solutions, receiving reports, and updating strategies. This teamwork-like structure prevents the 'premature convergence' seen in traditional models and enables reasoning that persistently examines problems from multiple angles.



Shocking Results: The Small Giant Overwhelms Large Models

A Feat in ARC-AGI

In the ARC test, which measures human abstract reasoning, HRM achieved 40.3%, surpassing OpenAI's o3-mini-high (34.5%) and Anthropic's Claude 3.7 Sonnet (21.2%).

Overwhelming Victory in Difficult Puzzles
• Sudoku-Extreme: LLMs have a 0% success rate. HRM achieved 55.0%.
• Maze-Hard (30x30 maze): LLMs have 0%. HRM achieved 74.5%.

Furthermore, it used only 1,000 samples of training data. This is the exact opposite approach to massive retail-type LLMs, making it truly a 'small giant.'



Why Can It Reason 'Smartly'?

HRM does not use CoT to express thoughts in language; instead, it performs 'Latent Reasoning' directly in the latent space.
The results:
• No language generation cost → Potential for up to 100x speedup.
• Maintains and verifies multiple hypotheses in parallel → Dramatic increase in robustness.

Visualization of the maze task showed that it initially explores multiple routes simultaneously, gradually eliminates dead ends, and finally converges on the optimal solution. This movement is extremely close to human 'trial and error.'



Expert Perspectives and Challenges

In a reproduction verification by Live Science, while the scores were confirmed, it was pointed out that 'the source of the performance is likely not just the hierarchical structure, but also the iterative improvement process during training.'

In other words, while HRM is innovative, there is still room for debate regarding what exactly was the breakthrough. This is a healthy progression of science, and further research is required.



Application Areas Opened by HRM
• Industry/Business: Powerful in fields where low-cost, high-speed reasoning is essential, such as logistics optimization, financial risk analysis, and factory anomaly detection.
• Scientific Research: Because it functions with small amounts of data, it is useful in pharmaceutical, climate, and materials research.
• Edge AI: Can be directly installed in autonomous vehicles, drones, and smart home appliances.

Ultimately, it could also serve as a stepping stone toward AGI (Artificial General Intelligence).



Conclusion: The Future of AI is from 'Quantity' to 'Structure'

The emergence of HRM has shifted the direction of AI evolution from 'scaling up' to 'structural intelligence.' By learning from the ultimate blueprint—the human brain—it extracts intelligence efficiently.

The future of AI development is not just about getting bigger, but about becoming more sophisticated, more efficient, and 'more humanly' intelligent.



References
• arXiv: Hierarchical Reasoning Model
• Sapient Intelligence: A New Paradigm of Scaling Law
• GitHub: sapientinc/HRM
• Live Science: Reproduction of ARC-AGI Results

いいなと思ったら応援しよう!