SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

How AI Becomes Smarter: From Knowledge to Application to Reasoning

The foundational technologies supporting the recent AI boom can be categorized into three areas: pre-training, post-training, and reasoning capabilities. Zach, who has served as a product manager at DeepMind for the longest time, systematically described the roles and challenges of each of these three elements, as well as market application trends, essentially detailing the 'process by which models become smarter.' In this article, using his lecture as a guide, we unravel how AI models acquire intelligence and are utilized in the real world.


1. Acquisition of General Knowledge through Pre-Training


1-1. Massive Datasets that Incorporate World Knowledge into Weights

Zach explains that 'pre-training is training to capture the world's knowledge and information into the model's weights.' By repeatedly learning the task of predicting the next token (such as words or pixels) using vast amounts of text, image, and audio data, the model acquires basic language understanding and factual recognition capabilities. This method was dramatically improved in efficiency by the Transformer architecture proposed in 2017, enabling training using large-scale clusters (hundreds of thousands of GPUs).

1-2. The 'Undergraduate' Analogy: Pre-Training as General Education

Zach compares a pre-trained model to 'a state of having acquired undergraduate-level general education.' This is a 'stage of learning very general knowledge broadly and shallowly,' and serves as the foundation for subsequent tuning. This metaphor intuitively conveys to not only AI researchers but also various business personnel 'why high-quality data is necessary.'

2. Functional Enhancement through Post-Training


2-1. SFT (Supervised Fine-Tuning) and RL (Reinforcement Learning)

In post-training, we mainly take two steps to elevate the model's style and response quality. First, in SFT (Supervised Fine-Tuning), we use high-quality input-output pairs to teach 'what constitutes a good response.' Subsequently, as RLHF (Reinforcement Learning from Human Feedback), we perform reinforcement learning based on human evaluation to efficiently improve quantitative evaluation metrics. Zach explained the learning process in an easy-to-understand way, stating, 'SFT is like training where you are taught by a senior colleague in your first job, and RL is closer to OJT through practice.'

2-2. The 'New Employee' Analogy: Task Adaptation and Style Adjustment

By comparing the model to a 'new employee assigned to their first workplace' and carefully teaching them how to proceed with specific tasks such as 'summarize this email' (whether to use bullet points or body text, how to emphasize important figures, etc.), useful results for practical work can be obtained.

3. Evolution of Reasoning Capabilities


3-1. Visualizing the Thought Process with Chain of Thought

In recent years, the ability to solve complex tasks has improved dramatically by utilizing 'Chain of Thought,' which sequentially shows the reasoning process as token output. Just like showing the intermediate steps of a problem-solving process at school, the model can also output 'why it reached that answer' step-by-step, making it easier for humans to verify and correct.

3-2. The Trade-off between Scale and Performance

On the other hand, Chain of Thought, which is used to increase reasoning accuracy, increases the computational cost during inference. Zach asked himself, 'Is it worth it?' regarding the benefits of performing multi-million dollar training using massive GPU clusters, and discussed the balance between cost and performance.

4. The Wall of Scaling and Field Insights


4-1. Expensive Cluster Training and Unexpected Debugging Challenges

In a training environment that bundles hundreds of thousands of GPUs, debugging challenges that were previously unimagined, such as damage from animals chewing on cables or minute hardware failures, occur. Real-world examples such as 'a raccoon chewed on a cable and caused a cell to break' were introduced, highlighting real-world operational risks.

4-2. The Creativity and Intuition of the '100x Engineer'

The existence of the "100x engineer", who goes beyond the "10x engineer", is the key to successfully designing models at an unprecedented scale. Zach stated that "research is as much an art as it is a science," emphasizing that human intuition and creativity support cutting-edge AI.

5. Open Source vs. Proprietary


5-1. The Rise of Small Models and Distillation Technology

Small to medium-sized open-source models from Moonshot and Meta offer the advantage of running in local environments while efficiently incorporating the knowledge of large-scale models through distillation technology. They are gaining attention for use cases such as execution on mobile and edge devices, as well as privacy preservation.

5-2. Large-Scale Model Competition and Technical Defenses

Meanwhile, large-scale models developed by companies like Google and OpenAI are adopting full-stack strategies to prevent distillation via APIs. Moves are underway to build technical barriers by strengthening RHF mechanisms and direct relationships with customers.

6. Application Areas and Market Trends


6-1. From Coding to Gaming, Finance, and Research

The most prominent success story is coding assistance tools, but diverse market needs are emerging, such as NPC generation for games, financial analysis, and the creation of consulting materials. Companies are focusing on industry verticals, such as Anthropic's model specialized for financial services and Mariner's UI control agents.

6-2. The Future of Task Delegation through Agentification

"Agents" are expected to be a new interface that allows specific tasks to be offloaded entirely to models, freeing humans from "routine work" such as research, document creation, and schedule management. It is expected that the level of automation will shift from Level 1 (assistance) to Level 5 (full autonomy) in the future.

7. Future Challenges and Outlook


7-1. Balancing Hallucinations and Factuality

As creativity increases, the risk of factual errors (hallucinations) rises, so developing methods that balance style and fact-checking is an urgent task.

7-2. Multimodal and UI/Control Interfaces

Multimodal models that handle video, audio, and sensor data in an integrated manner, as well as PC operation agents like Mariner, are the next frontier. The key will be improving "intent understanding capability" to accurately interpret user intent and complete tasks as intended.

Conclusion


What emerges from Zach's remarks is the perspective that "the evolution of AI models is not just a competition of scale, but a question of how to fuse human creativity with market needs." Refining knowledge gained in pre-training through post-training and opening up application markets with reasoning capabilities—the challenges taken on by companies like DeepMind will be the driving force that opens up a new dimension of AI. We will continue to keep a close eye on these developments.

Recommended Articles


Next Big Wave (Growth Stocks, Seeds of Ideas, Deep Dives into Trends)


Unraveling the secrets of growth, from seed startups to legendary companies


いいなと思ったら応援しよう!

この記事は noteマネー にピックアップされました

noteマネーのバナー