SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

The Three Frontiers Where AI Models Compete—Google Cloud's Vertex AI Lead Discusses a New Framework for Performance, Speed, and Cost

AI model capability comparisons are often discussed solely in terms of "intelligence." However, Michael Gerstenhaber, who oversees Vertex AI at Google, presents a completely different perspective. He argues that models are competing simultaneously across three distinct frontiers—"raw intelligence," "response speed," and "scalable cost"—each catering to different use cases. As AI adoption enters a mature phase, this framework offers practical insights for everyone from engineers to executive leadership.


1. The Strength of Google Cloud's Vertical Integration


Gerstenhaber became the head of Vertex AI at Google after spending a year and a half at Anthropic.When asked about his motivation for the move, he says:

“I believe Google is the only company in the world that owns everything from the interface to the infrastructure layer. We can build data centers, procure power, and even own power plants. We have proprietary chips, proprietary models, an inference layer we can control, an agent layer, APIs for memory, interleaved code writing capabilities, agent engines that ensure compliance and governance—and even conversational interfaces like Gemini Enterprise and Gemini Chat. I decided that this vertical integration would be a strength.”

Many Vertex AI users are engineers building their own applications. What they seek are agent-based patterns, agent-based platforms, and access to the world's most advanced model inference capabilities. The structure is such that companies like Shopify and Thomson Reuters build applications within their respective domains, with Vertex AI functioning as the underlying infrastructure.

2. The Three Frontiers: Performance, Speed, and Cost


2-1. The First Frontier: Raw Intelligence

The first frontier Gerstenhaber mentions is that of "pure intelligence." Models like Gemini Pro fall into this category.

“Think about writing code. You want the best code. It’s okay if it takes 45 minutes, because you have to maintain that code and deploy it into a production environment. You just want to do your best.”

In this domain, response time is a secondary issue, and output quality is the top priority. Complex system design and the generation of deliverables with long-term value are typical use cases for this frontier.

2-2. The Second Frontier: Latency

The second is maximum intelligence under latency constraints.

“If you need to know how to apply a return policy immediately in customer support, it’s meaningless if the correct answer comes after 45 minutes. The person will have hung up. That’s why you need the model with the highest intelligence within the latency budget.”

Handling airline seat upgrades, insurance claim decisions, and real-time price quotes—these are all scenarios that require a balance of both "speed" and "accuracy."

2-3. The Third Frontier: Scalable Cost

The third frontier is handling large-scale and unpredictable volume. Gerstenhaber suggests this is the most overlooked dimension.

“When a platform like Reddit tries to moderate content across the entire internet, even if they have a generous budget, they cannot take corporate risks when the scale is unknown. You don’t know how many harmful posts there are today, and you don’t know about tomorrow. So, they have no choice but to choose the model with the best intelligence within their budget in a way that is scalable for an unlimited target. That’s where cost becomes extremely important.”

This framework updates the conventional argument that simply "the smartest model wins." Depending on the use case, what needs to be maximized differs fundamentally, and the criteria for model selection change accordingly.

3. Why the Spread of Agentic AI is Lagging: Unprepared Infrastructure


3-1. The Reality That "Demos Are Great, But Production Is Different"

While AI agents are garnering attention, many feel that actual business adoption is not progressing as quickly as expected. Mr. Gerstenhaber speaks candidly about the reasons for this.

"This technology is fundamentally only two years old. Much of the infrastructure is still missing. Patterns for auditing what agents are doing have not been established. Patterns for data authorization for agents are also not yet in place. Work is required to deploy these patterns into production environments, and production environments are always a lagging indicator of technical capability."

3-2. Why Software Development is Leading the Way

On the other hand, agentic AI is rapidly taking hold in the field of coding assistance. Mr. Gerstenhaber analyzes the reasons for this structurally.

"It spread particularly quickly in software engineering because it fits well into the existing software development lifecycle. You can fail safely in a development environment. Then, you promote it to a test environment. At Google, code requires two people to review and approve it before it goes into production. Because there is a process where humans are in the loop, implementation risk is kept low."

To popularize agentic AI in fields other than coding, it is necessary to establish similar patterns of "safe spaces for failure" and "human approval processes" for each job function and industry.

When choosing an AI model, the question of "which model is the smartest" is no longer sufficient. Only after defining "which frontier to compete in" can one select the appropriate model and architecture—Mr. Gerstenhaber's framework helps organize that order of thinking. While the practical challenge of infrastructure development remains for the spread of agentic AI, its rapid adoption in software development is a leading example that demonstrates its potential for expansion.

Recommended Articles


Next Big Wave (Growth Stocks, Seeds of Ideas, Deep Dives into Trends)



いいなと思ったら応援しよう!