SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

From 'Using' to 'Managing' AI Agents: Two Concepts Executives Should Know

Continuing from last time.


Inference Economics

Why do companies that have introduced AI agents say it 'didn't go well'?

In many cases, it is not a technical problem.
It is a management problem. Whether you treat agents as 'tools' or as 'team members.'
This difference is determining the success or failure of AI utilization in 2026.

The first thing that surprises you when you introduce AI agents is the cost.

According to Gartner's March 2026 analysis, AI agents consume 5 to 30 times more tokens than regular chatbots. In multi-agent configurations where multiple agents work in coordination, this increases even further.

While the cost may be small for a single task, it can reach tens of millions of yen if executed tens of thousands of times per month.

Gartner also predicts that 'by 2027, 40% of AI agent projects will be canceled due to cost overruns.'

https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027

Overseas, this problem has been named 'Inference Economics'.
And a new professional field called 'Inference FinOps' has even emerged.

This applies the same concept as FinOps, which manages cloud costs, to AI inference costs. As of 2026, AWS and Azure have implemented dedicated Inference FinOps dashboards in their management consoles, effectively acknowledging that AI inference costs have become an 'item that should be managed as an expense' by the cloud giants themselves.

Think of cost in terms of 'optimization' rather than 'minimization'

What is important here is not to simply try to cut costs.

There is an optimization triangle.
Cost, Latency, and Quality—the balance of these three.

For example, in complex tasks, the accuracy of a single LLM call is moderate; to achieve accuracy, multiple inference loops are required, which inevitably increases cost and latency. This is why Deep Research and reasoning models take time and cost money.

Simple inquiries should be handled quickly by small models.
Tasks requiring complex judgment should be handled carefully by large models. This 'routing design' has become a core skill in AI agent management.

It is not about 'using it because it's cheap' or 'using it because it's high-performance,' but constantly asking, 'What is optimal for this task?' This is the job of an AI operator.

There is no improvement without evaluation. The advantage of Evals

Another thing being discussed overseas as the 'greatest competitive advantage' is 'Evals (Evaluation)'.

If you cannot systematically measure the quality of AI agents, you will not know what or how to improve.

In 2026 agent evaluations, more granular metrics are being used than the traditional 'accuracy rate'.

  • PlanQualityMetric: Evaluating the quality of task planning

  • ToolCorrectnessMetric: Whether the correct tool is being used with the correct parameters

  • Task Completion Rate: Whether it was able to execute autonomously to the end

  • Hallucination Rate: Whether it is generating misinformation

Evals are not just about 'measuring'.
Because you can measure, you can improve. Because you can improve, you can trust. Because you can trust, you can delegate.

Once this loop begins to turn, it creates a gap that competitors cannot close.

AI agents require 'hiring, evaluation, and development'

I will summarize what I wanted to convey throughout this series.

AI agents are no longer 'tools to be used,' but 'subjects to be managed'.

Just like human team members, agents need to be given roles and permissions, set with goals, evaluated for performance, and continuously improved.

'People who master AI' vs. 'People who build systems to keep AI running'
This gap is gradually widening in 2026. The divide is already visible overseas.

There are not many people in Japan who can speak to this yet.

いいなと思ったら応援しよう!