SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

The 'Future of Work' Revealed by Codex Usage Data: Why OpenAI Abandoned Chat

What is your current image of 'working with AI'?

Until recently, I thought it was normal to sit in front of a chat screen and have a one-on-one exchange, asking things like 'fix this text' or 'summarize these materials.' But lately, I feel that perception is gradually being updated.

Reading the economic research paper published by OpenAI on June 25th made that feeling clear. It is a report analyzing Codex usage data, co-authored by researchers from Columbia Business School, the Wharton School, and Duke University with OpenAI.

https://openai.com/index/how-agents-are-transforming-work/

The unit of work is shifting from conversation to delegation. The data proves it.



🔢 First, the case where the numbers speak too much

I will write down some of the numbers the paper presented as they are.

Within OpenAI, the ratio of tokens output by Codex is 99.8%. ChatGPT is used for only 0.2%.

The company that created ChatGPT barely uses its own flagship product. Instead, they use Codex—isn't that figure interesting?

The story is different when you look at individual users. Less than 1% of people have used Codex. But 'among those who do use it,' the output token ratio reaches 16.5%, which indicates a tendency to 'use it heavily once you start.'

Enterprise account users fall in between—17.3% of users utilize Codex, and 63.3% of the total output tokens for enterprises come from Codex.

This three-tier comparison was the most interesting part to me.


💬 The shift from 'asking' to 'doing'

There is an interesting description in the paper comparing it to past research.

In a usage survey of ChatGPT (Chatterji et al., 2025), nearly half of user actions were 'asking.' 'Summarize this,' 'Explain this concept.' Seeking information, learning, and advice were the central ways chat AI was used.

Codex is different. The paper defines Codex usage patterns not as 'consultation' but as 'production.' Debugging, refactoring, verifying changes, configuring applications, drafting documents, and data analysis. What is being done is having the AI perform the 'work itself.'

Reading this, I thought, 'Ah, this is the essential difference.'

Chat AI has been used as a 'smart advisor.' If you ask 'What do you think I should do?', it tells you the policy. But in the case of Codex, if you say 'Go ahead and do it,' it actually does it.

It sounds simple, but this is directly connected to the topic ofwhat humans spend their time on.


⏱ The fact that the delegation of '8-hour tasks' has increased tenfold

Specific figures are also available regarding the scale of tasks.

Between January and May 2026, the proportion of users who requested Codex to perform a 'task estimated to take an experienced human at least 8 hours' at least once increased by nearly tenfold.

70.2% of users have requested tasks taking over an hour, and 80.6% have requested tasks taking over 30 minutes.

Incidentally, the paper includes a note stating that these figures are based on the LLM reading conversation logs and estimating 'how long it would take a human.' In other words, it comes with the caveat that while not perfectly accurate, it should be viewed as a trend. It's an honest way of writing, isn't it?

However, the trend is clear. Users are increasingly handing over 'big jobs' to AI.


🏛 The Reversal Phenomenon: Legal Departments 'Writing Code'

This might be the most surprising part.

The paper includes a heat map visualizing 'what each department at OpenAI is doing with Codex.'

While it is natural for the engineering department to write code with Codex, what surprised me was that 25-31% of the work done by finance, marketing, and recruiting staff consisted of engineering or coding-related tasks.

In other words, people whose primary jobs should be finance or recruiting seem to be writing code or running data conversion and analysis tools through Codex.

I think the premise that 'technical work is done by technical staff' is slowly beginning to crumble.

Recruiters can ask Codex to 'convert this data into this format' without having to ask an engineer. That is the phenomenon actually happening within OpenAI.


🔄 'Power Users' Are Running Multiple Agents in Parallel

There is another figure that cannot be overlooked.

More than 10% of users, at some point during a week, are running three or more Codex agents simultaneously.

Furthermore, 26.6% of users are utilizing 'skills.' Skills refer to reusable instructions or complex workflow settings; in short, more than a quarter of all users are automating or templating 'how to operate Codex' itself.

In my view, this means that the number of users transitioning from the phase of 'using AI one-on-one' to the phase of 'managing it like a team' is definitely increasing.


📊 Explosive Increase in Output: 50x for Researchers, 13x for Lawyers

The figures written in the paper's abstract are also incredible.

Between November 2025 and June 2026, the median monthly output tokens generated by legal staff at OpenAI increased by 13 times. For research staff, it has increased by over 50 times.

I don't think it's quite right to interpret this as 'productivity increased 50-fold.' However, it is a fact that 'the volume of output that can be driven through AI has increased 50-fold.'

Thinking about what to do, making judgments, and verifying the results—these are becoming the human's job, while the model is shifting toward AI handling most of the actual hands-on work.


🤔 The meaning of 'mastering AI' has changed in the agent era

What struck me most after reading this paper is that the meaning of the phrase 'mastering AI' may have gradually changed.

In the chat era, 'mastering' meant the ability to write good prompts. It was a technique like 'be as specific as possible, assign a role, provide examples...'

But in the agent era, 'mastering' looks like a different set of skills.

Being able to design what to delegate and to what extent. Being able to monitor the agent's actions and correct its course at the right time. Running multiple agents in parallel and having humans handle the bottlenecks. And possessing an 'eye for quality' to evaluate the results achieved by AI.

To borrow the paper's terminology, these are the skills of 'delegation, monitoring, review, and coordination.'

It sounds like a manager, doesn't it? And that might be exactly what it is.


📌 What this change means

I believe the OpenAI Codex report showed something more than just numbers.

  • The fundamental difference: Chat AI is a 'conversation partner,' while Agent AI is a 'task executor'

  • Within OpenAI, the transition from ChatGPT to Codex is already nearly complete (99.8%)

  • A reversal is occurring where even non-technical staff are handling 'coding-related tasks' using Codex

  • What matters is not 'can you use it,' but 'can you design what to delegate'

It was impressive that the word 'Evidence' was included in the title of this paper. I think it should be read as OpenAI's assertion that we are no longer in the discussion phase, but that evidence has emerged.

Is a similar change starting in your workplace? If you are using it as a 'chat partner,' I encourage you to try an experiment where you 'hand over an entire task at once.'


📝 Thoughts while writing this article

The figure that surprised me most while reading the paper was that '10% of users run 3 or more agents in parallel at least once a week.' I don't have that sense yet. But the fact that the use case of 'managing multiple Codex agents' is already somewhat widespread made me want to rethink how I use it. Since the paper includes researchers from Columbia and Wharton, it was impressive that they included quite rigorous caveats. I appreciated the honesty in stating, 'The estimation of tasks exceeding 8 hours is based on LLM judgment, so please read it as a directional indicator.'


Miccell - Once you understand the mechanism, the world becomes more interesting.

いいなと思ったら応援しよう!