SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

AI Wins Through "Self-Made Data": Learning "Quality Over Quantity" and Data Sovereignty Strategies from Turing and Fyxer

The evolution of AI is no longer determined solely by model size or computational power.
In recent years, "how to independently collect and generate high-quality data" has become the primary axis of competition among AI startups.

In particular, companies developing multimodal AI, such as image, video, and audio, are moving away from relying on existing web scraping or inexpensive crowdsourcing, and are accelerating efforts to prioritize "primary data filmed and created in-house."

A symbolic example of this is the vision model development project being advanced by the AI company Turing.


1. Turing: An experiment to "make AI understand manual work"


At Turing, they are running a project where professionals who use their hands, such as artists, chefs, and electricians, wear GoPro cameras to film their daily tasks. Artist Taylor (a pseudonym) says:

"We wake up in the morning, put a camera on our heads, and start our daily routine. We make breakfast, wash dishes, and then each of us paints or creates sculptures. We were contracted to film five hours of footage every day, but in reality, it took seven hours including breaks."

The filming was physically demanding, to the point where "red marks were left on my forehead when I took it off."
However, the footage obtained in this way is valuable learning material for AI. Sudarshan Sivaraman, Chief AGI Officer at Turing, explains it as follows:

"To make AI understand 'how to perform a task,' we have no choice but to collect data from various professions. We film diverse data, focusing on blue-collar work, to let the AI learn continuous real-world actions."

2. From web scraping to "handmade data"


In the past, it was common in AI development to train models by scraping large amounts of images and text from the internet.
However, as copyright and quality issues have surfaced, "training with proprietary data" is becoming the mainstream trend since 2024.

About 80% of Turing's data is said to be synthetic data generated based on original footage.
But Sivaraman emphasizes:

"If the source data is low quality, no matter how much you synthesize it, the result will be bad. That is why the 'original footage' is the most important."

The reason AI companies "create data with their own hands" is not just for quality improvement.
It is because data that others cannot easily imitate becomes a "moat."

3. Fyxer: The value of "human-led learning" shown by email AI


The same trend can be seen in Fyxer, which develops email automated processing AI.
Founder Richard Hollingsworth says he realized this in early experiments.

"It is the 'quality' of data, not the 'quantity,' that determines performance."

At Fyxer, they hired many skilled executive assistants to train the AI.
There was even a time when there were four times as many assistants as engineers.
Human business judgment was essential to determine whether an email "should be replied to or not."

Through this approach, Fyxer's AI evolved from a mere classification model into a "judgment support AI" that understands context.
Hollingsworth says:

"Anyone can incorporate an AI model. However, data nurtured by humans cannot be easily reproduced. Our greatest strength is the high-quality data created by humans."

4. The essence of competitive advantage lies in "data collection capability"


Both Turing and Fyxer share a common culture of "crafting data in-house."
Now that model performance has reached a plateau in the AI industry, "data sovereignty" is being prioritized due to three factors:
・Maintaining the quality of synthetic data
・Avoiding copyright and ethical risks
・Proprietary data for model differentiation

Whether it is "AI that learns from video" like Turing, or "AI that mimics humans" like Fyxer, the underlying philosophy is the same.
That is—the future of AI will be determined by who possesses the highest quality data.

5. Conclusion — AI Development is Becoming a "Data Production Industry"


Competition in the AI era is shifting from algorithms to the "data production industry."
Artists filming videos and secretaries sorting emails are, in fact, "raising" the next generation of AI.

This change is not merely a technical trend, but symbolizes a new relationship between AI labor and creativity.
For AI to understand the world, we humans must continue to create "real-world data"—that is the beginning of the next AI revolution.

Recommended Articles


Next Big Wave (Growth Stocks, Seeds of Ideas, Deep Dives into Trends)



いいなと思ったら応援しよう!