Vertical AI②: The Evolution of Multimodal AI and Its Impact on Vertical AI
In recent years, Large Language Models (LLMs), which have played a central role in generative AI, have primarily focused on applications for text-based tasks. However, the emergence of "multimodal AI," capable of handling multiple modalities (data formats) such as audio, images, and video, is dramatically expanding the range of challenges that can be solved by Vertical AI (AI for specific industries).
Article:
This article provides an overview of how the advancement of new multimodal technologies, including audio, images, and video, is impacting the industrial sector and what kind of use cases are actually emerging. We will also provide concrete explanations of real-world implementation examples and the potential of AI agents, presenting new approaches to complex business operations.
1. The Evolution of Multimodal AI and Its Impact on Vertical AI
1-1. Expansion Brought by Multimodal AI
Conventional AI applications have relied primarily on text-based input, and the scope of generative AI utilization has been dominated by text-centric tasks such as "drafting emails," "creating website copy," and "drafting contracts." However, in recent years, AI models capable of handling diverse data formats such as "audio," "images," and "video" have been appearing one after another.
Multimodal models like GPT-4 have become capable of understanding not only text but also images and audio, and performing reasoning based on them. Furthermore, with the addition of "visual" and "audio" processing capabilities to these models, the range of tasks has expanded, setting the stage for Vertical AI (industry-specific AI) to enter larger markets.
1-2. Characteristics of Multimodal Architecture
Multimodal models that have appeared in the last year have improved significantly in areas such as "contextual understanding," "suppression of hallucinations (incorrect answers)," and "reasoning ability." For example, in fields such as speech recognition, image processing, and speech synthesis, there are instances where performance matches or exceeds human accuracy.
While conventional text-only Large Language Models primarily generated the next word by predicting patterns in training data, more models are now incorporating "chain-of-thought" to reason through more complex steps. Such architectural innovations have made complex tasks that process audio and images simultaneously possible.
2. New Use Cases Created by Audio-Enabled AI
2-1. Advances in Audio Models: From Cascaded to Native Audio Models
Previously, cascaded architectures were mainstream, involving "Automatic Speech Recognition (ASR)" → "Response generation by LLM" → "Text-to-Speech (TTS)." However, recently emerged "audio-native models" (e.g., OpenAI's Realtime API, Kyutai's Moshi, etc.) take input audio directly and generate responses while understanding the user's voice and even emotions.
Such models have very low latency, with cases reported of achieving responses in under 500 milliseconds. Furthermore, a major strength is their ability to respond while capturing the speaker's emotions, tone, and intent, making interactions more natural. It is expected that the quality and speed of conversational AI applications will improve significantly in the future.
2-2. Use Cases for Audio
There are four main use cases for utilizing audio as follows.
Call transcription and operational efficiency
・"Abridge" for the medical field records conversations between patients and doctors in real-time, and generates medical reports, prescriptions, and suggests billing codes as needed. This allows doctors to significantly reduce manual recording work and devote more time to patient care.
・"Rillavoice" provides a mechanism to automatically transcribe conversations in the home services sector and analyze sales representative calls for training purposes. This has made it possible to provide effective feedback without actually having to accompany the sales representative.Handling inbound calls
By using audio agents specialized for specific industries (e.g., car dealerships or home repair services), companies can handle customer inquiries outside of business hours or when staff are unavailable, automating appointment scheduling and price quotes. High-performance conversational AI is capable of responses that are incomparably more natural than previous "voice bots," and there are increasing cases where this leads to orders without missing customer intent.Streamlining customer support
Conventional IVR (Interactive Voice Response) tended to cause high user stress because it could only react to specific keywords. However, the latest audio agents have high natural language understanding capabilities and can flexibly handle complex inquiries. There is an increasing trend of automating FAQ-level responses and shifting customer support staff time toward high-level problem solving.Automating outbound calls
In sales and recruiting scenarios, mechanisms are being introduced where AI agents automatically contact potential customers or candidates and hand them over to staff once a response indicating interest or a desire to hear more is obtained. Such automation offers benefits like "being able to approach many leads at once" and "allowing staff to focus on necessary follow-ups." However, consideration for regulations and compliance (e.g., anti-spam and robocall regulations) is extremely important, and measures such as ensuring contact is made only with those who have opted in are required.
3. The Frontline of Vertical AI in the Image and Video Domain
3-1. Breakthroughs in Vision Models
As for models that handle images (still images), image understanding functions implemented in GPT-4 (e.g., GPT-4V) and Gemini 1.5 Pro (Google), which can process both images and video, have emerged. These models hold contexts of millions of tokens and are becoming capable of understanding and comparing multiple images across the board. If the inference cost and inference speed of models improve further in the future, it will be a significant tailwind for application developers.
3-2. Four Representative Use Cases
Data Extraction
This is for extracting necessary information from unstructured documents such as images and PDFs. For example, "Raft", provides functionality to extract critical information from invoices used in the logistics industry and input it into ERP systems. Furthermore, it automates the preparation of documents required for customs declarations. By not just transcribing text but utilizing the data to automate subsequent steps, it contributes to reducing work time and improving processing accuracy.Streamlining Visual Inspection
An example is the automation of construction site inspections using "xBuild". It analyzes photo data of roof or building damage to create reports on necessary repair scopes and methods. This shortens the process of sharing with insurance companies to obtain approval, thereby smoothing the insurance claim procedure. Solutions are also emerging that incorporate AI into quality checks for architectural drawings to reduce human oversight.Automation of Design (2D/3D Design)
AI platforms for the architecture and engineering industries are increasing rapidly. For instance, "Snaptrude", automates the creation of 3D architectural designs and handles detailed tasks such as piping and structural design. Engineers are freed from tedious work and can focus on more creative design and planning tasks. It is also said to contribute to higher order-winning rates by enabling speedy cost calculations and plan revisions during project proposals.Video Analytics
Although models capable of video understanding are less mature than those for still images, they are already being demonstrated in manufacturing and factory floor safety monitoring. Tracking and classifying specific objects, as well as searching within videos, are beginning to reach a practical stage, and more advanced applications in collaboration with robotics are expected in the future.
When launching image and video-based solutions, it is important to pay attention to their compatibility with existing workflows. Even if a system can generate advanced 3D models, if it is difficult to integrate with major software that users are accustomed to, such as Revit, it may not be adopted. Rather than attempting complex functions from the start, it is also important to adopt a strategy of providing products that balance implementation costs and ROI even with single functions, and then expanding them in stages.
4. The Potential of AI Agents
4-1. Automation Through Evolving Agents
While the term "AI agent" was temporarily overhyped, mechanisms to control the thought process are now being developed to execute complex multi-step tasks without errors. Reasoning-specialized models researched by OpenAI (such as o1) adopt a design that takes time for chain-of-thought reasoning to derive answers, rather than simply predicting the "next word" contained in training data. This is making it possible to respond more accurately and flexibly to multi-step problems.
4-2. Concrete Use Cases and Future Outlook
Sales and Marketing
Many solutions are emerging where AI agents research prospects (leads) and automatically generate personalized emails. By collecting corporate information and news via web searches and extracting appeal points tailored to the person in charge to make contact, the "research + outreach" portion can be significantly streamlined. Sales representatives can focus only on "warmed-up leads" to improve closing rates.Automation of Negotiation Tasks
As an example, "Pactum", provides a mechanism where AI agents negotiate supply chain contracts and business terms on behalf of the company. When a user sets key metrics important to their company (price, delivery date, contract period, etc.), the agent can negotiate with multiple suppliers in parallel to extract optimal conditions. Similar methods are expected to be applied to sales promotions and bulk purchase negotiations.Security and Audit Support
Corporate security teams are plagued by vast amounts of alerts and logs every day. By introducing AI agents here, it is possible to realize operations where low-risk alerts are first automatically investigated, and reports are summarized for the person in charge, detailing "what kind of methods are conceivable" and "possibility of damage." By having humans intervene in stages for critical events to make final decisions, the overall efficiency of security response can be increased.
In the future, agents capable of executing tasks across multiple modalities, including voice, images, and even video, will emerge. In particular, there is room to improve performance through optimal model combinations and design ingenuity, even in environments without large-scale data or massive computational resources. For startups, it is more effective to use such architectural differentiation and optimization for specific use cases as weapons rather than competing with major companies simply on the volume of learning data or computing power.
With the rapid development of multimodal AI that supports voice, images, and video, moving beyond text-centric applications, automation is progressing in a variety of business areas that were previously unimagined. As voice-native models have reached a practical level, front-office operations such as call centers, sales, and customer support are undergoing significant transformation. Meanwhile, due to the dramatic evolution of image recognition technology, industries that handle real assets, such as construction, manufacturing, and logistics, are also starting to utilize AI one after another.
Furthermore, the potential of AI agents is increasing in tasks that require "multi-step reasoning," such as complex negotiations and research work. Models that emphasize reasoning, which go beyond the framework of traditional LLMs, are emerging and are beginning to realize more flexible decision-making and task automation.
In any field, foundation models are expected to rapidly become commoditized. Therefore, vertically specialized unique datasets and integrated workflows hold the key to differentiation. This wave of Vertical AI is permeating industries in ways unimaginable a few years ago and is fundamentally changing how we work and how business is conducted.
Now, as we leap from the stage of "automating text generation" and transition to a new phase called multimodal AI, we are required to build Vertical AI with an eye toward the next innovation and adopt strategic approaches to maximize its potential.
Related Articles
