NVIDIA's Dilemma and Groq's Counterattack: New Trends in the AI Market
The recent generative AI boom has focused not only on "training," which requires massive computational resources, but also on "inference," which operates during actual user usage. Previously, it was often thought that "the computational volume required for training is overwhelmingly large," but as inference demands have expanded globally after training large-scale models, the view that "inference processing at the practical stage is likely to consume the most cost and energy" has been gaining strength.
In this article, we will organize the perspective of "GPU vs. LPU (an inference-oriented architecture proposed by Groq)" as told by Jonathan Ross (a former Google TPU development member), founder and CEO of Groq, the challenges faced by NVIDIA and other major players, and the large-scale infrastructure and social structural changes brought about by AI. Note that this article is a summary of the key points based on an interview with Mr. Ross.
1. Training and Inference: The Misunderstood Cost Structure
1-1. Is it true that "training is the most costly"?
Conventionally, in AI development, it was widely recognized that "the training phase of learning large-scale models consumes the most funds and computational resources." This is because there was an image that costs would be enormous due to using large-scale GPUs on the cloud to learn vast amounts of data over months.
However, Jonathan Ross says that even during his time at Google, he felt that "in fact, the computational volume required for inference is often 10 to 20 times that of training." For example, when services such as search engines and chatbots are actually provided to users, the number of times the trained model is called for "inference" becomes enormous.
"Your job is not to 'ride the wave,' but to get ahead of the next wave and position yourself."
(Jonathan Ross)
According to Mr. Ross, many companies still tend to think that "as long as you secure high-end GPUs for training, you will be fine," but he points out that the inference environment, where the number of users continues to increase, will bring the greatest cost burden in the future.
1-2. The dilemma NVIDIA faces
NVIDIA holds an overwhelming share of the GPU market, and in terms of training in particular, it still has performance and an ecosystem that "no one else can follow." On the other hand, if you try to handle inference usage with GPUs, there is a problem of high costs and power consumption due to an architecture that makes heavy use of expensive HBM (High Bandwidth Memory).
Mr. Ross points out, "For NVIDIA, it might be difficult to fully commit to low-cost solutions for inference because their business is already sufficiently successful by capturing training demand with high margins (profit rates)." In fact, looking at NVIDIA's business structure, they maintain a very high profit margin of 70-80%, and if they were to significantly lower prices specifically for inference, they might undermine their own high margins.
"In a sense, we at Groq might be the best thing for NVIDIA. They sell as many high-margin training GPUs as they want, and we take on the low-cost, high-volume inference demand—our interests do not conflict."
(Jonathan Ross)
2. The inference-specialized architecture that Groq aims for
2-1. What is an LPU?
The LPU (Language Processing Unit or Latency Processing Unit) architecture proposed by Groq has the following features:
Suppression of external memory bandwidth
A unique design that shares computational data between a large number of chips without using expensive HBM. As a result, communication between memories is minimized, and it is said that the power efficiency (power per token generated) is about 3 times better than that of a GPU.Design based on scale-out
"If necessary, connect many chips like a pipeline and flow data into them." While GPUs perform batch processing via external memory, etc., the Groq chip group performs calculations continuously like an "assembly line." This aims to achieve both high throughput and low latency.Overwhelmingly short introduction lead time
Conventional large-scale GPU cluster construction required advanced settings and tuning of network switches and racks. However, the LPU connection design is simplified, and the introduction is so rapid that there are cases where "tens of thousands of chips are set up in about 50 days from the contract."
2-2. Ingenuity in the business model
Groq is also unique in that it provides hardware for inference not by having "customers own it," but by providing it quickly through a "joint venture-like investment scheme." Large overseas capital, such as from Saudi Arabia, covers the cost of data center construction, and Groq prepares the hardware and takes the form of distributing profits according to inference usage.
"What is truly needed for inference is token generation efficiency when viewed in terms of 'performance x cost.' By specializing in that, we can lower the barrier to entry for users."
(Jonathan Ross)
3. Infrastructure Construction in the AI Era and the Fate of Capital
3-1. The Zigzag of Data Center Power Demand
Major IT companies such as Meta (Facebook), Google, and Microsoft are accelerating data center investments on the scale of tens to hundreds of billions of dollars annually. While moves to find regions with abundant power (in the GW range) and operate large-scale clusters are progressing, there are also frequent "all-show, no-go" projects where "land was bought, but upon opening the lid, necessary conditions like water resources were lacking, leaving the site non-operational."
Ross describes this as a "kind of bubble," while sounding an alarm: "In a few years, a massive amount of useless facilities that do not match actual demand will be created, but a few years after that, there is a possibility that true inference demand will explode and lead to shortages once again." His view is that this zigzag situation, where "investment alternates between excess and shortage," will continue for a while.
3-2. The Positioning of China and Europe
In China, AI is being promoted with large-scale national budgets, but Ross points out that "in China, where model and information regulations are strict, there is a high possibility of conflict with 'freedom of speech,' and there is an aspect where accelerating open foundations and R&D is difficult." On the other hand, because there is a vast amount of data and usage demand within the domestic market alone, and because they can invest heavily in hardware and power, the scenario of independent evolution cannot be ruled out.
Meanwhile, Europe has the challenge of a weak culture and ecosystem for startups to take risks, as it places too much emphasis on regulation and personal information protection. Ross states, "Even in Europe, if there were an environment like a 'special zone where it is easy to take bold risks,' there is plenty of room for talent and startups to concentrate and for innovation to blossom."
4. Organizational Building and Scaling: The Groq Way of Thinking
4-1. The Trade-off Between Growth Speed and Team Size
In its seventh year, Groq finally jumped on the wave of major demand and felt the "Product-Market Fit (PMF)." Ross looks back on past hardships as follows.
"(It took) 7 years (until the product fit). There were times when funds were running out, and I asked employees for a 'GROQ Bond' to exchange their salaries for stock. 80% of the employees agreed, and we held on just before running out of funds."
Currently, while achieving "exponential" scale growth, such as expanding from thousands to tens of thousands of LPUs at once, the team is kept to a minimum at the 300-person scale. This is because of the management decision that "everything from chip design, cloud, and compilers is developed in-house, but if you increase the number of people too much, communication costs will skyrocket and productivity will drop."
"Instead of increasing the number of people, we respond with process automation and softwareization. What is important is to maintain speed without sacrificing quality."
(Jonathan Ross)
4-2. Sharing Vision and 'Coins'
At Groq, they seem to distribute 'challenge coins' to employees, with 'Achieve 25 million tokens per second (25M TPS)' engraved on them to share explicit goals. It is a unique mechanism where if someone thinks during a meeting that 'that discussion does not contribute to achieving the goal,' they can express their opinion just by lightly tapping the coin on the desk. Such easy-to-understand symbols are said to enhance organizational cohesion and cut away miscellaneous discussions.
5. Future Outlook: What Inference Optimization Brings
5-1. The Era Where 'Inference' is the Protagonist
In the world of large-scale LLM (Large Language Model) models, cost and energy efficiency at the "time of actual operation on servers or in the cloud (inference)" will be the final deciding factor, rather than at the "time of developing the model (training)." Especially in the current situation where the number of users using generative AI is exploding, it can be said that inference optimization significantly influences "speed," "quality," and "price."
Groq's LPU design anticipates this wave of the "inference revolution" and is thoroughly aiming for "low cost and low power consumption per token generation." GPU forces like NVIDIA, which have tremendous strengths in training, will naturally challenge the inference domain as well, but it is not easy to lower costs while maintaining high margins. In the future, the complementary use of "GPU (learning) x LPU (inference)" may become the standard for practical work.
5-2. How to Protect "Human Decision-Making" in the AI Era
Another perspective Mr. Ross emphasizes is the "danger of humans surrendering their decision-making power to AI." If AI becomes capable of taking over too many tasks, we may lose the motivation to think or make decisions for ourselves, compromising by saying, "AI gives me good enough results, so that's fine." This is a major risk that could undermine not only individual creativity but also the vitality of society as a whole.
Of course, there are significant benefits to using AI in fields like medicine and law to reduce human error, but it will be necessary to reset the boundaries of "how much to leave to AI and where human responsibility begins." In fact, considering the current situation where AI model "hallucinations" (incorrect answers) and ethical issues have not yet been resolved, it is true that human checking and final judgment remain indispensable.
In this article, based on an interview with Jonathan Ross, CEO of Groq, we have summarized the importance of the AI inference field and the potential of new chip business models. The key points are as follows.
The era has arrived where inference requires massive computational power
Inference, the phase of practical use for large-scale models, requires far more computational resources and costs than training.NVIDIA maintains an advantage in "training," but faces a dilemma in inference specialization
While high-margin GPUs are in excess demand for "training purposes," there is a strong demand for cost reduction in inference applications, making the same architecture potentially disadvantageous.Groq's LPU architecture and sense of speed
Achieving high throughput and low power consumption with a multi-chip configuration without using HBM. Furthermore, rapid deployment through innovative investment schemes.Infrastructure investment bubble and the wave of scale
While investment in data centers and chips is overheating globally, constraints and imbalances in power and water resources are beginning to emerge. There is a risk of alternating surpluses and shortages in the coming years.Rethinking the human role in the AI era
The more inference efficiency increases, the higher the possibility that all work and decision-making will be replaced by AI. While convenience increases, there is a concern that humans will lose their agency.
Mr. Ross repeatedly emphasizes that "while people's lives will become 'richer' due to the large-scale development of AI technology, there is also a risk that the motivation to think and act for oneself will be eroded." The future where inference-specialized hardware is ready and anyone can use large-scale models is just around the corner. At that time, shouldn't what we cherish be the agency to continue making "choices that are meaningful to ourselves"? While enjoying the efficiency brought by next-generation chips that handle massive inference requests and the latest cloud infrastructure, how will we maintain the sense of "deciding with human hands" in the end? A major question for the co-evolution of technology and society is now before us.
Related Articles

