When a user types a sentence into a chat window, a fluent response appears on screen within tens of milliseconds. This seemingly lightweight process, however, turns into a bewildering string of numbers at the billing page: how many million input tokens, how many million output tokens, how long the context window is, whether the cache was hit. Why does an intangible answer get billed by volume? Why can’t AI companies sell it like traditional software—one license, unlimited copies?
There are at least three different ledgers tangled together here: production cost—what a token actually consumes in electricity, chips, data center space, and engineering resources; market price—the per-call rate that model vendors quote externally, which also bakes in competition, margins, and service tiering; and task cost—how many tokens a user actually burns to complete a job. When the three are conflated, confusion follows: “Electricity isn’t that expensive, so why is AI still so pricey?” or “Models got cheaper, so why did our enterprise bill go up?”
AI certainly has a software side: algorithms can iterate, models can be compressed, services can be distributed through APIs. But it has another side that is becoming increasingly conspicuous: every generation requires real computation, depending on real electricity, VRAM, networks, and data center space. Typical SaaS costs get amortized rapidly as copies multiply; AI inference, by contrast, always carries a recurring marginal cost. A token looks like software, but underneath, it increasingly resembles heavy industry.
The Five-Layer Structure: Where the Money Goes
In January 2026, Nvidia CEO Jensen Huang first described AI infrastructure as a “five-layer cake” at the World Economic Forum in Davos: from bottom to top—energy, chips, infrastructure, models, and applications. He has since used this framework repeatedly at CES 2026 and GTC 2026. In the book The Token Economy, the authors describe this system as a “real-time cognitive production line”: user inputs, context, tool instructions, and enterprise knowledge bases serve as raw materials; electricity, chips, and inference time serve as energy; model architectures, caching mechanisms, routing strategies, and workflow orchestration serve as the manufacturing process. Unlike an industrial product that can be built and warehoused, a token is typically produced, delivered, and consumed at the instant of invocation.
Breaking down costs from a large-scale commercial inference perspective, the five layers distribute roughly as follows:
| Layer | Core Components | Share of Total Token Production Cost |
|---|---|---|
| Energy | Electricity, cooling, network transmission | 10%–20% |
| Chips | GPU/AI accelerators, HBM, chip-to-chip interconnect, CPU | 40%–55% |
| Infrastructure | Data center facilities, power distribution, liquid cooling, optical modules, InfiniBand networking, scheduling and operations | 15%–20% |
| Model | Training cost amortization, inference framework optimization, model serving orchestration | 10%–15% |
| Application | Interface, integration, compliance, content safety | 5%–10% |
Note: These percentages are not standard financial statements from any particular company, but rather a decomposition perspective on large-scale commercial inference scenarios. Whether you build or rent, use frontier or economy models, train or infer, run long or short tasks—each shifts the weighting of every layer. But one ranking remains relatively stable: the three hardware-related layers typically constitute the bulk of costs, roughly 70% or more.
The energy layer, while intuitive, is not the largest item. At U.S. industrial electricity rates of $0.06–$0.08 per kilowatt-hour, a 500-megawatt data center runs an annual power bill exceeding $300 million. Yet SemiAnalysis’s inference TCO model shows that facility leasing and electricity combined typically account for less than 20% of total cost of ownership.
The chip layer is where the real cost sits. A single Nvidia B200 costs $30,000–$40,000, which typically already includes the cost of high-bandwidth memory (HBM) in the package. From a bill-of-materials perspective, HBM and advanced packaging together account for roughly two-thirds of high-end chip manufacturing costs. Model parameters must reside in HBM to enable efficient inference; HBM capacity and bandwidth directly determine inference speed and concurrency ceilings. And global HBM capacity is controlled by an oligopoly of three players: South Korea’s SK Hynix holds roughly 57%, South Korea’s Samsung Electronics holds 22%, and Micron Technology holds 21%. The AI demand surge combined with this oligopolistic structure has driven memory’s share of hyperscaler capital expenditure from roughly 8% in 2023–2024 to approximately 30% by 2026—almost entirely driven by HBM.
The model layer is the “softest” layer and the one falling fastest. Take MoE (Mixture of Experts) architectures as an example: DeepSeek V4 has a total of 1.6 trillion parameters, yet each inference activates only about 49 billion parameters (roughly 3%)—equivalent to a giant consulting firm with 16,000 employees dispatching only a ~500-person elite team for each engagement. Distillation techniques allow small models to replicate large-model capabilities at roughly 1% of the inference cost. The common thread across these algorithmic optimizations: no hardware replacement needed—just software and framework updates, and costs drop another notch.
Dual Clocks: Algorithms Fast, Physics Slow
Once the five-layer cost structure is understood, the most paradoxical aspect of the token economy becomes clear: algorithm-side prices are falling rapidly, while physical-side supply remains tight. The Token Economy calls this the “dual clock”—algorithms and software run on a fast clock; chips, data centers, and power run on a slow clock.
How fast is the fast clock? Over the past three years, inference costs at equivalent capability levels have fallen 99.5%. Stanford’s 2026 AI Index Report estimates a 250-fold reduction over three years. MoE, attention optimization, distillation, speculative decoding—these technologies iterate on a daily and weekly cadence, cutting costs another tier without any hardware changes.
How slow is the slow clock? The physical world imposes four constraints. First, data center expansion takes 18 to 36 months, and many projects are running late. In 2026, the U.S. had roughly 16 gigawatts of data center capacity planned for delivery across about 140 projects; as of April, only about 5 gigawatts were under construction. The most emblematic case is OpenAI’s Stargate project: announced with great fanfare at $500 billion, construction has stalled due to grid interconnection queues and transformer delays.
Second, transformer delivery cycles average 2.5 years. As of Q2 2025, the average lead time for standard power transformers in the U.S. was 128 weeks, compared with just 6–8 months before the pandemic. These bottleneck components account for less than 10% of total data center construction costs—a $2 billion campus could sit idle for over a year because of a $40 million transformer order. In some ways, money isn’t the bottleneck; time is.
Third, HBM supply competition and labor shortages. The three oligopolists’ combined capacity is growing at only 30%–50% annually, unable to keep pace with AI demand that is doubling. Fourth, structural power shortages. International Energy Agency data shows global data center electricity consumption will double between 2024 and 2029, from roughly 415 terawatt-hours to over 900 TWh. Goldman Sachs projects the U.S. will face a 45-gigawatt power deficit for data centers by 2028. As of end-2025, the U.S. grid interconnection queue still stood at 2,060 gigawatts, with median interconnection timelines for new projects approaching five years—Google has even reported some sites facing 12-year waits.
With two clocks running simultaneously, the token market produces a unique landscape: prices can fall quickly due to algorithmic breakthroughs while physical supply remains tight; utilization gains can amortize costs, but once capacity hits its ceiling, new GPU, facility, and power contracts cause costs to jump in step-function increments.
Nvidia’s Three-Front Offensive
It is precisely within this structural tension—slow hardware, fast software—that Nvidia has chosen to consolidate its core position in the AI industry chain with unusual intensity of action. Under multiple pressures—chip price increases, rising data center construction costs, and increasingly negative public sentiment—Jensen Huang has opted to accelerate rather than hold position.
According to The Information, Nvidia recently announced a lead investment in a multi-billion-dollar funding round for AI search and agent company Perplexity, committed $6 billion to license and acquire portions of AI coding startup Poolside’s software and talent, and continues to inject capital into data center land and power developers. Three fronts advancing simultaneously, all pointing to one strategic objective: locking in Nvidia’s hardware-software bundled solutions during the early stages of AI infrastructure buildout and seizing first-mover advantage.
At the data center level, Nvidia made a minority equity investment in Cloverleaf Infrastructure, a company focused on coordinating land, power, and construction resources to provide “build-ready” sites for data centers. Additionally, Nvidia ultimately invested in Lancium—the power developer behind OpenAI’s Stargate facility in Abilene, Texas—after OpenAI had previously explored investing in or acquiring Lancium itself. The investment logic is clear: embed Nvidia’s hardware-software bundle into data center construction processes as early as possible, while hedging against potential chip oversupply risks arising from power delivery delays and cost overruns.
At the model level, Nvidia has agreed to pay $6 billion by the end of next year to license portions of Poolside’s software and recruit 100 employees from the company. The deal signals Huang’s commitment to building Nemotron into a frontier-grade open-source model, positioning Nvidia not merely as a chip supplier but as a significant participant in the AI model ecosystem—digging a deeper moat across the entire AI value chain.
Nvidia’s dominant position is also reflected in its pricing power. Server vendors have informed customers that certain flagship Blackwell and Rubin chip systems will see price increases of roughly 17% upon delivery next year. That increase would add at least $5 billion to the construction cost of a gigawatt-scale data center, with cost pressure passing directly through to cloud service providers and data center developers.
From the perspective of the South Korean market, this wave of price increases is a near-term positive for memory chip makers. South Korea’s SK Hynix and South Korea’s Samsung Electronics see their pricing leverage strengthened accordingly—Nvidia’s willingness to absorb server price increases signals that AI demand is robust enough to digest higher component costs. In the medium to long term, however, a data center deploying thousands of servers will see 15%–17% price increases compound into significantly higher total cost of ownership, and margin pressure on cloud providers may resurface.
Nvidia’s expansion is not without headwinds. U.S. public sentiment toward AI and the data centers behind it is turning increasingly negative, and the economic arguments advanced by anti-data-center activists have proven difficult to refute effectively. Meanwhile, the pace of AI commercialization has been slower than expected. OpenAI CEO Sam Altman recently admitted he underestimated the difficulty of AI disrupting the software industry: “The economy has too much inertia; people keep doing the same things.” He compared the current phase to the eve of the iPhone’s debut, calling it largely a “product-level failure.” By contrast, Anthropic’s growth momentum is stronger—its banking advisors have told potential investors that the company’s IPO valuation could reach $2 trillion, with a raise potentially exceeding $100 billion, and the IPO could launch as early as this fall.
Returning to the original question: where exactly does the money go when producing a token? The answer will likely disappoint anyone seeking a “standard answer.” Token market prices will change, models will be replaced, and cost structures will shift depending on task type and deployment configuration—there will never be a “fixed price” for a single token. But one thing is certain: every token requires real energy, real chips, and real infrastructure to produce.
People have grown accustomed to the internet narrative: write the code once, serve unlimited users, marginal cost approaching zero. The token economy drags the physical world back to the center of the story—when intelligence services become measurable, callable commodities, the future competition is no longer just about who writes better code, but about who can more efficiently and sustainably organize energy, chips, models, and human demand. For model companies, the competition is converting every kilowatt-hour and every GPU into more high-value tokens; for cloud providers, it is running chips, networks, and facilities at higher utilization; for application companies, it is proving that these tokens actually deliver results users are willing to pay for. When intelligence services begin to be metered, invoked, and priced, AI is entering not just the chat window, but a new industrial order.
Source link
Author

- Ytv Market News
- Share-market news writer and analyst with deep experience covering equities, commodities, forex, and cryptocurrencies for readers in the USA, UK, Canada, and Australia. Ytv Market News delivers timely market updates, practical trading insights, and clear explanations of macro and company-level catalysts that move prices. Combines on-the-ground financial reporting with technical analysis, using concise charts and actionable ideas to help investors and traders make smarter decisions.
Latest entries
Commodities NewsAugust 25, 2026US Antimony Doubles Ore Extraction at Stibnite HillIndiaAugust 25, 2026Samudra Manthan: India prepares year-wise drilling, rig deployment roadmap
Company Press ReleasesAugust 25, 2026Invesco Ltd: Form 8.3 – Prologis Inc; Public dealing disclosure
Stock Market VideosAugust 25, 2026New Carrier To Replace USS Lincoln in Middle East | Balance of Power 08/13/2026
