The Pricing Trap That Made Nvidia Vulnerable

Nvidia’s monopoly was built on two pillars: CUDA software lock-in and a pricing strategy that forces customers to overpay for the combination of fast compute and large memory. The host, speaking on Theo – t3.gg, breaks down this strategy with a kitchen analogy: you can have a fast chip or lots of memory, but paying for both means paying a massive premium.

The numbers tell the story. The RTX 5090, originally priced at $2,000 and now fetching around $4,500 due to demand, offers 32 GB of GDDR7 memory and roughly 1,800 GB/s of bandwidth. The RTX Pro 6000, priced between $12,000 and $16,000, offers 96 GB of memory — but it’s the same speed or slower on raw compute than the 5090. The DGX Spark, at $4,000, offers 128 GB of unified LPDDR5 memory but runs on a 20-core ARM chip with only 6,000 CUDA cores — a fraction of the 24,000+ on the 5090 — and uses memory that is roughly one-seventh the speed of GDDR7.

Product Price Memory Memory Bandwidth Compute
RTX 5090 ~$4,500 (street) 32 GB GDDR7 ~1,800 GB/s 24,000+ CUDA cores
RTX Pro 6000 $12,000–$16,000 96 GB GDDR7 Similar to 5090 Same or slower than 5090
DGX Spark $4,000 128 GB LPDDR5 ~1/7th of GDDR7 6,000 CUDA cores, 20-core ARM
Mac Studio M5 Ultra ~$10,000 256 GB unified 1.2 TB/s 80-core GPU, 36-core CPU

The host’s point is sharp: Nvidia has engineered a market where the cheap option is always crippled in one dimension, and the option that satisfies both needs costs six to ten times more. The DGX Spark, in particular, is described as “my least favorite computer in this apartment” — it ships with a botched Ubuntu install, runs ARM Linux which the host calls “hell,” and is only useful for inference, not as an actual computer. The 200 Gbit/s NIC for connecting Sparks is technically there but “hell to set up.”

Apple’s M5 Ultra: Killing a Whole Category of Nvidia Devices

Apple’s surprise announcement of the Mac Studio M5 Ultra directly attacks Nvidia’s pricing leverage. The M5 Ultra offers 256 GB of unified memory at 1.2 TB/s bandwidth for roughly $10,000 — more memory than any Nvidia consumer option, at a fraction of the RTX Pro’s price, with bandwidth only 30–40% slower than the 5090 and nearly six times faster than the DGX Spark. The M5 Max variant offers 128 GB at 614 GB/s.

The host notes that Apple had gimped the M4 generation by never releasing an M4 Ultra, leaving Mac Studio buyers stuck on the M3-era chip for over three years; the M5 Ultra fixes that gap. The time-to-first-token is 10x faster than the M1 Ultra Mac Studio, driven largely by memory improvements.

See also  India Courts Japanese Capital With $1 Billion Deep-Tech Fund Pitch and Startup Corridor Proposal — BigGo Finance

“This kills almost all of the reasons you could ever justify buying the DGX Spark, which is a whole category of Nvidia devices killed.”

The host frames this as Apple’s first real play in the AI space — a direct challenge to Nvidia’s pricing model rather than a niche product. “This is Apple’s first real play in the AI space. Them coming in and saying, ‘Sorry guys, you’re messing around too much. We’re going to put an end to that.'” For local inference workloads, the M5 Ultra now offers a no-compromise option where Nvidia forces a choice between speed and capacity.

OpenAI’s Jalapeno: Beating Nvidia at Its Own Game

The most consequential development covered is OpenAI’s Jalapeno chip, developed in partnership with Cerebras and detailed by SemiAnalysis after an early-access review. The headline numbers are striking: 216 GB of HBM4 memory, 700 watts power draw, 13.4 petaflops for FP4, and performance-per-watt that beats every Nvidia, AMD, and Google chip tested across multiple open-source models. SemiAnalysis explicitly notes that Jalapeno is a generalized inference chip, not a narrow ASIC — OpenAI even ported Doom to it as a demonstration.

Metric Jalapeno Nvidia Blackwell Nvidia Rubin (not yet shipping)
Memory 216 GB HBM4 Less, HBM3e HBM4
Power 700 W Higher Higher
FP8 performance Slightly behind GB200/3000 Baseline Ahead
Performance per watt Beats all tested Lower Unknown
FP4 13.4 petaflops Lower Unknown

The efficiency numbers are the real story. At concurrency 1 on DeepSeek R1, Jalapeno hits over 700 tokens per second per user; on Kimi K2.5 and GPT-OSS, it reaches 1,400 tokens per second per user — all with single-token prediction, no speculative decoding, and no prefill-decode disaggregation. SemiAnalysis confirmed the GSM-8K results match the reference chip, so the speed is not coming from nerfed models.

Why does OpenAI care so much about efficiency? Because the company is “limited by data center power, not by budget or floor space, and thus tokens per megawatt is paramount.” This is the same framing Jensen Huang himself conceded at Computex 2026: “If you have one gigawatt of power, then throughput per watt is revenue.”

The host flags one caveat: all tested models are relatively small, so large-model performance (e.g., Kimi K3 or DeepSeek Pro) remains unverified. But the strategic implication is clear — OpenAI is building its own escape hatch from Nvidia dependency, and the first-generation chip is already competitive on the metric that matters most for scaling.

See also  Water suppliers on Iran hacking alert

China’s Chip Independence: GLM-5.3 Flash on Huawei Silicon

The episode opens with a concrete demonstration that Nvidia-free AI is already viable at scale. An anonymous model called OX Alpha appeared on OpenRouter and OpenCode with absurd free throughput — 100 trillion tokens per day — and turned out to be GLM-5.3 Flash, served entirely on Huawei chips. The host notes this was a genuine surprise: he had assumed the traffic was running on Nvidia hardware, but the Chinese lab achieved per-token costs and hardware efficiency comparable to Nvidia GPUs without using any Nvidia silicon.

This matters because US export controls have largely banned Nvidia’s most powerful GPUs from China, forcing Chinese labs to rely on domestic alternatives like Huawei. The GLM-5.3 Flash deployment proves that constraint is not fatal — it is a forcing function. The host connects this to the broader pattern: every major AI player is now motivated to reduce Nvidia dependency, and China has the strongest motivation of all.

The CUDA Moat and Nvidia’s Defensive Moves

The host identifies CUDA as the second pillar of Nvidia’s monopoly, alongside chip pricing. CUDA is the language and system of choice for the vast majority of AI research, and it is the reason Nvidia’s software lock-in persists even as hardware alternatives emerge. Nvidia’s reported bid to acquire Hugging Face for $12.9 billion is read as a defensive play: by funding open-source AI development, Nvidia encourages more training and fine-tuning — which still happens on CUDA — thereby extending its saturation.

The host also flags Jensen Huang’s appearance on Jim Cramer’s show, where he dismissed the Jalapeno threat with characteristic confidence: “There’s so many XPUs that are being announced… lots of projects get started, lots of projects get cancelled.” The host’s read is that this bravado masks genuine concern, and that Nvidia’s market share growth is real but increasingly contested.

Terrafab: The Long-Shot That Could Change Everything

Elon Musk’s Terrafab is described as “the most epic chip building effort ever” — a chip fab that will be 25 times the size of Apple Park and 20 times the size of the Pentagon. Musk is building this despite having some of the largest Nvidia contracts in existence; he reportedly has more GPUs than Anthropic, and Anthropic is now renting GPUs from him. The host notes that Intel is reportedly involved in Terrafab 2. The strategic logic is the same as OpenAI’s: even the biggest Nvidia customer is trying to build an exit ramp.

See also  CLS Classmate: China's Auto Sales Plunge 20%; If Beijing Rolls Out Home Purchase Subsidies, the New Home Market Could "Completely Reverse"

The Rubin Window and What Comes Next

The host predicts that Nvidia’s Rubin line using HBM4 will not meaningfully matter for two to three years, as Blackwell orders are still being fulfilled, giving OpenAI’s Jalapeno a window to compete. This is a critical timing question: can Nvidia’s Rubin-generation HBM4 chips arrive in time to defend the franchise, or will the customers now building alternatives — OpenAI, Apple, Huawei, Musk — collectively erode the monopoly before Rubin ships at scale?

The host is an AMD investor and acknowledges AMD is “very behind,” but the broader point stands: the era of Nvidia as the default, unavoidable supplier is ending. The open question is whether the CUDA moat and Rubin’s ramp can hold the line.

The through-line of this episode is that Nvidia’s monopoly is being attacked from four directions simultaneously — OpenAI on performance-per-watt, Apple on memory pricing, China on hardware independence, and Musk on manufacturing scale. The future of AI compute will be fought on electricity, not chips: whoever delivers the most tokens per megawatt wins, and Nvidia’s current pricing model is vulnerable to exactly that metric. For investors, the implication is clear: Nvidia’s growth is real, but the structural cracks in its pricing power and software lock-in are no longer hypothetical. The question is no longer whether Nvidia will face competition, but whether the competition will arrive before Rubin ships at scale.


Source link

Author

Shin John
Shin JohnYtv Market News
Share-market news writer and analyst with deep experience covering equities, commodities, forex, and cryptocurrencies for readers in the USA, UK, Canada, and Australia. Ytv Market News delivers timely market updates, practical trading insights, and clear explanations of macro and company-level catalysts that move prices. Combines on-the-ground financial reporting with technical analysis, using concise charts and actionable ideas to help investors and traders make smarter decisions.