The AI infrastructure bubble may stay contained while chips, power, and construction capacity remain scarce. In July 2026, NVIDIA CEO Jensen Huang told Axios that a bubble was highly unlikely within five years. His claim rests on physical supply constraints. It offers no guarantee that AI stock valuations will hold.

Scarcity is doing more work than optimism
Huang describes the current buildout as a new industrial infrastructure layer. He expects the semiconductor industry to expand five to ten times over the next decade. Foundry capacity, advanced packaging, memory, storage, and optical networking all have to grow with it.
The International Energy Agency gives the constraint argument real weight. It projects global data center electricity use to rise from about 485 TWh in 2025 to roughly 950 TWh in 2030. The agency also expects high-bandwidth memory shortages through the end of 2027.
The Federal Reserve's July 2026 Beige Book recorded strong data center orders and construction. It also found persistent shortages of skilled technicians and tradespeople. Those delays stretch the time between approved capex and productive compute capacity.
Scarcity can postpone overbuilding. It cannot turn a low-return facility into a good investment.
Kimi K3 compresses model rents and expands the workload
Moonshot AI released Kimi K3 on July 16, 2026. The model has 2.8 trillion parameters and a one-million-token context window. Its sparse architecture activates 16 of 896 experts per token.
Moonshot estimates a 2.5-fold scaling-efficiency gain over Kimi K2. Its API costs \(3 per million cache-miss input tokens and \)15 per million output tokens. That price puts more pressure on the scarcity premium charged by closed-model providers.
Huang's hardware thesis follows a price-elasticity argument. Cheaper AI lets more developers and companies add inference to everyday work. Total compute rises when workload growth outruns efficiency gains.
If cost per completed task falls 50 percent and usage more than doubles, aggregate compute demand still increases. The actual result depends on token growth, cache rates, and hardware utilization.

Investors need three separate ledgers
One market trade often combines model competition, data center construction, and equity valuation. Each ledger has a different cash-flow test.
| Ledger |
Verifiable 2026 evidence |
The pressure test |
| Model rents |
K3 API: $3 input and $15 output per million tokens |
Cost per completed task, retention, cache-hit rate |
| Physical compute |
IEA: data center electricity use rises from 485 TWh to 950 TWh |
Grid connection time, HBM lead times, accelerator utilization |
| Equity valuation |
NVIDIA fiscal Q1 2027 revenue: $81.6 billion |
AI revenue conversion, depreciation, cash return on capex |
NVIDIA reported $81.6 billion in quarterly revenue on May 20, 2026, up 85 percent from a year earlier. Data center revenue reached $75.2 billion. Networking revenue rose 199 percent to $14.8 billion.
The company guided to $91 billion for the following quarter and assumed no China data center compute revenue. Huang also described current China revenue as approximately zero during the Axios interview.
The numbers show powerful demand and growing concentration. Data centers now drive most of NVIDIA's revenue. Future returns require high utilization and customer revenue after the hardware ships.
Capacity relief starts the valuation stress test
SEMI expects spending on 300mm memory fab equipment to reach $52 billion in 2026, up 29 percent. Its forecast rises to $57 billion in 2027. New capacity still needs construction, tool installation, qualification, and production ramp.
Shorter HBM lead times, faster grid connections, and higher completion rates will reveal the true supply gap. High utilization would convert new capacity into revenue. Falling utilization would expose depreciation, electricity, and financing costs.
GPU shipments alone provide an incomplete scorecard. Investors also need power per trillion tokens, data center utilization, cloud AI revenue per dollar of capex, and cost per completed enterprise task.
Huang's five-year view may prove directionally right because physical shortages delay excess supply. That mechanism says little about the price investors should pay for the eventual cash flows.

Does Kimi K3 mean U.S. closed models have lost their edge?
Current evidence falls short of that claim. Kimi says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall. K3 proves that open models can approach the frontier, while adoption still depends on product quality, security, compliance, and service reliability.
Is Huang's five-year bubble claim a forecast?
It is Huang's July 2026 judgment, grounded in shortages of chips, memory, land, power, and labor. He also said a bubble will likely form someday. The five-to-ten-year window carries much more uncertainty.
Why can cheaper models increase GPU demand?
Lower unit costs make more workloads economical. Aggregate compute rises when token growth exceeds efficiency gains. Kimi recommends supernodes with at least 64 accelerators, which shows that large sparse models still require memory capacity and fast interconnects.
Which metrics could reveal an AI infrastructure bubble early?
Start with data center utilization and cloud AI revenue relative to capex. Then watch HBM lead times, grid connections, and cost per completed task. Rapid capacity growth paired with falling utilization would be a clearer warning than one model benchmark.
Sources
Author Insight
I read Huang's five-year claim as a supply clock, rather than a valuation target. Kimi K3 lowers the price of model capability while physical constraints postpone oversupply. When both forces ease, utilization and cash returns will face their real test.