NVIDIA B200 Compute Pricing: Dec 31, 2026 Prediction Market & Spot Analysis
- David Rogers
- Technology Prediction Markets
- 2026-08-20
NEED TO KNOW
- Settled Baseline Spot Rate: NVIDIA B200 compute trades at $6.40/hr on the Ornn settled spot index (as of September 1, 2026), consolidating near its post-summer recovery level.
- Prediction Market Distribution: Robinhood event contracts project an implied median between $7.50 and $8.00/hr by Dec 31, 2026, with Above $7.37 priced at 73¢, Above $7.97 at 49¢, and Above $13.07 down to 11¢.
- Unrivaled Throughput Economics: MLPerf benchmarks published in Ornn Data's September 2026 whitepaper show B200 generating 10,756 server tok/s on dense Llama-2-70B and 8,949 server tok/s on sparse
gpt-oss-120b, delivering the lowest compute cost in the industry ($0.15 to $0.16 per million tokens at full utilization). - Forward Curve Dichotomy: Ornn forward term marks demonstrate exceptional near-term scarcity (1-year term retains 97.9% of spot at $6.27/hr), but steep long-term compression (3-year falls to 69.1%, 5-year to 53.8%, and implied forward for months 37–60 drops to 31%) as B300 and Rubin approach.
- Power Cost De-coupling: At 1,000 W nameplate and PUE 1.3, electricity costs are just $0.104/GPU-hr (1.6% of spot rent), confirming that secondary pricing is governed by capital amortization and supply-chain delivery schedules rather than utility bills.
- Base Equilibrium Target: The consensus price channel for secondary B200 compute by year-end 2026 remains centered in the $7.20 to $8.10/hr range, supported by near-term term-contract firmness.
NVIDIA B200 spot compute pricing on secondary market indexes settled at $6.40 per hour on September 1, 2026, holding steady after rebounding from early summer lows near $4.30 per hour /Ornn/ /Ornn Data/. Robinhood prediction order books reflect a more balanced sentiment: contracts trade at 73 cents for prices above $7.37, 60 cents above $7.67, 49 cents above $7.97, and just 11 cents above $13.07 /Robinhood/.
The market hourly price is framed as:
- Base Hardware & Operating Floor: ~$3.80 – $4.50 / hr (based on an amortized ~$300k–$400k HGX chassis over 3–4 years at ~80% utilization plus base datacenter power).
- Current Spot Baseline (Ornn): $6.40 / hr (reflecting a ~$1.90–$2.40 scarcity margin).
| Scenario | Modeled Price Target | Probability Weight (Robinhood) | Primary Market Drivers |
|---|---|---|---|
| Bear Breakdown (< $6.00) | $4.20 – $5.80 / hr | ~27% (Implied by 100% − 73¢ at $7.37) | • CoWoS packaging glut & volume deliveries. • Internal enterprise migration to TPU/Trainium. • Distillation/quantization reducing per-token GPU demand. • Mega-campus grid energizations relieving power premiums. |
| Base Case Equilibrium ($6.00 – $10.00) | $7.20 – $8.10 / hr | ~60% (Centered at $7.67–$7.97 bracket) | • Steady demand growth balanced by ongoing hardware shipments. • Liquid cooling infrastructure constraints keep floor firm. • Agentic fine-tuning absorbs new hardware capacity smoothly. |
| Bull Scarcity Surge (> $13.07) | $13.50 – $16.00 / hr | ~11% (11¢ at $13.07) | • Dual-die CoWoS-L packaging halts or HBM3e yield crises. • Sovereign AI and tier-1 lab monopolies starving spot clouds. • Test-time compute/reasoning inference explosion. • Grid interconnection moratoria and Pacific tariff escalations. |
Empirical Throughput & Token Economics
Granular benchmark derivations published by Ornn Data in The Economics of Open-Weight Inference /Ornn Data/ demonstrate the raw performance advantage of the Blackwell architecture. In MLPerf v4.1 inference audits, an HGX B200 generates 10,756 server output tokens per second per GPU on dense Llama-2-70B—nearly four times the throughput of an H100 SXM (2,701 tok/s) and three times an H200 (3,654 tok/s). On sparse mixture-of-experts workloads like OpenAI’s open-weight gpt-oss-120b (5.1B active parameters), Red Hat’s MLPerf v6.0 submissions clock the B200 at 8,949 server tokens per second per GPU.
This performance multiplier delivers the lowest compute cost per token in the industry:
- Full Utilization ($u = 1, h = 0$): $0.16 per million tokens on dense Llama-2-70B; $0.15 per million tokens on sparse
gpt-oss-120b. - Base Utilization ($u = 0.5, h = 0.15$): $0.39 per million tokens on dense; $0.47 per million tokens on sparse MoE.
Crucially, operating electricity costs are largely decoupled from spot rental spreads. At 1,000 W nameplate power, a PUE of 1.3, and commercial power rates of $0.08/kWh, energy accounts for just $0.104 per GPU-hour—representing only 1.6% of B200’s $6.40 spot price. Consequently, spot pricing is governed by capital amortization and supply-chain delivery cadence rather than utility power bills.
The Forward Curve Dichotomy: 1-Year Scarcity vs. 5-Year Replacement
While prediction markets focus on short-term price discovery through December 2026, Ornn Data’s forward term curves reveal a stark divergence across contract durations:
- 6-Month Term: 99.1% of one-month mark
- 1-Year Term: 97.9% of one-month mark ($6.27/hr implied)
- 3-Year Term: 69.1% of one-month mark ($4.42/hr implied)
- 5-Year Term: 53.8% of one-month mark ($3.44/hr implied; months 37–60 implied forward drops to 31%)
This forward curve structure reflects extreme near-term supply tightness: enterprise demand for high-throughput reasoning models, synthetic data pipelines, and agentic rollouts absorbs nearly every delivered B200 tray, keeping one-year forward commitments at parity with spot. However, the long end of the curve steepens sharply as the market prices in replacement cycles from B300 (Blackwell Ultra, whose term marks were initiated by Ornn on September 1, 2026) and the upcoming Rubin architecture.
To anchor long-term software ecosystem demand for its hardware fleet, NVIDIA announced the acquisition of Hugging Face on September 3, 2026 /NVIDIA Blog/. By directly controlling the open-weight model hub, NVIDIA is securing the software pipeline that routes open-source inference, agentic sandboxes, and post-training workloads onto Blackwell infrastructure.
Ultimately, prediction markets are pricing neither a catastrophic supply collapse nor a commoditized race to the bottom. Instead, the contracts show strong support around the $7.50 to $8.00 per hour median, treating the NVIDIA B200 /NVIDIA/ as a valuable, steadily utilized enterprise asset rather than an unconstrained commodity.
Disclaimer: All forecasts, probability models, and price target estimates are independent projections for educational and research purposes only. Prediction markets carry financial risk and high volatility. This is not investment or financial advice; participate at your own risk.