Frontier Inference Margins

Range explorer ↓ Full report ↓

An interactive model of what it costs to serve one token of a frontier LLM — and the modeled unit direct-serving contribution margin that implies (a list-price metric only when the batch and discount sliders are 0%; not a company gross margin). Prices and inputs are current as of July 15, 2026 unless a different effective date is shown. The live v3.0 path resolves a per-hardware roofline from a declared serving regime, traffic-derived sequence length, per-row precision tuple and reviewed calibration; it visibly caps memory-limited operating points and withholds numbers for infeasible ones. TPU v7 uses an analyst-set same-platform bridge; Trainium 2/3 use an out-of-family joint coefficient without matched serving anchors; Rubin remains a shape-only projection. Every slider documents its sources. Adjust the assumptions — nothing here is any provider's actual ledger.

What this calculator does — and deliberately does not — model (read before quoting numbers)
· expand for the perspective, its source, evidence label, and what would falsify it
Serving contribution margin
Not a company gross margin. Unit direct-serving contribution margin only — see the methods box.

The two independent estimates — and the stress test under them

GPT-5.6 Pro

83.1%

68 – 92%

undiscounted list · judgment range

Explain this estimate

The face carries the round-3 self-authored revision — 83.1% at the undiscounted list tariff, span 68–92%, with a lead-only diagnostic of 79.65–85.89% across 0–4 months. Its author stated round 3 at list only, which is why this face's basis line differs from the card beside it. The two bases differ only in the denominator — effective billings is this page's own illustrative mix (15% Batch share, 5% negotiated discount) applied to the published $5/$25 tariff with cache reads at 10%; list is that tariff undiscounted.

Estimand: Claude Opus 4.x unit direct-serving contribution margin — accelerator capacity, occupancy slack, direct host/network/control-plane burden and production serving overhead, at the Reference traffic mix (15:1 input:output, 60% cache), mid-2026. Not a company gross margin (§7 explains why the two diverge).

Its load-bearing assumptions: 65% paid-capacity occupancy — its stated midpoint over a 50–80% band, explicitly an analyst judgment and not a source; list-only billing on its own vector (0% batch, 0% discount), with this page's 15/5 mix asked for as a companion denominator rather than baked in; per-leg strategic procurement — NVIDIA ×0.95 (range 0.90–1.00), TPU ×0.50 (0.30–0.70, where it says the regime disagreement is concentrated), Trainium ×0.85 (0.70–1.00) at zero weight, with Trainium excluded from the point until the per-chip-versus-replica form defect is repaired; no algorithmic-lead prior; and no fleet-wide speculative-decoding credit (0 pp in the headline, an eligible-leg sensitivity instead).

The honest gap. Its declared vector computes about 80% at list and about 77% on the illustrative mix in this engine — roughly four points below what its author states, because the midpoints above are the middles of ranges it declined to collapse. It asks explicitly that the vector not be tuned until it reproduces the number, so it is not. What does reproduce is its downside: its public-rate stress case computes 64.13% here against the 64.1% it states.

Where it diverges from the Fable estimate: occupancy (65 vs 60 — worth about 1.9 points at the other posture), and a more strategic fleet and serving-form synthesis: its $0.595/M judgment cost sits below anything this engine computes at the NA-blend fleet without either the lead prior or near-owned-TCO economics. Both hold the algorithmic-lead prior at zero. The page does not adjudicate between them.

Not the same object as §6. §6 below is this adjudicator's July 2026 independent consult, at 92–94% list on a mature-fleet strategic-contract reading; the number on this card is its round-2 re-run of 2026-08-06, made with the findings that postdated that consult and on the estimand stated above. Both are preserved rather than harmonized.

Which vintage is where. The face figures and the calculator's opening state are the same round-3 self-authored reading, so the headline and the opener agree. The assumption, gap and divergence paragraphs above are the fuller published reasoning from this adjudicator's contextual review of 2026-08-06.

Provenance: face figures quoted from the round-3 self-authored revision; the analysis above from the round-2 contextual review of 2026-08-06. Its declared dials are itemized in the adopted grounding ledger, and the July consult is published in full as a verbatim original with an adopted-findings summary.

Reach this operating point in the calculator: Load this estimate's own vector ↑ (round-3 self-authored, on Opus — the calculator's opening state)

Fable 5

≈77%

65 – 82%

effective billings · judgment range

Explain this estimate

The face carries the round-3 self-authored reading — ≈77% effective billings, span 65–82%, ≈80% at list. Same two denominators as the card beside it — effective billings is this page's 15% Batch / 5% discount mix on the $5/$25 tariff, list is that tariff undiscounted.

Estimand: identical to the card beside it by construction — Opus 4.x unit direct-serving margin at the Reference 15:1 / 60% mix, published tariff, mid-2026 — which is what makes the pair comparable at all.

Its load-bearing assumptions: 60% paid-capacity occupancy (band 50–70), held deliberately unmoved rather than raised to preserve altitude after the lead prior was withdrawn; procurement per leg — TPU ×0.30, Trainium ×0.70, and NVIDIA at no discount, one it explicitly declined to claim; 2.5T total / ~300B active; algorithmic-lead prior off, with a separately declared +2.0-point serving-stack efficiency credit that is not the lead prior and must never be stacked with it.

The honest gap. This engine, at that posture with the lead at zero, computes 75.1% effective billings / 78.2% at list; the stated central is the +2.0 credit above it. That credit is bounded by two first-party anchors — a 1.15× decode-only speculative-decoding gain (+0.9 here) and a 20% all-in serving-cost reduction (+5.0 here) — and has no legal dial on this engine, because the speculative-decode lever is gate-blocked at this stack setting precisely to stop anchor-embedded efficiency being credited twice. So the estimate ships as vector plus declared increment rather than as a solved dial, and its author is explicit that the lead must not be switched on to reach the headline (lead-on is a separate sensitivity, at 81.1% / 83.4%).

Where it diverges from the GPT-5.6 Pro estimate: occupancy 60 against 65 — and it holds that 65 is exactly as unmeasured as its own 60, since the public-evidence sweep returned a rigorous negative for both — plus the rate treatment above and the declared credit in place of a lead prior.

Produced without sight of the other. A Fable 5 session derived the round-2 estimate from this page's own numbers and hash-committed it before any GPT-Pro round-2 output existed. The face carries its later round-3 self-authored reading. The assumption and gap paragraphs above are the fuller published reasoning from its independent estimate of 2026-08-06.

Provenance: face figures quoted from the round-3 self-authored reading; the analysis above from the round-2 independent estimate of 2026-08-06, computed against this site's deployed engine. Its declared dials are itemized in the adopted grounding ledger.

Reach this operating point in the calculator: Load this estimate's own vector ↑ (round-3 self-authored, on Opus)

Judgment ranges across regimes and structural uncertainty — not statistical confidence intervals. Both adjudicators say so in their own words.

Stress test ~59% this page's own conservative policy-labeled low/committed-rate planning scenario, at the public-evidence reference (algorithmic lead 0 months) — a scenario of this page, not a third adjudicator

Explain this scenario

Every judgment dial at its unmoved value: the registered heterogeneous low/committed planning-rate vector at 1.0×, 50% fleet occupancy, open-source-level serving stack, balanced latency, no algorithmic-lead prior. It reads about 59% on this page's 15% Batch / 5% discount mix and about 64% at the undiscounted list tariff.

It is a reproducibility floor — what a reader can rebuild from public rate cards without trusting a private-economics judgment — and deliberately not an estimate of likely actual economics. It is the only number on this page a reader can check from the outside. The full reasoning behind it, including where it is fragile, is the section body below.

Reach this operating point in the calculator: Load the public-rate reproducibility floor ↑ (on Opus)

The answer
The full reading — scope, bases, bands and spans
Why not the higher numbers?
What would have to be true to reach the higher readings

What would have to be true for a ~90% margin? Or ~45%?

These are “what would it take?” questions — not this page's estimate. Each range below is a hypothesis to stress-test: pick one to see the assumptions it would require. The page's own modeled result is shown only in the calculator below, and never takes its value from this explorer.

Pick a margin range you've heard claimed and see the assumptions it would take to land there. Each range with a claimed position behind it carries a page-authored route — what would have to be true of procurement, occupancy and serving efficiency — that you can load into the calculator as an explicit counterfactual. The headline estimate (derived from the Model / Traffic-mix selections at the central cost lens) never takes its value from this explorer.

Saved scenarios

Scenarios saved in this browser. Selecting one loads it as a modified state — the saved numbers, never the currently-selected lens. Save new ones from “Save scenario” at the end of the adjustment sections below.

Evidence catalog — sourced claims per range, badged by relation (verbatim where archived; reported figures labeled as such), plus every page-authored route

This grouping of claims into ranges is this page's organization, not the claimants'. A claim renders under a range only with its typed relation badged — asserts, locates-within, conditional transition, unnamed subject, different metric, disclosure anchor — and a floor claim never renders as interval membership: it is compatible with its range and every higher one. Company-GM and reported figures are different objects from this calculator's unit metric (§7). Routes are ordered by fewest changed registry fields from the central configuration (stable-id tie-break) — an edit count, not a measure of which route is right.

Blended cost
per 1M tokens served (traffic mix)
Blended effective price
per 1M tokens billed (after cache mix)
Cost per 1M output
input: —
Serving feasibility
Save scenario

Saved to this browser only; it reloads as a modified state (never inheriting a lens). Load, rename or delete under “Saved scenarios” in the range explorer above.

Shared links and saved scenarios are stamped with a defaults epoch (currently v22) and the headline margin at share time. Links or presets minted before v2.2 (2026-08-06) no longer resolve — they show a deprecation notice rather than a silently wrong number.

Margin if served entirely on each accelerator
Where the cost of a token goes click any bar to open that accelerator's cost control · owned/strategic-TCO decomposition per 1M blended tokens, by accelerator — a counterfactual build-up under rent lenses (see chip)
Margin sensitivity — active parameters holding everything else at current settings
Cost per 1M output tokens across hardware generations current model settings; margin at current price labeled on each bar. Rubin is shown as an unpriced placeholder (no public rack rate, benchmark or purchase price) — it is not plotted on the $ scale.
Subscription economics is the $200 plan underwater? Depends whether you price usage at list or at marginal cost.

Full report: how fat are frontier inference margins?

Scope note: this is first and foremost an investigation of Anthropic's serving margins — that is where the strongest 90–95%-range claims cluster (as conditional and floor claims, not an unconditional "Anthropic = 90–95%" statement — see the §1 source audit), where the public evidence concentrates, and where the two independent research runs went deep. The other providers in the calculator are included for comparison at materially lower evidence density; their presets are best-effort estimates (marked * where speculative). Each non-Anthropic provider has since received its own independent GPT-5.6 Pro deep dive — §10 audits what is actually known about each one, with judgmental margin ranges and the reasons for their width, and links the full public research reports.

1 · The claim and the claimants

Expand this sectionCollapse this section

The framing under examination: that Anthropic's marginal cost of serving a Claude token sits so far below its API list price that serving gross margins reach the 90–95% range — the loudest figure the discourse repeats. This is a claim about unit serving economics — one token, on a warm GPU, at scale, in the claimants' framing — not about Anthropic's income statement (see §7 for why those diverge). Note this page's own metric is stricter than the warm-GPU framing: it allocates paid idle capacity through the utilization divisor (methods box). Read what follows as an audit of that range: as this section documents, the cited corpus carries the 90–95% zone as conditional and floor claims, not as an unconditional "Anthropic = 90–95%" statement.

The loudest and most quantitative proponent is @teortaxesTex:

His cited record, summarized (verbatim posts in the annex sweep): an Anthropic-specific serving-cost claim — DeepSeek-style math prices serving Opus at at most $4/Mtok (token class unspecified), implying very high unit margins even on subscriptions except at full utilization (Jun 28, 2026) — and a CONDITIONAL claim, made in a western-provider context:

"No, they'll just increase the batch size, have the same speed, and drive margins from 90% to 95%. You're welcome" — Jun 27, 2026

Earlier and more conservative from the same account: "if we exclude R&D and look at inference alone, Anthropic and OpenAI are making like 80% margins" (Mar 2025). The cited corpus does not contain an unconditional claim that Anthropic's margin is exactly 90–95%.

@zephyr_z9 (Zephyr) makes a related pricing allegation from the semis side: a Jul 8, 2026 post says xAI is not "juicing up the gross margins to 90%-95%" — but it names no comparison lab and does not define the metric as unit serving margin (both are this page's readings). Separately, a Jun 24 post explicitly places Anthropic at ~70% company-level gross margin with 15–20% FCF margin. The two figures coexist in one account's posts; this corpus does not establish that the 90–95% allegation is Anthropic-specific.

One reported-margin objection: @fleetingbits — "we know approximately what frontier lab inference margins are; it's like 40-50%; it's been reported a bunch of times. anthropic labels cloud provider commissions as a sales and marketing expense; so the gross margins are mostly inference compute costs." And in the middle, a Jul 1, 2026 PodcastAlphaX clip-account post featuring Dylan Patel of SemiAnalysis quotes him putting Anthropic's margin on an Opus API token "north of 80%" — quoted-secondary evidence, not a post from his or SemiAnalysis's account.

2 · The evidence base, verified

Expand this sectionCollapse this section

2a · DeepSeek's serving disclosure — the one primary source in the field

On March 1, 2025 (Open Source Week "Day 6"), DeepSeek published actual production serving statistics for V3/R1 on H800 clusters: 73.7k input / 14.8k output tokens per second per 8-GPU H800 node, 56.3% input cache-hit rate, ~$87,072/day GPU cost (at $2/H800-hr) against $562,027/day theoretical revenue at R1 list prices — the famous "545% cost-profit ratio," which is a markup; as a margin it is 84.5% (a distinction TeorTaxes himself insisted on). Caveats DeepSeek listed: much traffic was free web/app users and off-peak V3 pricing, so realized revenue was materially lower than the theoretical figure.

Why it matters for Anthropic: DeepSeek's disclosure implies an 84.5% theoretical margin had all observed traffic been billed at R1 list prices — DeepSeek itself stated actual realized revenue was materially lower — on export-restricted hardware, in early 2025, at prices far below Anthropic's on a like-for-like token-class basis: 3.6× cheaper for cache reads, 9.1× for fresh input, 11.4× for output. The hardware and serving-software inputs to that calculation have since improved (§4), though tariffs, fleet constraints and latency targets can move the other way. The a-fortiori argument — if DeepSeek could get to ~85% theoretical at $2.19/Mtok output on H800s, what does $25/Mtok output on GB300s and TPUv7 imply? — is the strongest single piece of evidence in the bull case. It is an anchor for what optimized serving can cost, not proof of what Anthropic's margin is.

2026-07-15 update — a 2026 sequel to the same argument. The Information reported (Jul 14–15, 2026, via @jukan05/@jingyanghk relays) that DeepSeek's annualized revenue is nearing ~$500M, with a ~¥50B (~$7.4B) raise at a ~¥500B (~$74B) valuation and STAR Market IPO prep — and, in a follow-up, that V4 API gross margin is 70–80% (up from an earlier ">50%" figure). This is a reported, realized figure at V4's much lower list prices, and it corroborates the direction of the argument above even though it is a different object: "gross margin" here carries no disclosed accounting definition or company confirmation, so it is not directly comparable to the disclosed 545%/84.5% unit-serving arithmetic — the metric distinction matters as much now as it did then. This page's own model already puts V4 Pro below 90% at list on any fleet (§10), so a reported 70–80% company-side figure sits inside, not above, what the unit math would suggest.

2b · The Musk size leak — total vs active, a distinction that matters

The leak is widely retold as "Musk revealed Opus is much smaller than expected" — but on total parameters it says the opposite. What Musk actually posted (Apr 9, 2026): "0.5T total. Current Grok is half the size of Sonnet and 1/10th the size of Opus. Very strong model for its size." The community deduction (e.g.): Sonnet ≈ 1T, Opus ≈ 5T total parameters — Opus is big.

The "smaller than expected" intuition belongs to active parameters: TeorTaxes, observing Fable served at ~90 tok/s, concluded it had "shockingly FEW active parameters for what it was", and Zephyr puts OpenAI's frontier at "~100B active range" with Opus/Fable "the highest active parameters" among peers. For serving cost, active parameters are the dominant modeled driver (compute scales with active-parameter FLOPs/token, Φ ≈ 2 × active plus a context-length-dependent attention term); the activated roofline separately resolves compute, HBM/KV-traffic and fabric/communication-bound terms as functions of context length, batch and sharding topology rather than collapsing them into one scalar. Total size mostly sets the HBM footprint and the feasible sharding topology. The current default is a 2.5T-total/~300B-active MoE; the 5T Musk-relative reading is preserved as a separate legacy scenario — and the active number is the single most uncertain, most consequential slider on this page. Even the total is contested: the "incompressible knowledge probes" estimation method (arXiv:2604.24827) carries a ~threefold prediction interval (per the paper's v2 revision of Jul 5, 2026: median fold-error 1.48×, with 86% of models within 3×), and a methodological re-analysis lands nearer 1.1T total for Opus 4.7. GPT Pro's sanity check cuts the other way: a dense 5T Opus would cost more than its own $25/Mtok list price to serve — so if Musk's number is right, heavy sparsity isn't optional, it's implied.

2c · GLM 5.2 on GB300 — the ncode/Noumena deployment by @_xjdr

The researcher is @_xjdr; the product is ncode on the Noumena platform (code.noumena.com), served on GB300 NVL72 racks from Prime Intellect. His free-week postmortem (Jun 30, 2026) is the best public look at frontier-style serving on Blackwell Ultra:

"final GLM 5.2 served stats: ~12000 unique api keys served ~300B tokens total 232 tok/s/gpu output average 431 tok/s/gpu output max sustained 2.1 sec TTFT overage [sic] (1M ctx) 61 sec p95 TTFT (1M ctx) 81k tok average input size 41% cache hit rate 0 chat logs kept (dont be evil)"

Configuration: 60× B300 ("15 trays"), bf16 attention with fp8 experts/KV ("virtually 0 measurable quality difference"), no MTP, no Eagle — a custom Rust stack. A third-party estimate on his thread put the implied cost at ~$0.35/Mtok in, ~$1.50/Mtok out at $6/GPU-hr. Note his 232 tok/s/GPU average is output-only bookkeeping, and the same post reports 81k-token average inputs at a 1M context limit — but it does not report average output length, request rate, phase-level GPU allocation, or an input:output split, so it does not establish how GPU time divided between prefill and decode; it is not comparable to SGLang's 12k tok/s/GPU throughput records (different model, batch regime, and MTP). The original labels the 2.1-second TTFT statistic "overage"; that field remains uninterpreted. (Erratum: an earlier revision of this page silently normalized "overage" to "average"; the token is now preserved verbatim as "overage [sic]" and left uninterpreted — see the changelog.) The separately reported 61-second p95 TTFT at 1M context documents a long tail, but the corpus does not identify its cause.

Adjacent results worth separating from the ncode story (GPT Pro's browsing initially merged them, §6): GLM-5.2's architecture is public via Baseten's serving writeup744B total / 40B active, NVFP4-clean, 280+ tok/s/user on Blackwell (a per-user speed record, not aggregate throughput). And the headline GB300 aggregate number belongs to SGLang + NVIDIA serving DeepSeek-V4: 2,200 → 11,200 tok/s/GPU from April to June 2026 on the same racks — five-fold, from software alone.

2d · What the plans actually hand out

Covered in §8 — the short version: heavy Claude Max users have documented API-equivalent usage far above the subscription price. That is consistent with low marginal serving cost, but it does not identify the subscriber distribution or exclude breakage, throttling, routing or cross-subsidy; §8 treats it as tail evidence only.

3 · The cost model and its calibration

Expand this sectionCollapse this section

The calculator above prices a token from an explicit cost identity:

output $/Mtok = ($/device-hr ÷ 3600 ÷ rendered tokens/s/device) × 10⁶ ÷ utilization. Decode throughput is the inverse of the slowest reviewed roofline term — compute, HBM or fabric — divided by eta_eff. The selected serving regime resolves a declared per-row batch; memory feasibility may cap that batch or mark the leg infeasible. Traffic fixes output length and therefore context length; precision selects a per-row tuple of FLOPS and byte widths; q=1 and a=1 on every render path. Fresh prefill is a separate compute/fabric roofline with the universal reviewed prefill calibration. Cache reads cost a selected fraction of fresh prefill. The traffic mix blends those costs, and the same mix prices revenue at list minus cache/batch/negotiated discounts.

Calibration points, not a validated predictive model. Only rows whose live coefficient reproduces a matched observation are labeled fitted. H20 and Ascend observations informed the frozen joint calibration, but their deployed per-row coefficients are separate source-informed neutral choices that do not reproduce those observations; H100/H200 are explicit H800-family transfers, GB300 is analyst-set at an assumed operating point, TPU v7 is an analyst-set same-platform aggregate-form bridge, Trainium2/3 use an out-of-family joint coefficient without a matched serving anchor, and Rubin is projection-only. Reproducing a calibration point is an identity, not a validation; the anchored observations are all ≤~50B-active while flagship scenarios assume 120–300B active. The previous single-scalar transfer test failed (mean error 37%, worst 59%), while the physically informed follow-up still failed the whole-platform gate (worst −40%); see the LOAO methods note. The live engine therefore exposes each row's operating-point basis and feasibility state rather than implying one scalar transfers everywhere:

Anchor (published)PublishedLive engine at listed pointDeployed default
DeepSeek disclosure, H800 decode (37B active, FP8, production average)1,850 tok/s/GPU1,873 at b=96, L=4,989fitted per-row decode calibration
DeepSeek disclosure, H800 input flow — includes the 56.3% disk-cache-hit share, so it is not a fresh-prefill benchmark9,212 tok/s/GPU aggregatenot used directlyfresh-prefill reconstruction calibrates the universal prefill transfer
vLLM GB200, R1 decode (source precision basis not fully pinned — possibly NVFP4)~10,100 tok/s/GPU10,114 at b=128, L=3,000 on the registered FP4 tuplefitted per-row calibration; within 0.1% of the ~10,108 anchor
SGLang GB300 record, V4 Pro 1.6T FP4 + MTP (49B active, disclosed)>12,000 tok/s/GPU7,460 at declared b=128, L=2,740analyst-set scenario-only row; measured value is null and the source batch/MTP basis is not reconstructable, so this is not a fit
Ant Group/SGLang production, H20 decode (R1, relaxed <70 ms tier; full weight-precision path not disclosed)714 tok/s/GPU (675 at <50 ms; 423 at <30 ms)680 at the source-published b=48, L=4,096 MTP pointsource-informed neutral adjustment; does not reproduce the observation and earns no fitted-cluster credit
CloudMatrix-Infer, Ascend 910C decode (R1, INT8; vendor-measured, optimized 384-NPU supernode)1,943 tok/s/NPU (538 at the separate b=8 / 14.9 ms fit input)1,422.7 at the source-published b=96, L=4,096, MTP-70% pointsource-informed neutral adjustment on the registered INT8 tuple; does not reproduce the observation and earns no fitted-cluster credit

Decode is often memory- or interconnect-bound; this is why "inference is memory" and why bigger NVLink domains and HBM can beat raw FLOPS. The live roofline path makes that binding term explicit at each rendered operating point. Domain of validity: the registered traffic lengths and declared serving regimes; the selector does not promise a TTFT/TPOT target, and KV-cache lifecycle outside the stated memory term remains unmodeled.

Registry/engine identity — reconciled (schema rev 2.2, reviewed 2026-07-27). The table above describes what the live engine actually computes with (per-row decode/prefill roofline η at each listed basis). The internal evidence registry separates source observations from deployed analyst identities and binds every deployed identity by test equality to the engine's calibration registry. Its closed classes are fitted, fitted-inherited, family-transfer, platform-native aggregate bridge, joint-fit, source-informed neutral, analyst-set assumed operating point, and projection. Retired scalar decode/prefill efficiency fields (effDec/effPre) remain archived as historical metadata, not a second live identity. The H20/Ascend relabel does not change their per-row throughput math; it does remove them from the typed throughput-eligible fleet set, so the named anchored-only counterfactual now contains only GB200 and H800.

4 · Chip efficiency: H100 → GB300 → Rubin

Expand this sectionCollapse this section
PlatformDense FP8 (PF)Dense FP4 (PF)HBMBWTDPRental (Jul 2026)
H100 SXM (2022)1.9880 GB3.35 TB/s700 W$1.99–3.90/hr, spot ≈ $2.40
H200 (2024)1.98141 GB4.8 TB/s700 W$2.45–4.50/hr
B200 HGX, per GPU (2024)4.59180 GB8 TB/s1.0 kW$6.69/hr on-demand (Lambda)
GB200 NVL72, per GPU (2024-25)5.010186 GB8 TB/s1.2 kW$3.50–6.00/hr; rack ≈ $3–3.5M; AWS Capacity Block rack rate $761.904/hr ($10.582/GPU-hr) — cleanest full-rack market datum
GB300 NVL72 "Blackwell Ultra" (2025-26)5.015288 GB8 TB/s1.4 kWNo numeric rack rate public on any major provider (Jul 2026); B300 node anchor: Nebius $7.85/GPU-hr ⇒ $0.289/M output (§4)
TPU v7 Ironwood (GA Mar 2026)4.61192 GB7.37 TB/s~1 kW$12/chip-hr on-demand, $5.40 3-yr (Jul 2026); Anthropic deal: up to 1M chips; rental-floor anchor: $6.42/M output tok on-demand (§10) — a vendor saturation run at infinite offered load; no achieved TTFT/TPOT published
Trainium2 (GA Dec 2024) / Trainium3 (GA Dec 2025)1.3 / 2.5196 / 144 GB2.9 / 4.9 TB/s~0.5–0.8 kWTrn2 Capacity Block ≈$2.235/chip-hr (Jul 2026), engineering anchor $68.90–98.65/M output tok; Rainier launched with ~500k Trainium2 confirmed running Claude inference (Nov 2025), Anthropic reported >1M in use by Apr 2026, zero Trn3/Rainier economics disclosed
H800 — China export SKU (2023)1.9880 GB3.35 TB/s700 WIDC annual-commit $1.47–2.06/hr (mid-2026, +30% post-Spring-Festival); DeepSeek's disclosure assumed $2
H20 — China-legal SKU (2024)0.29696 GB4.0 TB/s400 W~$10–12k chip / ~$20k installed; rental class spans ~$0.76 (IDC annual) to $7+ (cloud on-demand)
Huawei Ascend 910C (2024–25)1.50 (INT8; no FP8)128 GB3.2 TB/s~0.6 kW~$23k/chip installed; Huatai procurement $1.71–2.25/hr; CloudMatrix 384 ≈ RMB 60M ($8.2M); CM384 throughput anchored (§4), cost unanchored; 910B (older chip) cost proxy ≈$2.09/M output
Vera Rubin NVL72, per GPU (preliminary 2026 specs)17.5 (dense FP8/FP6)50 sparse NVFP4 inference (35 dense-class training)288 GB HBM422 TB/s~1.8 kWproduction power, throughput and cost TBD — 2026-07-15 dive: no MLPerf, no InferenceX result, no public price; only a relative Kimi-K2-Thinking claim vs GB200 (10× tok/s/MW, ~1/10 cost)

Measured, not marketing: SemiAnalysis InferenceX (the successor to InferenceMAX, the field's independent benchmark) finds the most-optimized GB300 NVL72 delivers ~17× the best H100 config in FP8 and ~32× in FP4 on DeepSeek R1 (Jun 27, 2026) — and, critically, that software alone was a 14× gain on the same silicon (baseline FP8 ~1k → wideEP+disagg ~8k → +MTP ~14k tok/s/GPU). SGLang's reproducible benchmark runs agree: ~11–12k tok/s/GPU on V4 Pro 1.6T (an InferenceX/SGLang benchmark, not production telemetry), 6.5× over B200 with Dynamo disaggregation. At rack level: an H100 rack two years ago did ~8.8k tok/s; a GB300 NVL72 does ~370k — 42× in two years (8× of it HBM growth).

2026-07-15 — GB300 is now audited-throughput-anchored, but its rack price is still not public. A round-2 dive found MLPerf Inference v6.0 (Apr 1, 2026) contains valid, reproducible single-rack GB300 NVL72 results on DeepSeek-R1 FP4 — refuting any "GB300 has no public serving benchmark" framing: NVIDIA Interactive 250,634 / Server 400,437 / Offline 647,076 generated tok/s; Nebius submitted the strongest one-rack results, Server 575,580 / Offline 673,936. These are generated-output tokens (LoadGen sums completed-output-token counts), not input+output totals, and NVIDIA's headline 2.49M tok/s Offline figure is four racks (288 GPUs), not one. What remains missing: no numeric GB300 rack rental rate is public on AWS, CoreWeave, Nebius, GCP, Azure, OCI or Crusoe's pricing pages as of Jul 15, 2026 — so the honest model is GB300 $/M-output as a function of rack-hour price, not a point estimate (SemiAnalysis's own $190.94/rack-hr GB300 figure is explicitly labeled a temporary 1.2×-GB200 placeholder in its source code, not an observed rate). Note on the live calculator: that "function of price" framing describes the honest representation of the newly-documented bridge anchor — the deployed calculator still prices GB300 from an analyst-estimated $6/GPU-hr scenario (no observed market rate), which is 12% of the activated NA-blend default fleet; treat that $6 as a labeled scenario input, not a market anchor. The cleanest present Blackwell-Ultra anchor instead comes from the same-provider B300 node: Nebius's 8-GPU MLPerf Server result (60,413 gen tok/s) paired with its own public $7.85/GPU-hr rate gives $0.289/M generated output tokens — an accelerator-rental floor, before non-GPU serving overhead and utilization reserve (B200 comparably: $0.307/M — B300 is only ~6% cheaper despite ~17% more throughput, because its hourly price is higher). A GB200 rack bridge exists via AWS's Capacity Block rate ($761.904/rack-hr, the cleanest explicit full-rack market datum — though a reserved/upfront Capacity Block price, not ordinary on-demand): $0.881/$0.630/$0.435 per M output at Interactive/Server/Offline. Rubin remains absolute-economics-unanchored — no MLPerf submission, no InferenceX result (listed "Coming Soon"), no public rental or purchase price for Rubin or Rubin Ultra; NVIDIA's only public claim is relative (Kimi-K2-Thinking, 32K-in/8K-out): up to 10× tok/s/MW and ~1/10 cost per M tokens vs GB200 NVL72 — not an absolute anchor, and not to be confused with the different NVL144/Rubin CPX/R100/VR200 configurations. None of this changes the deployed GB200/GB300/Rubin roofline parameters above — it upgrades the evidence quality behind GB300 and documents, rather than fits, the new bridge anchors. (Full NVIDIA forward inference-economics dive →)

The China stack is its own cost universe. Under export controls the Chinese labs serve on three tiers: hoarded pre-ban H800s (H100 compute, capped NVLink — where DeepSeek's disclosure happened), the deliberately compute-starved but bandwidth-rich H20 (the main legal SKU since mid-2025; decode-friendly because decode is bandwidth-bound — Ant Group's production SGLang deployment is the best public anchor), and Huawei's Ascend 910C (Huawei's CloudMatrix-Infer paper reports 1,943 tok/s/NPU decode on R1, while DeepSeek's internal evaluation put the chip at ~60% of H100 — vendor and customer numbers disagree, so the calculator's Ascend row carries wide error bars). H200s were license-cleared for ten named Chinese firms in early 2026 but essentially none had shipped as of May; Beijing reportedly moved to allow those purchases this week (Jul 7–8). Chip scarcity also cuts the other way on price: Chinese H-series lease rates rose 20–30% over early 2026 on a >1,000× surge in national token volume (TrendForce) — Chinese margins are being squeezed from the cost side at exactly the moment Western $/token falls. One more trap: "the China rental rate" is a category error — the same H20 spans roughly ¥5–72 per card-hour (>10×) between annual IDC bare-metal leases and hyperscaler on-demand instances. The calculator's China rows use annual-commit rates; the GPU-hour cost multiplier is the dial for other rental classes.

2026-07-15 — CloudMatrix 384 gets a real generated-throughput anchor; production cost stays unmatched. A round-2 dive found the strongest public CM384 result to date: FlexNPU (arXiv:2606.04415, Jun 3, 2026) served DeepSeek-R1 W8A8 on a full 384-card 910C system at ≈633,000 generated tok/s system-wide (≈1,646/card) under stated TTFT ≤1s / TPOT ≤50ms constraints — end-to-end, SLO-constrained, and a different (full-system) measurement from the isolated-decode figure already in the row above. No public CM384 hourly rental price accompanies it, so throughput is anchored but cost is not. This dive also nails down a total-vs-generated conflation in the widely-repeated CloudMatrix numbers: the "6,688 tok/s/NPU" figure is PREFILL/INPUT throughput and an idealized perfect-expert-balancing projection (the directly measured default is 5,655 prefill tok/s/NPU), while the "1,943 tok/s/NPU" figure already cited above is DECODE at batch 96 (49.4 ms TPOT) — only ≈20 generated tok/s per active sequence, not an interactive per-user rate; these two figures must never be blended as simultaneous end-to-end throughput. A cross-source proxy exists for the older 910B chip: JD's xLLM paper (709 generated tok/s/card on 16 cards) paired with China Telecom CTyun's public RMB 38.45/hr 910B instance rate gives ≈$2.09/M generated output tokens (public-evidence range $1.61–2.80/M) — a cross-source reconstruction, not audited provider COGS, and not the 910C row above. Ascend 920 has no official product page, benchmark, deployment, cloud SKU or price — Huawei's own Sept-2025 roadmap goes 910C → 950PR (Q1'26) → 950DT (Q4'26) → 960 (Q4'27) → 970 (Q4'28), with no 920 — so any "920" cost figure would be fabricated from rumor and stays excluded from this model. None of this changes the deployed Ascend 910C roofline parameters above: the 1,943 and 538 isolated-decode observations remain historical inputs to the frozen joint calibration, while the live row is the separate source-informed neutral η=0.299324 identity and does not reproduce either observation. The FlexNPU full-system result remains an evidence annotation, not a pending automatic refit. (Full Ascend inference-economics dive →)

But cost per token falls slower than throughput rises, because rental prices track capability: $2.40/hr H100 → $6/hr GB300 eats ~2.5× of the 17×. Net hardware-economics gain per generation in $/Mtok: H100→H200 ~1.2×, H200→GB200 ~2–3×, GB200→GB300 ~1.1–1.5× (GB300's edge is HBM capacity for reasoning/long-context, not FLOPS), Rubin ~2–3× again. Cumulative 2024→2027: roughly 5–16× cheaper per token at the hardware level (the product of the per-generation ranges just listed; the commissioned fixed-model projection lands in the same band — ~4.5–10× by the 2027 Rubin ramp, ~7–17× by mature Rubin), before model-side efficiency (sparser MoEs, MTP, quantization) which historically contributed as much again. This is why margins at constant list prices would, on these hardware-cost assumptions alone, trend toward 95%+ — a mechanical consequence of the cost model under a fixed-price counterfactual, not a prediction, and one the market visibly pre-empts, which is why in practice prices fall instead (Opus's 2025-11 cut to $5/$25 was 3×; fast mode — $10/$50 on Opus 4.8, down from 4.7's $30/$150 tier retiring July 24 — shows the latency premium being monetized separately, and itself falling).

5 · The two independent estimates

The two estimate cards and the stress test under them are at the top of this page, beside the answer tile. The reasoning that stood in this position is in the expander below.

Expand this sectionCollapse this section

Verdict: conditional, not blanket. A 90–95% list-price serving contribution margin requires assumptions beyond this page's conservative planning vector. At the page-adopted 2.5T size, the Opus-class policy-labeled scenario lands at ~59% AT THE PUBLIC-EVIDENCE REFERENCE (algorithmic lead 0 months, family multipliers 1.0×); its declared compatible cost lenses span ~59–83% at that same reference. The calculator's own default state additionally carries the owner-ratified +3-month algorithmic-lead prior — a labeled scenario prior, not a measurement — and reads ~69% (lenses ~69–87%). Every CALCULATOR figure in this verdict is the reference reading unless it says otherwise; the 90–95% zone is the external claim under examination, not a reading of this page. All 7 of 7 declared fleet legs render under the loaded-bytes planning policy, but placement remains unverified, so this is not a central or verified estimate. The more serious limit is model form: at the public-evidence reference, holding η fixed while re-expressing legacy traffic produces a 53.24–64.17% calibration-debt span, and the Trainium batch interpretation may make affected legs ~15× wrong. The claim's truth is a function of the invoice, operating point and unresolved form, and it sits atop a stack of margins that shrinks at every step toward the income statement.

My planning scenarios for mid-2026 — balanced latency, 50% fleet utilization, low/committed planning rates, the deployed billing defaults (15% batch share, 5% discount; a pure list-price run lands ~4.9 points higher than the ~59% public-evidence reference reading):

  • Opus-class (assume ~2.5T total / ~300B active): at the public-evidence reference the calculator produces an output-token cost of ≈$6.37/Mtok against $25 list (≈75% output-token margin); blended across the calculator's Reference traffic convention (15:1 I/O, 60% cache; not a measured operating point), the selected cost-lens alternatives span ~59–83% and the policy-labeled scenario is ~59% — both at that reference; with the ratified prior the default reads ~69%. The cited "$4/1Mt" ceiling is not a validation target because the post does not identify input, output or blended tokens.
  • Sonnet-class (~1T total / ~120B active): at Sonnet 5's current $2/$10 introductory tariff (through Aug 31, 2026) the central scenario lands at ~56% blended at the public-evidence reference (algorithmic lead 0 months); changing only to the $3/$15 standard tariff from Sep 1 lands at ~71% on that same reference. With the ratified algorithmic-lead prior the two readings are ~67% and ~78%. All seven declared fleet legs render under the policy, though placement remains unverified. The earlier 90–95% zone is therefore not reproduced by this activated operating point. (Sonnet 5's tokenizer also emits ~30% more tokens per text than 4.6 — per-token and per-text economics diverge across generations.)
  • If active params are lower than my assumption (TeorTaxes's reading of Fable's serving speed), Opus margins move up 5–10 points. The claim's truth is nearly monotonic in one number nobody outside Anthropic knows — but only locally, at fixed fleet membership: this holds within a feasibility boundary, not across one. A total-parameter change that crosses a boundary changes which legs are renderable and can move the displayed margin discontinuously, non-monotonically — but that is a different axis from the one below: the calculator's own sensitivity chart sweeps active parameters only, at a fixed total, so it does not cross or display this total-parameter boundary (at the activated default, all 121 swept points share one fleet membership).

Three things the bull case gets right: (1) the DeepSeek a-fortiori argument is the strongest first-party-grounded argument identified in this research (an anchor for what optimized serving can cost, not proof of any provider's margin — §2); (2) caching can be a margin machine under favorable billing and reuse conditions — cache reads bill at 10% of input price but cost ~1–5% of prefill to serve, so agentic traffic stays high-margin even at deep effective discounts (assuming reuse is billed at the cache tariff; the billable-share caveat in the methods box spans 75.7%→32.3% at the public-evidence reference on that assumption, and cache storage/retention/replication are unmodeled); (3) each hardware generation adds margin at constant prices, and Anthropic sits on three platforms (GB-class NVIDIA, up to 1M TPU v7, >1M Trainium2) with pricing leverage none of its open-weight competitors have.

2026-07-15 — a separate blinded model-generated cross-check lands in a broad adjacent band. A GPT-5.6 Pro dive rebuilt this page's own metric — unit serving margin, defined precisely, not company gross margin — bottom-up from public anchors only, under an explicit instruction not to consult this site or its repository; it disclosed at the end that neither ever appeared in its search results and neither was opened. For an equal-dollar basket of GPT-5.5 / Gemini 3.1 Pro / DeepSeek V4 Pro on a 4M-in/1M-out workload it landed on a central 73.3% (cross-model median 71.8%), with a modeled scenario range of 47–89% (not a statistical interval). This page's reference reading of ~59% (algorithmic lead 0 months) sits below that central figure but inside the range; its ratified-prior default of ~69% sits above the central figure and also inside the range. This is a robustness comparison, not an independent empirical measurement or a matched-estimand validation: it is the same model family operating over the same public source universe, estimating a different model basket, workload and hardware-cost path — and a 47–89% band is wide enough that containing the ~59% reference reading is weak corroboration. Blinding reduces direct copying; it does not create independent evidence. Its top sensitivity driver matches this page's own: delivered tokens/sec/chip at the actual workload (halving it moves the basket to 46.6%; doubling it, to 86.7%). (Full blinded cross-check dive →)

Three things it elides: (1) utilization — fleets are provisioned for peak; changing only utilization from 50% to 35% moves the Opus policy-labeled scenario from 59.18% to 41.69% at the public-evidence reference, a 17.49-point drop (from the calculator's ratified-prior default the same change moves 68.98% to 55.69%, a 13.29-point drop — the lever is real on either basis); (2) the latency premium is not free — interactive serving at user-acceptable speed costs ~1.4–3× throughput-optimal serving (fast mode's price premium exists for a reason — 2× on Opus 4.8, and it launched at 6×); (3) subscriptions and whales — documented tail users extract 15–40× their subscription in list-value tokens (§8); how much of the subscriber base is underwater is unknowable without a usage distribution nobody publishes, but the leakage direction is real and the API numbers never show it.

Forward hardware sensitivity scenario (not a forecast): the GB300→Rubin cost path would imply 92–96% Opus-class output margins by 2027 only if list prices and every non-hardware assumption held fixed (they will not) — a conditional sensitivity output of the cost model, not a prediction, with no probability assigned to future list prices, price cuts, token volume, or realized margins. The open question the scenario sharpens: how long can list prices sit this far above a falling cost floor while open-weight models (DeepSeek V4, GLM 5.2, Kimi) sell adequate quality at 10–20× less?

6 · GPT-5.6 Pro's analysis and verdict

Expand this sectionCollapse this section

Reach this operating point in the calculator: Load the strategic-partner fleet ↑ (GPT-5.6 Pro lens on Opus)

Model-generated scenario analysis; the spans below are uncalibrated sensitivity ranges, not confidence intervals.

GPT-5.6 Pro's verdict: "Standard-list-price marginal serving margin: approximately 92–94% for Opus and 94–96% for Sonnet on a mature 2026 fleet." And: "Is 95% an upper bound? No." — Sonnet at normal $3/$15 exceeds it, Opus on strategic TPU contracts exceeds it (97.2%) — but 95% is "not conservative for every token": at public cloud rates the same math yields only 55–85%.

Where its 93% and this page's ~59% reference-reading policy-labeled scenario actually differ — assumption by assumption, not a single "disagreement":

Assumption§5 (this page's policy-labeled scenario)§6 (GPT-5.6 Pro)Why it matters
Estimand / metricModeled unit direct-serving contribution margin at modeled effective billings, blended across the traffic mix"Economic marginal cost" margin; input and output legs quoted separately, no cache blending in the headlineDifferent metrics can differ by points before any input differs
Traffic mixReference 15:1 / 60% cachePer-leg costs (cache economics treated separately)Cache-heavy mixes raise blended margin at Anthropic's 10% cache tariff
ProcurementRegistered low/committed planning rates (1.0×; a scenario vector, not one observed invoice)Anthropic-scale strategic contracts (SemiAnalysis's Nov 2025 estimate for the 600,000 GCP-rented TPU v7: ≈$1.60/TPU-hr inclusive of Google's margin — an estimate, not a disclosed contract price)The dominant term, alongside differing fleet, architecture and metric assumptions, in the 59-vs-93 gap
Utilization50%75%Linear divisor on all costs
ArchitectureOpus ≈ 2.5T total / 300B active; Φ ≈ 2 × active + L-dependent attention (compute/HBM/fabric roofline, not a flat scalar)Historical 5T total / 300B active; 2.3 × activeThe active count agrees, but total size does not; total geometry changes residency, feasible width and fleet membership
FleetMixed 7-leg default blend (Hopper/Blackwell/TPU/Trainium); all 7 render, placement unverified40% TPU v7 · 25% GB300 · 15% GB200 · 15% Trn2 · 5% H200TPU-heavy blends are cheaper under strategic rates
OverheadStack factor 1.0, composed outside both the per-row decode calibration and universal prefill calibration; cluster overhead modeled in TCO mode+10% CPU/network/reliability; explicitly not electricity-only, not full COGSSmall beside procurement

Its dominant sensitivity is the same as this page's: active parameters (Opus at 150B active → 95–97%; at 600B → 81–87%). The activated GPT-5.6-Pro cost lens lands at ~83%, with all 5 of 5 declared fleet legs numerically renderable; that remains a page-authored scenario, not convergence with the external 93% analysis. The honest summary is therefore conditional on procurement, unverified placement and the open calibration debt — while architecture, traffic mix, latency targets and utilization remain first-order unknowns in their own right. Its other unique contributions are placed where they belong: the PitchBook/Morningstar 44% gross-margin estimate (§7), the GB300 5×-from-software-alone progression (§2c), and its own 93.5%→44% bridge allocation (verbatim annex).

Attribution divergence, preserved: GPT Pro's browsing never located the ncode/Noumena deployment and instead attributed the GB300 GLM-5.2 story to a merger of the SGLang/NVIDIA DeepSeek-V4 result with Baseten's GLM-5.2 work. The Grok X sweep, however, found the actual account — @_xjdr's ncode/Noumena, with primary post URLs (§2c). Where the two engines' account attributions conflicted we kept only what primary post URLs confirm; the direct post links in §1–2 are the ground truth we could verify.

A rare piece of direct evidence on the invoice question surfaced in the xAI dive (§10): the SpaceXAI prospectus discloses that Anthropic pays xAI $1.25B per month for ~325,000 GPUs plus supporting CPUs, storage and networking — about $5.27 per bundled GPU-hour, a real, dated, arm's-length price for capacity Anthropic actually buys. Caveats: it is a bundled service price, not a bare chip-hour, and contracted reserve capacity need not price the marginal fleet. But it brackets one tranche of the debate from above: a disclosed, dated, arm's-length price for capacity Anthropic actually buys. One bundled contract need not bound the blended invoice across all partners; it can be compared with, but does not validate, either this page's registered low/committed planning-rate vector or its strategic-partner scenarios.

2026-07-15 — Trainium gets a narrow public $/token candidate; Project Rainier itself still discloses none. A dedicated dive found AWS publishes two reproducible named-model Neuron benchmarks on a full trn2.48xlarge (16 Trainium2 chips) which, paired with the current Ohio Capacity Block rate (≈$2.235/chip-hr), give infrastructure-only anchors: Llama 3.3 70B (speculative decoding) $68.90/M output tokens; Llama 3.1 405B (FP8-rescaled target weights in a BF16 execution setting, plus speculative decoding) $98.65/M — engineering reference points (batch=1, concurrency=1, 10K-token prompt), not production TCO; the dive warns that calling the 405B run simply "FP8 inference" would overstate the disclosure. The negatives are the more load-bearing finding: no public Trainium3 anchor exists (GA'd with a GPT-OSS-120B recipe but no published throughput and no public instance/UltraServer price — neither side of $/token is public); Project Rainier is confirmed running Claude inference on ~500,000 Trainium2 chips (AWS, Nov 2025) but discloses no inference/training allocation, utilization, token volume, precision or internal rate — production-validation evidence, not a serving-economics anchor; there is no Trainium MLPerf submission through v6.0; and AWS's 30–40% price-performance claim over P5e/P5en remains unreproducible from any public matched benchmark. None of this changes the fleet-blend assumptions in the table above. (Full Trainium inference-economics dive →)

7 · Why reported gross margins are so much lower

Expand this sectionCollapse this section

Reported/leaked figures for Anthropic's business-level gross margin paint a different picture from the unit math. The Information (Jan 2026, from people with knowledge of its financials): Anthropic lowered its 2025 gross-margin projection to 40% (down 10 points from the earlier 50% plan) because inference costs on Google/Amazon servers ran 23% higher than anticipated; including free-tier chatbot inference the margin would be ≈38%. Same reporting: 2025 revenue ≈ $4.5B (~12× 2024's $381M), 86% of it API. SemiAnalysis put the 2024 accounting gross margin at −94% (yes, negative). Zephyr's mid-2026 read is ~70% GM with 15–20% FCF margin; a detailed July 2026 thread claims quarterly gross profit swung from −$55M to ~$453M with inference cost/token down ~40× since early 2024 (@IvanaSpear). The freshest independent estimate is PitchBook/Morningstar (Jun 2026): gross margin ≈ 44%, with compute spend of $0.71 per revenue dollar in Q1 2026, projected $0.56 in Q2. A related forward-looking read from the lease-economics side: Anthropic converting roughly $5B of compute spend into an expected $15B of ARR (May 2026) — a spend-to-revenue multiple, not a margin. Inverting those reported margins back through this calculator does not yield a unique cost story: many combinations of hyperscaler markup, utilization, discounting and free-traffic share rationalize a 40–44% book margin equally well — which is why the reported-margin camp appears here as a diagnostic (the skeptic lens) rather than a claimed parameter set. These observations are not a time series and should not be read as one common-perimeter progression: −94% is SemiAnalysis's estimate of 2024 company gross margin; 40% is a reported internal projection for 2025; 44% is a PitchBook/Morningstar estimate tied to projected compute spend (which is not necessarily accounting cost of revenue); 70% is one analyst's mid-2026 social-media claim. Differences among them can reflect perimeter, method and forecast status as much as operating improvement — and The Information ran a follow-up on why the labs kept missing their own gross-margin forecasts.

2026-07-26 update — a claimed realized-profit data point. At the RAISE Summit (Paris, recorded Jul 9, published Jul 16), SemiAnalysis's Dylan Patel said on stage that Anthropic turned its first gross profit in Q2 (June 2026) and will book slightly over $1B of operating profit in Q3 — figures he described as coming from the financials Anthropic is preparing to disclose in its IPO. This is Patel's own characterization of pre-IPO financials he says he has reviewed, not a company disclosure or filing, and it stays unverified until Anthropic's actual IPO/S-1-equivalent materials surface. If accurate, it would be the first realized-profit claim in this section — every other figure above is a projection (the 40% 2025 GM plan), a modeled estimate (PitchBook/Morningstar's 44%) or a social-media analyst read (Zephyr's ~70%), none of which asserts an actual profit was booked.

Sensitivity and perimeter map — not an accounting reconciliation. The rows below identify one-at-a-time sensitivities and items outside the calculator, roughly in order of size. Several are already embedded in the selected unit result; the ranges must not be added or subtracted from the headline, and they do not reconcile unit contribution margin to company gross margin:

ItemMechanismRough magnitudeAlready in the calculator?
Compute procurement markupAnthropic buys most compute from AWS/GCP, who take their own margin; partner-committed economics can also sit well below public cloud rental. DeepSeek/High-Flyer appear to control substantial compute (the $2/hr figure was a costing assumption; the current owned/rented mix is undisclosed).−10 to −20 ptsYes — the cost lens / rent multiplier
Marketplace / channel revenue shareSome channel revenue shares 30–40% with the clouds (Bedrock/Vertex economics) — a revenue-side item, distinct from procurement cost.−3 to −10 ptsNo — outside the calculator
Utilization & peak provisioningCapacity sized for Monday-morning peak. 35% vs 70% utilization is a 2× on cost.−5 to −15 ptsYes — the utilization divisor (the default already allocates 50% slack)
Subscription over-consumptionFlat-fee plans where the tail extracts 15–40× the fee (§8); weekly caps (Aug 2025) exist to clamp exactly this.−5 to −10 ptsNo — API-billing perimeter only
Free tier & internal inferenceclaude.ai free traffic, evals, RL/synthetic-data generation all burn serving compute against zero revenue (some booked as R&D, treatment varies).−5 to −10 ptsNo — billed traffic only
Discounts & mixEnterprise/committed-use discounts, 50%-off batch tier, and cache-heavy traffic. SemiAnalysis estimated ~$0.99/Mtok for its Opus 4.7 agentic workload at ~300:1 input:output and >90% cache hits — a workload-specific blended billed price, not Anthropic-wide realized revenue per token.−3 to −8 ptsPartly — batch/discount sliders and cache tariffs
Unreconciled accounting perimeterWhat lands in COGS vs S&M vs R&D differs by lab — fleetingbits' point that Anthropic books cloud commissions under sales & marketing cuts the other way, flattering GM.±No — accounting policy, not unit economics

Dollars make the map legible in a way percentage points hide: on the Opus policy-labeled scenario (Reference 15:1/60% traffic, 15% batch share, 5% average discount), 1M blended tokens bill $3.71875 before discounting and $3.26785 realized; modeled direct-serving cost is $1.33392 at the public-evidence reference (algorithmic lead 0 months) across all 7 declared fleet legs, and $1.01356 under the calculator's ratified-prior default (+3 months for the closed labs — a labeled scenario prior, not a measurement) (placement remains unverified). At that fixed denominator one percentage point ≈ $0.03268/Mtok, so a notional 10–20-point cost increase is $0.32679–$0.65357 per million tokens — about 24.5–49.0% of that modeled direct-serving cost. These are scenario conversions at the deployed defaults, not an accounting reconciliation.

Training compute — the industry's historically heavy cash burn — sits below gross margin in R&D — it explains why the company loses money overall (training reportedly falling from 400%+ of revenue toward ~36% by end-2026 per the IvanaSpear thread), not why gross margin is below the unit margin.

8 · The subscription-token investigations

Expand this sectionCollapse this section

Selection rule for this section: multi-month, tool-tracked (ccusage-class) longitudinal records only — single-session anomalies and one-off screenshots are excluded as tail-of-the-tail. The two that qualify: ksred's Claude Code pricing guide — ~10B tokens over 8 months: ~$15,000 at API list prices against ~$800 of Max subscription fees (~19×) — and btcbigd's 60-day record — 8.6B tokens / ~$8.5k at list (~21×). Shorter-window corroborators (the widely-shared melvynx $3,200-capacity calculation, olofj's 2.5-month track, and others) point the same direction and are archived in the annex sweep; the subscription card below defaults to the melvynx capacity figure.

What these investigations do and do not establish: heavy subscribers demonstrably extract API-list-value equivalents many times their fee — $3,200 of sticker sold for $200 in the melvynx case, which at this page's modeled direct serving cost is perhaps $300–900 to actually serve depending on mix. They characterize the tail, not the median — no representative usage distribution, mean or median is public, so nothing here identifies whether the typical subscriber is profitable. What the plans' generosity is consistent with is low marginal serving cost; it does not independently prove it — breakage, throttling, workload routing, acquisition subsidy and cross-subsidy can all contribute. The weekly limits (Aug 2025) and repeated 5-hour-window retuning (May 2026 doubling) show the tail is real enough to clamp. The subscription card in the calculator computes break-even for whatever usage level you set — it says nothing about the distribution.

9 · The most comprehensive public analyses

Expand this sectionCollapse this section

10 · The other providers: what's known, what isn't

Expand this sectionCollapse this section

The Anthropic verdict above rests on an unusual density of evidence: a primary serving disclosure from a direct competitor, a parameter leak from a rival CEO, leaked business financials, and a live X-sphere argument between named analysts. None of the other providers has all of that, and some have almost none of it. Read the cards below as provider-native case studies, not a like-for-like ranking: each numeric headline reports the activated shared-roofline replay — OpenAI, Google, xAI, DeepSeek and Zhipu are workload-specific blended list-price serving contribution margins, while Moonshot is an output-token margin only — and the traffic mixes, service tiers and cost lenses differ materially across cards, so rank order would encode the analyst's assumptions as much as the providers' economics. All current card fleets produce numeric outputs, but that is implementation feasibility, not empirical validation or verified placement. The displayed ranges are analyst-elicited scenario ranges spanning selected downside and upside cases; no statistical probability is attached to them. The badge grades public observability of the evidence base, not numeric reliability, and each card carries a five-dimension evidence profile inside. A deterministic test suite cross-checks every card headline and membership statement against the activated engine. These replays translate each dive into one shared roofline; they do not reproduce each dive's own component model.

OpenAI — GPT-5.6 Sol evidence: medium unit CM ~94% (scenario range 86–97%; uncalibrated; all 4 of 4 declared fleet legs renderable at declared serving topology)

Evidence profile — architecture: medium · pricing: high · fleet & TCO: medium · production throughput: low · financial perimeter: medium

Metric: blended list-price serving contribution margin · lens: partner-fleet economics (dive: 0.85× rental, 73% utilization) · traffic: OpenAI dive mix 9:1 / 78% cache. Provider-native case — not cross-provider comparable.

The closest analogue to the Anthropic case: a partner-hosted fleet (OpenAI discloses 3 GW of dedicated inference capacity on Hopper/Blackwell across Microsoft, OCI and CoreWeave — contractually controlled but not owned), premium list prices (Sol $5/$30 short-context), and an active-parameter estimate (~100B) that is only a community figure — though one independently corroborated by Epoch's inference-economics work. The most interesting tension: the activated ~94% modeled list-price margin coexists with a reported 70% "compute margin for paying users" (The Information, Oct 2025) and a 33% company adjusted gross margin — the gap is take-or-pay reserved capacity, subscriptions and free traffic, i.e. the same margin-stack that separates Anthropic's unit math from its books (§7).

Known knowns (top 3 — full ledger in the grounding ledger and dive)
  • Sol/Terra/Luna list prices: $5/$30, $2.50/$15, $1/$6 (in/out, short-context); batch/flex exactly −50%; cache reads −90% (disclosed)
  • 3 GW dedicated inference capacity; Stargate Abilene runs GB200 via OCI; strategy explicitly "partner-centric" (disclosed)
  • 70% compute margin on paying users (Oct 2025) and 33% adjusted gross margin for 2025, vs a 46% forecast (The Information)
Known unknowns (top 3)
  • Parameters and routing: ~100B active is an estimate; effective-capacity methods allow 3–29T total
  • The Azure/Stargate transfer price ($2.2–4.8/GPU-hr modeled band) and Microsoft's revenue-share percentage
  • Fleet occupancy and aggregate tok/s/GPU — the dominant driver: a >4× cost span on its own

2026-07-26 update — a later, differently-cut margin trajectory (analyst characterization, not company disclosure). At the RAISE Summit (Paris, recorded Jul 9, published Jul 16), SemiAnalysis's Dylan Patel said OpenAI's company-wide gross margin has moved from ~30% (late 2025) to ~55% today, and — stripping away free users — from ~50% to ~65%. This is a differently-cut, more recent figure than the 33%/39% adjusted-GM figures above (The Information/Reuters, FY2025/Q1-2026): Patel's cut explicitly separates paying-user economics from the free-tier drag, the same distinction this page's own unit-vs-accounting-margin framing needs (§7). The figure carries the same caveat as the Anthropic remark from the same appearance (§7): Patel's own characterization of financials he says he has seen, not an OpenAI disclosure.

Why the scenario range is 86–97%: throughput assumptions alone span >4× in cost; reserved take-or-pay capacity can make slack-period allocated cost several times the engineering marginal cost; and the true transfer price is bracketed only by market comparables and Oracle-offtake arithmetic. Would update on: a production tok/s/GPU datapoint, a Stargate/Azure transfer-price disclosure, or a credible parameter leak.

Reach this operating point in the calculator: Reproduce this card ↑ (§10 dive replay)

Full GPT-5.6 Pro report →

Google — Gemini 3.1 Pro evidence: low (params) unit CM ~84% (activated replay; uncalibrated; all 1 of 1 declared fleet legs renderable at declared serving topology)

Evidence profile — architecture: low · pricing: high · fleet & TCO: medium · production throughput: low · financial perimeter: medium

Metric: blended list-price serving contribution margin · lens: internal-cost path (dive: derived ≈$1.28/Ironwood-hr, 75% utilization, throughput batching) · traffic: Reference 15:1 / 60% (no Google-native mix is public). Provider-native case — not cross-provider comparable.

The Gemini margin is not publicly identifiable; the activated calculator now produces a conditional ~84% replay. Its single TPU v7 leg is numerically renderable at the capacity-solved operating width, but the internal TPU price, production placement and Gemini architecture are not public. Google still has the most direct internal-cost path on the page: it designs the TPU stack and serves Gemini on Google-operated infrastructure — though model-to-fleet routing and the internal capacity charge are undisclosed, so the derived ≈$1.28/Ironwood-hour remains an unanchored scenario input rather than an identified margin. It also carries the weakest architecture evidence of any provider: no credible leak of Gemini's total or active parameters exists at all, so the 120B-active/3T-total inputs are scenario midpoints, not estimates. The disclosed facts are all about scale and trajectory: serving unit costs down 78% during 2025, ~3.2 quadrillion tokens/month across surfaces, ~19B API tokens/minute.

Known knowns (top 3 — full ledger in the grounding ledger and dive)
  • List: $2/$12 with $0.20 cache reads ≤200K (higher tier above); no free API tier for 3.1 Pro (disclosed)
  • Serving unit cost −78% during 2025; a further >30% cut on Search AI responses after Gemini 3 (disclosed, earnings calls)
  • Full-stack energy telemetry: accelerators are only 58% of per-prompt energy — the 1.72× stack-overhead anchor (disclosed, paper)
Known unknowns (top 3)
  • Total and active parameters — the dominant driver; genuinely unknown outside Google; a 60–240B active bracket is a 4× cost swing
  • Internal transfer price for TPU time — no credible report exists
  • Production throughput at latency (decode MFU modeled at 4–15%)

2026-07-15 — TPU now has real cost-per-token anchors, but they still don't close on Gemini. A dedicated dive found public named-model, named-precision serving benchmarks now exist for every current TPU generation — the best is Qwen3-Coder-480B-A35B-Instruct-FP8 on four Ironwood chips, 518.86 output tokens/sec/chip at 1K-in/8K-out, which at current list ($12/chip-hr on-demand, $5.40 at 3-yr) derives $6.42/M output tokens on-demand, $2.89/M at 3-yr commitment; supporting anchors include v6e Llama 2 70B at ~$0.94/M and v5e Llama 2 70B at ~$1.77/M. This is real progress over "peak FLOPS only" — but it does not change this card's own dominant unknown. The load-bearing negative the dive was explicit about: there is still no public basis to compute Gemini's margin from TPU economics — the rental prices above are Google's external customer list prices, not its internal fleet cost, and Google has never published a mapping from a named Gemini API SKU to a TPU generation, topology, precision, throughput or utilization. The accelerator-rental anchors are a hardware-platform upgrade, not a Gemini-specific one.

Why the scenario range remains unquantified: the replay's ~84% is a page-authored point, while the two dominant unknowns (active parameters and effective decode efficiency) each span 3–4× in cost and no probability model is justified. Vertical integration would only put a floor under the answer if Google's internal TPU cost and Gemini's architecture were known, and neither is. Would update on: any credible Gemini parameter information, an internal TPU-rate datapoint, or a public Ironwood serving benchmark at stated latency.

Reach this operating point in the calculator: Reproduce this card ↑ (§10 dive replay)

Full GPT-5.6 Pro report → · TPU inference-economics dive →

xAI — Grok 4.5 evidence: medium unit CM ~64% full-cycle-TCO lens (scenario range 10–85%; uncalibrated; all 4 of 4 declared fleet legs renderable at declared serving topology)

Evidence profile — architecture: medium · pricing: high · fleet & TCO: high · production throughput: low · financial perimeter: high

Metric: blended list-price serving contribution margin under the full-cycle-TCO lens (cash-marginal ≈91% and Anthropic-contract opportunity ≈20% are selectable as valuation replays in the calculator) · traffic: Uncached 3:1 / 0% (the dive convention). Provider-native case — not cross-provider comparable.

The one lab that operationally controls a purpose-built fleet — much of it finance-leased, so “owned” is not the relevant cost classification — and the widest judgmental range on the page, because "what does a GPU-hour cost xAI?" has three defensible answers. The SpaceXAI prospectus disclosed cluster counts that sum to >440k accelerators (200k+ H100/H200/GB200 at Colossus, 110k GB200 + 110k GB300 at Colossus II — the total is a derivation from those disclosed counts) — but $20.2B of it sits on finance leases, and AI capex ran $12.7B in 2025 alone. Grok 4.5 (launched Jul 8 at $2/$6) is disclosed at 1.5T total parameters; active count is anyone's guess (100–500B). The decisive fact: Anthropic pays xAI $1.25B/month for ~325k GPUs (≈$5.27/GPU-hr) — so, to the extent that capacity is fungible with the Anthropic contract, a GPU-hour spent serving Grok instead of billed to Anthropic forgoes a disclosed wholesale price more than twice xAI's estimated full-cycle cost. (This opportunity cost is bounded by the ~325k contracted GPUs; it does not apply uniformly to every GPU-hour across the whole >440k fleet.) In the activated replay, cash-marginal cost lands at ~91%, full-cycle TCO at ~64%, and the Anthropic-contract opportunity value at ~20%; all four declared fleet legs now render numerically, though H100 is batch-capped and placement remains unverified. Zephyr's "xAI isn't juicing margins" is true or false depending entirely on which lens you pick.

Known knowns (top 3 — full ledger in the grounding ledger and dive)
  • Grok 4.5 = 1.5T-total MoE ("V9 foundation model", Musk; MoE per Cursor); 80 tok/s user streams; $2/$0.50/$6 list (disclosed)
  • Fleet: >440k accelerators across Colossus I/II — but $20.2B sits on finance leases: "owned" ≠ paid-for (prospectus, disclosed)
  • Anthropic capacity contract: $1.25B/mo for ~325k GPUs ≈ $5.27/GPU-hr; Google (from Oct 2026): $920M/mo for ~110k ≈ $11.45/hr (disclosed/derived)
Known unknowns (top 3)
  • Saturated aggregate throughput — 80 tok/s is per stream; the dominant driver (~8× cost swing)
  • Active parameters and routing on the 1.5T MoE (100–500B — the Grok-2 lineage ran unusually dense at 42.7%)
  • How capacity is allocated among Grok, Anthropic, Google, Cursor, training, and idle

Why the scenario range is 10–85%: unknown batch throughput spans ~8× in cost, and the GPU-hour can be honestly valued anywhere from $0.60 (cash) through $2.40 (full-cycle) to $5.27 (contracted opportunity cost) — the margin question dissolves into the valuation question. Would update on: any saturated-throughput datapoint, new capacity-contract terms in SpaceXAI filings, or an active-parameter disclosure.

Reach these operating points in the calculator: Reproduce this card ↑ (full-cycle TCO) · Cash-marginal valuation ↑ (≈91%) · Anthropic-contract opportunity cost ↑ (≈20%)

Full GPT-5.6 Pro report →

DeepSeek — V4 Pro evidence: mixed unit CM ~74% (scenario range 45–83%; uncalibrated; all 1 of 1 declared fleet legs renderable at declared serving topology)

Evidence profile — architecture: high · pricing: high · fleet & TCO: low · production throughput: low · financial perimeter: low

Metric: blended list-price serving contribution margin · lens: dive central scenario (≈$1.50/H800-equivalent-hr, their own stack) · traffic: DeepSeek disclosure 4:1 / 56% cache. Provider-native case — not cross-provider comparable.

The inversion of the Anthropic story. DeepSeek has unusually strong public evidence — disclosed architecture (1.6T total / 49B active, selective FP4), disclosed pricing, and the clearest production serving disclosure identified in this research — yet one of the lower provider-native estimates (the cards are not directly rankable), because it chose to convert its efficiency into price: a permanent 75% cut took V4 Pro to $0.435/$0.87, roughly 60% below R1's old output price. The activated ~74% replay is still only a sensitivity surrogate: the engine applies its FP8 precision tuple globally and does not model V4's component-selective FP4/indexer/KV precisions, so the open checkpoint does not make this particular number exact. Repricing the old disclosed workload at today's tariff turns the famous 84.5% margin into ~67% before any cost improvement — and V4-Flash would be underwater on the old cost structure.

Known knowns (top 3 — full ledger in the grounding ledger and dive)
  • V4 Pro: 1.6T/49B; V4-Flash: 284B/13B; FP4 is selective (routed experts + indexer; KV stays BF16/FP8) (disclosed)
  • List (Jul 9): $0.003625 cache-hit / $0.435 miss / $0.87 out — after a permanent 75% cut (disclosed)
  • The 2025 H800 production trace: $2/hr assumption, 84.5% theoretical margin at R1 list (disclosed)
Known unknowns (top 3)
  • Current fleet mix (H800 vs H20 vs Ascend 950 vs internal) — the dominant driver
  • V4 production throughput — no V4 equivalent of the 2025 73.7k/14.8k tok/s disclosure
  • The 2026 owned/rented split and true chip-hour basis

2026-07-15 update. The Information reported (Jul 14–15, via @jukan05/@jingyanghk relays) DeepSeek's annualized revenue nearing ~$500M, a ~¥50B (~$7.4B) raise at a ~¥500B (~$74B) valuation with STAR Market IPO prep, and — in a follow-up — a 70–80% V4 API gross margin (up from an earlier ">50%" figure; The Information). This is the first press-reported (The Information, via relays) post-V4 API-margin datapoint identified in this research — not company-confirmed, and carrying no disclosed accounting definition — corroborating in direction, though not metric-identical to, the disclosure-era a-fortiori argument (§2a). TeorTaxes's fleet backsolves from the $500M figure bracket a wide range depending on utilization assumptions: at V4 list output prices and 100% utilization, ~$88.3K/GPU/yr implies ~5,662 GPUs; at more realistic bundle-cost utilization, $28K–$55.7K/GPU/yr implies ~9,000–18,000 GPUs — a useful independent check against this dive's own dominant unknown (current fleet mix), though TeorTaxes's own caveat applies: the math covers paid API only, not the consumer app.

2026-07-15 — a non-NVIDIA production data point, corrected. DigitalOcean and RadixArk publicly reported serving DeepSeek V4 Pro (FP4) on AMD Instinct at 3,500+ aggregate tok/s/GPU (crediting HIP graphs for a claimed ~10× gain over an undisclosed baseline), attributed to a June 2026 InferenceX result. A follow-up dive traced this to primary source: the actual InferenceX benchmark is MI355X-specific — MI350X is not an InferenceX-benchmarked SKU at all, so DigitalOcean's "MI350X/MI355X" phrasing is marketing shorthand, not a second distinct result. The underlying InferenceX article (ISL 8192/OSL 1024, single MI355X 8-GPU node) independently confirms the headline is genuinely reproducible progress — 2,256 aggregate tok/s/GPU at 9.4 generated tok/s/user on 2026-05-21, a 110.5× gain over the 2026-04-25 first-light point — and specifically describes extending the concurrency sweep to 1,024 as drawing out "the high-throughput, low-interactivity end of the frontier." That confirmed mechanism means DigitalOcean's later 3,500+ figure, from that same low-interactivity end of the curve, should not be read as an interactive per-user generation rate. A more granular generated-throughput/TTFT split and a derived TensorWave $/M-output cost proxy also appeared in the same dive, but they came from a GPT-5.6 Pro run that hit a 120-minute hard timeout and returned only a reasoning-summary with no pinned source URLs; those exact figures could not be independently verified against InferenceX's chart-rendered data, are shown nowhere on this page, and are excluded from the analysis pending source recovery. This remains a third-party deployment of the open-weight model, not DeepSeek's own production fleet, so it still does not close the "V4 production throughput" known-unknown above. (Full AMD inference-economics dive (partial) →)

Why the scenario range is 45–83%: the estimate scales the disclosed 2025 baseline by two unknown factors — hardware-hour cost (0.5–1.1× the old basis) and per-token work for the bigger-but-sparser V4 (1.05–1.5×) — which compound to a 3× cost span. Would update on: a V4-era serving disclosure (the 2025 one set the standard), realized peak-surcharge data after mid-July, or fleet-mix reporting.

Reach these operating points in the calculator: Reproduce this card ↑ (dive replay) · DeepSeek Feb-2025 disclosure replay ↑ (≈87%, on R1) · Ant-informed H20 scenario ↑ (live 649 tok/s/GPU on R1; no SLO enforcement) · China public-cloud on-demand lens ↑ (on R1)

Full GPT-5.6 Pro report →

Zhipu / Z.ai — GLM 5.2 evidence: mixed unit CM ~−166% (activated dive replay; historical analyst range 35–77%; all 3 of 3 declared fleet legs renderable at declared serving topology; every leg batch-capped and policy-unclean)

Evidence profile — architecture: high · pricing: high · fleet & TCO: low · production throughput: low · financial perimeter: high

Metric: blended list-price serving contribution margin · lens: audit-implied procurement haircut (dive: 1.9× rent, 75% utilization, 0.60 stack) · traffic: site-assumed 8:1 I/O + the ncode-week 41% cache observation. Provider-native case — not cross-provider comparable.

The rare provider with audited numbers: Zhipu's Hong Kong listing (Jan 2026) forces disclosure nobody else makes. Its cloud/API segment ran a −0.4% gross margin in H1 2025 (price-war casualty) recovering to 18.9% for FY2025. The activated dive replay does not reproduce the earlier mid-50s price-only counterfactual: it lands near −166%, with all three fleet legs numerically renderable only after being capped to substantially smaller batches; all three are policy-unclean, so the number is not a policy-clean result. The architecture is fully disclosed (744B/40B, open weights), and third-party reseller prices give a conditional cost ceiling (valid only if the reseller serves at positive contribution margin — the same conditionality noted on the Kimi card): DeepInfra sells GLM-5.2 below Z.ai's own list.

Known knowns (top 3 — full ledger in the grounding ledger and dive)
  • GLM-5.2: 744B/40B, BF16+FP8 open checkpoints, 1M context (disclosed)
  • List: $1.40 / $0.26 cached / $4.40 out (disclosed)
  • Audited FY2025: cloud/API revenue RMB 190.4M at 18.9% GM (H1: −0.4%); company net loss RMB 4.7B (HKEX filings)
Known unknowns (top 3)
  • Traffic share by chip platform — nine named, zero percentages
  • The confidential compute-service rate — with decode throughput, the dominant drivers (public Chinese card-hour prices span >4×)
  • Coding Plan consumption distribution — claimed 15–30× (even "~100×") API-value leverage vs actual breakage

Why the scenario range is 35–77%: decode throughput and the confidential compute price each move the answer ±10–20 points; the audited 18.9% segment-margin anchor and the conditional reseller price ceiling are what keep the interval from being wider still. Would update on: the H1 2026 interim filing (does cloud/API GM keep climbing from 18.9%?), any platform-share disclosure, or reseller-economics evidence.

Reach this operating point in the calculator: Reproduce this card ↑ (§10 dive replay)

Full GPT-5.6 Pro report →

Moonshot — Kimi K2.7 Code evidence: medium output-token CM ~84% OUTPUT-TOKEN margin (scenario range 55–91%; uncalibrated; all 2 of 2 declared fleet legs renderable at declared serving topology)

Evidence profile — architecture: high · pricing: high · fleet & TCO: low · production throughput: medium · financial perimeter: low

Metric: OUTPUT-TOKEN margin only (a blended margin needs Moonshot's undisclosed traffic mix) · lens: dive replica (H200-class decode floor, their stack) · traffic: n/a for the output metric; the calculator's blended headline uses the Kimi dive mix 8:1 / 40%. Provider-native case — not cross-provider comparable.

An output-token margin, not blended — so despite its large headline it cannot be ranked against DeepSeek’s and Zhipu’s blended figures. K2.7 kept a premium output price ($4.00/M — 4.6× DeepSeek V4 Pro's post-cut $0.87) on a 1T-total, 32B-active, selectively INT4 architecture. The activated calculator does not implement that component-level INT4 layout: its Kimi preset uses the global FP8 tuple as a sensitivity surrogate, so the ~84% headline must not be read as an exact open-checkpoint reconstruction. A reproducible 128×H200 benchmark still supplies an external short-context decode-accelerator floor at ~$0.21/M output, while SemiAnalysis's same-lineage runs span $0.14–$1.00/M by hardware and speed.

Known knowns (top 3 — full ledger in the grounding ledger and dive)
  • K2.7 Code / K2.6: 1T total, 32B active, 384 experts, 256K context; selective weight-only INT4 (595GB); no native MTP (disclosed, open config)
  • List: $0.95 / $0.19 cached / $4.00 out; HighSpeed exactly 2× price for 5–6× user speed; batch = 60% of standard (disclosed)
  • Reproducible K2 decode floor ~$0.21/M output (128×H200); same-lineage cost curves $0.14–$1.00/M (LMSYS, SemiAnalysis)
Known unknowns (top 3)
  • Current accelerator mix — A800/H800 verified only historically; H20/Ascend production use unproven
  • Production throughput/batching at their latency target — the dominant driver (32→90 tok/s/user alone moves cost 2.4×)
  • Owned-vs-rented split and effective GPU-hour basis (China retail proxies $1.01–$2.01/hr)

Why the scenario range is 55–91%: the open weights pin the architecture but not the deployment — per-user speed targets and the actual fleet each move serving cost 2–3×; the $0.21 decode floor and the $3.50 reseller price bracket the answer from both sides. Note this is an output-token margin; a blended margin needs Moonshot's undisclosed traffic mix. Would update on: a K2.7-specific serving benchmark, any fleet disclosure, or revenue-mix updates against the >$300M ARR report.

Reach this operating point in the calculator: Reproduce this card ↑ (output-token dive replay)

Full GPT-5.6 Pro report →

Same-assumption scenario outputs — one normalized lens (not a ranking)

Read this as scenario outputs, not estimates or a leaderboard: these are deterministic results of one arbitrary shared lens, they propagate no input uncertainty, and they replace each provider's own assumptions — so the row order encodes this page's chosen lens, not the providers' relative economics. For readers who want one comparable row per provider anyway, this table is computed live by the calculator with the lens held fixed (the registered low/committed planning-rate vector, 50% utilization, balanced latency — this page's central read) and the traffic mix pinned to the Reference profile (15:1 input:output, 60% cache hits — regardless of the interactive traffic selector above), while keeping each provider's own prices, cache tariff and fleet. This is deliberately not each provider's operating point — it prices everyone through one page-authored procurement vector — and DeepSeek V4's negative number is the honest consequence of its post-price-war tariff under those scenario economics (its own operating point is the ~74% card above). Margins are rounded to whole points. These are deterministic outputs of one normalized scenario; the provider-native §10 ranges do NOT apply after changing the lens and traffic mix, and this table does not propagate input uncertainty. Excluded tariff-only scenarios with unidentified architecture: GPT-5.6 Terra, GPT-5.6 Luna, Gemini 3.5 Flash.

Method note: X posts located and quoted via a Grok 4.5 agent sweep (2026-07-09); hardware/TCO data via Parallel deep research + Exa across SemiAnalysis, MLCommons, NVIDIA/Google/AWS primary pages; cross-model verification via an independent GPT-5.6 Pro research run. The per-provider audits in §10 come from one dedicated GPT-5.6 Pro deep dive per provider plus two Chinese-accelerator hardware sweeps (all 2026-07-09, archived in the research annex; provenance headers condensed for public release, conclusions unchanged). Attributions (the direction of the Musk leak, the ncode/Noumena identification) were verified against primary posts; where the research engines disagreed, both readings are shown rather than harmonized. Model sizes for closed models remain estimates — the calculator exists so you can disagree with a slider instead of a screenshot.

Feedback

Feedback on this public research page — not a support channel; please don't submit sensitive personal data. Submissions are reviewed manually and periodically purged.

Cloudflare's anti-abuse challenge loads only after you interact with this form.