← Frontier Inference Margins · all research reports
Research reports
The full public research artifacts behind Frontier Inference Margins: one GPT-5.6 Pro deep dive per provider, hardware sweeps, and the X-post hunts. Bulky by design — they carry the citations. Provenance headers were condensed for public release (conversation IDs and raw interface citation markers removed); conclusions are unchanged and pre-edit copies are archived.
- Anthropic — GPT-5.6 Pro independent consult — The original 54-minute Pro research run behind report §6: Opus 92–94% / Sonnet 94–96% at strategic rates.
- OpenAI — GPT-5.6 Pro deep dive — GPT-5.6 family serving economics: Azure/OCI/CoreWeave fleet, pricing, margin verdict with a judgmental uncertainty range.
- Google (Gemini) — GPT-5.6 Pro deep dive — TPU vertical integration, Ironwood internal TCO, Gemini pricing, margin verdict with a judgmental uncertainty range.
- DeepSeek — GPT-5.6 Pro deep dive — V4/V4 Pro economics, post-disclosure margin evidence, China fleet reality, margin verdict with a judgmental uncertainty range.
- Zhipu / Z.ai (GLM) — GPT-5.6 Pro deep dive — GLM-5.2 economics, HK IPO financial disclosures, domestic fleet, margin verdict with a judgmental uncertainty range.
- Moonshot (Kimi) — GPT-5.6 Pro deep dive — K2.x economics, reseller price floors, fleet evidence, margin verdict with a judgmental uncertainty range.
- xAI (Grok) — GPT-5.6 Pro deep dive — Grok 4.x economics on the owned Colossus fleet, pricing, margin verdict with a judgmental uncertainty range.
- Chinese accelerators — GPT-5.6 Pro deep dive — Ascend 910C/950, H20/H800, CloudMatrix: specs, costs, throughput anchors, export-control state.
- Huawei Ascend — web sweep — Cross-check sweep: 910B/910C/950 specs, CloudMatrix 384 pricing, MFU anchors, who serves on Ascend.
- H800 / H20 / export controls — web sweep — Cross-check sweep: specs, China pricing, the 2025–26 export-control timeline, fleet reality, H20 throughput anchors.
- X-sphere margin claims — Grok 4.5 sweep — The primary-post hunt behind §1–2: TeorTaxes, Zephyr, Jukan, fleetingbits, DeepSeek disclosure threads.
- GLM-on-GB300 & subscription plans — Grok 4.5 sweep — The ncode/Noumena GB300 deployment posts and the Claude plan-tokenomics investigations.
- The final answer — rationale and evidence chain — The load-bearing rationale behind the landing FINAL-ANSWER block: the estimand, every evidence link in the chain, the policy identities, and what disclosure would move the answer.
- Methods: leave-one-anchor-out validation — The falsification test behind methodology v2: a single scalar MFU fails to transfer across platforms (mean error 37%), and the roofline follow-up's negative result.
- External review #1 — four-persona council (Opus synthesis) — Unedited adversarial pre-publication review: skeptic, architect, risk analyst, empiricist + synthesis. Drove methodology v2.
- External review #2 — GPT-5.6 Pro — Independent adversarial review with exact replacement wording; confirmed the six §10 numbers transferred faithfully. Drove methodology v2.
- Design consultation — four-persona council (Opus synthesis) — The typed-ontology, attribution-honesty and ship-list adjudications behind methodology v2.1's preset expansion.
- Preset grounding pack — GPT-5.6 Pro — First-party tariff verifications and locked parameter decisions for the v2.1 presets.
- Roofline model consultation — GPT-5.6 Pro — The negative result that keeps anchor fits: a physically-informed roofline fails the whole-platform gate and the preregistered test is formally unrunnable on public data.
- Final review — four-persona council on v2.1 (Opus synthesis) — The NO-SHIP gate that produced v2.1.1: lens-range membership, xAI operating-point, attribution and permalink-identity repairs.
- Final review — GPT-5.6 Pro on v2.1 — Independent final gate on the finished product.
- v2.1.2 plan review — four-persona council (Opus synthesis) — Pre-implementation review of the traffic-mix/reception-audit plan: seven P0 conditions, the xAI 26.95/36.52 conflict proof, and two live attribution defects found.
- v2.1.2 plan review — GPT-5.6 Pro — Pre-implementation review: traffic-mix state contract, provenance rules for regenerated artifacts, and the redesign of named-person reception testing into corpus-bounded source audits.
- Changelog — Dated revision history — what changed in each methodology revision and why.
- Anthropic consult — verbatim original (recovered) — The complete 54-minute GPT-5.6 Pro response, original bytes recovered from the receiving session transcript, SHA-256-stamped.
- Roofline consultation — verbatim full derivation (recovered) — The complete roofline derivation behind the adopted negative result — original bytes, SHA-256-stamped.
- Adopted grounding ledger (machine-generated) — Every preset parameter with value, source and evidence label, generated from the deployed registry — the authoritative parameter record.
- Preset pack re-audit — GPT-5.6 Pro re-emission + delta — Dated author-model re-emission of the expired 192-row pack, with a checked delta table against the adopted values (the ledger wins).
- Reception phase — synthesis & resolution ledger — Nine simulated-reader audits (five corpus-bounded source-faithfulness, four audience archetypes) against the frozen v2.1.2 preview: 16 P0s and ~26 P1s found, dispositioned finding-by-finding. A simulated-reader exercise, not validation.
- Reception phase — corpus manifest — The fixed source corpus behind the faithfulness audits, with its stated circularity limit.
- Google TPU inference economics — GPT-5.6 Pro deep dive — Recovered Targeted-4 dive: named-model, named-precision serving anchors for TPU v5e/v6e/v7 (Ironwood), paired with current GCP rental prices — plus the rigorous negative that no public path exists from TPU rental prices to Gemini's internal margin.
- AWS Trainium inference economics — GPT-5.6 Pro deep dive — Recovered Targeted-4 dive: a narrow engineering-only Trainium2 $/token anchor from two AWS Neuron tutorials, plus rigorous negatives for Trainium3 and Project Rainier (500K+ Trn2 chips confirmed running Claude inference, zero economics disclosed).
- Blinded unit-margin cross-check — GPT-5.6 Pro deep dive — A from-scratch bottom-up unit-serving-margin estimate built under an explicit instruction not to consult this site or its repository; disclosed blind held. A robustness comparison, not an independent empirical measurement or matched-estimand validation. Central 73.3%, cross-model median 71.8%, range 47–89%.
- AMD Instinct inference economics — GPT-5.6 Pro deep dive (partial, 120-min timeout) — Recovered reasoning-summary only (the MCP's 120-minute hard timeout hit before a full report was emitted): refines the DigitalOcean/RadixArk MI350X claim to MI355X and a total-vs-generated-throughput caveat; AMD anchor-quality negatives (MLPerf, Azure, Oracle).
- NVIDIA forward (GB300/Rubin) inference economics — GPT-5.6 Pro deep dive — Round-2 dive: MLPerf v6.0 audited GB300 NVL72 generated-throughput anchors, a clean B300 $/token pair, a GB200 rack-price bridge — and a rigorous negative that Rubin's economics remain fully unanchored.
- Huawei Ascend / CloudMatrix 384 inference economics — GPT-5.6 Pro deep dive — Round-2 dive: a full-system CM384 generated-throughput anchor and a cross-source 910B $/token proxy — plus rigorous negatives on CM384 cost, the 6,688-tok/s prefill/decode conflation, and the nonexistent Ascend 920.