| H100 SXM |
flops.bf16 |
990000000000000 FLOP/s |
published — 989 TF dense BF16 (NVIDIA datasheet); equals fp8/2, consistent with fallback rule C10 |
| H100 SXM |
flops.fp8 |
1980000000000000 FLOP/s |
published — 1,979 TF dense FP8 (NVIDIA datasheet; = engine.js flopsFp8 1.98 carried) |
| H100 SXM |
hbmBytes |
85520809984 B (85.521 decimal GB) |
observed framebuffer — nvidia-smi total 81,559 MiB × 2^20 = 85,520,809,984 B for 'NVIDIA H100 80GB HBM3' (verified 2026-07-18 across 6+ independent pasted outputs: scaleway.com MIG guide, github.com/wilicc/gpu-burn#114, docs.databricks.com H100 starter, support.crusoecloud.com, github.com/ray-project/ray#53266). The '80 GB' label × 1e9 undercounts this by ~6.9% (R5 finding). |
| H100 SXM |
bwHBM |
3350000000000 B/s |
published — H100 SXM HBM3 3.35 TB/s (NVIDIA datasheet; = engine.js bw 3.35 in SI B/s) |
| H100 SXM |
fabric |
900000000000 B/s |
published — NVIDIA H100 datasheet: NVLink (4th gen) 900 GB/s aggregate bidirectional per GPU (nvidia.com/en-us/data-center/h100, verified 2026-07-18). Same convention as the frozen h800 400e9 (the export-capped variant of this same fabric). |
| H100 SXM |
nShard |
144 |
N_shard; analyst-declared strong topology carry — 144, copied from the sourced DeepSeek H800 EP144/DP144 decode deployment unit. H100 and H800 have identical registered compute, HBM bandwidth, and framebuffer capacity; their fabric difference does not enter MoE feasibility, and N_shard does not enter MoE decode throughput. No H100-specific EP144 deployment receipt exists: this is a physically coherent same-capacity deployment scenario, not an observed H100 width. |
| H100 SXM |
etaDec |
0.313491 |
FITTED-inherited (h800 fit; 'H800-class compute', memo §2); evidence=fitted-inherited; inherits h800's deployed η; reproduces h100's current deployed 1,872.9730 tok/s exactly (identical FLOPS/HBM constants; binding t_H unaffected by the h100/h800 fabric difference — shown in the derivations file) |
| H100 SXM |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=144; sources=inherits CALIBRATION.h800 (FITTED-inherited; memo §2) |
| H100 SXM |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| H100 SXM |
rent |
$2.4/hr |
price evidence=analyst-set; economic basis note below |
| H100 SXM |
capex |
$25000 |
engine.js HW economic input; economic basis note below |
| H100 SXM |
tdp |
0.7 kW |
engine.js HW economic input; economic basis note below |
| H100 SXM |
economicBasisNote |
— |
2022. Anchor: DeepSeek served V3/R1 on H800 (H100-class compute). Prefill MFU is reconstructed FRESH-only (the disclosed 73.7k/node input flow includes the 56.3% disk-cache-hit share). |
| H200 |
flops.bf16 |
990000000000000 FLOP/s |
published — 989 TF dense BF16 (H200 datasheet); = fp8/2 (C10-consistent) |
| H200 |
flops.fp8 |
1980000000000000 FLOP/s |
published — 1,979 TF dense FP8, same Hopper compute as H100 (= engine.js flopsFp8 1.98) |
| H200 |
hbmBytes |
150754820096 B (150.755 decimal GB) |
observed framebuffer — nvidia-smi total 143,771 MiB × 2^20 = 150,754,820,096 B for 'NVIDIA H200' (verified 2026-07-18, 5+ independent pasted outputs: github.com/NVIDIA/cuda-samples#311, forums.developer.nvidia.com peermem thread, github.com/axolotl-ai-cloud/axolotl#2688, userguide.ncshare.org, github.com/NVIDIA/cccl#1672). Note the '141 GB' label reads as ~GiB: observed FB ≈ 150.75e9 B. |
| H200 |
bwHBM |
4800000000000 B/s |
published — H200 HBM3e 4.8 TB/s (NVIDIA datasheet; = engine.js bw 4.80) |
| H200 |
fabric |
900000000000 B/s |
published — NVIDIA H200 datasheet: NVLink 900 GB/s (4th gen, SXM; PNY-hosted NVIDIA H200 NVL/SXM datasheet, verified 2026-07-18) |
| H200 |
nShard |
144 |
N_shard; analyst-declared weak topology transfer — 144, copied from the sourced DeepSeek H800 EP144/DP144 decode deployment unit. H200 has the same registered compute but 1.76× the framebuffer and higher HBM bandwidth; DeepSeek states large EP is chosen partly to enlarge aggregate/per-expert batch, so EP144 remains plausible, but the added memory also permits a rational narrower deployment. No H200-specific width receipt exists. This topology assumption is independent of, and stacked with, the separate H800-η family transfer; η inheritance is not evidence for N_shard. |
| H200 |
etaDec |
0.313491 |
family-transfer (in-family extrapolation off the H800/H100-class fit — tier-(b), memo §1; NOT in the joint-η bucket); evidence=family-transfer; h800's deployed η applied to h200 hardware constants (memo §1/§2). NOT tuned to reproduce the current hand-set effDec 0.085: at F1-op-point conditions the roofline gives ≈2,683.7 tok/s vs the current 2,274.3 (+18.0%) — an EXPECTED family-transfer movement, deltas tabulated at slice 2/3 (derivations file, finding 4) |
| H200 |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=144; sources=family-transfer off CALIBRATION.h800 (memo §1; derivations finding 4) |
| H200 |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| H200 |
rent |
$2.9/hr |
price evidence=analyst-set; economic basis note below |
| H200 |
capex |
$32000 |
engine.js HW economic input; economic basis note below |
| H200 |
tdp |
0.7 kW |
engine.js HW economic input; economic basis note below |
| H200 |
economicBasisNote |
— |
Same compute as H100, 1.76× HBM capacity/1.4× bandwidth → better batching. |
| GB200 NVL72 |
flops.bf16 |
2500000000000000 FLOP/s |
C10 fallback — fp8/2 = 2.5 PF dense; consistent with NVIDIA's published precision ladder |
| GB200 NVL72 |
flops.fp8 |
5000000000000000 FLOP/s |
published — 5.0 PF dense FP8 per GPU (engine.js HW note: NVIDIA sparse/dense split verified; the widely-copied 4.5 PF is B200's) |
| GB200 NVL72 |
flops.fp4 |
10000000000000000 FLOP/s |
frozen — 10 PF dense NVFP4 (d2 §3.1 F4 recipe FLOPS=10e15 / roofline-diagnostic GB200 case; = 2× dense FP8 per NVIDIA ladder) |
| GB200 NVL72 |
hbmBytes |
198674743296 B (198.675 decimal GB) |
observed framebuffer — nvidia-smi total 189,471 MiB × 2^20 = 198,674,743,296 B for 'NVIDIA GB200' per GPU (verified 2026-07-18: docs.nvidia.com NIM openfold3 arm64 prerequisites, support.crusoecloud.com NVBandwidth-on-GB200 guide, github.com/sgl-project/sglang#8911 and #7504). Distinct from standalone B200 180 GB; ≈185 GiB per GPU, the 186-GB-class B200-in-GB200 part. |
| GB200 NVL72 |
bwHBM |
8000000000000 B/s |
published — 8 TB/s HBM3e per GPU (frozen F4 recipe hbmBps 8.0e12; = engine.js bw 8.00) |
| GB200 NVL72 |
fabric |
1800000000000 B/s |
frozen — NVLink5 1.8 TB/s per GPU (d2 §3.1 F4 recipe / roofline-diagnostic.mjs GB200 case; memo §6 known value) |
| GB200 NVL72 |
nShard |
8 |
N_shard; declared analyst default (unchanged from the pre-existing engine value), precision-conditioned — multiple Class C/D anchors support 8 as a genuinely deployed low-precision replica width, but per the 2026-07-20 GPT Pro evidence consult, the public record does not establish a specific required or minimum replica width for a multi-trillion-parameter closed-weight model at unknown precision (the actual Opus/5T scenario this engine models); see the declared-sensitivity case list on this row and research/gptpro-reports/2026-07-20-replica-width-consult.md. |
| GB200 NVL72 |
etaDec |
0.315997 |
FITTED (deployed; computed per §2 rule — slice-1a derivation; b9 M1: precision double-credit REMOVED); evidence=fitted; 0.315997 IS THE FP4-BASIS COEFFICIENT, and its use in the FP8 default render is a DECLARED CONSERVATIVE TRANSFER, not a bridge (owner ruling q-im-fp4-gb200-eta-basis, 2026-08-02, accepted default; relabel only — no number moves). Executed through this repo's own roofline at F4's operating point, 0.315997 reproduces the 10,108 tok/s anchor to 0.06% ON THE FP4 TUPLE (10,114.4) and yields 6,408.5 on the FP8 tuple; the coefficient that reproduces the anchor in the FP8 basis is 0.498413. So at the FP8 default this row under-predicts its own best Blackwell measurement by ~1.58× — conservative, which is why every overclaim-tuned review passed it. THE RETIRED CLAIM, recorded so it cannot come back: b9 M1 (r4 defect D3, run B §B4/§C1) derived 0.315997 = 0.585795 / 1.8538 to de-embed the retired NVFP4 precision scalar, and called the result 'the FP8 bridge value', defended as equalling F4's own implied η 0.316 'an independent confirmation, not a second fit'. It is not independent: research/im3-slice1a-derivations.md:69-70 records that 0.585795 was CONSTRUCTED as 0.316 × 1.8538, so dividing it back cannot fail to return 0.316. The de-embedding correctly removed a double-counted scalar; it did not convert the basis. Slice-1a authored-decision item h (:204-207) registered the FP4 calibration basis as a choice FOR REVIEW with the FP8 alternative pre-computed (10,135.1 / 8,581.1 tok/s) — REOPENED, and the vLLM source-dtype resolution is a registered evidence task (research/update-queue.md). Three incompatible senses of 'the FP8 basis' are live in the documents (the legacy ladder rung 10,135.1, this de-embedded coefficient, and the live tuple render 6,408.5) and what ships is none of them; that is why this field no longer uses the phrase. Archived NVFP4 replay basis: engine deployed effDec 0.150 × PRECISION_MULT.fp4 1.85 ⇒ 18,750.000 tok/s (5.0e15 × 1.85 × 0.150 / 74e9) at the F4 operating point. |
| GB200 NVL72 |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=8; sources=research/evidence-instances-v22.json#gb200-vllm-r1; frozen d2 §3.1 F4 (vLLM R1 decode observation); research/reviews/im-adv-r4-runB-internal-raw.md §B4/§C1 (the de-embedding; its 'FP8 bridge value' label is RETIRED per q-im-fp4-gb200-eta-basis — see deployedBasis) |
| GB200 NVL72 |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| GB200 NVL72 |
rent |
$4.5/hr |
price evidence=observed-source-named; economic basis note below |
| GB200 NVL72 |
capex |
$45000 |
engine.js HW economic input; economic basis note below |
| GB200 NVL72 |
tdp |
1.2 kW |
engine.js HW economic input; economic basis note below |
| GB200 NVL72 |
economicBasisNote |
— |
72-GPU NVLink domain (~$3-3.5M/rack ⇒ ~$44-49k/GPU). Per-GPU dense FP8 = 5.0 PF (NVIDIA sparse/dense split, verified — the widely-copied 4.5 PF is B200's). The vLLM ~10.1k R1 decode observation anchors F4 at b=128/rank and L=3,000. The live FP8-basis coefficient is η=0.315997 after de-embedding the retired FP4 precision scalar; the source run's precision basis (possibly NVFP4) is not fully pinned, so treat it as an upper anchor rather than reapplying an FP4 gain. Neocloud rates $3.50-6/hr (Jul 2026). 2026-07-15 dive: AWS Capacity Block $761.904/rack-hr ($10.582/GPU-hr) is the cleanest explicit full-rack market datum; paired with audited MLPerf v6.0 rack throughput it derives $0.881/$0.630/$0.435 per M output tok (Interactive/Server/Offline) — a rack-scale bridge anchor, not fitted into this row's rent (left at the existing neocloud estimate). |
| GB300 NVL72 |
flops.bf16 |
2500000000000000 FLOP/s |
C10 fallback — fp8/2 = 2.5 PF dense |
| GB300 NVL72 |
flops.fp8 |
5000000000000000 FLOP/s |
carried — 5.0 PF dense FP8 (engine.js flopsFp8 5.00; Blackwell Ultra FP8 unchanged from GB200 basis in the engine row) |
| GB300 NVL72 |
flops.fp4 |
15000000000000000 FLOP/s |
frozen — C2: NVFP4 active path FLOPS = 15e15 (d2 §4.2; Blackwell Ultra 15 PF dense FP4, 1.67× B200) |
| GB300 NVL72 |
hbmBytes |
298013687808 B (298.014 decimal GB) |
vendor-documented framebuffer — 284,208 MiB × 2^20 = 298,013,687,808 B per GPU: NVIDIA DGX GB300 NVL72 release notes (docs.nvidia.com/dgx/dgxgb300nvl72-release-notes, v1.0.0/1.0.1/1.0.6, known-issue 14) state 'NSM Type 3 (0x0C) and nvidia-smi (-q -d MEMORY) return 284,208 MiB, while Redfish TotalMemorySizeMiB returns 285,324 MiB'. The nvidia-smi value adopted (feasibility-conservative of the two); FLAGGED single-source-type — NVIDIA's own docs, no third-party pasted output found (verified 2026-07-18). ≈3.6% below the '288 GB'-as-GiB reading. |
| GB300 NVL72 |
bwHBM |
8000000000000 B/s |
published — 288 GB HBM3e @ 8 TB/s per GPU (frozen §5 recipe 1-2; = engine.js bw 8.00) |
| GB300 NVL72 |
fabric |
1800000000000 B/s |
memo §6 known value — NVLink5 1.8 TB/s per GPU (gb300 1.8e12, same NVLink5 domain as gb200; frozen §5 recipe 1-2) |
| GB300 NVL72 |
nShard |
8 |
N_shard; declared analyst default (unchanged from the pre-existing engine value), precision-conditioned — multiple Class C/D anchors support 8 as a genuinely deployed low-precision replica width, but per the 2026-07-20 GPT Pro evidence consult, the public record does not establish a specific required or minimum replica width for a multi-trillion-parameter closed-weight model at unknown precision (the actual Opus/5T scenario this engine models); see the declared-sensitivity case list on this row and research/gptpro-reports/2026-07-20-replica-width-consult.md. |
| GB300 NVL72 |
etaDec |
0.258295 |
analyst-set at a DECLARED assumed operating point (b9 M1 relabel — the prior FITTED label was FALSE: the calibration observation carries measured:null and an assumed batch; precision double-credit also REMOVED); evidence=analyst-set-assumed-op; 0.258295 IS AN FP4-BASIS COEFFICIENT used in the FP8 render as a DECLARED CONSERVATIVE TRANSFER, exactly as on gb200 and derived identically (owner ruling q-im-fp4-gb200-eta-basis, 2026-08-02; relabel only — no number moves). b9 M1 (r4 defects D3 + D6, run B §B4/§B12-2/§C1): 0.258295 = 0.477845 / 1.85, de-embedding the retired NVFP4 precision scalar from a value built at the NVFP4 tuple (research/im3-slice1a-derivations.md:72-78). The magnitude class matches gb200's: on the V4 Pro geometry this row renders 7,460.3 tok/s at FP4 — matching the site's own validation row — against 4,411.6 at FP8, a 1.69× gap. UNLIKE gb200 there is no measured anchor to under-predict (calObs.measured is null), so 'conservative by 1.69×' is a magnitude statement, not a demonstrated error. The row is ALSO relabelled from 'fitted' to analyst-set: measured:null and an ASSUMED batch, so no fit was ever performed here (run B: 'indefensible as fitted'). Archived NVFP4 replay basis: engine deployed effDec 0.127 × 1.85 ⇒ 15,875.000 tok/s (5.0e15 × 1.85 × 0.127 / 74e9) at the declared operating point L=2,740 (the C5 gb300 value). SCENARIO-ONLY (§C3): both the batch and the η are assumed; upgrading requires a reconstructable non-MTP serving curve — a registered evidence task (research/update-queue.md). |
| GB300 NVL72 |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=8; sources=research/evidence-instances-v22.json#gb300-analyst-set-dec; research/evidence-instances-v22.json#gb300-sglang-v4 (RETRO observation; excluded from fit set); memo §2 declared assumed operating point (source batch not reconstructable; excluded from fit set); research/reviews/im-adv-r4-runB-internal-raw.md §B1/§B4/§B12-2/§C1 |
| GB300 NVL72 |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| GB300 NVL72 |
rent |
$6/hr |
price evidence=analyst-set; economic basis note below |
| GB300 NVL72 |
capex |
$55000 |
engine.js HW economic input; economic basis note below |
| GB300 NVL72 |
tdp |
1.4 kW |
engine.js HW economic input; economic basis note below |
| GB300 NVL72 |
economicBasisNote |
— |
Blackwell Ultra: 15 PF dense FP4 (1.67× B200), 288GB HBM. The SGLang >12k tok/s/GPU V4 Pro observation (FP4+MTP, 49B active) is not a fit: its batch and MTP acceptance are not reconstructable. The live FP8-basis η=0.258295 is analyst-set at a declared b=128/rank, L=2,740 operating point after de-embedding the retired FP4 precision scalar. InferenceX ~17× H100 FP8. xjdr served GLM 5.2 on these ($4-7/hr early rates). 2026-07-15 dive: audited MLPerf v6.0 single-rack results now confirm generated-throughput at scale (NVIDIA Interactive 250,634 / Server 400,437 / Offline 647,076 gen tok/s; Nebius Server 575,580 / Offline 673,936 gen tok/s) — but NO numeric GB300 rack rental rate is public on any major provider checked (AWS/CoreWeave/Nebius/GCP/Azure/OCI/Crusoe, Jul 2026): model GB300 $/M-output as a function of rack-hour price, not a point estimate. Cleanest current Blackwell-Ultra pair is same-provider B300 (node-scale, not rack): Nebius 8-GPU MLPerf Server 60,413 gen tok/s at $7.85/GPU-hr ⇒ $0.289/M output. |
| H800 (China) |
flops.bf16 |
990000000000000 FLOP/s |
published — 989 TF dense BF16 (H100-class datasheet value); = fp8/2 (C10-consistent) |
| H800 (China) |
flops.fp8 |
1980000000000000 FLOP/s |
published — 1.98 PF dense FP8, H100-class compute (frozen F1 recipe flops 1.98e15; = engine.js flopsFp8 1.98) |
| H800 (China) |
hbmBytes |
85520809984 B (85.521 decimal GB) |
observed framebuffer (H100-80GB-class carry) — nvidia-smi total 81,559 MiB × 2^20 = 85,520,809,984 B, verified for 'NVIDIA H100 80GB HBM3' (see h100.prov.hbmBytes sources); H800 is the same 80 GB SXM HBM3 die/SKU class and no H800-specific pasted output was found — the carry is labeled, not observed-on-H800 (R5 verification note). |
| H800 (China) |
bwHBM |
3350000000000 B/s |
published — 3.35 TB/s HBM3 (H100-class; frozen F1 recipe hbmBps 3.35e12; = engine.js bw 3.35) |
| H800 (China) |
fabric |
400000000000 B/s |
frozen — NVLink export-capped 400 GB/s aggregate bidirectional per GPU (H800 SKU; roofline-diagnostic.mjs H800 case; d2 receipt pack) |
| H800 (China) |
nShard |
144 |
N_shard; PUBLISHED (R5 AMENDMENT: was analyst-declared 8, STALE — 'F1's 8-GPU node' is the anchor's per-node throughput-normalization basis, not the deployment width) — 144: DeepSeek Day-6 disclosure, verbatim: 'Decoding Phase [Routed Expert EP144, MLA/Shared Expert DP144]: Each deployment unit spans 18 nodes' (×8 GPUs = 144). Prefill unit = EP32/DP32 over 4 nodes (32 GPUs) — prefill width, NOT used by decode feasibility. Source 'deployment unit' adopted as memo §5 'replica width'. Cached: research/primary-sources/deepseek-day6-inference-2026-07-18/ (sha256 3b145f12…). Amends the memo §5 declared cell; adjudication in im3-slice1a-derivations.md §6. |
| H800 (China) |
etaDec |
0.313491 |
FITTED (deployed; computed per §2 rule — slice-1a derivation); evidence=fitted; current engine deployed effDec 0.070 ⇒ 1,872.9730 tok/s (1.98e15 × 0.070 / 74e9) at the F1 operating point — the snapshot-pin ≈1,873 basis, which relates to the 1,850 source anchor through the engine's documented rounding conventions (memo §2) |
| H800 (China) |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=144; sources=research/evidence-instances-v22.json#h800-deepseek-prod-dec; frozen d2 §3.1 F1 (DeepSeek Feb 2025 production disclosure) |
| H800 (China) |
etaPre |
0.17581 |
FITTED (identity, F7); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| H800 (China) |
rent |
$1.75/hr |
price evidence=observed-source-named; economic basis note below |
| H800 (China) |
capex |
$40000 |
engine.js HW economic input; economic basis note below |
| H800 (China) |
tdp |
0.7 kW |
engine.js HW economic input; economic basis note below |
| H800 (China) |
economicBasisNote |
— |
H100 compute with NVLink capped at 400GB/s (export SKU, finite pre-ban stock — capex carries the scarcity premium). THE H800/H100 DIFFERENTIAL, STATED (d-im-h800, owner note aca09d 2026-08-18; grounded on the NVIDIA H800 datasheet 2631447 vs the H100 SXM datasheet, Lenovo Press LP1814, Tencent Cloud HCCPNV5): the export SKU differs from the H100 SXM in NVLink (400 vs 900 GB/s aggregate bidirectional) and FP64 (1 vs 34 TF) ONLY — FP8/BF16 tensor rate, 80 GB HBM3, 3.35 TB/s and 700 W are identical, so nothing else that reaches a serving number differs. CORRECTION to the 2026-08-16 annotation on this row (owner verbatim: 'H800 numbers are clearly broken'), which said the cap is modelled nowhere and that this engine has no fabric term on these rows: it does — HW_ROOFLINE carries 400e9 here and 900e9 on h100, and the live roofline consumes both (decode t_N = b·D/fabric, prefill t_fabric = D/fabric). What is true is that at every expert-parallel operating point this page ships the fabric term is SLACK under the frozen max() form (≈3% of the binding memory term at the page default, ≈13% on kimi), so this row and h100 render identical throughput, and that h100/h200 INHERIT the efficiency FITTED on this row (F1: DeepSeek's production disclosure ran on H800s). This row is therefore the MEASURED part; the borrowing rows are where the assumption lives, and the page's implicit assumption has been that the cap costs nothing. That assumption is now a NAMED, ADJUSTABLE fit-transfer assumption — nvlinkCapMinRatio, default 1.00 = this historical model — the assumed MINIMUM ratio of a borrowing row over its own capped counterfactual, applied per phase to the borrowing rows only, never here, never stacking on an advantage the roofline already renders, with the engine's own serial-exposure counterfactual (≈1.005× at the page default) and the sensitivity computed beside it. On the dense tensor-parallel donor the fabric term already BINDS prefill and the H100 renders ≈2.25× this row's prefill throughput with no lever at all. Why this row still renders a HIGHER margin than the H100 at the default: identical throughput at $1.75/hr against $2.40/hr — a rent difference the page states, not a finding about export SKUs. Registered evidence task E-2026-08-16-c (a matched capped-vs-uncapped serving observation) stays open; no such observation is public. THE DeepSeek V3/R1 workhorse. Prefill MFU = FRESH-only reconstruction (~4,026 tok/s/GPU net of the 56.3% disk-cache share; the raw 9,212 aggregate includes cache hits). Annual-commit IDC rate $1.47-2.06/hr mid-2026; DeepSeek's 2025 disclosure assumed $2/hr. |
| H20 (China) |
flops.bf16 |
148000000000000 FLOP/s |
C10 fallback — fp8/2 = 148 TF dense; matches NVIDIA's published H20 BF16 148 TF |
| H20 (China) |
flops.fp8 |
296000000000000 FLOP/s |
published — 296 TF dense FP8 (frozen F2/F3 recipes flops 0.296e15; = engine.js flopsFp8 0.296) |
| H20 (China) |
hbmBytes |
102625181696 B (102.625 decimal GB) |
observed framebuffer — nvidia-smi total 97,871 MiB × 2^20 = 102,625,181,696 B for 'NVIDIA H20' (verified 2026-07-18: github.com/NVIDIA/TensorRT-LLM#8023, forums.developer.nvidia.com cudaLaunchHostFunc thread; same total documented for GH200's 96 GB HBM3 partition, docs.nvidia.com grace-perf-tuning-guide). |
| H20 (China) |
bwHBM |
4000000000000 B/s |
published — 4.0 TB/s HBM3 (frozen F2/F3 recipes hbmBps 4.0e12; = engine.js bw 4.00) |
| H20 (China) |
fabric |
900000000000 B/s |
frozen — 900 GB/s NVLink (H20 keeps full NVLink4; roofline-diagnostic.mjs H20 cases fabricBps 900e9) |
| H20 (China) |
nShard |
16 |
N_shard; PUBLISHED (R5 AMENDMENT: was analyst-declared 8 'deployment-family convention', STALE — the F2/F3 anchor source states the width directly) — 16: LMSYS/Ant Group post (lmsys.org/blog/2025-09-26-sglang-ant-group), verbatim: 'The Decode instance is deployed on a 2-node setup (16× H20 GPUs)' / 'All Decode instances are deployed with a dual-node setup: Attention-DP16 + MoE-EP16'; prefill = single-node TP8 (prefill width, not used by decode feasibility). The repo already carried this at research/consultation-2026-07-10-roofline-verbatim.md:154 without it reaching the normative data. Cached: research/primary-sources/sglang-ant-h20-2026-07-18/ (sha256 53c49762…). Amends the memo §5 declared cell; adjudication in im3-slice1a-derivations.md §6. |
| H20 (China) |
etaDec |
0.217022 |
analyst-set (source-informed neutral adjustment: deployed 680 tok/s differs from the 714 tok/s observation; not a fit); evidence=source-informed-neutral; legacy effDec 0.170 defines the neutral 680.0000 tok/s throughput target (0.296e15 × 0.170 / 74e9) at the F3 operating point; the executable roofline mapping is etaDec 0.217022. This is the hardware dive's neutral recommendation (≈680 basis), NOT the 714 anchor-reproducer (memo §2) |
| H20 (China) |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=16; sources=research/evidence-instances-v22.json#h20-ant-sglang-pro (F2 input observation); research/evidence-instances-v22.json#h20-ant-sglang (F3 input observation); research/evidence-instances-v22.json#h20-neutral-live-dec (live analyst-set identity); frozen d2 §3.1 F2/F3 |
| H20 (China) |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| H20 (China) |
rent |
$1/hr |
price evidence=observed-source-named; economic basis note below |
| H20 (China) |
capex |
$20000 |
engine.js HW economic input; economic basis note below |
| H20 (China) |
tdp |
0.4 kW |
engine.js HW economic input; economic basis note below |
| H20 (China) |
economicBasisNote |
— |
The China-legal NVIDIA SKU: only 296 TF dense FP8 but 4.0 TB/s HBM — decode is bandwidth-bound, so it serves far better than its FLOPS suggest (hence the high effective MFU vs a tiny denominator). This row's executable roofline coefficient is the source-informed-neutral η_dec=0.217022; it maps the legacy effDec=0.170 throughput basis to the hardware dive's ~680 tok/s recommendation and does NOT fit Ant Group's relaxed <70ms production observation (714 tok/s at b=48, L=4,096, one-step/two-draft-token MTP with ~1.8-1.9 accepted tokens). Rental class matters: same chip spans ~$0.76 (IDC annual) to $7+ (hyperscaler on-demand). |
| TPU v7 Ironwood |
flops.bf16 |
2307000000000000 FLOP/s |
C10 fallback — fp8/2 = 2.307 PF dense (no published Ironwood dense-BF16 figure in the project source base; analyst-set via the C10 rule, labeled) |
| TPU v7 Ironwood |
flops.fp8 |
4614000000000000 FLOP/s |
published — 4,614 TF FP8 per Ironwood chip (frozen recipe 5 / C9-adjacent [RP §2b]; = engine.js flopsFp8 4.61) |
| TPU v7 Ironwood |
hbmBytes |
206158430208 B (206.158 decimal GB) |
published, unit EXPLICIT at source — 192 GiB × 2^30 = 206,158,430,208 B: Google Cloud TPU docs table 'HBM capacity per chip (GiB) … 192' (docs.cloud.google.com/tpu/docs/tpu7x, also compute/docs/tpus/tpu-machines), corroborated by the Hot Chips 2025 Ironwood deck ('capacity 192 GiB, 8 stacks HBM3E') and arXiv:2606.15870 (verified 2026-07-18). No reserved-carve-out figure published; raw capacity adopted, labeled. |
| TPU v7 Ironwood |
bwHBM |
7370000000000 B/s |
published — 7.37 TB/s HBM per chip (frozen recipe 5 'HBM 192 GB @ 7.37e12'; = engine.js bw 7.37) |
| TPU v7 Ironwood |
fabric |
1200000000000 B/s |
frozen — ICI 1.2 TB/s per chip (d2 §5 recipe 5; memo §6 known value) |
| TPU v7 Ironwood |
nShard |
4 |
N_shard; frozen — 4 (d2 §1.3: TP8 across 8 tensorcores = the 4-chip host, one replica) |
| TPU v7 Ironwood |
etaDec |
0.55 |
analyst-set (platform-native aggregate-form bridge: two same-platform diagnostics 0.528/0.574, midpoint 0.55 — b9 M1; VALID ONLY in this row's declared decodeTrafficBasis); evidence=platform-native-aggregate-bridge; b9 M1 (r4 defect D2, run B §A4/§B1/§C1): the joint fleet fit 0.36142 is REJECTED as evidence for this row — it is a geometric mean over six NVIDIA/Ascend observations containing ZERO TPU data, against the project's own LOAO record of 37% average / 59% worst-case cross-platform transfer error. Replaced by the mean of two SAME-PLATFORM aggregate-form diagnostics: η ≈ 518.86 × 480 GB / (64 × 7.37 TB/s) = 0.528 from the rental anchor, and η ≈ 677 × 400 GB / (64 × 7.37 TB/s) = 0.574 from Google's July 2026 Ironwood playbook; midpoint 0.55, declared band 0.528–0.574. These are weight-only, representation-specific diagnostics (they ignore KV and recurrent-state traffic), so the coefficient is bound to decodeTrafficBasis 'replica-resident-distinct' and is NOT a universal platform efficiency. The retired ≈0.199 value is NOT the alternative: run B §A4 shows it is the coefficient required to reproduce the old anchor INSIDE the malformed per-device identity, absorbing the batch/weight unit mismatch rather than measuring efficiency. |
| TPU v7 Ironwood |
decodeTrafficBasis |
replica-resident-distinct |
paired eta representation=replica-resident-distinct; nPhysDeclared=16; sources=research/reviews/im-adv-r4-runB-internal-raw.md §A2/§A4/§B1/§C1; Google Ironwood Qwen 3.5 serving playbook 2026-07-14 (677 tok/s/chip, concurrency 64, 4 chips, 400 GB replica footprint — 79.6% of the discounted HBM roofline); TPU v7 Qwen3-Coder-480B rental anchor (518.86 tok/s/chip, stated concurrency 64 over 4 chips) |
| TPU v7 Ironwood |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| TPU v7 Ironwood |
rent |
$5.4/hr |
price evidence=observed-source-named; economic basis note below |
| TPU v7 Ironwood |
capex |
$35000 |
engine.js HW economic input; economic basis note below |
| TPU v7 Ironwood |
tdp |
1 kW |
engine.js HW economic input; economic basis note below |
| TPU v7 Ironwood |
economicBasisNote |
— |
Google's inference TPU (GA Mar 31, 2026): 4,614 TF FP8 ≈ B200-class, 9,216-chip pods. Anthropic committed up to ~1M TPUs (Oct 2025). Rent = Google's PUBLISHED 3-year committed rate $5.40/chip-hr (b9 M1, r4 defect D4: the retired $4.20 sat BELOW every public comparator — 3-yr $5.40, DWS Flex $6, on-demand $12 — so it was never a purchasable market rate; this is a low/committed PLANNING rate, and it is labeled as one). 2026-07-15 dive: named-model accelerator-rental anchors exist (Qwen3-Coder-480B, 4 chips, 518.86 tok/s/chip ⇒ $6.42/M output on-demand, $2.89/M 3-yr) — barred from calibration as a RETRO evidence annotation (memo §2), though b9 M1 admits its aggregate-form weight-only efficiency reading (0.528) as one endpoint of this row's platform-native η bridge. The live decode path uses η_dec=0.55, the midpoint of that bridge (0.528 rental anchor / 0.574 Google Ironwood playbook), replacing the joint fleet fit that contained ZERO TPU observations. Google-internal fleet cost still unknown: no public Gemini-SKU→TPU mapping exists — estimates. |
| Trainium2 |
flops.bf16 |
650000000000000 FLOP/s |
frozen — C10 convention value 0.65e15/chip (d2 §4.2 C10; the dossier-stated 667 TF/chip is within 3% and is reported as sensitivity in the frozen C10 note, convention retained) |
| Trainium2 |
flops.fp8 |
1300000000000000 FLOP/s |
published — ~1.3 PF dense FP8 per chip (AWS Trainium2 spec; = engine.js flopsFp8 1.30) |
| Trainium2 |
hbmBytes |
103079215104 B (103.079 decimal GB) |
frozen, unit explicit — the frozen d2 §5 recipe 6-7 states '96 GiB @ 2.9e12 per chip': 96 GiB × 2^30 = 103,079,215,104 B. No AWS framebuffer receipt; the frozen source's own GiB statement adopted. |
| Trainium2 |
bwHBM |
2900000000000 B/s |
published — 96 GiB @ 2.9 TB/s per chip (frozen recipe 6-7; = engine.js bw 2.90) |
| Trainium2 |
fabric |
1280000000000 B/s |
PUBLISHED (b9 M1 AMENDMENT: was frozen 1.024e12, STALE — run B §B1/§C1) — 1.28 TB/s NeuronLink per Trainium2 device (AWS specification). The frozen d2 §5 recipe 6-7 value of 1,024 GB/s materially understated the published hardware. NOTE: this term has NO effect on the displayed FP8/HBM-binding midpoint (the decode path binds on t_H, not t_N, at every default operating point) — it corrects the registry for subsequent communication-modeling scenarios (M2). |
| Trainium2 |
nShard |
16 |
N_shard; frozen — 16 (d2 §1.3: TP16 = the 16-chip trn2.48xlarge instance, one replica) |
| Trainium2 |
etaDec |
0.36142 |
analyst-set (joint fleet fit — out-of-family tier-(c)); evidence=joint-fit; joint fleet fit (frozen d2 §3.2); dense-TP row — t_cc applies per TCC_CONSTANTS (frozen §1.5), outside η. SCENARIO-ONLY (r4 §C3, b9 M1): no model-, topology- and SLO-matched Trainium serving throughput observation exists in the public record, so this coefficient CANNOT support a central Trainium margin — run B §A4: 'the joint η may remain only as a visibly marked sensitivity parameter'. Upgrading it requires a named model, topology, batch, precision, SLO and achieved output throughput. |
| Trainium2 |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=16; sources=frozen d2 §3.2 joint fleet fit |
| Trainium2 |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| Trainium2 |
rent |
$2.235/hr |
price evidence=observed-source-named; economic basis note below |
| Trainium2 |
capex |
$15000 |
engine.js HW economic input; economic basis note below |
| Trainium2 |
tdp |
0.5 kW |
engine.js HW economic input; economic basis note below |
| Trainium2 |
economicBasisNote |
— |
GA Dec 2024. Project Rainier launched with ~500k Trainium2 for Anthropic (activated ~Nov 2025, confirmed running Claude inference alongside training); Anthropic reported >1M Trainium2 in use across AWS by Apr 2026 — inference/training allocation, utilization and internal rate undisclosed. Rent = AWS Capacity Blocks PUBLISHED $2.235/chip-hr (b9 M1, r4 defect D4: the retired $1.50 sat below the only public comparator). 2026-07-15 dive: a narrow public ENGINEERING anchor exists (AWS Neuron tutorials, batch=1/concurrency=1: Llama 3.3 70B spec-decode $68.90/M output tokens, Llama 3.1 405B $98.65/M) — not a production-TCO measurement, not fitted into this roofline. b9 M1 operating point: the AWS Qwen3-235B recipe on one trn2.48xlarge states replica-global batch 16 online / 64 offline (16 chips, tp_degree 64, attention-DP8, MoE EP32/TP2); the surrogate midpoint 32 replaces the retired b=4, which was an AWS tutorial demo value read as a production point. Throughput remains UNVERIFIED — no matched serving anchor exists — so this leg is scenario-only. |
| Trainium3 |
flops.bf16 |
671000000000000 FLOP/s |
PUBLISHED (post-review adjudication 2026-07-27) — the AWS Neuron architecture docs state BF16/FP16/TF32 dense = 671 TFLOPS/chip (the 2,517 figure is SPARSE) (https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/arch/neuron-hardware/trainium3.html). Replaces the C10 fp8/2 fallback of 1.255 PF, which overstated published dense BF16 by 1.87×. NOTE the same docs state Trainium2 BF16 = 667 TFLOPS; the trn2 row deliberately retains its frozen C10 convention value 0.65e15 (labeled, within 3%) — moving it is a frozen-recipe identity decision, flagged in the adjudication report. |
| Trainium3 |
flops.fp8 |
2510000000000000 FLOP/s |
published — 362 PF FP8 per 144-chip UltraServer ⇒ 2.51 PF/chip (engine.js trn3 note, GA Dec 2025); Neuron architecture docs concur: 2,517 MXFP8/MXFP4 TFLOPS/chip |
| Trainium3 |
flops.fp4 |
2510000000000000 FLOP/s |
published NATIVE (b9 M1 — run B §B4/§C1) — AWS states 2.517 PF for MXFP8 AND MXFP4 per Trainium3 device, i.e. ONE figure covering both formats. Registered EQUAL to the fp8 basis: native FP4 arithmetic exists, but no FP4 compute doubling is published, so the row takes the capacity/weight-loading benefit (PRECISION_TUPLES.trn3.fp4 sW 0.5) and NO throughput credit. |
| Trainium3 |
hbmBytes |
154618822656 B (154.619 decimal GB) |
PUBLISHED (post-review adjudication 2026-07-27) — the AWS Neuron architecture docs are unit-explicit: 'HBM Capacity (GiB): 96 → 144' and '144 GiB of device memory' (https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/arch/neuron-hardware/trainium3.html, retrieved 2026-07-27). 144 GiB × 2^30 = 154,618,822,656 B. This REVERSES the external review's SI re-reading (144e9), which took the marketing page's loose '144 GB' label literally against the unit-explicit engineering docs; HBM3e stack capacities are physically binary quantities, and the trn2 row's frozen source states GiB for the same family. |
| Trainium3 |
bwHBM |
4900000000000 B/s |
PUBLISHED — 144 GiB HBM3e at 4.9 TB/s per chip (AWS Neuron architecture docs; SI carry of bw 4.90). |
| Trainium3 |
fabric |
2560000000000 B/s |
PUBLISHED (post-review adjudication 2026-07-27) — the AWS Neuron architecture docs state Inter-chip Interconnect 2,560 GB/sec/chip for Trainium3, in the same table and convention as Trainium2's 1,280 (already registered as 1.28e12) (https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/arch/neuron-hardware/trainium3.html). CONFLICT NOTED: the marketing page (https://aws.amazon.com/ec2/instance-types/trn3/) says '2TB/s of bandwidth per chip' for NeuronLink-v4 — likely a different direction/aggregation convention; the engineering docs' same-convention family figure is adopted. The displayed FP8/HBM-binding midpoint does not bind on t_N. |
| Trainium3 |
nShard |
16 |
N_shard; analyst-declared — 16 (trn2 carry, memo §5) |
| Trainium3 |
etaDec |
0.36142 |
analyst-set (joint fleet fit — out-of-family tier-(c); trn2-carry platform constants, labeled); evidence=joint-fit; joint fleet fit; no serving anchor of any kind exists for Trn3 (engine.js note: confirmed negative). SCENARIO-ONLY (r4 §C3, b9 M1): both the coefficient and the operating point are Trn2-derived carries — run B §B1: 'no central Trainium3 throughput should be represented as observed'. The row's rent is separately scenario-only (no public instance/UltraServer rate exists). |
| Trainium3 |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=16; sources=frozen d2 §3.2 joint fleet fit (trn2-carry platform constants) |
| Trainium3 |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| Trainium3 |
rent |
$2.2/hr |
price evidence=analyst-set; economic basis note below |
| Trainium3 |
capex |
$20000 |
engine.js HW economic input; economic basis note below |
| Trainium3 |
tdp |
0.8 kW |
engine.js HW economic input; economic basis note below |
| Trainium3 |
economicBasisNote |
— |
GA Dec 2025: 144-chip UltraServers, 362 PF FP8 ⇒ 2.51 PF/chip, 144GB HBM3e. AWS's aggressive cost-per-token play — pricing estimates. 2026-07-15 dive: AWS has published a GPT-OSS-120B inference recipe but no achieved tok/s and no public Trn3 instance/UltraServer price — neither side of $/token is public; NO public serving anchor of any kind (confirmed negative). |
| Ascend 910C |
flops.bf16 |
752000000000000 FLOP/s |
ANALYST-SET, C10-ANALOG (flagged for 1a review) — 1.504e15/2 = 0.752e15. C10 as frozen reads 'bf16 = fp8_dense/2 where unpublished'; ascend has no fp8, so the rule is applied to its INT8 8-bit dense basis. This exactly carries the current engine's PRECISION_MULT.bf16 = 0.5 rendering behavior forward. Declared convention, not a published value. |
| Ascend 910C |
hbmBytes |
128000000000 B (128.000 decimal GB) |
ANALYST CONVENTION (labeled) — 128e9 B, the SI reading of the '128 GB' label. R5 verification (2026-07-18) found NO Huawei datasheet with unit precision and NO pasted npu-smi output; third-party sources state a round '128 GB' (one outlier claims 96 GB) with the HBM generation itself contested (HBM2e vs HBM3). The SI reading is the conservative choice and matches the frozen d2 harness's own hbmCap 128e9 for the cm384 rows. Revisit if an npu-smi receipt surfaces. |
| Ascend 910C |
bwHBM |
3200000000000 B/s |
published — 128 GB @ 3.2 TB/s per NPU (frozen recipe 3-4; = engine.js bw 3.20) |
| Ascend 910C |
fabric |
784000000000 B/s |
frozen — UB (unified bus) 784 GB/s per NPU (d2 §5 recipe 3-4 / roofline-diagnostic.mjs Ascend cases) |
| Ascend 910C |
nShard |
128 |
N_shard; declared — 128, the cm384 CO-LOCATION instance width (memo §5; frozen §1.3 co-loc 128). The 6P2D decode-pool width 144 (frozen §1.3) is an observation-level constant of the cm384-6p2d recipe, NOT this row's replica width. |
| Ascend 910C |
etaDec |
0.299324 |
analyst-set (source-informed neutral adjustment: deployed 1,422.7 tok/s differs from the 1,943 tok/s observation; not a fit); evidence=source-informed-neutral; legacy effDec 0.070 defines the neutral 1,422.7027 tok/s throughput target (1.504e15 × 0.070 / 74e9) at the F5 operating point; the executable roofline mapping is etaDec 0.299324. This is the hardware dive's neutral recommendation, NOT the 1,943 anchor-reproducer (memo §2). 8-bit basis is INT8 (W8A8) |
| Ascend 910C |
decodeTrafficBasis |
active-parameter-surrogate |
paired eta representation=active-parameter-surrogate; nPhysDeclared=128; sources=research/evidence-instances-v22.json#ascend-cminfer-decode (F5 input observation); research/evidence-instances-v22.json#ascend-cminfer-15ms (F6 input observation); research/evidence-instances-v22.json#ascend-neutral-live-dec (live analyst-set identity); frozen d2 §3.1 F5/F6 |
| Ascend 910C |
etaPre |
0.17581 |
extrapolated (single-anchor transfer); frozen F7 identity fit: η_pre = 0.17581 (H800 fresh-prefill reconstruction ~4,026 tok/s/GPU at L_in = 4,989, FP8; d2 §3.1-§3.2). EVERY row's prefill leg computes the frozen E2 roofline with this single value (memo §8 universal-transfer rule). Prefill never carries a measured basis (IM5-5 boundary; no two-sided prefill evidence exists — BLOCK-2 item 9). |
| Ascend 910C |
rent |
$1.95/hr |
price evidence=observed-source-named; economic basis note below |
| Ascend 910C |
capex |
$23000 |
engine.js HW economic input; economic basis note below |
| Ascend 910C |
tdp |
0.6 kW |
engine.js HW economic input; economic basis note below |
| Ascend 910C |
economicBasisNote |
— |
Huawei's dual-die flagship (SMIC 7nm). NO native FP8 — 8-bit here means INT8 (1.504 PF per Huawei's Atlas spec; a widely-quoted 1,054 figure is a typo in the CloudMatrix paper). This row's executable roofline coefficient is the source-informed-neutral η_dec=0.299324; it maps the legacy effDec=0.070 throughput basis to the hardware dive's ~1,420 tok/s recommendation on an R1-class workload and does NOT fit Huawei's optimized source observation (1,943 tok/s at b=96, L=4,096, one speculative token at assumed 70% acceptance, q=2/a=1.7). DeepSeek's internal '60% of H100' eval implies a lower independent efficiency estimate (~5.5-6.5% of peak FLOPs). Rent = Huatai procurement award ($1.71-2.25/hr); CloudMatrix 384 ≈ RMB 60M. 2026-07-15 dive: a full 384-card CloudMatrix system (FlexNPU, arXiv:2606.04415) serves DeepSeek-R1 W8A8 at ≈633,000 generated tok/s system-wide under TTFT≤1s/TPOT≤50ms. CLOSED (owner ruling 2026-07-18, memo §9 'CM384 ruling'): these are EVIDENCE-RECORD ANNOTATIONS — never a selectable scenario, an operating point, or a calibration input — delivered full-system SLO point: 1,648 tok/s/card ⇒ 8.1%, mixed prefill+decode, all 384 cards charged; decode-pool standalone ceiling: 2,885 tok/s/decode-card ⇒ 14.2%, capacity ceiling, prefill hardware excluded. The full-system point brackets (corroborates) this row's deployed 7% default; the decode-pool ceiling does not (a different standalone measurand, ~1.49× above the full-system point) — neither is fitted into this row. No public CM384 hourly rental price exists — throughput anchored, cost unanchored. Cross-source proxy for the OLDER 910B chip (not this row): JD xLLM 709 gen tok/s/card × CTyun RMB 38.45/hr ⇒ ≈$2.09/M output (range $1.61-2.80/M), medium-low confidence. Ascend 920 has no official SKU, benchmark, deployment or price (Huawei's roadmap goes 910C→950PR/950DT→960→970, no 920) — excluded from this model. |