Versioned benchmark collection
Space Flight v2
Harness v2 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- Space Flight Stage 2 — Landfall
Model coverage
- Claude Opus 5 · Max
- Claude Fable 5 · Max
- GPT-5.6 Sol · Ultra
- Kimi K3 · Max
Evidence available
- Four measured run ledgers
- Four public presentation recordings
- Cost, token, workflow timing, and artifact receipts
Still missing
- Repeat-run reliability evidence
- Standardized gameplay and physical-input verification
- Comparable local FPS measurements
Public receipts
4 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01Claude Opus 5Max reasoning · Space Flight Stage 2 — Landfall$133.83 estimated API-equivalent costrecorded estimate2h 31m 33.5s wall-clockEnd-to-end workflowOpen 02Claude Fable 5Max reasoning · Space Flight Stage 2 — Landfall$280.14 estimated API-equivalent costrecorded estimate2h 05m 20.4s wall-clockEnd-to-end workflowOpen 03GPT-5.6 SolUltra reasoning · Space Flight Stage 2 — Landfall$20.84 estimated API-equivalent costrecorded estimate1h 29m 57.7s wall-clockEnd-to-end workflowOpen 04Kimi K3Max reasoning · Space Flight Stage 2 — Landfall$15.58 estimated API-equivalent costrecorded estimate2h 23m 18.1s wall-clockEnd-to-end workflowOpen
