Versioned benchmark collection

Space Flight v2

Harness v2 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Astronaut approaching a landed spacecraft in the Space Flight v2 collection
Astronaut approaching a landed spacecraft in the Space Flight v2 collection
4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Space Flight Stage 2 — Landfall

Model coverage

  • Claude Opus 5 · Max
  • Claude Fable 5 · Max
  • GPT-5.6 Sol · Ultra
  • Kimi K3 · Max

Evidence available

  • Four measured run ledgers
  • Four public presentation recordings
  • Cost, token, workflow timing, and artifact receipts

Still missing

  • Repeat-run reliability evidence
  • Standardized gameplay and physical-input verification
  • Comparable local FPS measurements
Public receipts

4 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.