Versioned benchmark collection

Space Flight v1

Harness v1 · 8 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Space-flight benchmark output in the Space Flight v1 collection
Space-flight benchmark output in the Space Flight v1 collection
8Measured runs
1Frozen targets
6Model families
7Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Space Flight Game

Model coverage

  • GPT-5.6 Terra · Ultra
  • GPT-5.6 Sol · Ultra
  • Claude Fable 5 · Medium
  • GPT-5.6 Luna · xhigh
  • Claude Fable 5 · Max
  • Kimi K3 · Max
  • Qwen3.8 Max Preview · Unreported

Evidence available

  • Six measured run ledgers
  • Public browser builds
  • Cost, token, and workflow timing

Still missing

  • Browser/hardware environment
  • Local FPS
  • Final capture metadata
  • Minimum blind-vote sample
Public receipts

8 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.