Versioned benchmark collection

Red Sands v1

Harness v1 · Builder projects v1 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Western frontier town with a player character and saddled horse in Red Sands v1
Western frontier town with a player character and saddled horse in Red Sands v1
4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Red Sands v1

Model coverage

  • GPT-5.6 Sol · Max
  • Kimi K3 · Max
  • Claude Opus 5 · Max
  • Qwen3.8 Max · xhigh

Evidence available

  • Four selected measured run ledgers
  • Exact frozen prompt identity
  • Per-run quest source and generated voice artifact identities
  • Four reconstructed projects with clean install, production build, and browser smoke receipts

Still missing

  • Human start-to-finish playthrough for every run
  • Blind evaluation and quality scoring
  • Repeat-run reliability evidence
Public receipts

4 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.