Versioned benchmark collection
Red Sands v1
Harness v1 · Builder projects v1 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- Red Sands v1
Model coverage
- GPT-5.6 Sol · Max
- Kimi K3 · Max
- Claude Opus 5 · Max
- Qwen3.8 Max · xhigh
Evidence available
- Four selected measured run ledgers
- Exact frozen prompt identity
- Per-run quest source and generated voice artifact identities
- Four reconstructed projects with clean install, production build, and browser smoke receipts
Still missing
- Human start-to-finish playthrough for every run
- Blind evaluation and quality scoring
- Repeat-run reliability evidence
Public receipts
4 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01GPT-5.6 SolMax reasoning · Red Sands v1$19.89 estimated API-list-price equivalentrecorded estimate1h 0m 43.1s wall-clockEnd-to-end workflowOpen 02Kimi K3Max reasoning · Red Sands v1≥$5.68 estimated API-list-price equivalent (lower bound)recorded lower bound1h 0m 6.6s (measured window only) wall-clockEnd-to-end workflowOpen 03Claude Opus 5Max reasoning · Red Sands v1$129.47 estimated API-list-price equivalentrecorded estimate1h 30m 54.0s wall-clockEnd-to-end workflowOpen 04Qwen3.8 Maxxhigh reasoning · Red Sands v1¥235.23 estimated API-list-price equivalent (upper bound)recorded upper bound1h 5m 52.3s wall-clockEnd-to-end workflowOpen
