Versioned benchmark collection
Red Sands v1
Harness v1 · Builder projects v1 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- Red Sands v1
Model coverage
- GPT-5.6 Sol · Max
- Kimi K3 · Unreported
- Claude Opus 5 · Max
- Qwen3.8 Max · xhigh
Evidence available
- Four selected measured run ledgers
- Exact frozen prompt identity
- Per-run quest source and generated voice artifact identities
- Four reconstructed projects with clean install, production build, and browser smoke receipts
Still missing
- Human start-to-finish playthrough for every run
- Blind evaluation and quality scoring
- Repeat-run reliability evidence
Public receipts
4 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01GPT-5.6 SolMax reasoning · Red Sands v1$19.89 API-equivalent estimate from official OpenAI API Standard pricing for the measured token mix; NOT the actual subscription-backed Codex chargerecorded estimate1h 0m 43.1s wall-clockEnd-to-end workflowOpen 02Kimi K3Unreported reasoning · Red Sands v1≥$5.68 Independently recomputed API-list-price equivalent for the measured token mix; not an itemized subscription charge (lower bound)recorded lower bound1h 0m 6.6s (measured window only) wall-clockEnd-to-end workflowOpen 03Claude Opus 5Max reasoning · Red Sands v1$129.47 First-party Claude API list-price equivalent for the measured token mix; not an itemized subscription chargerecorded estimate1h 30m 54.0s wall-clockEnd-to-end workflowOpen 04Qwen3.8 Maxxhigh reasoning · Red Sands v1¥235.23 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound)recorded upper bound1h 5m 52.3s wall-clockEnd-to-end workflowOpen
