Measured harness ledgerPublic result
GPT-5.6 SolRed Sands v1 — GPT-5.6 Sol Max
Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.
Max reasoningHeadline result
- Workflow cost
- $19.89
- Wall-clock
- 1h 0m 43.1s wall-clock
- Processed tokens
- 29.52M processed
- Record state
- ledger
Public summary
GPT-5.6 Sol Max public partial ledger: 1h 0m 43.1s wall-clock, 29.52M processed, and $19.89 estimated API-list-price equivalent.
Cost basis
- Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold; the largest was 242,915 tokens.
Visible evidence gaps
- any evidence that the quest can be completed
- a human playthrough, start to finish
- blind-evaluation record — the primary instrument for this task
- the handback note the prompt asks for
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Red Sands v1 · Builder projects v1
