Measured harness ledgerPublic result
GPT-5.6 Sol

Red Sands v1 — GPT-5.6 Sol Max

Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.

Max reasoningHeadline result
Workflow cost
$19.89
Wall-clock
1h 0m 43.1s wall-clock
Processed tokens
29.52M processed
Record state
ledger
Public summary

GPT-5.6 Sol Max public partial ledger: 1h 0m 43.1s wall-clock, 29.52M processed, and $19.89 estimated API-list-price equivalent.

Cost basis
  • Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold; the largest was 242,915 tokens.
Visible evidence gaps
  • any evidence that the quest can be completed
  • a human playthrough, start to finish
  • blind-evaluation record — the primary instrument for this task
  • the handback note the prompt asks for
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Red Sands v1 · Builder projects v1
RemakeBenchResearch console