Measured harness ledgerPublic result
GPT-5.6 Sol

Red Sands v1 — GPT-5.6 Sol Max

Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.

Max reasoningHeadline result
Workflow cost
$19.89
Wall-clock
1h 0m 43.1s wall-clock
Processed tokens
29.52M processed
Record state
artifact_operator_static_checks_ledger_no_playthrough_evidence
Public summary

GPT-5.6 Sol Max artifact_operator_static_checks_ledger_no_playthrough_evidence ledger: 1h 0m 43.1s wall-clock, 29.52M processed, and $19.89 API-equivalent estimate from official OpenAI API Standard pricing for the measured token mix; NOT the actual subscription-backed Codex charge.

Run identity and stack
  • Result ID: red-sands-quest-threejs-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: Provider not separately recorded
  • Client: Codex
  • Stack: Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Three.js quest-authoring fixture
  • Stack: Harness v1 frozen prompt
Cost basis
  • Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold; the largest was 242,915 tokens.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/red-sands-quest-threejs-gpt-5.6-sol-max/source/src/quest/quests/red-ink.js
  • SHA-256: c3881a7413268e96d8c3d4b0c2b666d12bf91882325a257454806a01bde8f5d3
Recorded caveats
  • This task has NO automated completion gate, by design.
  • No playthrough evidence of any kind was supplied, and no operator playthrough exists. Nothing shows this quest has been completed.
  • npm run smoke runs under ?capture=1, which never constructs the quest system.
  • The prompt's required handback note was not delivered; no stated-gaps section exists in the submission.
  • Cache-creation is recorded as zero because the transcript exposes no cache-write counter; it is a logging absence, not a measured zero.
  • The receipt records 2 user turns where the prompt is a single fixed brief.
  • Wall-clock is end-to-end latency including tool execution and ElevenLabs synthesis, not model-only compute.
  • Output tokens include the 28,866-token reasoning subset, code, and tool-call JSON.
  • Cache reads are discounted, so the 29.5M total-processed figure overstates cost.
  • The cost is an API-list-price equivalent, not the actual subscription-backed Codex charge, and excludes tool charges.
  • The archived smoke report and capture are the operator's rerun, which overwrote the model's equivalent copy.
  • The voice audio was not listened to; delivery, count and casting are confirmed, performance is not.
  • This is a quest-authorship task over a fixed game, so its figures are not comparable to the archive's from-scratch build rows.
  • The archive is a delta over the fixture, not a standalone runnable tree.
  • No quality score and no blind-evaluation record exist — which for this task is where the entire measurement lives.
Visible evidence gaps
  • any evidence that the quest can be completed
  • a human playthrough, start to finish
  • blind-evaluation record — the primary instrument for this task
  • the handback note the prompt asks for
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Red Sands v1 · Builder projects v1
RemakeBenchResearch console