Measured harness ledgerPublic result
GPT-5.6 SolRed Sands v1 — GPT-5.6 Sol Max
Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.
Max reasoningHeadline result
- Workflow cost
- $19.89
- Wall-clock
- 1h 0m 43.1s wall-clock
- Processed tokens
- 29.52M processed
- Record state
- artifact_operator_static_checks_ledger_no_playthrough_evidence
Public summary
GPT-5.6 Sol Max artifact_operator_static_checks_ledger_no_playthrough_evidence ledger: 1h 0m 43.1s wall-clock, 29.52M processed, and $19.89 API-equivalent estimate from official OpenAI API Standard pricing for the measured token mix; NOT the actual subscription-backed Codex charge.
Run identity and stack
- Result ID: red-sands-quest-threejs-gpt-5.6-sol-max
- Technical model: gpt-5.6-sol
- Provider: Provider not separately recorded
- Client: Codex
- Stack: Codex
- Stack: Technical model/configuration: gpt-5.6-sol
- Stack: Three.js quest-authoring fixture
- Stack: Harness v1 frozen prompt
Cost basis
- Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold; the largest was 242,915 tokens.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/red-sands-quest-threejs-gpt-5.6-sol-max/source/src/quest/quests/red-ink.js
- SHA-256: c3881a7413268e96d8c3d4b0c2b666d12bf91882325a257454806a01bde8f5d3
Recorded caveats
- This task has NO automated completion gate, by design.
- No playthrough evidence of any kind was supplied, and no operator playthrough exists. Nothing shows this quest has been completed.
- npm run smoke runs under ?capture=1, which never constructs the quest system.
- The prompt's required handback note was not delivered; no stated-gaps section exists in the submission.
- Cache-creation is recorded as zero because the transcript exposes no cache-write counter; it is a logging absence, not a measured zero.
- The receipt records 2 user turns where the prompt is a single fixed brief.
- Wall-clock is end-to-end latency including tool execution and ElevenLabs synthesis, not model-only compute.
- Output tokens include the 28,866-token reasoning subset, code, and tool-call JSON.
- Cache reads are discounted, so the 29.5M total-processed figure overstates cost.
- The cost is an API-list-price equivalent, not the actual subscription-backed Codex charge, and excludes tool charges.
- The archived smoke report and capture are the operator's rerun, which overwrote the model's equivalent copy.
- The voice audio was not listened to; delivery, count and casting are confirmed, performance is not.
- This is a quest-authorship task over a fixed game, so its figures are not comparable to the archive's from-scratch build rows.
- The archive is a delta over the fixture, not a standalone runnable tree.
- No quality score and no blind-evaluation record exist — which for this task is where the entire measurement lives.
Visible evidence gaps
- any evidence that the quest can be completed
- a human playthrough, start to finish
- blind-evaluation record — the primary instrument for this task
- the handback note the prompt asks for
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Red Sands v1 · Builder projects v1
