Measured harness ledgerPublic result
GPT-5.6 LunaOff-road Mud Game — GPT-5.6 Luna Max
Build a one-request procedural off-road driving game with vehicle physics, independent suspension, streamed terrain, mud feedback, and multiple cameras.
Max reasoningHeadline result
- Workflow cost
- $0.73
- Wall-clock
- 51m 53.5s wall-clock
- Processed tokens
- 25.51M processed
- Record state
- partial_token_timing_source_build_ledger
Public summary
GPT-5.6 Luna Max partial_token_timing_source_build_ledger ledger: 51m 53.5s wall-clock, 25.51M processed, and $0.73 Token-only Standard API-equivalent estimate, re-priced at the official GPT-5.6 Luna rates verified during archival; not the actual Codex subscription charge..
Run identity and stack
- Result ID: off-road-driving-game-threejs-gpt-5.6-luna-max
- Technical model: gpt-5.6-luna
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: gpt-5.6-luna
- Stack: Three.js vehicle workflow
- Stack: Harness v1 off-road prompt
Cost basis
- The official Luna model page applies long-context rates only when a call exceeds 272,000 input tokens. The receipt reports zero such calls.
- The submitted receipt used five-times-higher rates than the official GPT-5.6 Luna model page. Its tokens are retained but its cost is re-priced above.
Primary artifact integrity
- Kind: ledger-checksummed-source-file
- Path: artifacts/off-road-driving-game-threejs-gpt-5.6-luna-max/source/src/main.js
- SHA-256: 9d328fefcab954cc44abee962dab04c6577765add39a1b290c0954d8739921dc
Recorded caveats
- Wall-clock is end-to-end workflow latency, including tools and idle gaps; it is not model-only compute time.
- The post-generation metrics request is included in the recorded wall-clock and token total.
- Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
- Cache-read tokens are discounted, so total processed tokens substantially overstate cost.
- The archive contains a production build created during archival verification, not a recorded final gameplay capture.
- No independent browser acceptance, sustained-drive, local-FPS, generation-seam, suspension-articulation, or blind-evaluation receipt is archived.
Visible evidence gaps
- independent browser acceptance run
- local FPS, browser, viewport, and hardware receipt
- multi-minute sustained-drive and generation-seam receipt
- four-corner suspension-articulation receipt
- final gameplay capture
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
