Measured harness ledgerPublic result
GPT-5.6 SolThe Last of Us Main Menu — GPT-5.6 Sol Max
Recreate the pinned The Last of Us-style main menu reference as an interactive browser experience.
Max reasoningHeadline result
- Workflow cost
- $6.18
- Wall-clock
- 28m 14.7s wall-clock
- Processed tokens
- 9.57M processed
- Record state
- partial_static_build_and_sanitized_token_ledger
Public summary
GPT-5.6 Sol Max partial_static_build_and_sanitized_token_ledger ledger: 28m 14.7s wall-clock, 9.57M processed, and $6.18 Verified API-equivalent token estimate; not an invoice or actual subscription charge..
Run identity and stack
- Result ID: last-of-us-main-menu-browser-gpt-5.6-sol-max
- Technical model: gpt-5.6-sol
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: gpt-5.6-sol
- Stack: Browser recreation
- Stack: Pinned visual reference
- Stack: Harness v1 menu prompt
Primary artifact integrity
- Kind: interactive-threejs-main-menu-scene
- Path: artifacts/last-of-us-main-menu-browser-gpt-5.6-sol-max/source/components/ruin-scene.tsx
- SHA-256: f0e6ceea300e2e44a9538d2279500a841a6bdc10b499e7b77d8793848a47dce4
Recorded caveats
- Wall-clock and output throughput were recovered from the archived session's response-item/event payload schema: the first benchmark user message was recorded at 2026-09-01T20:45:27.485Z and task_complete at 2026-09-01T21:13:42.202Z. The legacy parser missed these fields because it only read an older message schema.
- The API-equivalent estimate is not an invoice or actual charge for the subscription-backed session.
- Output tokens include hidden reasoning, code, visible prose, and tool-call JSON.
- Cache reads are discounted, so total processed tokens overstate effective cost.
- The source archive excludes the original raw metrics file because it contains a private local transcript path and task identifier.
- The production build passed, but no browser/WebGL replay, local FPS measurement, hardware/viewport record, final capture, or blind evaluation was supplied.
Visible evidence gaps
- verifiable candidate-prompt delivery and protocol evidence
- browser, hardware, viewport, and local-FPS receipt
- independent browser/WebGL replay and final-capture metadata
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
