Measured harness ledgerPublic result
GPT-5.6 Sol

The Last of Us Main Menu — GPT-5.6 Sol Max

Recreate the pinned The Last of Us-style main menu reference as an interactive browser experience.

Max reasoningHeadline result
Workflow cost
$6.18
Wall-clock
28m 14.7s wall-clock
Processed tokens
9.57M processed
Record state
partial_static_build_and_sanitized_token_ledger
Public summary

GPT-5.6 Sol Max partial_static_build_and_sanitized_token_ledger ledger: 28m 14.7s wall-clock, 9.57M processed, and $6.18 Verified API-equivalent token estimate; not an invoice or actual subscription charge..

Run identity and stack
  • Result ID: last-of-us-main-menu-browser-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 menu prompt
Primary artifact integrity
  • Kind: interactive-threejs-main-menu-scene
  • Path: artifacts/last-of-us-main-menu-browser-gpt-5.6-sol-max/source/components/ruin-scene.tsx
  • SHA-256: f0e6ceea300e2e44a9538d2279500a841a6bdc10b499e7b77d8793848a47dce4
Recorded caveats
  • Wall-clock and output throughput were recovered from the archived session's response-item/event payload schema: the first benchmark user message was recorded at 2026-09-01T20:45:27.485Z and task_complete at 2026-09-01T21:13:42.202Z. The legacy parser missed these fields because it only read an older message schema.
  • The API-equivalent estimate is not an invoice or actual charge for the subscription-backed session.
  • Output tokens include hidden reasoning, code, visible prose, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate effective cost.
  • The source archive excludes the original raw metrics file because it contains a private local transcript path and task identifier.
  • The production build passed, but no browser/WebGL replay, local FPS measurement, hardware/viewport record, final capture, or blind evaluation was supplied.
Visible evidence gaps
  • verifiable candidate-prompt delivery and protocol evidence
  • browser, hardware, viewport, and local-FPS receipt
  • independent browser/WebGL replay and final-capture metadata
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console