Measured harness ledgerPublic result
GPT-5.6 Sol

ECHO Main Menu — GPT-5.6 Sol Max

Recreate the pinned ECHO-style cinematic main menu reference as an interactive browser experience.

Max reasoningHeadline result
Workflow cost
$11.35
Wall-clock
45:46.1 (supplemental schema-derived) wall-clock
Processed tokens
19.59M processed
Record state
partial_static_build_and_supplemental_timing_ledger
Public summary

GPT-5.6 Sol Max partial_static_build_and_supplemental_timing_ledger ledger: 45:46.1 (supplemental schema-derived) wall-clock, 19.59M processed, and $11.35 Verified API-equivalent token estimate; not an invoice or actual subscription charge..

Run identity and stack
  • Result ID: echo-main-menu-browser-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Browser recreation
  • Stack: Pinned visual references
  • Stack: Harness v1 ECHO prompt
Primary artifact integrity
  • Kind: interactive-browser-title-menu
  • Path: artifacts/echo-main-menu-browser-gpt-5.6-sol-max/source/components/echo-eye-experience.tsx
  • SHA-256: 626030381bd521bd5e29d98d191ab075740390cf7da6b754b60eaa76a00259c3
Recorded caveats
  • Raw parser timing is null because its user-message selector did not match this transcript schema; wall-clock and throughput use the separately labeled supplement.
  • The supplement identifies three response-item user messages, so one-shot compliance is not established.
  • The API-equivalent estimate is not an invoice or actual charge for the subscription-backed session.
  • Output tokens include hidden reasoning, code, visible prose, and tool-call JSON; cache reads are discounted, so processed-token volume overstates cost.
  • The source archive excludes raw metrics because the supplied original includes a private local transcript path and task identifier.
  • The production build passed, but no browser/WebGL replay, keyboard/menu interaction check, reference-state capture, local FPS, hardware/viewport record, or blind evaluation was supplied.
Visible evidence gaps
  • verifiable candidate-prompt delivery
  • one-shot protocol evidence or explicit protocol-deviation adjudication
  • independent browser/WebGL replay with keyboard/menu interaction checks
  • reference-state candidate captures
  • browser, hardware, viewport, and local-FPS receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console