Measured harness ledgerPublic result
GPT-5.6 Sol

Sekiro Main Menu — GPT-5.6 Sol Max

Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.

Max reasoningHeadline result
Workflow cost
$15.96
Wall-clock
2h 19m 44.1s wall-clock
Processed tokens
22.54M processed
Record state
partial_metrics_and_artifact_ledger
Public summary

GPT-5.6 Sol Max partial_metrics_and_artifact_ledger ledger: 2h 19m 44.1s wall-clock, 22.54M processed, and $15.96 API-equivalent token estimate; not an actual Codex subscription charge.

Run identity and stack
  • Result ID: sekiro-main-menu-browser-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 Sekiro prompt
Cost basis
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: browser-title-menu-source-package
  • Path: artifacts/sekiro-main-menu-browser-gpt-5.6-sol-max/deliverable/sekiro-cinematic-menu.zip
  • SHA-256: b0154e7c095af67f3dfbad347e1b8df34d2f29a84935d7e873b1b079bf77ce70
Recorded caveats
  • The supplied protocol contains two user turns; it is not a one-shot benchmark run.
  • Wall-clock is end-to-end latency, not model-only compute, and includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session; separately priced tools and non-token services are excluded.
  • The source raw receipt is not published because it contains private local-session information; the published record binds its SHA-256 and provides a sanitized metrics derivative.
  • No independent browser replay, local FPS with browser/hardware/viewport receipt, RTX final capture, or blind-evaluation record was supplied.
Visible evidence gaps
  • One-shot protocol evidence
  • Independent browser interaction, audio, keyboard, and gamepad replay
  • local FPS with browser, hardware, and viewport receipt
  • RTX final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console