Measured harness ledgerPublic result
Claude Opus 5Sekiro Main Menu — Claude Opus 5 Max
Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.
Max reasoningHeadline result
- Workflow cost
- ≥$23.38
- Wall-clock
- 4h 50m 04.2s wall-clock
- Processed tokens
- 12.49M processed
- Record state
- partial_multi_agent_root_metrics_and_artifact_ledger
Public summary
Claude Opus 5 Max partial_multi_agent_root_metrics_and_artifact_ledger ledger: 4h 50m 04.2s wall-clock, 12.49M processed, and ≥$23.38 Claude API list-price equivalent for the measured root orchestrator session only; lower bound, not an itemized Claude Code charge or total multi-agent workflow cost (lower bound).
Run identity and stack
- Result ID: sekiro-main-menu-browser-opus-5-max
- Technical model: claude-opus-5
- Provider: Anthropic Claude Code
- Stack: Anthropic Claude Code
- Stack: Technical model/configuration: claude-opus-5
- Stack: Browser recreation
- Stack: Pinned visual reference
- Stack: Harness v1 Sekiro prompt
Primary artifact integrity
- Kind: interactive-browser-title-menu
- Path: artifacts/sekiro-main-menu-browser-opus-5-max/source/index.html
- SHA-256: b173000abcb4da87ad6b48fbf98b786e0f0bd5fc7dae630301089c939690d536
Recorded caveats
- This is a multi-agent workflow. The supplied root-session cost omits separate subagent cost, so $23.378917 is a lower bound rather than total workflow cost.
- The user-turn count is unreported; this is not validated as a one-shot benchmark run.
- The primary wall-clock reporting value is active-only time: 4h 50m 04.2s. The original script-reported raw span (11h 39m 31.2s) is retained for reproducibility and includes a documented user-idle interval; neither duration is model-only compute.
- The later one-hour cache-TTL audit reports a larger cache-write count from a later transcript moment. The primary cost follows the internally self-consistent snapshot; it should not be treated as an exact total workflow invoice.
- Included images are model-supplied project captures, not independent RemakeBench replay captures.
- No local FPS receipt, browser/hardware/viewport evidence, RTX capture, or blind-evaluation record was supplied.
Visible evidence gaps
- Complete multi-agent cost accounting
- Recorded user-turn / one-shot protocol evidence
- Independent browser interaction and audio replay
- local FPS with browser, hardware, and viewport receipt
- RTX final capture
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
