Measured harness ledgerPublic result
Claude Opus 5

Sekiro Main Menu — Claude Opus 5 Max

Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.

Max reasoningHeadline result
Workflow cost
≥$23.38
Wall-clock
4h 50m 04.2s wall-clock
Processed tokens
12.49M processed
Record state
partial_multi_agent_root_metrics_and_artifact_ledger
Public summary

Claude Opus 5 Max partial_multi_agent_root_metrics_and_artifact_ledger ledger: 4h 50m 04.2s wall-clock, 12.49M processed, and ≥$23.38 Claude API list-price equivalent for the measured root orchestrator session only; lower bound, not an itemized Claude Code charge or total multi-agent workflow cost (lower bound).

Run identity and stack
  • Result ID: sekiro-main-menu-browser-opus-5-max
  • Technical model: claude-opus-5
  • Provider: Anthropic Claude Code
  • Stack: Anthropic Claude Code
  • Stack: Technical model/configuration: claude-opus-5
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 Sekiro prompt
Primary artifact integrity
  • Kind: interactive-browser-title-menu
  • Path: artifacts/sekiro-main-menu-browser-opus-5-max/source/index.html
  • SHA-256: b173000abcb4da87ad6b48fbf98b786e0f0bd5fc7dae630301089c939690d536
Recorded caveats
  • This is a multi-agent workflow. The supplied root-session cost omits separate subagent cost, so $23.378917 is a lower bound rather than total workflow cost.
  • The user-turn count is unreported; this is not validated as a one-shot benchmark run.
  • The primary wall-clock reporting value is active-only time: 4h 50m 04.2s. The original script-reported raw span (11h 39m 31.2s) is retained for reproducibility and includes a documented user-idle interval; neither duration is model-only compute.
  • The later one-hour cache-TTL audit reports a larger cache-write count from a later transcript moment. The primary cost follows the internally self-consistent snapshot; it should not be treated as an exact total workflow invoice.
  • Included images are model-supplied project captures, not independent RemakeBench replay captures.
  • No local FPS receipt, browser/hardware/viewport evidence, RTX capture, or blind-evaluation record was supplied.
Visible evidence gaps
  • Complete multi-agent cost accounting
  • Recorded user-turn / one-shot protocol evidence
  • Independent browser interaction and audio replay
  • local FPS with browser, hardware, and viewport receipt
  • RTX final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console