Measured harness ledgerPublic result
Claude Fable 5.1

Sekiro Main Menu — Claude Fable 5.1 Max

Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.

Max reasoningHeadline result
Workflow cost
$26.06
Wall-clock
70m (rounded provider display) wall-clock
Processed tokens
Not recorded
Record state
partial_artifact_build_and_rounded_usage_ledger
Public summary

Claude Fable 5.1 Max partial_artifact_build_and_rounded_usage_ledger ledger: 70m (rounded provider display) wall-clock, Not recorded, and $26.06 provider-reported session cost; not independently recomputable from rounded display values.

Run identity and stack
  • Result ID: sekiro-main-menu-browser-fable-5.1-max
  • Technical model: claude-fable-5.1
  • Provider: Anthropic Claude Code
  • Stack: Anthropic Claude Code
  • Stack: Technical model/configuration: claude-fable-5.1
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 Sekiro prompt
Cost basis
  • No raw session export, exact token counts, cache TTL, or per-line price receipt was supplied.
Primary artifact integrity
  • Kind: interactive-browser-title-menu
  • Path: artifacts/sekiro-main-menu-browser-fable-5.1-max/source/index.html
  • SHA-256: e01aec65526763bde15549666230cfcbbef7a10f84a3fffcb581dc2b3b9aa33d
Recorded caveats
  • Wall-clock and API duration are whole-minute values displayed by the provider, not raw timestamps.
  • The displayed input and cache values are abbreviated; the recorded 61,313,398 processed-token total is an approximation from those displayed components.
  • The $26.06 total is provider-reported and is not a verified API-equivalent cost calculation.
  • The supplied report does not establish one-shot protocol compliance or session-specific assistant/model-call counts.
  • Rolling local-activity requests are excluded because they are not session-attributed benchmark metrics.
  • Build/typecheck success does not establish runtime interaction, audio, WebGL visual quality, local FPS, RTX capture, or blind-vote outcomes.
  • The supplied usage report identifies a session but does not provide user-turn, assistant-turn, or raw transcript evidence.
Visible evidence gaps
  • raw Claude Code session export or exact machine-readable token receipt
  • per-line pricing and cache-TTL receipt sufficient to recompute cost
  • recorded one-shot protocol evidence
  • browser interaction and audio replay receipt
  • local FPS with browser, hardware, and viewport receipt
  • RTX final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console