Measured harness ledgerPublic result
Claude Opus 5

SR-71 Cinematic — Claude Opus 5 Max

Create a complete SR-71 cinematic scene in Blender through the pinned MCP workflow.

Max reasoningHeadline result
Workflow cost
≥$290.01
Wall-clock
16h 22m 10.1s wall-clock
Processed tokens
408.05M processed
Record state
complete_artifact_verified_parent_cost_lower_bound
Public summary

Claude Opus 5 Max complete_artifact_verified_parent_cost_lower_bound ledger: 16h 22m 10.1s wall-clock, 408.05M processed, and ≥$290.01 API-equivalent list-price accounting for the parent transcript only; lower bound, not an itemized Claude Code plan invoice (lower bound).

Run identity and stack
  • Result ID: sr71-cinematic-blender-claude-opus-5-max
  • Technical model: claude-opus-5
  • Provider: Anthropic
  • Client: Claude Code
  • Stack: Anthropic
  • Stack: Claude Code
  • Stack: Technical model/configuration: claude-opus-5
  • Stack: Blender MCP
  • Stack: Cinematic aircraft scene
  • Stack: Pinned Harness prompt
  • Stack: Requested tool profile: blender-mcp-reference-grounded
Cost basis
  • All-5-minute cache-write alternative: $268.71.
Primary artifact integrity
  • Kind: blender-cinematic-aerospace-scene
  • Path: artifacts/sr71-cinematic-blender-claude-opus-5-max/sr71_blackbird_cinematic_final.blend
  • SHA-256: 42567b39a13949fdbb1c52c70f44af272793e3f1ef0a995aaa3d0a96ac9e9e24
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including Blender rendering, tool work, autonomous iteration, and possible idle intervals; it is not model-only compute time.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON, not just visible messages.
  • Cache reads are discounted, so the 408M processed-token total substantially overstates cost relative to fresh input and output pricing.
  • The $290.01+ figure is a parent-transcript lower bound. Judge-subagent usage is reported but not billing-reconciled, so no complete all-agent total is asserted.
  • The automated technical validator was rerun during archival and passed. The 91.5 agent-fidelity score is reported source evidence, not a blind-voter result or an independently replayed judge transcript.
  • No blind human evaluation, RTX final-render run, or hardware-FPS result is currently supplied.
  • The benchmark reference images are checksum-pinned but withheld from the public fixture pending documented redistribution authority.
Visible evidence gaps
  • rights-cleared public distribution of every canonical reference-image payload, or an operator-access protocol that can be independently audited
  • reconciled billing receipt for the independent judge subagents
  • blind human evaluation
  • RTX final render and practical FPS evidence if those measurements are required for the scoreboard
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console