Measured harness ledgerPublic result
Claude Opus 5

F-117 Stealth Jet — Claude Opus 5 Max

Build an interactive Three.js F-117 stealth-jet experience under the pinned Harness task and fixture contract.

Max reasoningHeadline result
Workflow cost
$46.23
Wall-clock
61m 18.5s wall-clock
Processed tokens
73.65M processed
Record state
partial_metrics_source_build_ledger
Public summary

Claude Opus 5 Max partial_metrics_source_build_ledger ledger: 61m 18.5s wall-clock, 73.65M processed, and $46.23 Token-only API-equivalent estimate, not the actual Claude Code subscription charge.

Run identity and stack
  • Result ID: f117-stealth-jet-threejs-opus-5-max
  • Technical model: claude-opus-5
  • Provider: Claude Code
  • Stack: Claude Code
  • Stack: Technical model/configuration: claude-opus-5
  • Stack: Three.js flight experience
  • Stack: Harness v1 F-117 prompt
Cost basis
  • No separately billed server-side web-search or web-fetch cost is included; the source receipt reports zero server-tool counters.
Primary artifact integrity
  • Kind: interactive-threejs-f117-model
  • Path: artifacts/f117-stealth-jet-threejs-opus-5-max/source/src/main.js
  • SHA-256: cd5225416f7320a12c2b533e9e682584a3de1bc03f8abbae27b0de2d52688ad7
Validation evidence
  • Path: artifacts/f117-stealth-jet-threejs-opus-5-max/validation/build-and-source-check.public.json
  • SHA-256: a24b86137bee1d52ed2d7708fe91927ea89cd8845364892fbddc1ca2eb0539fc
Recorded caveats
  • The source receipt identifies one measured Claude Code session but does not record user-turn count, so this is not claimed as a clean one-shot result.
  • Wall-clock is end-to-end workflow latency, including tool execution, dev-server restarts, browser round trips, and waits; it is not model-only compute time.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache-read tokens are discounted, so processed-token volume substantially overstates cost.
  • The raw receipt's cache-write prose says 394,711 tokens whereas its numerical ledger and calculation use 409,868; the latter is used and the mismatch is disclosed.
  • The cost is API-list-price equivalent, not the actual subscription-backed Claude Code charge.
  • The production build passed, but no test script, independent browser acceptance run, local FPS measurement, hardware/viewport receipt, or blind evaluation was supplied.
Visible evidence gaps
  • explicit user-turn protocol receipt or clean one-user-turn rerun
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • independent final-capture receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console