Measured harness ledgerPublic result
Claude Fable 5.1

CARVE — Claude Fable 5.1 Max

Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.

Max reasoningHeadline result
Workflow cost
$492.04
Wall-clock
2h 20m 30.3s wall-clock
Processed tokens
279.60M processed
Record state
partial_user_turn_count_unverified_token_timing_cost_build_and_operator_webgpu_replay_ledger
Public summary

Claude Fable 5.1 Max partial_user_turn_count_unverified_token_timing_cost_build_and_operator_webgpu_replay_ledger ledger: 2h 20m 30.3s wall-clock, 279.60M processed, and $492.04 Token-only API-equivalent estimate, not the actual Claude Code subscription charge.

Run identity and stack
  • Result ID: carve-webgpu-snowboarding-fable-5.1-max
  • Technical model: claude-fable-5-1
  • Provider: Anthropic Claude Code
  • Stack: Anthropic Claude Code
  • Stack: Technical model/configuration: claude-fable-5-1
  • Stack: Babylon.js WebGPU
  • Stack: Supplied snowboarder fixture
  • Stack: Harness v1 CARVE prompt
Cost basis
  • No separately priced tool or non-token cost is included because no verifiable usage counts were supplied.
  • All-5-minute cache-write alternative: $415.31.
Primary artifact integrity
  • Kind: interactive-webgpu-snowboarding-game
  • Path: artifacts/carve-webgpu-snowboarding-fable-5.1-max/source/src/main.js
  • SHA-256: ef4acc44c77386cf23d5693692e6de1997855348f75cd7a7d5080f892f2d9f0e
Validation evidence
  • Path: artifacts/carve-webgpu-snowboarding-fable-5.1-max/validation/build-and-replay-check.public.json
  • SHA-256: 7aeba30cd18ef8249c4aff6d2abed16a6f5447d8259a965f9367838ca730b0cc
Recorded caveats
  • The supplied metrics receipt does not expose a verified benchmark user-turn count, so this row is not labeled clean one-shot.
  • Wall-clock is end-to-end workflow latency including tool execution and waits, not model-only compute time.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are deeply discounted, so total processed tokens substantially overstate cost.
  • The raw receipt says its metric recomputation includes the metrics-reporting turns themselves.
  • The source tree is provenance-preserving but not byte-identical to the private source capture: eight Blender-only absolute paths were made portable and fully disclosed.
  • The fixed snowboarder is a shared, immutable benchmark input; its prior generation is excluded from this model's benchmark cost and output.
  • The operator replay passed the scripted interaction but retained a nonterminal WebGPU warning. Its in-app FPS values are not a fixed-hardware, cross-run performance result.
  • No blind evaluation is recorded.
Visible evidence gaps
  • verified user-turn count for the supplied session
  • standardized fixed-hardware FPS measurement
  • final 1080p render or capture receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console