Measured harness ledgerPublic result
Grok 4.6CARVE — Grok 4.6 xhigh
Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.
xhigh reasoningHeadline result
- Workflow cost
- $12.66
- Wall-clock
- 25m 24.387s wall-clock
- Processed tokens
- 15.12M processed
- Record state
- partial_metrics_source_and_production_build_ledger
Public summary
Grok 4.6 xhigh partial_metrics_source_and_production_build_ledger ledger: 25m 24.387s wall-clock, 15.12M processed, and $12.66 Provider-recorded usage estimate, not an itemized subscription cash charge.
Run identity and stack
- Result ID: carve-webgpu-snowboarding-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok Build
- Client: Grok 1.0.3
- Stack: xAI Grok Build
- Stack: Grok 1.0.3
- Stack: Technical model/configuration: Grok 4.6
- Stack: Babylon.js WebGPU
- Stack: Supplied snowboarder fixture
- Stack: Harness v1 CARVE prompt
Cost basis
- The exact billed grok-4.6-build alias has no official catalog row. The $12.71 figure uses published grok-4.6 rates as a labeled selected-model API equivalent, not as a grok-4.6-build list price.
Primary artifact integrity
- Kind: interactive-webgpu-snowboarding-game-source
- Path: artifacts/carve-webgpu-snowboarding-grok-4.6-xhigh/source/src/main.js
- SHA-256: 1fc1cd0ba174e6a4fe98ef3d03daeee9b3d683e92340547b8888c08feca9f622
Recorded caveats
- This is a three-user-turn protocol deviation, not a comparable one-shot run.
- The canonical extractor result remains FAIL. The partial aggregate relies on a separately disclosed isolated-window reconciliation and must not be relabeled PASS.
- Wall-clock is end-to-end workflow latency including tool execution, installs, tests, browser checks, and idle gaps; it is not model-only compute.
- Reported API duration is summed across model calls and can exceed wall-clock when subagents overlap.
- Generated output includes visible output and reasoning output as provider-reported.
- Cache-read input is discounted in provider billing, so processed-token totals should not be treated as cost equivalents.
- The $12.66 is a provider-recorded usage estimate, not proof of a marginal cash charge. The $12.71 selected-model equivalent is separately labeled because the billed alias lacks a public rate.
- The build and static asset inventory pass, but browser/WebGPU behavior, audio response, gameplay loop, persistent deformation, frame rate, final capture, and visual quality were not independently measured.
- No blind evaluation is recorded.
Visible evidence gaps
- clean one-user-turn rerun
- WebGPU browser runtime check
- timed-run, checkpoint, wipeout/respawn, and run-summary check
- persistent snow deformation and re-ridable track check
- audio response check
- 1080p browser, GPU, viewport, and local FPS receipt
- final capture
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
