Measured harness ledgerPublic result
Grok 4.6

CARVE — Grok 4.6 xhigh

Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.

xhigh reasoningHeadline result
Workflow cost
$12.66
Wall-clock
25m 24.387s wall-clock
Processed tokens
15.12M processed
Record state
partial_metrics_source_and_production_build_ledger
Public summary

Grok 4.6 xhigh partial_metrics_source_and_production_build_ledger ledger: 25m 24.387s wall-clock, 15.12M processed, and $12.66 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: carve-webgpu-snowboarding-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Client: Grok 1.0.3
  • Stack: xAI Grok Build
  • Stack: Grok 1.0.3
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Babylon.js WebGPU
  • Stack: Supplied snowboarder fixture
  • Stack: Harness v1 CARVE prompt
Cost basis
  • The exact billed grok-4.6-build alias has no official catalog row. The $12.71 figure uses published grok-4.6 rates as a labeled selected-model API equivalent, not as a grok-4.6-build list price.
Primary artifact integrity
  • Kind: interactive-webgpu-snowboarding-game-source
  • Path: artifacts/carve-webgpu-snowboarding-grok-4.6-xhigh/source/src/main.js
  • SHA-256: 1fc1cd0ba174e6a4fe98ef3d03daeee9b3d683e92340547b8888c08feca9f622
Recorded caveats
  • This is a three-user-turn protocol deviation, not a comparable one-shot run.
  • The canonical extractor result remains FAIL. The partial aggregate relies on a separately disclosed isolated-window reconciliation and must not be relabeled PASS.
  • Wall-clock is end-to-end workflow latency including tool execution, installs, tests, browser checks, and idle gaps; it is not model-only compute.
  • Reported API duration is summed across model calls and can exceed wall-clock when subagents overlap.
  • Generated output includes visible output and reasoning output as provider-reported.
  • Cache-read input is discounted in provider billing, so processed-token totals should not be treated as cost equivalents.
  • The $12.66 is a provider-recorded usage estimate, not proof of a marginal cash charge. The $12.71 selected-model equivalent is separately labeled because the billed alias lacks a public rate.
  • The build and static asset inventory pass, but browser/WebGPU behavior, audio response, gameplay loop, persistent deformation, frame rate, final capture, and visual quality were not independently measured.
  • No blind evaluation is recorded.
Visible evidence gaps
  • clean one-user-turn rerun
  • WebGPU browser runtime check
  • timed-run, checkpoint, wipeout/respawn, and run-summary check
  • persistent snow deformation and re-ridable track check
  • audio response check
  • 1080p browser, GPU, viewport, and local FPS receipt
  • final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console