Measured harness ledgerPublic result
Grok 4.6Off-road Mud Game — Grok 4.6 xhigh
Build a one-request procedural off-road driving game with vehicle physics, independent suspension, streamed terrain, mud feedback, and multiple cameras.
xhigh reasoningHeadline result
- Workflow cost
- $23.95
- Wall-clock
- 1h 00m 18.521s wall-clock
- Processed tokens
- 28.28M processed
- Record state
- partial_multi_user_turn_metrics_artifact_build_and_runtime_ledger
Public summary
Grok 4.6 xhigh partial_multi_user_turn_metrics_artifact_build_and_runtime_ledger ledger: 1h 00m 18.521s wall-clock, 28.28M processed, and $23.95 Provider-recorded usage estimate, not a subscription cash-charge invoice.
Run identity and stack
- Result ID: off-road-driving-game-threejs-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok Build
- Client: Grok 1.0.3
- Stack: xAI Grok Build
- Stack: Grok 1.0.3
- Stack: Technical model/configuration: Grok 4.6
- Stack: Three.js vehicle workflow
- Stack: Harness v1 off-road prompt
- Stack: Requested tool profile: Grok Build CLI
Cost basis
- The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred.
Primary artifact integrity
- Kind: interactive-threejs-off-road-driving-game
- Path: artifacts/off-road-driving-game-threejs-grok-4.6-xhigh/source/src/main.js
- SHA-256: 354ab4c9cb1105047119b83f3b43858037dbbb227b14a7e4438076fc6ea3d5cf
Recorded caveats
- The run used two benchmark user turns, including one operator follow-up, so it is not a one-shot result.
- Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool calls and waits.
- Two tool calls reported failures; they are retained in the metrics rather than hidden.
- Generated output includes visible output and reasoning output as provider-reported.
- Cache reads are discounted in provider billing, so total processed tokens overstate cost.
- The $23.95 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
- The published supplied captures lack browser, hardware, control, timing, and reproducible-procedure metadata.
- A clean production build does not establish browser FPS or driving quality; the archive's scripted high-speed probe ended with no wheel contacts, so no sustained-handling claim is made.
- No final RTX Pro 6000 render or blind-evaluation record is archived.
Visible evidence gaps
- one-shot-compliant run
- browser and local-hardware identity
- sustained driving and physics validation
- local FPS measurement
- final capture metadata
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
