Measured harness ledgerPublic result
Grok 4.6

Off-road Mud Game — Grok 4.6 xhigh · Attempt 2

Build a one-request procedural off-road driving game with vehicle physics, independent suspension, streamed terrain, mud feedback, and multiple cameras.

xhigh reasoningHeadline result
Workflow cost
$3.57
Wall-clock
45m 52.619s wall-clock
Processed tokens
22.61M processed
Record state
partial_multi_user_turn_metrics_source_build_ledger
Public summary

Grok 4.6 xhigh partial_multi_user_turn_metrics_source_build_ledger ledger: 45m 52.619s wall-clock, 22.61M processed, and $3.57 Provider-recorded usage estimate, not a subscription cash-charge invoice.

Run identity and stack
  • Result ID: off-road-driving-game-threejs-grok-4.6-xhigh-attempt-2
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Attempt: 2
  • Stack: xAI Grok Build
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Attempt 2
  • Stack: Three.js vehicle workflow
  • Stack: Harness v1 off-road prompt
  • Stack: Requested tool profile: Grok Build CLI
Cost basis
  • The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred.
Primary artifact integrity
  • Kind: interactive-threejs-off-road-driving-game
  • Path: artifacts/off-road-driving-game-threejs-grok-4.6-xhigh-attempt-2/source/src/main.js
  • SHA-256: 6b7cb82f461cb72ca334c2057a5fa7a0337ad200106d579e4d7eb31aaecdba76
Recorded caveats
  • This is attempt 2; the existing NORTHLINE KESTREL Grok result remains a separate first attempt.
  • The run used two benchmark user turns, including one operator follow-up, so it is not a one-shot result.
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool calls and waits.
  • Four tool calls reported failures; they are retained in the metrics rather than hidden.
  • Generated output includes visible output and reasoning output as provider-reported.
  • Cache reads are discounted in provider billing, so total processed tokens overstate cost.
  • The $3.57 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
  • The global extractor failed closed because a sibling session reused a static marker; the source reports a passing session-scoped extraction for the live generation session.
  • The supplied captures lack browser, hardware, control, timing, and reproducible-procedure metadata. The local production build and static JavaScript checks pass, but no independent driving run or FPS measurement is claimed.
  • No final RTX Pro 6000 render or blind-evaluation record is archived.
  • The first user turn is the benchmark generation prompt. The second is an operator background-task-notice follow-up. This is not a one-shot result.
Visible evidence gaps
  • one-shot-compliant run
  • globally unique extractor marker or global discovery isolation
  • browser and local-hardware identity
  • sustained driving and physics validation
  • local FPS measurement
  • final capture metadata
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console