Measured harness ledgerPublic result
Grok 4.6

Neon Gunner — Grok 4.6 xhigh

Build a playable destructible neon run-and-gun game in the browser under the pinned Harness task contract.

xhigh reasoningHeadline result
Workflow cost
$6.73
Wall-clock
32m 15.261s wall-clock
Processed tokens
9.00M processed
Record state
partial_multi_user_turn_metrics_artifact_and_supplied_validator_ledger
Public summary

Grok 4.6 xhigh partial_multi_user_turn_metrics_artifact_and_supplied_validator_ledger ledger: 32m 15.261s wall-clock, 9.00M processed, and $6.73 Provider-recorded usage estimate, not a subscription cash-charge invoice.

Run identity and stack
  • Result ID: neon-gunner-canvas-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Client: Grok 1.0.3
  • Stack: xAI Grok Build
  • Stack: Grok 1.0.3
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Browser canvas game
  • Stack: Harness v1 Neon Gunner prompt
  • Stack: Requested tool profile: Grok Build CLI
Cost basis
  • The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred.
Primary artifact integrity
  • Kind: interactive-vanilla-canvas-run-and-gun
  • Path: artifacts/neon-gunner-canvas-grok-4.6-xhigh/source/index.html
  • SHA-256: 01f8201cabc05d4f85e50942dfaf0899cd358ae9d0d7819eb183c930f7c34f7a
Validation evidence
  • Result: PASS
  • Path: artifacts/neon-gunner-canvas-grok-4.6-xhigh/provenance/validation.public.json
  • SHA-256: e0d2ff294986de65f4161dbd157a0bfe303a93be9ff74e98a5a9b9d20044b40b
  • Validator SHA-256: fa5ec13e9ef2f5f96949283077b7c6278bf9c9f14032817c21743ab973d06090
Recorded caveats
  • The run used three benchmark user turns, including two operator follow-ups, so it is not a one-shot result.
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool calls and waits.
  • Two tool calls reported failures; they are retained in the metrics rather than hidden.
  • Generated output includes visible output and reasoning output as provider-reported.
  • Cache reads are discounted in provider billing, so total processed tokens overstate cost.
  • The $6.73 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
  • The supplied 26/26 validator log was not independently replayed during archival; it is marked as supplied evidence rather than an independent validation claim.
  • The archived browser screenshot verifies title-screen boot only; no local FPS/hardware measurement or final gameplay capture was supplied.
  • No blind-evaluation record is archived.
Visible evidence gaps
  • one-shot-compliant run
  • independent browser replay of the validator
  • local FPS measurement and hardware receipt
  • browser and viewport final gameplay capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console