Measured harness ledgerPublic result
Grok 4.6

Red Sands v1 — Grok 4.6 xhigh

Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.

xhigh reasoningHeadline result
Workflow cost
$5.07
Wall-clock
21m 43.319s wall-clock
Processed tokens
7.42M processed
Record state
partial_one_shot_metrics_artifact_model_smoke_without_quest_playthrough
Public summary

Grok 4.6 xhigh partial_one_shot_metrics_artifact_model_smoke_without_quest_playthrough ledger: 21m 43.319s wall-clock, 7.42M processed, and $5.07 Provider-recorded usage estimate, not a subscription cash-charge invoice.

Run identity and stack
  • Result ID: red-sands-quest-threejs-v2-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Client: Grok 1.0.3
  • Stack: xAI Grok Build
  • Stack: Grok 1.0.3
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Three.js quest-authoring fixture
  • Stack: Harness v1 frozen prompt
  • Stack: Requested tool profile: Grok Build CLI
Cost basis
  • The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/red-sands-quest-threejs-v2-grok-4.6-xhigh/source/src/quest/quests/pikes-star.js
  • SHA-256: ca48c9351379009f33bb07a7ccec1c9c595eae8daa59b6ac0895813cacbbeb69
Recorded caveats
  • Red Sands v2 is behaviorally distinct from v1; its results should not be compared one-to-one with v1 submissions.
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool calls and waits.
  • Three tool calls reported failures; they are retained in the metrics rather than hidden.
  • Generated output includes visible output and reasoning output as provider-reported.
  • Cache reads are discounted in provider billing, so total processed tokens overstate cost.
  • The $5.07 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
  • The candidate is purely additive and the protected fixture files match their v2 hashes, but no end-to-end quest playthrough, completion trace, or independent visual review was supplied.
  • No blind evaluation is archived.
Visible evidence gaps
  • Independent end-to-end playthrough or replay showing start, objective progression, gunfight, return, and completion
  • Quest-system runtime capture outside capture mode
  • Blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console