Measured harness ledgerPublic result
Grok 4.6

Space Flight Stage 2 — Landfall — Grok 4.6 xhigh

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

xhigh reasoningHeadline result
Workflow cost
$4.32
Wall-clock
26m 36.958s wall-clock
Processed tokens
7.01M processed
Record state
partial_one_shot_token_timing_production_build_and_scripted_acceptance_ledger
Public summary

Grok 4.6 xhigh partial_one_shot_token_timing_production_build_and_scripted_acceptance_ledger ledger: 26m 36.958s wall-clock, 7.01M processed, and $4.32 Provider-recorded usage estimate, not a subscription cash-charge invoice.

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Client: Grok CLI 1.0.3
  • Stack: xAI Grok Build
  • Stack: Grok CLI 1.0.3
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Cost basis
  • The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred. Aggregate metrics also cannot apply a per-request long-context threshold.
Primary artifact integrity
  • Kind: archive-operator-rebuilt-static-production-entry
  • Path: artifacts/space-flight-game-threejs-stage-2-landfall-grok-4.6-xhigh/production-build/index.html
  • SHA-256: fce6a96f5598e73eca257482382857cddda4a32000eab1d1481da8ef2572e296
Validation evidence
  • Result: production_build_and_scripted_two_cycle_check_passed
Recorded caveats
  • The one-shot Stage 2 metric is a continuation of this model's separate Stage 1 STARWAKE project, which retained a three-user-turn protocol deviation.
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle intervals.
  • One tool failure is retained in the reported metrics rather than hidden.
  • Generated output includes visible output and provider-reported reasoning output.
  • Cache reads are discounted in provider billing, so total processed tokens overstate cost.
  • The $4.32 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
  • The supplied checker is replayed against a fresh production build, but it uses application test hooks for scenario setup. It does not independently establish an uninterrupted manually controlled no-cut loop, benchmark FPS, or transition hitch behavior.
  • The shipped infantry asset is hashed but no frozen input fixture or distributable license record was supplied for it.
  • No RTX Pro 6000 1080p60 render or blind-evaluation record is archived.
  • This Stage 2 request is one genuine benchmark user turn. It deliberately continues the user-provided Grok STARWAKE Stage 1 project; the inherited Stage 1 result remains separately disclosed as a three-user-turn protocol deviation.
Visible evidence gaps
  • independent manual controls acceptance run through the uninterrupted landing and ascent loop
  • local FPS and transition-hitch receipt with hardware, viewport, and graphics settings
  • frozen fixture manifest and licensing record for the provided infantry asset
  • RTX Pro 6000 1080p60 render
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console