Measured harness ledgerPublic result
Grok 4.6Space Flight Stage 2 — Landfall — Grok 4.6 xhigh
Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.
xhigh reasoningHeadline result
- Workflow cost
- $4.32
- Wall-clock
- 26m 36.958s wall-clock
- Processed tokens
- 7.01M processed
- Record state
- partial_one_shot_token_timing_production_build_and_scripted_acceptance_ledger
Public summary
Grok 4.6 xhigh partial_one_shot_token_timing_production_build_and_scripted_acceptance_ledger ledger: 26m 36.958s wall-clock, 7.01M processed, and $4.32 Provider-recorded usage estimate, not a subscription cash-charge invoice.
Run identity and stack
- Result ID: space-flight-game-threejs-stage-2-landfall-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok Build
- Client: Grok CLI 1.0.3
- Stack: xAI Grok Build
- Stack: Grok CLI 1.0.3
- Stack: Technical model/configuration: Grok 4.6
- Stack: Three.js Stage 2 workflow
- Stack: Harness v1 Landfall prompt
Cost basis
- The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred. Aggregate metrics also cannot apply a per-request long-context threshold.
Primary artifact integrity
- Kind: archive-operator-rebuilt-static-production-entry
- Path: artifacts/space-flight-game-threejs-stage-2-landfall-grok-4.6-xhigh/production-build/index.html
- SHA-256: fce6a96f5598e73eca257482382857cddda4a32000eab1d1481da8ef2572e296
Validation evidence
- Result: production_build_and_scripted_two_cycle_check_passed
Recorded caveats
- The one-shot Stage 2 metric is a continuation of this model's separate Stage 1 STARWAKE project, which retained a three-user-turn protocol deviation.
- Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle intervals.
- One tool failure is retained in the reported metrics rather than hidden.
- Generated output includes visible output and provider-reported reasoning output.
- Cache reads are discounted in provider billing, so total processed tokens overstate cost.
- The $4.32 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
- The supplied checker is replayed against a fresh production build, but it uses application test hooks for scenario setup. It does not independently establish an uninterrupted manually controlled no-cut loop, benchmark FPS, or transition hitch behavior.
- The shipped infantry asset is hashed but no frozen input fixture or distributable license record was supplied for it.
- No RTX Pro 6000 1080p60 render or blind-evaluation record is archived.
- This Stage 2 request is one genuine benchmark user turn. It deliberately continues the user-provided Grok STARWAKE Stage 1 project; the inherited Stage 1 result remains separately disclosed as a three-user-turn protocol deviation.
Visible evidence gaps
- independent manual controls acceptance run through the uninterrupted landing and ascent loop
- local FPS and transition-hitch receipt with hardware, viewport, and graphics settings
- frozen fixture manifest and licensing record for the provided infantry asset
- RTX Pro 6000 1080p60 render
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
