Measured harness ledgerPublic result
Grok 4.6

Courtyard Blossom Stage 2 — Grok 4.6 xhigh

Create a second cherry-blossom temple courtyard arena for the fixed fighting game under the pinned reference-guided Stage 2 contract.

xhigh reasoningHeadline result
Workflow cost
$9.22
Wall-clock
22m 49.670s wall-clock
Processed tokens
10.24M processed
Record state
delivered_playable_stage_2_artifact_non_scoring
Public summary

Grok 4.6 xhigh delivered_playable_stage_2_artifact_non_scoring ledger: 22m 49.670s wall-clock, 10.24M processed, and $9.22 Provider-recorded usage estimate, not a subscription cash-charge invoice.

Run identity and stack
  • Result ID: courtyard-arena-blossom-stage-2-reference-exposed-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Client: Grok CLI 1.0.3
  • Stack: xAI Grok Build
  • Stack: Grok CLI 1.0.3
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Blender / browser arena workflow
  • Stack: Reference-guided Stage 2 fixture
Cost basis
  • The exact billed alias grok-4.6-build has no official public catalog price, and per-request long-context buckets are unavailable.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/courtyard-arena-blender-mcp-grok-4.6-xhigh-reference-exposed/source/web/models/arena_source.blend
  • SHA-256: 69a7870328c80774060eeb20b1e5e0d68f3ea2fc1222446fbf6b16b532aa61c6
Validation evidence
  • Result: PASS_WITH_NONCOMPARABILITY
Recorded caveats
  • This record stores a completed playable Stage 2 Blossom Temple artifact; it is not a failure record.
  • It is non-scoring historical evidence rather than a clean Stage 2 leaderboard entry because the original workspace had additional withheld Stage 1 material outside the new fixture's explicit inputs.
  • Wall-clock is end-to-end workflow latency, not model-only compute.
  • Eight tool failures are retained in the reported metrics rather than hidden.
  • Generated output includes visible output and provider-reported reasoning output.
  • Cache reads are discounted in provider billing, so total processed tokens overstate cost.
  • The $9.22 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or subscription cash charge.
  • Runtime replay proves boot and deterministic stepping only; it is not an FPS, sustained-play, visual-quality, or blind-evaluation measurement.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console