Measured harness ledgerPublic result
Grok 4.6Courtyard Blossom Stage 2 — Grok 4.6 xhigh
Create a second cherry-blossom temple courtyard arena for the fixed fighting game under the pinned reference-guided Stage 2 contract.
xhigh reasoningHeadline result
- Workflow cost
- $9.22
- Wall-clock
- 22m 49.670s wall-clock
- Processed tokens
- 10.24M processed
- Record state
- delivered_playable_stage_2_artifact_non_scoring
Public summary
Grok 4.6 xhigh delivered_playable_stage_2_artifact_non_scoring ledger: 22m 49.670s wall-clock, 10.24M processed, and $9.22 Provider-recorded usage estimate, not a subscription cash-charge invoice.
Run identity and stack
- Result ID: courtyard-arena-blossom-stage-2-reference-exposed-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok Build
- Client: Grok CLI 1.0.3
- Stack: xAI Grok Build
- Stack: Grok CLI 1.0.3
- Stack: Technical model/configuration: Grok 4.6
- Stack: Blender / browser arena workflow
- Stack: Reference-guided Stage 2 fixture
Cost basis
- The exact billed alias grok-4.6-build has no official public catalog price, and per-request long-context buckets are unavailable.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/courtyard-arena-blender-mcp-grok-4.6-xhigh-reference-exposed/source/web/models/arena_source.blend
- SHA-256: 69a7870328c80774060eeb20b1e5e0d68f3ea2fc1222446fbf6b16b532aa61c6
Validation evidence
- Result: PASS_WITH_NONCOMPARABILITY
Recorded caveats
- This record stores a completed playable Stage 2 Blossom Temple artifact; it is not a failure record.
- It is non-scoring historical evidence rather than a clean Stage 2 leaderboard entry because the original workspace had additional withheld Stage 1 material outside the new fixture's explicit inputs.
- Wall-clock is end-to-end workflow latency, not model-only compute.
- Eight tool failures are retained in the reported metrics rather than hidden.
- Generated output includes visible output and provider-reported reasoning output.
- Cache reads are discounted in provider billing, so total processed tokens overstate cost.
- The $9.22 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or subscription cash charge.
- Runtime replay proves boot and deterministic stepping only; it is not an FPS, sustained-play, visual-quality, or blind-evaluation measurement.
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
