Measured harness ledgerPublic result
Grok 4.6

F-117 Stealth Jet — Grok 4.6 xhigh

Build an interactive Three.js F-117 stealth-jet experience under the pinned Harness task and fixture contract.

xhigh reasoningHeadline result
Workflow cost
Not recorded
Wall-clock
24m 52.094s wall-clock
Processed tokens
7.71M processed
Record state
partial_runtime_evidence
Public summary

Grok 4.6 xhigh partial_runtime_evidence ledger: 24m 52.094s wall-clock, 7.71M processed, and Not recorded.

Run identity and stack
  • Result ID: f117-stealth-jet-threejs-grok-4.6-xhigh
  • Technical model: grok-4.6
  • Provider: xAI Grok Build
  • Stack: xAI Grok Build
  • Stack: Technical model/configuration: grok-4.6
  • Stack: Three.js flight experience
  • Stack: Harness v1 F-117 prompt
Cost basis
  • No measured or estimated run cost is recorded.
Primary artifact integrity
  • Kind: interactive-threejs-f117-model
  • Path: artifacts/f117-stealth-jet-threejs-grok-4.6-xhigh/source/src/main.js
  • SHA-256: 8fcaaf4dc62404f2dfe08e346fe4de3cc99bc0480c851eec9d4a9e983e369c15
Validation evidence
  • Path: artifacts/f117-stealth-jet-threejs-grok-4.6-xhigh/validation/build-and-source-check.public.json
  • SHA-256: 3e17b42ffb71a7161ef826a1c4c2fdbb21dc53cb8eedaa3d84440f9d13379a44
Recorded caveats
  • The benchmark window is one user turn with zero pre-marker follow-ups, but the canonical global extractor discovery failed closed because an unrelated historical session reused a static marker and had compaction. The current session's manual isolation reconciles and is disclosed rather than silently substituting the failed discovery output.
  • Wall-clock is end-to-end latency, including tools, installs, browser checks, and waits; it is not model-only compute time.
  • Visible and reasoning output are separate receipt fields but are both included in output_generation and normally output-priced.
  • Cache reads are discounted, so processed-token volume overstates effective cost.
  • The $4.92 recorded Grok Build cost is not proof of a marginal cash charge under a plan. The $4.91 API-equivalent applies public grok-4.6 rates to expected grok-4.6-build billing and excludes $0.035 in observed search costs.
  • No quota before/after receipt, independent browser run, local FPS, hardware/viewport receipt, or blind evaluation was supplied.
Visible evidence gaps
  • extractor rerun with a unique marker or global discovery isolation
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console