Measured harness ledgerPublic result
Grok 4.6

The Last of Us Main Menu — Grok 4.6 xhigh

Recreate the pinned The Last of Us-style main menu reference as an interactive browser experience.

xhigh reasoningHeadline result
Workflow cost
Not recorded
Wall-clock
43m 54s wall-clock
Processed tokens
12.51M processed
Record state
verified_exact_metrics_runtime_replay_qualified_no_blind_evaluation
Public summary

Grok 4.6 xhigh verified_exact_metrics_runtime_replay_qualified_no_blind_evaluation ledger: 43m 54s wall-clock, 12.51M processed, and Not recorded.

Run identity and stack
  • Result ID: last-of-us-main-menu-browser-grok-4.6-xhigh
  • Technical model: grok-4.6
  • Provider: xAI via grok-sub
  • Stack: xAI via grok-sub
  • Stack: Technical model/configuration: grok-4.6
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 menu prompt
Cost basis
  • No measured or estimated run cost is recorded.
Primary artifact integrity
  • Kind: interactive-browser-title-menu
  • Path: artifacts/last-of-us-main-menu-browser-grok-4.6-xhigh/source/index.html
  • SHA-256: ad651f42b48ad92795c3a420f64869b895d7541ee78f9970877c77037c0e5c9e
Validation evidence
  • Result: PASS_WITH_QUALIFICATION
Recorded caveats
  • Wall-clock time includes model reasoning, tool execution, asset generation, installation, rendering, visual iteration, and verification; it is not model-only compute time.
  • Reasoning tokens are reported separately but are already included in provider-accounted output tokens, so they are not added again to the total.
  • The $1.36086734 value is the exact recorded grok-sub total for this run; no alternate catalog-price inference is substituted.
  • Independent replay measured 20.88 FPS, below the 30 FPS target, and observed two non-functional favicon 404 requests.
  • Model-supplied captures used an offscreen WebGL1 fallback; live Chrome WebGL2 replay is archived separately.
  • Independent replay verifies runtime behavior on one machine/browser, not visual-fidelity scoring or cross-device performance.
Visible evidence gaps
  • blind visual-fidelity evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console