Measured harness ledgerPublic result
Grok 4.6

ECHO Main Menu — Grok 4.6 xhigh

Recreate the pinned ECHO-style cinematic main menu reference as an interactive browser experience.

xhigh reasoningHeadline result
Workflow cost
$2.71
Wall-clock
1h 09m 03s wall-clock
Processed tokens
20.04M processed
Record state
verified_exact_metrics_runtime_replay_qualified_protocol_deviation_no_blind_evaluation
Public summary

Grok 4.6 xhigh verified_exact_metrics_runtime_replay_qualified_protocol_deviation_no_blind_evaluation ledger: 1h 09m 03s wall-clock, 20.04M processed, and $2.71 Exact final-segment CLI receipt plus deterministic reconstruction from the native per-inference ledger for the interrupted segments.

Run identity and stack
  • Result ID: echo-main-menu-browser-grok-4.6-xhigh
  • Technical model: grok-4.6
  • Provider: xAI via grok-sub
  • Stack: xAI via grok-sub
  • Stack: Technical model/configuration: grok-4.6
  • Stack: Browser recreation
  • Stack: Pinned visual references
  • Stack: Harness v1 ECHO prompt
Primary artifact integrity
  • Kind: interactive-browser-title-menu
  • Path: artifacts/echo-main-menu-browser-grok-4.6-xhigh/source/index.html
  • SHA-256: b2a46457eabf789114150415c6b37ade55f622ba8a2cf1646d47783f1b242052
Validation evidence
  • Result: FUNCTIONAL_PASS_WITH_QUALIFICATION
Recorded caveats
  • The run is not one runner process: a local one-hour watchdog interrupted active work, and the preserved Grok session was continued until it stopped itself.
  • Wall-clock time includes model reasoning, tool execution, asset generation, installation, rendering, visual iteration, and verification; it is not model-only compute time.
  • Reasoning tokens are reported separately but are already included in provider-accounted output tokens, so they are not added again to the total.
  • The 2.712231 USD total combines one exactly recorded final segment with ledger reconstruction for the interrupted segments using the CLI's per-call rates and long-context rule.
  • Independent replay verifies runtime behavior on one browser and machine, not visual-fidelity scoring or cross-device performance.
  • A trustworthy numeric FPS value was not captured; continuous motion was verified through changing live frames instead.
  • The initial runner imposed a 3600-second local watchdog and exited while Grok was still working. The preserved conversation and workspace were continued. A short intermediate continuation was interrupted only to remove the practical cutoff; the final continuation ran until Grok returned end_turn.
Visible evidence gaps
  • blind visual-fidelity evaluation record
  • trustworthy local FPS measurement
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console