Measured harness ledgerPublic result
Grok 4.6

Sekiro Main Menu — Grok 4.6 xhigh

Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.

xhigh reasoningHeadline result
Workflow cost
$1.15
Wall-clock
51m 59.543s wall-clock
Processed tokens
9.42M processed
Record state
partial_two_turn_metrics_source_build_and_static_serve_ledger
Public summary

Grok 4.6 xhigh partial_two_turn_metrics_source_build_and_static_serve_ledger ledger: 51m 59.543s wall-clock, 9.42M processed, and $1.15 Provider-recorded usage estimate, not a subscription cash-charge invoice.

Run identity and stack
  • Result ID: sekiro-main-menu-browser-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok Build
  • Client: Grok 1.0.5
  • Stack: xAI Grok Build
  • Stack: Grok 1.0.5
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 Sekiro prompt
  • Stack: Requested tool profile: Grok Build CLI
Cost basis
  • The exact billed alias grok-4.6-build has no official public catalog price, so no API-equivalent cost is inferred.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/sekiro-main-menu-browser-grok-4.6-xhigh/production-build/index.html
  • SHA-256: 61d77b2e5a8836a79f9a530bd0bb4dc7b131a03a4f4dc1f4d48a4910e29b4a1c
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tools, installs, browser checks, and idle intervals.
  • The benchmark protocol includes two genuine ACP user turns, including a host task-completed notification. It is a disclosed protocol deviation, not a one-shot result.
  • Generated output includes visible output and reasoning output. Cache reads are discounted, so total processed tokens overstate effective cost.
  • The $1.15 figure is the source ledger's provider-recorded usage estimate, not an inferred API-equivalent rate or a subscription cash charge.
  • The fresh build and static HTTP checks validate basic runnability only; they are not browser interaction, keyboard navigation, audio, local-FPS, or final-render measurements.
  • No RTX Pro 6000 final render, local FPS/environment record, supplied final capture, or blind evaluation is archived.
Visible evidence gaps
  • local FPS measurement with browser, hardware, and viewport metadata
  • keyboard navigation and menu interaction record
  • RTX Pro 6000 final render or capture metadata
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console