Measured harness ledgerPublic result
GPT-6 Astra

Sekiro Main Menu — GPT-6 Astra Max

Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.

Max reasoningHeadline result
Workflow cost
$11.65
Wall-clock
44m 15.4s (qualified snapshot span) wall-clock
Processed tokens
7.23M processed
Record state
partial_post_task_metrics_snapshot_with_build_and_live_load
Public summary

GPT-6 Astra Max partial_post_task_metrics_snapshot_with_build_and_live_load ledger: 44m 15.4s (qualified snapshot span) wall-clock, 7.23M processed, and $11.65 Standard API-equivalent estimate from a qualified cumulative root-session snapshot, not a subscription invoice.

Run identity and stack
  • Result ID: sekiro-main-menu-browser-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Browser recreation
  • Stack: Pinned visual reference
  • Stack: Harness v1 Sekiro prompt
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/sekiro-main-menu-browser-gpt-6-astra-max/production-build/index.html
  • SHA-256: 02744bff4305c32586efe4a29986a100d15cfd6b623a942c514f51d6a58c7946
Recorded caveats
  • The supplied token and cost totals are cumulative root-session snapshot values that include any initial metrics-follow-up usage already recorded; they are not a clean isolated generation-turn score.
  • The raw parser did not recognize response_item user records, so the displayed wall-clock is a schema-aware recovered span rather than the parser's raw duration.
  • Wall-clock includes tools and idle gaps, not model-only compute.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates cost.
  • Supplied Playwright input-verification material is retained, but the archive operator only independently rebuilt and live-loaded the project; no fresh direct-input replay was run.
  • Blind visual evaluation and independent fidelity scoring remain outstanding.
Visible evidence gaps
  • clean isolated generation-turn token and timing receipt
  • fresh independent keyboard and mouse replay
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console