Measured harness ledgerPublic result
GPT-6 AstraSekiro Main Menu — GPT-6 Astra Max
Recreate a high-end cinematic Sekiro-style main menu in the browser from the pinned reference and task contract.
Max reasoningHeadline result
- Workflow cost
- $11.65
- Wall-clock
- 44m 15.4s (qualified snapshot span) wall-clock
- Processed tokens
- 7.23M processed
- Record state
- partial_post_task_metrics_snapshot_with_build_and_live_load
Public summary
GPT-6 Astra Max partial_post_task_metrics_snapshot_with_build_and_live_load ledger: 44m 15.4s (qualified snapshot span) wall-clock, 7.23M processed, and $11.65 Standard API-equivalent estimate from a qualified cumulative root-session snapshot, not a subscription invoice.
Run identity and stack
- Result ID: sekiro-main-menu-browser-gpt-6-astra-max
- Technical model: gpt-6-astra
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: gpt-6-astra
- Stack: Browser recreation
- Stack: Pinned visual reference
- Stack: Harness v1 Sekiro prompt
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/sekiro-main-menu-browser-gpt-6-astra-max/production-build/index.html
- SHA-256: 02744bff4305c32586efe4a29986a100d15cfd6b623a942c514f51d6a58c7946
Recorded caveats
- The supplied token and cost totals are cumulative root-session snapshot values that include any initial metrics-follow-up usage already recorded; they are not a clean isolated generation-turn score.
- The raw parser did not recognize response_item user records, so the displayed wall-clock is a schema-aware recovered span rather than the parser's raw duration.
- Wall-clock includes tools and idle gaps, not model-only compute.
- Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
- Cache reads are discounted, so processed-token volume overstates cost.
- Supplied Playwright input-verification material is retained, but the archive operator only independently rebuilt and live-loaded the project; no fresh direct-input replay was run.
- Blind visual evaluation and independent fidelity scoring remain outstanding.
Visible evidence gaps
- clean isolated generation-turn token and timing receipt
- fresh independent keyboard and mouse replay
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.