Measured harness ledgerPublic result
GPT-6 AstraNeon Gunner — GPT-6 Astra Max
Build a playable destructible neon run-and-gun game in the browser under the pinned Harness task contract.
Max reasoningHeadline result
- Workflow cost
- $9.90
- Wall-clock
- 44m 55.2s (supplemental snapshot span) wall-clock
- Processed tokens
- 4.72M processed
- Record state
- partial_post_task_metrics_snapshot_with_supplied_validation
Public summary
GPT-6 Astra Max partial_post_task_metrics_snapshot_with_supplied_validation ledger: 44m 55.2s (supplemental snapshot span) wall-clock, 4.72M processed, and $9.90 API-equivalent estimate from a qualified post-task snapshot, not a subscription invoice.
Run identity and stack
- Result ID: neon-gunner-canvas-gpt-6-astra-max
- Technical model: gpt-6-astra
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: gpt-6-astra
- Stack: Browser canvas game
- Stack: Harness v1 Neon Gunner prompt
- Stack: Requested tool profile: vanilla-canvas
Validation evidence
- Validator SHA-256: fa5ec13e9ef2f5f96949283077b7c6278bf9c9f14032817c21743ab973d06090
Recorded caveats
- The fixed metrics snapshot includes the beginning of the later metrics follow-up and an idle gap after game delivery, so its timing, token total, and cost are not a clean completed-run score.
- The supplied parser reported zero user turns and null timing because this transcript stored user text in a different schema; the displayed timing is a supplemental schema-aware derivation.
- Cache-creation tokens are unavailable in the transcript schema; zero is not proof of no billable cache writes.
- The supplied 26/26 validator and automated browser traversal are model-supplied evidence, not independent operator replay.
- Archive verification independently checked JavaScript syntax, frozen-fixture identity, and title-screen boot only; no independent keyboard/mouse gameplay replay is claimed.
- Blind evaluation and a standardized local-FPS run remain outstanding.
Visible evidence gaps
- clean completed-run timing and token receipt
- independent real-input gameplay replay
- standardized local FPS measurement
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.