Measured harness ledgerPublic result
GPT-6 Astra

Neon Gunner — GPT-6 Astra Max

Build a playable destructible neon run-and-gun game in the browser under the pinned Harness task contract.

Max reasoningHeadline result
Workflow cost
$9.90
Wall-clock
44m 55.2s (supplemental snapshot span) wall-clock
Processed tokens
4.72M processed
Record state
partial_post_task_metrics_snapshot_with_supplied_validation
Public summary

GPT-6 Astra Max partial_post_task_metrics_snapshot_with_supplied_validation ledger: 44m 55.2s (supplemental snapshot span) wall-clock, 4.72M processed, and $9.90 API-equivalent estimate from a qualified post-task snapshot, not a subscription invoice.

Run identity and stack
  • Result ID: neon-gunner-canvas-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Browser canvas game
  • Stack: Harness v1 Neon Gunner prompt
  • Stack: Requested tool profile: vanilla-canvas
Validation evidence
  • Validator SHA-256: fa5ec13e9ef2f5f96949283077b7c6278bf9c9f14032817c21743ab973d06090
Recorded caveats
  • The fixed metrics snapshot includes the beginning of the later metrics follow-up and an idle gap after game delivery, so its timing, token total, and cost are not a clean completed-run score.
  • The supplied parser reported zero user turns and null timing because this transcript stored user text in a different schema; the displayed timing is a supplemental schema-aware derivation.
  • Cache-creation tokens are unavailable in the transcript schema; zero is not proof of no billable cache writes.
  • The supplied 26/26 validator and automated browser traversal are model-supplied evidence, not independent operator replay.
  • Archive verification independently checked JavaScript syntax, frozen-fixture identity, and title-screen boot only; no independent keyboard/mouse gameplay replay is claimed.
  • Blind evaluation and a standardized local-FPS run remain outstanding.
Visible evidence gaps
  • clean completed-run timing and token receipt
  • independent real-input gameplay replay
  • standardized local FPS measurement
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console