Measured harness ledgerPublic result
GPT-6 Astra

Courtyard Blossom Stage 2 — GPT-6 Astra Max

Create a second cherry-blossom temple courtyard arena for the fixed fighting game under the pinned reference-guided Stage 2 contract.

Max reasoningHeadline result
Workflow cost
$74.50
Wall-clock
Not recorded
Processed tokens
53.63M processed
Record state
partial_root_snapshot_metrics_with_independent_in_app_browser_replay_and_unverified_model_geometry_evidence
Public summary

GPT-6 Astra Max partial_root_snapshot_metrics_with_independent_in_app_browser_replay_and_unverified_model_geometry_evidence ledger: Not recorded wall-clock, 53.63M processed, and $74.50 OpenAI Standard API-list-price equivalent, not an itemized Codex subscription charge.

Run identity and stack
  • Result ID: courtyard-arena-blossom-stage-2-reference-exposed-gpt-6-astra-max
  • Technical model: GPT-6 Astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: GPT-6 Astra
  • Stack: Blender / browser arena workflow
  • Stack: Reference-guided Stage 2 fixture
Cost basis
  • Separately priced tools and non-token services are excluded.
Recorded caveats
  • The raw parser has no timing result because its user-message-schema expectation did not match this transcript. The displayed 6h 28m 15.8s figure is a qualified supplemental root-session span, not a clean generation-only duration.
  • The cumulative root-task snapshot includes tool time and idle gaps, and does not establish a clean public one-user-turn boundary.
  • Output tokens include reasoning, visible prose and code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates effective cost.
  • This is an API-equivalent estimate; separately priced tools and non-token services are excluded.
  • Archive-operator replay establishes functional routes and input response only. It does not establish standardized FPS, transition hitch, cross-hardware performance, blind preference, or independent textures-off geometry quality.
  • No blind evaluation is recorded.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console