Measured harness ledgerPublic result
GPT-6 Astra

CARVE — GPT-6 Astra Max

Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.

Max reasoningHeadline result
Workflow cost
$39.12
Wall-clock
1h 52m 02.7s wall-clock
Processed tokens
26.89M processed
Record state
partial_model_supplied_runtime_evidence_with_independent_clean_build
Public summary

GPT-6 Astra Max partial_model_supplied_runtime_evidence_with_independent_clean_build ledger: 1h 52m 02.7s wall-clock, 26.89M processed, and $39.12 Standard API-equivalent estimate from the supplied Codex receipt, not an actual subscription charge..

Run identity and stack
  • Result ID: carve-webgpu-snowboarding-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Babylon.js WebGPU
  • Stack: Supplied snowboarder fixture
  • Stack: Harness v1 CARVE prompt
Primary artifact integrity
  • Kind: interactive-webgpu-snowboarding-game
  • Path: artifacts/carve-webgpu-snowboarding-gpt-6-astra-max/source/src/main.js
  • SHA-256: 37507556ab69cb32728c3e2d8cbb88d40134eebfac73905eb679a7573ff83d79
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool execution and waits.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates effective cost.
  • The archive operator independently reproduced a clean build but did not independently play the WebGPU game; gameplay, asset-review, and performance reports are retained as model-supplied evidence.
  • The source performance report records a native-1080p 60 FPS target miss on Apple M1 Pro. Its reduced-resolution figures are not standardized desktop-GPU results.
  • No portable Blender builder source for the submitted generated environment assets was supplied; the runtime GLBs, asset manifests, and close-up evidence are archived instead.
  • The shared snowboarder fixture is an immutable supplied input; its prior generation is excluded from this contestant's cost and output.
  • No blind evaluation is recorded.
Visible evidence gaps
  • independent WebGPU gameplay replay through real input
  • standardized fixed-hardware FPS measurement
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console