Measured harness ledgerPublic result
GPT-5.6 Sol

CARVE — GPT-5.6 Sol Max

Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.

Max reasoningHeadline result
Workflow cost
$27.40
Wall-clock
2h 19m 55s wall-clock
Processed tokens
37.02M processed
Record state
partial_two_user_turn_token_timing_source_build_ledger
Public summary

GPT-5.6 Sol Max partial_two_user_turn_token_timing_source_build_ledger ledger: 2h 19m 55s wall-clock, 37.02M processed, and $27.40 Token-only standard API-equivalent estimate, not the actual Codex subscription charge.

Run identity and stack
  • Result ID: carve-webgpu-snowboarding-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Babylon.js WebGPU
  • Stack: Supplied snowboarder fixture
  • Stack: Harness v1 CARVE prompt
Cost basis
  • Prompts with more than 272,000 input tokens use the published long-context rate for the whole request. The receipt reports zero calls over that threshold.
  • No separately priced tool or non-token cost is included because no verifiable usage counts were supplied.
Primary artifact integrity
  • Kind: interactive-webgpu-snowboarding-game
  • Path: artifacts/carve-webgpu-snowboarding-gpt-5.6-sol-max/source/src/main.js
  • SHA-256: 27d8a14a063715f864d94df57cc9dc90d5e1203f822be4dc40e1b04859488d1d
Validation evidence
  • Path: artifacts/carve-webgpu-snowboarding-gpt-5.6-sol-max/validation/build-and-source-check.public.json
  • SHA-256: ec393a1784baa8f0e2fe6db7f22d6875494da3122e8e25453383e158f0e4d409
Recorded caveats
  • The root-task receipt records two user turns where the frozen benchmark protocol expects one; it is disclosed rather than treated as a clean one-shot result.
  • Wall-clock is end-to-end workflow latency including tools and idle gaps, not model-only compute time.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens substantially overstate cost.
  • Cache creation is recorded as zero because the transcript schema did not expose a cache-write field; no estimate was inferred.
  • The cost is an API-list-price equivalent, not the actual subscription-backed Codex charge, and separately priced tools are excluded.
  • The source notices identify the player GLB as user-provided and do not assert redistribution rights. The source archives the asset because it is a required supplied input, without relicensing it.
  • The production build passed, but no independent browser acceptance run, local FPS measurement, hardware/viewport receipt, or blind evaluation was supplied.
  • The supplied root-task receipt reports two user turns. This result is not presented as a clean one-shot run.
Visible evidence gaps
  • clean one-user-turn rerun
  • independent WebGPU browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • independent final-capture receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console