Measured harness ledgerPublic result
GPT-6 Sol

CARVE — GPT-6 Sol Max

Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.

Max reasoningHeadline result
Workflow cost
$13.16
Wall-clock
Not recorded
Processed tokens
51.52M processed
Record state
complete_token_timing_artifact_and_validation_ledger
Public summary

GPT-6 Sol Max complete_token_timing_artifact_and_validation_ledger ledger: Not recorded wall-clock, 51.52M processed, and $13.16 API-equivalent list price, not a marginal subscription cash charge.

Run identity and stack
  • Result ID: carve-webgpu-snowboarding-gpt-6-sol-max
  • Technical model: GPT-6 Sol
  • Provider: OpenAI
  • Client: OpenAI Codex CLI
  • Stack: OpenAI
  • Stack: OpenAI Codex CLI
  • Stack: Technical model/configuration: GPT-6 Sol
  • Stack: Babylon.js WebGPU
  • Stack: Supplied snowboarder fixture
  • Stack: Harness v1 CARVE prompt
Recorded caveats
  • Wall-clock is end-to-end generation latency and includes tool time and idle gaps.
  • Output tokens include hidden reasoning, visible output, and tool-call JSON.
  • Raw local transcripts and private session identifiers are not published.
  • The unchanged operator parser returned null duration because this Codex transcript stores user prompts as response_item message records. The published 3983.6 seconds derive from the first such user record to the final usage record in the same transcript.
  • The generated asset batch failed the independent clay-geometry judge; this is an overall asset-quality failure despite completed gameplay and performance checks.
  • The supplied rider rendered as a black silhouette during gameplay.
  • Browser GPU timestamp timing was unavailable; the FPS receipt uses frame-time samples.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console