Measured harness ledgerPublic result
GPT-5.6 Sol

Campfire under the stars — GPT-5.6 Sol xhigh

Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.

xhigh reasoningHeadline result
Workflow cost
$2.26
Wall-clock
28:52.6 wall-clock
Processed tokens
2.43M processed
Record state
partial_token_timing_and_source_ledger
Public summary

GPT-5.6 Sol xhigh partial_token_timing_and_source_ledger ledger: 28:52.6 wall-clock, 2.43M processed, and $2.26 API-equivalent estimate, not a subscription invoice.

Run identity and stack
  • Result ID: campfire-threejs-gpt-5.6-sol-xhigh
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Three.js interactive scene
  • Stack: Harness v1 campfire prompt
Cost basis
  • Requests with more than 272,000 input tokens use long-context pricing; all 44 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: interactive-threejs-campfire-scene
  • Path: artifacts/campfire-threejs-gpt-5.6-sol-xhigh/source/app/components/CampfireScene.tsx
  • SHA-256: 8f65e9feca6a28510cc0f2a27ec4220aa9c31b442daa3a4677ac38d68879d7ba
Recorded caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • Subagent logs are excluded to avoid double-counting inherited parent context.
Visible evidence gaps
  • browser and hardware environment
  • local FPS
  • final capture
  • blind-evaluation record
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console