Measured harness ledgerPublic result
GPT-5.6 Sol

High-voxel-density Shrine Village diorama — GPT-5.6 Sol Ultra

Create a complete inspectable Blender voxel diorama of a shrine village at a fixed 0.05-unit voxel resolution through the Blender MCP workflow.

Ultra reasoningHeadline result
Workflow cost
$1.98
Wall-clock
20:36.2 wall-clock
Processed tokens
1.53M processed
Record state
partial_token_timing_and_artifact_ledger
Public summary

GPT-5.6 Sol Ultra partial_token_timing_and_artifact_ledger ledger: 20:36.2 wall-clock, 1.53M processed, and $1.98 API-equivalent estimate, not a subscription invoice.

Run identity and stack
  • Result ID: shrine-village-gpt-5.6-sol-ultra
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Blender MCP
  • Stack: 0.05-unit voxel grid
  • Stack: Inspectable .blend artifact
Cost basis
  • Requests with more than 272,000 input tokens use long-context pricing; all 29 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: blender-scene
  • Path: artifacts/shrine-village-gpt-5.6-sol-ultra/shrine_village.blend
  • SHA-256: b788c87e172e4a3577875c0e495339499c76296bb943284ce871abdfc7132cb2
Recorded caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • Subagent logs are excluded to avoid double-counting inherited parent context.
Visible evidence gaps
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Builder test available

This result is part of Builder tests. Open them for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Tech Review 001 · v1
  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console