Measured harness ledgerPublic result
GPT-6 Astra

High-voxel-density Jungle Temple diorama — GPT-6 Astra Max

Create a complete inspectable Blender voxel diorama of a jungle temple at a fixed 0.05-unit voxel resolution through the Blender MCP workflow.

Max reasoningHeadline result
Workflow cost
$16.35
Wall-clock
Not recorded
Processed tokens
10.06M processed
Record state
partial_post_task_metrics_timing_unavailable_with_blender_scene_open_inspection
Public summary

GPT-6 Astra Max partial_post_task_metrics_timing_unavailable_with_blender_scene_open_inspection ledger: Not recorded wall-clock, 10.06M processed, and $16.35 Standard API-equivalent estimate from the supplied Codex receipt, not an actual subscription charge..

Run identity and stack
  • Result ID: jungle-temple-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Blender MCP
  • Stack: 0.05-unit voxel grid
  • Stack: Inspectable .blend artifact
Primary artifact integrity
  • Kind: Blender scene
  • Path: artifacts/jungle-temple-gpt-6-astra-max/Jungle_Temple.blend
  • SHA-256: 54ab44f70f3c5a896915228d16a60e0ad007cfa11015b421a12fafc18cd6bc83
Recorded caveats
  • The unchanged parser found no qualifying user timestamp, so timing and output throughput are unavailable and not inferred.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed-token volume overstates effective cost.
  • The Blender scene was independently opened without rendering or modification; that does not establish visual quality.
  • Supplied viewport reviews and geometry audits are retained as supplied evidence. Remote final-render acceptance and blind evaluation remain unverified.
Visible evidence gaps
  • isolated generation-turn timing receipt
  • remote final-render acceptance frames with renderer, resolution, and hardware receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console