Measured harness ledgerPublic result
GPT-6 Astra

MacBook-class cinematic ad scene — GPT-6 Astra Max

Create one polished MacBook-class product-ad shot in Blender with a modeled device, legible industrial detail, intentional materials, lighting, camera movement, and a validator-ready scene.

Max reasoningHeadline result
Workflow cost
$20.35
Wall-clock
65m 04s (qualified snapshot span) wall-clock
Processed tokens
11.79M processed
Record state
partial_post_task_metrics_snapshot_with_independent_validation
Public summary

GPT-6 Astra Max partial_post_task_metrics_snapshot_with_independent_validation ledger: 65m 04s (qualified snapshot span) wall-clock, 11.79M processed, and $20.35 API-equivalent estimate from a qualified post-task snapshot, not a subscription invoice.

Run identity and stack
  • Result ID: macbook-cinematic-blender-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Blender MCP
  • Stack: Cinematic product scene
  • Stack: Harness v1 MacBook validator
  • Stack: Requested tool profile: blender-mcp
Primary artifact integrity
  • Kind: blender-cinematic-product-ad-scene
  • Path: artifacts/macbook-cinematic-blender-gpt-6-astra-max/macbook_cinematic_final.blend
  • SHA-256: 9f8bba5b2c1e9d4c5b08d57540da4fc98c886cd59df7ac4c90716e31103f78d6
Validation evidence
  • Result: 49/49 checks recorded
  • Path: artifacts/macbook-cinematic-blender-gpt-6-astra-max/validation_log.json
  • SHA-256: e6ed20d8a1d5d15955a4c3f8c2d6ed409c09f1e28bc43a536a86d0ad4cd78ec6
  • Validator SHA-256: 54591de8d4ef2953c6f7e2fe3e4f4b3d9407b953049cc858b1012afaf9424d51
Recorded caveats
  • The supplied metrics snapshot includes the beginning of a later metrics follow-up, so its timing, token total, and cost are not a clean completed-run score.
  • The raw parser did not recognize user-message payloads; the displayed 65m 04s span is a schema-aware derivation and user-turn count remains unavailable.
  • Wall-clock includes tools and idle gaps, not just model compute.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates cost.
  • Cache-creation tokens are unavailable in the transcript schema; zero is not proof of no billable cache writes.
  • The canonical validator and headless reopen confirm configured scene properties, not visual quality, final-animation rendering, render performance, or retained Blender MCP tool use.
  • Blind evaluation remains outstanding.
Visible evidence gaps
  • clean completed-run token/timing receipt
  • benchmark user-turn count
  • Blender MCP tool-use transcript or disclosure
  • render-hardware performance measurement
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console