Measured harness ledgerPublic result
Qwen3.8 Max

MacBook-class cinematic ad scene — Qwen3.8 Max xhigh

Create one polished MacBook-class product-ad shot in Blender with a modeled device, legible industrial detail, intentional materials, lighting, camera movement, and a validator-ready scene.

xhigh reasoningHeadline result
Workflow cost
$24.99
Wall-clock
2h 51m 49.9s wall-clock
Processed tokens
82.89M processed
Record state
artifact_validator_pass_metrics_partial_quota_interrupted
Public summary

Qwen3.8 Max xhigh artifact_validator_pass_metrics_partial_quota_interrupted ledger: 2h 51m 49.9s wall-clock, 82.89M processed, and $24.99 API-equivalent accounting only, not an itemized Token Plan cash charge.

Run identity and stack
  • Result ID: macbook-cinematic-blender-qwen3.8-max-xhigh
  • Technical model: Qwen 3.8 Max
  • Provider: Alibaba Cloud Model Studio Token Plan via Qwen Code
  • Client: Qwen Code 0.21.7 installed / 0.21.6 transcript-recorded
  • Stack: Alibaba Cloud Model Studio Token Plan via Qwen Code
  • Stack: Qwen Code 0.21.7 installed / 0.21.6 transcript-recorded
  • Stack: Technical model/configuration: Qwen 3.8 Max
  • Stack: Blender MCP
  • Stack: Cinematic product scene
  • Stack: Harness v1 MacBook validator
Cost basis
  • Qwen Code records no itemized per-run cash cost under this Token Plan.
Primary artifact integrity
  • Kind: blender-cinematic-scene
  • Path: artifacts/macbook-cinematic-blender-qwen3.8-max-xhigh/macbook_cinematic_final.blend
  • SHA-256: f1d7a5e482656e3d7440638326bb74aee83457b8a358c5921cb1ea1bf809af43
Validation evidence
  • Result: 49/49 checks recorded
  • Path: artifacts/macbook-cinematic-blender-qwen3.8-max-xhigh/validation_log.json
  • SHA-256: c6636a210a52408248a0157f71833ce80adbef7760dd98a4c56b2531ac0222c3
Recorded caveats
  • The final blend and 49/49 supplied validator log are present, but the run cannot be represented as a fully canonical metrics window because quota exhaustion forced a synthetic EOF measurement boundary.
  • The objective validator was not independently rerun in this archival pass.
  • Wall-clock is end-to-end latency, not model-only compute; it includes tools, installs, render work, browser checks, and idle gaps.
  • One benchmark user turn generated 237 API requests, including five same-model visual-inspector agents.
  • No cache-write counter, quota delta, blind-evaluation record, or standardized runtime-performance measurement is available.
  • Visual quality is deliberately not scored in this ledger.
Visible evidence gaps
  • natural post-run metrics boundary
  • independent validator rerun with Blender version and render hardware receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console