Measured harness ledgerPublic result
Qwen3.8 Max Preview

MacBook-class cinematic ad scene — Qwen3.8 Max Preview Mandatory thinking enabled · Attempt 3

Create one polished MacBook-class product-ad shot in Blender with a modeled device, legible industrial detail, intentional materials, lighting, camera movement, and a validator-ready scene.

Mandatory thinking enabled reasoningHeadline result
Workflow cost
$29.27
Wall-clock
7h 00m 56.4s wall-clock
Processed tokens
18.43M processed
Record state
partial_objective_pass_agent_nontermination
Public summary

Qwen3.8 Max Preview Mandatory thinking enabled partial_objective_pass_agent_nontermination ledger: 7h 00m 56.4s wall-clock, 18.43M processed, and $29.27 Qualified third-party NanoGPT API-list-price scenario; not an Alibaba Token Plan cash charge.

Run identity and stack
  • Result ID: macbook-cinematic-blender-qwen3.8-max-preview-attempt-3-partial
  • Technical model: qwen3.8-max-preview
  • Provider: Alibaba ModelStudio Token Plan International
  • Client: Qwen Code 0.20.1
  • Attempt: 3
  • Stack: Alibaba ModelStudio Token Plan International
  • Stack: Qwen Code 0.20.1
  • Stack: Technical model/configuration: qwen3.8-max-preview
  • Stack: Attempt 3
  • Stack: Blender MCP
  • Stack: Cinematic product scene
  • Stack: Harness v1 MacBook validator
Cost basis
  • Qualified API-list-price equivalent.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/macbook-cinematic-blender-qwen3.8-max-preview-attempt-3-validator-pass-model-noncompletion/artifact/macbook_cinematic_final.blend
  • SHA-256: 762c826ba6bee691c9bbbea6724122bd4d8f38dea0e28e7009e659d45e9b21e0
Validation evidence
  • Result: 49/49 checks recorded
  • Path: artifacts/macbook-cinematic-blender-qwen3.8-max-preview-attempt-3-validator-pass-model-noncompletion/artifact/validation_log.json
  • SHA-256: 5ee3736d1e6e8b67e1588f90a167aa33bb7b44ddf3a55f5405383b85e4bce550
Recorded caveats
  • This is a dual result: the core construction objective passed, but the agent workflow did not terminate normally. It is visible as PARTIAL, not as a completed or blind-vote-ready benchmark result.
  • The 38m 10.5s number is time to the independent 49/49 objective gate; the full agent attempt lasted 7h 00m 56.4s including the quota pause.
  • 61.42% of the API-equivalent cost was incurred after the objective gate had already passed.
  • Output includes model-accounted thoughts; the thoughts figure is a subset of output and is not additive.
  • The API-equivalent estimate is not Alibaba pricing, provider cost, or an itemized subscription charge.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console