Measured harness ledgerPublic result
Qwen3.8 Max PreviewHigh-voxel-density Jungle Temple diorama — Qwen3.8 Max Preview Unreported
Create a complete inspectable Blender voxel diorama of a jungle temple at a fixed 0.05-unit voxel resolution through the Blender MCP workflow.
Unreported reasoningHeadline result
- Workflow cost
- $6.77
- Wall-clock
- 48m 49.9s wall-clock
- Processed tokens
- 4.35M processed
- Record state
- partial_token_timing_artifact_validation_ledger
Public summary
Qwen3.8 Max Preview Unreported partial_token_timing_artifact_validation_ledger ledger: 48m 49.9s wall-clock, 4.35M processed, and $6.77 Third-party NanoGPT API-list-price equivalent for the measured token mix; not an Alibaba Token Plan cash charge.
Run identity and stack
- Result ID: jungle-temple-qwen3.8-max-preview
- Technical model: qwen3.8-max-preview
- Provider: Alibaba Cloud Model Studio
- Client: Qwen Code 0.20.0
- Stack: Alibaba Cloud Model Studio
- Stack: Qwen Code 0.20.0
- Stack: Technical model/configuration: qwen3.8-max-preview
- Stack: Blender MCP
- Stack: 0.05-unit voxel grid
- Stack: Inspectable .blend artifact
Cost basis
- NanoGPT published Qwen3.8 Max Preview rates of $1.50/M input and $5.00/M output on 2026-07-20. It did not publish a separate cache-read discount, so all 4,277,044 measured input tokens are priced at the same input rate. The actual run used Alibaba Token Plan Personal Lite; $6.77 is a third-party API-list-price equivalent, not a first-party Alibaba price, an itemized provider charge, or provider cost.
- Qualified API-list-price equivalent.
Primary artifact integrity
- Kind: blender-scene
- Path: artifacts/jungle-temple-qwen3.8-max-preview/jungle_temple.blend
- SHA-256: 6e568833a95517ad9d8b7a3f0fca115a6207566b1542dfb20228c9ea0c3c17a6
Recorded caveats
- Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
- Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 45,797 thought tokens are a subset of the 70,825 output tokens.
- The post-generation 3.931-second render is excluded from model wall-clock.
- No separate cache-read discount was available in the cited third-party rate card, so cached input is priced at its ordinary input rate for the API-equivalent calculation.
- The $6.77 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
- The final capture has a pale, low-contrast palette and frames the temple lower-right rather than the prompt's upper-right requirement; no quality score is assigned.
- The task crossed a five-hour quota reset boundary, so five-hour percentages cannot be subtracted. The weekly 12% to 16% change supports a nominal 100 Credits with a whole-percent display range of 75–125, not an invoice-grade exact charge.
- No blind evaluation was supplied.
Visible evidence gaps
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
