Measured harness ledgerPublic result
Qwen3.8 MaxCampfire under the stars — Qwen3.8 Max xhigh
Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.
xhigh reasoningHeadline result
- Workflow cost
- ¥20.79 First-party API-list-price-equivalent upper bound, not a token-plan cash charge (upper bound)
- Wall-clock
- 15m 33.8s wall-clock
- Processed tokens
- 1.62M processed
- Record state
- artifact_model_runtime_metrics_ledger_missing_fps_and_blind_evaluation
Public summary
Qwen3.8 Max xhigh artifact_model_runtime_metrics_ledger_missing_fps_and_blind_evaluation ledger: 15m 33.8s wall-clock, 1.62M processed, and ¥20.79 First-party API-list-price-equivalent upper bound, not a token-plan cash charge (upper bound).
Run identity and stack
- Result ID: campfire-threejs-qwen3.8-max-xhigh
- Technical model: Qwen 3.8 Max
- Provider: Alibaba Cloud Model Studio token plan via Qwen Code
- Client: Qwen Code 0.21.6
- Stack: Alibaba Cloud Model Studio token plan via Qwen Code
- Stack: Qwen Code 0.21.6
- Stack: Technical model/configuration: Qwen 3.8 Max
- Stack: Three.js interactive scene
- Stack: Harness v1 campfire prompt
Cost basis
- Qwen Code records no itemized per-run cash cost under this token plan.
- No model-specific cached-input rate was published for qwen3.8-max; the headline prices cache reads at the fresh-input rate.
- The official page's general 10% cache-hit rule would yield ¥5.55. It is reference-only and not the displayed headline.
Primary artifact integrity
- Kind: interactive-threejs-campfire-scene
- Path: artifacts/campfire-threejs-qwen3.8-max-xhigh/source/index.html
- SHA-256: 1f4032c100aa8710e695d056dd084db9f93cc3258adef831d16f026b007cf398
Validation evidence
- Result: PASS
- Path: artifacts/campfire-threejs-qwen3.8-max-xhigh/verification/runtime.public.json
- SHA-256: 4b73f3dcd6a56ececb80fa9745f9aa51f266a34ec4a8fd8f05438591eee95a62
Recorded caveats
- Wall-clock is end-to-end latency, including tool execution and model-run visual inspection, rather than model-only compute.
- The single benchmark user turn generated 26 API requests, including six same-model visual-inspector agents.
- Cache reads are discounted under the general cache rule, so processed-token volume overstates cost. ¥20.79 is explicitly an upper bound; ¥5.55 is a secondary reference calculation under the general 10% cache-hit rule.
- No cache-write counter, quota delta, standardized local FPS measurement, RTX Pro 6000 render, or blind-evaluation record was supplied.
- Visual quality is deliberately not scored here; it belongs to blind evaluation.
Visible evidence gaps
- standardized local FPS measurement
- RTX Pro 6000 1080p60 final render or capture metadata
- blind-evaluation record
- independent non-model-run runtime reproduction
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
