Measured harness ledgerPublic result
Qwen3.8 Max

Campfire under the stars — Qwen3.8 Max xhigh

Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.

xhigh reasoningHeadline result
Workflow cost
¥20.79 First-party API-list-price-equivalent upper bound, not a token-plan cash charge (upper bound)
Wall-clock
15m 33.8s wall-clock
Processed tokens
1.62M processed
Record state
artifact_model_runtime_metrics_ledger_missing_fps_and_blind_evaluation
Public summary

Qwen3.8 Max xhigh artifact_model_runtime_metrics_ledger_missing_fps_and_blind_evaluation ledger: 15m 33.8s wall-clock, 1.62M processed, and ¥20.79 First-party API-list-price-equivalent upper bound, not a token-plan cash charge (upper bound).

Run identity and stack
  • Result ID: campfire-threejs-qwen3.8-max-xhigh
  • Technical model: Qwen 3.8 Max
  • Provider: Alibaba Cloud Model Studio token plan via Qwen Code
  • Client: Qwen Code 0.21.6
  • Stack: Alibaba Cloud Model Studio token plan via Qwen Code
  • Stack: Qwen Code 0.21.6
  • Stack: Technical model/configuration: Qwen 3.8 Max
  • Stack: Three.js interactive scene
  • Stack: Harness v1 campfire prompt
Cost basis
  • Qwen Code records no itemized per-run cash cost under this token plan.
  • No model-specific cached-input rate was published for qwen3.8-max; the headline prices cache reads at the fresh-input rate.
  • The official page's general 10% cache-hit rule would yield ¥5.55. It is reference-only and not the displayed headline.
Primary artifact integrity
  • Kind: interactive-threejs-campfire-scene
  • Path: artifacts/campfire-threejs-qwen3.8-max-xhigh/source/index.html
  • SHA-256: 1f4032c100aa8710e695d056dd084db9f93cc3258adef831d16f026b007cf398
Validation evidence
  • Result: PASS
  • Path: artifacts/campfire-threejs-qwen3.8-max-xhigh/verification/runtime.public.json
  • SHA-256: 4b73f3dcd6a56ececb80fa9745f9aa51f266a34ec4a8fd8f05438591eee95a62
Recorded caveats
  • Wall-clock is end-to-end latency, including tool execution and model-run visual inspection, rather than model-only compute.
  • The single benchmark user turn generated 26 API requests, including six same-model visual-inspector agents.
  • Cache reads are discounted under the general cache rule, so processed-token volume overstates cost. ¥20.79 is explicitly an upper bound; ¥5.55 is a secondary reference calculation under the general 10% cache-hit rule.
  • No cache-write counter, quota delta, standardized local FPS measurement, RTX Pro 6000 render, or blind-evaluation record was supplied.
  • Visual quality is deliberately not scored here; it belongs to blind evaluation.
Visible evidence gaps
  • standardized local FPS measurement
  • RTX Pro 6000 1080p60 final render or capture metadata
  • blind-evaluation record
  • independent non-model-run runtime reproduction
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console