Measured harness ledgerPublic result
Qwen3.8 Max

Explorable space-flight game — Qwen3.8 Max xhigh

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

xhigh reasoningHeadline result
Workflow cost
¥49.71 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound)
Wall-clock
30m 19.1s wall-clock
Processed tokens
3.97M processed
Record state
artifact_runtime_metrics_ledger_no_blind_evaluation
Public summary

Qwen3.8 Max xhigh artifact_runtime_metrics_ledger_no_blind_evaluation ledger: 30m 19.1s wall-clock, 3.97M processed, and ¥49.71 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound).

Run identity and stack
  • Result ID: space-flight-game-qwen3.8-max-xhigh
  • Technical model: qwen3.8-max
  • Provider: Alibaba Cloud Model Studio token plan
  • Client: Qwen Code 0.21.6 (installed) / 0.21.6 (transcript-recorded)
  • Stack: Alibaba Cloud Model Studio token plan
  • Stack: Qwen Code 0.21.6 (installed) / 0.21.6 (transcript-recorded)
  • Stack: Technical model/configuration: qwen3.8-max
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
Cost basis
  • Qwen Code records no per-request cost field; qwen_recorded_usd is null by construction.
  • No model-specific cached-input rate is published for qwen3.8-max, so cache reads are priced at the fresh-input rate.
  • Reference only: applying the pricing page's general cache-hit rule (10% of the input rate) gives ¥9.94. It is not the reported figure.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/space-flight-game-qwen3.8-max-xhigh/source/index.html
  • SHA-256: fee68057865e13558823242d34c3849c790d4bdff908722a46319f65954e6ca9
Validation evidence
  • Result: PASS
Recorded caveats
  • Wall-clock (1,819.108 s) is end-to-end workflow latency including tool execution, browser captures, and idle gaps — not model-only compute. Summed API duration (2,035.016 s) is a separate figure and exceeds wall-clock because subagents ran concurrently.
  • One Qwen Code agent turn fans out to many billable requests and subagent sessions under a single session id (31 root + 12 subagent here).
  • Visible output and reasoning are separate fields; both are output-priced on this stack.
  • inputTokens includes cached tokens and outputTokens includes reasoning tokens, so the ledger's totalTokens and total_processed are different quantities by definition and must never be swapped; they coincide numerically here only by algebra.
  • Cache reads are normally discounted; pricing them at the fresh-input rate makes ¥49.71 an upper bound.
  • Qwen Code records no cache-write counter, so that line item is unavailable rather than zero.
  • Context occupancy is a single-request footprint, not cumulative usage; quota fields are null with no explicit before/after evidence.
  • Ledger layouts are version-specific; the extractor is validated against Qwen Code 0.21.x and fails closed on an incompatible schema.
  • The 120.1 FPS sample is Apple M1 Pro / ANGLE Metal evidence at 1920x993, not a standardized cross-machine comparison.
  • The cost figure is in CNY from the official Alibaba rate card and is not directly comparable to the third-party USD estimates used for other Qwen rows in this archive.
  • The sanitized session export named in receipt.sha256 is withheld from the public archive by its own privacy label.
Visible evidence gaps
  • blind-evaluation record
  • independent (non-model-run) runtime and FPS receipt
  • standardized cross-machine performance measurement
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console