Measured harness ledgerPublic result
Qwen3.8 Max

Space Flight Stage 2 — Landfall — Qwen3.8 Max xhigh

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

xhigh reasoningHeadline result
Workflow cost
¥55.93 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash.
Wall-clock
1h 50m 56.0s wall-clock
Processed tokens
33.80M processed
Record state
artifact_runtime_ledger_partial_metrics_no_blind_evaluation
Public summary

Qwen3.8 Max xhigh artifact_runtime_ledger_partial_metrics_no_blind_evaluation ledger: 1h 50m 56.0s wall-clock, 33.80M processed, and ¥55.93 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash..

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-qwen3.8-max-xhigh
  • Technical model: qwen3.8-max
  • Provider: Alibaba Cloud Model Studio token plan
  • Client: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
  • Stack: Alibaba Cloud Model Studio token plan
  • Stack: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
  • Stack: Technical model/configuration: qwen3.8-max
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Cost basis
  • This row reports the cache-hit-rule figure as its headline. The Stage 1 row for the same model reports its full-input-rate UPPER BOUND (¥49.71) as its headline. On a like-for-like basis Stage 1 is ¥9.94 under the cache-hit rule, or Stage 2 is ¥410.98 at the full rate. Do not compare the two headline numbers directly.
  • Qwen Code records no per-request cost field; null by construction.
  • ¥410.98 if the 32,874,891 cache-read tokens were priced at the full ¥12/M input rate instead of the 10% cache-hit rule.
  • The 华北2 (Beijing) listing was used: input ¥12/M, output ¥36/M. The Singapore listing is ¥14.988/¥44.965. The run's own region is not identifiable from the ledger, which is one of the reasons the metrics status is PARTIAL.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/space-flight-game-threejs-stage-2-landfall-qwen3.8-max-xhigh/source/index.html
  • SHA-256: 15d759eaf2f3778e8ab66a45200ba20c365091ad9b15257b55cd52e5571b28c1
Validation evidence
  • Result: PASS
  • Path: artifacts/space-flight-game-threejs-stage-2-landfall-qwen3.8-max-xhigh/verification/landfall-qa-report.json
  • SHA-256: a1a174792d765780394f2919351ce489d48528de592532e4a7419d1db4b9268b
Recorded caveats
  • Metrics status is PARTIAL. The root ledger and root transcript token sums disagree by 26,237 fresh input, 198,249 cache read, 9,285 visible output, and 4,948 reasoning tokens; the ledger figures are reported and the gap is disclosed, not reconciled.
  • The pricing region is not identifiable from the ledger, and the cached-input rate comes from a generic published rule rather than a model-specific price.
  • Wall-clock (6,656.0 s) is end-to-end workflow latency including tool execution, browser checks, and idle gaps — not model-only compute. Summed API duration (6,303.7 s) is a separate figure.
  • One Qwen Code agent turn fans out to many billable requests and subagent sessions under a single session id (158 root + 12 subagent here).
  • Visible output and reasoning are separate fields; both are output-priced on this stack.
  • inputTokens includes cached tokens and outputTokens includes reasoning tokens, so the ledger's totalTokens and total_processed are different quantities and must never be swapped.
  • Qwen Code records no cache-write counter, so that line item is unavailable rather than zero.
  • Context occupancy is a single-request footprint, not cumulative usage; quota fields are null with no explicit before/after evidence.
  • The client version is reported as-is: installed and transcript-recorded both read 0.21.6 while the expected stack listed 0.21.7.
  • All runtime evidence is the model's own hardware-backed Metal session; no independent operator rerun was performed for this result.
  • The 120 FPS figure is Apple M1 Pro / ANGLE Metal evidence at 1920x993, not a standardized cross-machine comparison, and the QA loop logs 13 frame-time spikes above 35 ms.
  • The sanitized session export named in receipt.sha256 is withheld from the public archive by its own privacy label.
Visible evidence gaps
  • blind-evaluation record
  • independent (non-model-run) runtime and FPS receipt
  • reconciliation of the root ledger-vs-transcript token gap
  • region-identified pricing and a model-specific cached-input rate
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console