Measured harness ledgerPublic result
Qwen3.8 Flash

Eiffel Tower Drawing — Qwen3.8 Flash Highest

Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.

Highest reasoningHeadline result
Workflow cost
$0.13
Wall-clock
19m 51.2s wall-clock
Processed tokens
5.18M processed
Record state
judge_pass_independently_replayed
Public summary

Qwen3.8 Flash Highest judge_pass_independently_replayed ledger: 19m 51.2s wall-clock, 5.18M processed, and $0.13 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-draw-mujoco-qwen3.8-flash-max-opencode-openrouter-2026-09-16
  • Technical model: Qwen 3.8 Flash
  • Provider: OpenRouter
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenRouter
  • Stack: OpenCode 1.18.23
  • Stack: OpenCode → OpenRouter
  • Stack: Technical model/configuration: Qwen 3.8 Flash
  • Stack: Requested variant: max
  • Stack: MuJoCo pencil-drawing fixture
  • Stack: Harness drawing prompt
  • Stack: Requested tool profile: python-mujoco-headless
Cost basis
  • actual marginal cash charged usd: $0.00 USD.
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-draw-mujoco-qwen3.8-flash-max-opencode-openrouter-2026-09-16/evidence/model-supplied/ctrl.npy
  • SHA-256: a6b6e37e3f6ce85a6471fb4c3d7d1c3c06eea57fabc6b313f27478b8c040869d
Validation evidence
  • Result: pass
  • Path: artifacts/dex-draw-mujoco-qwen3.8-flash-max-opencode-openrouter-2026-09-16/validation/judge-replay.public.json
  • SHA-256: 7de4e2851fb307c55b726629afec73cd8b9ade755e4ee3c78e7d098cef9005f9
  • Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
  • The judge measures task completion only (contact-gated ink against the given polyline, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay and progression clips are rendered by the operator's viewer from the judge's trajectory export and ink stream with a fixed camera; they are presentation evidence, not a measurement.
  • usedThreeFingers / threeFingerFraction are reported, never gating.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute, and all four runs of this batch shared one Mac.
  • The given harness exposes the judge's own simulate/rasterise/score pipeline, so the solver could self-score in-process; wall clock and cost are NOT comparable with the dex-cube-turn-mujoco cohorts.
  • This run is on fixture revision v1.1 (harness simulation archiving only). Verdicts stay comparable with the three v1 runs; scene_sha256 differs because it hashes harness.py.
  • One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
  • The quoted cost is OpenCode's catalog-rate estimate: the OpenRouter account routes this model through a BYOK Alibaba key, so OpenRouter billed $0.
  • The counted run is the second launch; the first died in about 5 s on OpenCode SQLite contention with zero model calls and is disclosed, not counted.
  • This workspace keeps the variants.max.max_tokens=131072 configuration fix carried over from dex-cube-turn.v1 (provenance/CONFIG_NOTE.md); the instant-`length`-abort retry loop never fired.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console