Measured harness ledgerPublic result
Qwen3.8 Flash

Dex Cube Turn — Qwen3.8 Flash Highest

Turn one face of a Rubik's cube a quarter turn by finger contact with two dexterous hands in MuJoCo.

Highest reasoningHeadline result
Workflow cost
$0.26
Wall-clock
48m 34.9s wall-clock
Processed tokens
10.49M processed
Record state
judge_fail_independently_replayed
Public summary

Qwen3.8 Flash Highest judge_fail_independently_replayed ledger: 48m 34.9s wall-clock, 10.49M processed, and $0.26 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-qwen3.8-flash-max-opencode-openrouter-2026-09-16
  • Technical model: Qwen 3.8 Flash
  • Provider: OpenRouter
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenRouter
  • Stack: OpenCode 1.18.23
  • Stack: OpenCode → OpenRouter
  • Stack: Technical model/configuration: Qwen 3.8 Flash
  • Stack: Requested variant: max
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn prompt
  • Stack: Requested tool profile: python-mujoco-headless
Cost basis
  • actual marginal cash charged usd: $0.00 USD.
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-qwen3.8-flash-max-opencode-openrouter-2026-09-16/evidence/model-supplied/ctrl.npy
  • SHA-256: 6ee869f2518e97fda2480e41b894d9dfce7be019ecd6cf31717e802a2ef5d643
Validation evidence
  • Result: pass
  • Path: artifacts/dex-cube-turn-mujoco-qwen3.8-flash-max-opencode-openrouter-2026-09-16/validation/judge-replay.public.json
  • SHA-256: 23ed4dab0f1130cc30157749a8ebe6619ba51809d089fe0b6ce3f63637f5a0fe
  • Validator SHA-256: 8425328aed969dfd11cc1e5107b1e0c83b888c59675bdf2ce86c4bc2878dcee9
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • Wall-clock includes tools, simulation and judge runs, and idle gaps; it is not model-only compute.
  • One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
  • OpenRouter billed $0 (BYOK Alibaba key); the OpenCode-recorded $0.2616 is a catalog-rate estimate. The receipt's own billing block wrongly equates it with marginal cash because the extractor assumes pay-as-you-go credits.
  • The counted run is a clean single session; three earlier launches aborted instantly on the max_tokens cap before the project opencode.json fix and are not counted (archived privately).
  • The solver submitted a gravity-compensated 3 s hold after its turn attempts failed; the log is violation-free and fails only the target-state checks.
  • Attempt archiving was not yet in this tier's judge for this run (fixture revision v1); the three reconstructed attempts under evidence/reconstructed-attempts/ were re-run by the operator after the fact from the solver's own scripts, judged FAIL, and are not scored.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • judge-time attempt archive (only the after-the-fact reconstruction exists)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console