Measured harness ledgerPublic result
GLM 5.3 Flash

Dex Cube — 3 Move — GLM 5.3 Flash Highest

Complete the three-move Rubik's cube sequence D, F', U2 with regrasp, using two dexterous hands in MuJoCo.

Highest reasoningHeadline result
Workflow cost
$1.44
Wall-clock
1h 31m 15.1s wall-clock
Processed tokens
41.18M processed
Record state
judge_fail_independently_replayed
Public summary

GLM 5.3 Flash Highest judge_fail_independently_replayed ledger: 1h 31m 15.1s wall-clock, 41.18M processed, and $1.44 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-stage-2-three-moves-glm-5.3-flash-max-opencode-go-2026-09-16
  • Technical model: GLM 5.3 Flash
  • Provider: OpenCode Go
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenCode Go
  • Stack: OpenCode 1.18.23
  • Stack: Technical model/configuration: GLM 5.3 Flash
  • Stack: Requested variant: max
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn stage-2 prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-glm-5.3-flash-max-opencode-go-2026-09-16/evidence/model-supplied/ctrl.npy
  • SHA-256: 45fc9bc920f23b61cbc30f709bdb889bd8d95b9d288fcfba90f3ef7c706a861f
Validation evidence
  • Result: pass
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-glm-5.3-flash-max-opencode-go-2026-09-16/validation/judge-replay.public.json
  • SHA-256: cd77183ba3ba6f16ca0d10e0fabae18362170aa88a61fb16ec39341a42842107
  • Validator SHA-256: 329a7f94ec9d9a37edbb94b6fcad5a2e48e32aca160e93e683f5de2763f4ab34
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • This is a judge FAIL with no partial credit: movesReached 0 of 3, the cube's facelet state unchanged from the initial scramble. The log is violation-free, so the failure is the absent moves, not a safety breach.
  • Most of the run's budget went into rotating the cube by writing orientation into qpos directly - a channel the re-simulating judge cannot see. The solver established that itself and did not submit it.
  • Wall-clock includes tools, simulation and judge runs, and idle gaps; it is not model-only compute. The solver used about 61% of its 150-minute budget and ended on its own final stop finish; it was never interrupted, nudged or resumed.
  • One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
  • The OpenCode-recorded figure is API-equivalent usage accounting under OpenCode Go, not an itemized marginal cash charge, and no quota delta was exposed.
  • The run used the v1.1 harness with the v1 TASK.md; validate.py is byte-identical across the revisions, so verdicts remain comparable with the v1 and v1.1 stage-2 runs. Disclosed in the input lineage rather than corrected after the fact.
  • The progression clip's segments are the solver's own sim-attempts, not judge attempts: all three judge attempts are the same log as the submission and were deduped out.
  • The replay clips are rendered by the operator's viewer from the judge's trajectory exports with a fixed camera; they are presentation evidence, not measurements.
  • The session id is withheld per precedent; the raw metrics receipts, the sanitized session export and the launcher transcripts are committed by hash only, and provenance/run-notes.md and the sim-attempt records are sanitized.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • quota delta (OpenCode Go exposes none)
  • the solver's /tmp scratch scripts (written outside the workspace and not preserved; only lib.py and flick.py remain)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console