Measured harness ledgerPublic result
GLM 5.3 FlashDex Cube Turn — GLM 5.3 Flash Highest
Turn one face of a Rubik's cube a quarter turn by finger contact with two dexterous hands in MuJoCo.
Highest reasoningHeadline result
- Workflow cost
- $0.34
- Wall-clock
- 40m 47.8s wall-clock
- Processed tokens
- 7.87M processed
- Record state
- judge_pass_independently_replayed
Public summary
GLM 5.3 Flash Highest judge_pass_independently_replayed ledger: 40m 47.8s wall-clock, 7.87M processed, and $0.34 Provider-recorded usage estimate, not an itemized subscription cash charge.
Run identity and stack
- Result ID: dex-cube-turn-mujoco-glm-5.3-flash-max-opencode-go-2026-09-16
- Technical model: GLM 5.3 Flash
- Provider: OpenCode Go
- Client: OpenCode 1.18.23
- Variant: max
- Stack: OpenCode Go
- Stack: OpenCode 1.18.23
- Stack: Technical model/configuration: GLM 5.3 Flash
- Stack: Requested variant: max
- Stack: MuJoCo dexterous-hand fixture
- Stack: Harness cube-turn prompt
- Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
- Kind: open-loop-1khz-actuator-log
- Path: artifacts/dex-cube-turn-mujoco-glm-5.3-flash-max-opencode-go-2026-09-16/evidence/model-supplied/ctrl.npy
- SHA-256: ad856ae4ddce86c5c0142a9dda985f89f86a3f2fa5700817e6e29c47905b40cc
Validation evidence
- Result: pass
- Path: artifacts/dex-cube-turn-mujoco-glm-5.3-flash-max-opencode-go-2026-09-16/validation/judge-replay.public.json
- SHA-256: 40a0971834df71bd641f2cb323fd351a227233d3bf5e4dd7f75971a771653964
- Validator SHA-256: 8425328aed969dfd11cc1e5107b1e0c83b888c59675bdf2ce86c4bc2878dcee9
Recorded caveats
- The judge measures task completion only; there is no visual quality or blind evaluation for this task.
- Wall-clock includes tools, simulation and judge runs, and idle gaps; it is not model-only compute.
- One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
- Attempt archiving was not yet in this tier's judge for these runs, so no judge attempt records exist; the progression published here is six unscored reconstructed attempts recovered after the fact - see reconstructed_attempts.
- The OpenCode-recorded figure is API-equivalent usage accounting under OpenCode Go, not an itemized marginal cash charge.
- Reconstructed attempts are this run's own progression, recovered and judged after the fact by the archive operator. They are unscored evidence: they do not change this run's verdict or cohort standing and are excluded from every comparison, including any of them that the judge passes.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.