Measured harness ledgerPublic result
GLM 5.3 FlashApple Stem — GLM 5.3 Flash Highest
Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.
Highest reasoningHeadline result
- Workflow cost
- $0.56
- Wall-clock
- 89m 33.2s wall-clock
- Processed tokens
- 11.99M processed
- Record state
- judge_pass_independently_replayed
Public summary
GLM 5.3 Flash Highest judge_pass_independently_replayed ledger: 89m 33.2s wall-clock, 11.99M processed, and $0.56 Provider-recorded usage estimate, not an itemized subscription cash charge.
Run identity and stack
- Result ID: dex-apple-stem-superdex-glm-5.3-flash-max-opencode-go-2026-09-16
- Technical model: GLM 5.3 Flash
- Provider: OpenCode Go
- Client: OpenCode 1.18.23
- Variant: max
- Stack: OpenCode Go
- Stack: OpenCode 1.18.23
- Stack: Technical model/configuration: GLM 5.3 Flash
- Stack: Requested variant: max
- Stack: SuperDex fingertip fixture
- Stack: Harness apple-stem prompt
- Stack: Requested tool profile: python-superdex-headless
Primary artifact integrity
- Kind: open-loop-200hz-joint-target-log
- Path: artifacts/dex-apple-stem-superdex-glm-5.3-flash-max-opencode-go-2026-09-16/evidence/model-supplied/ctrl.npy
- SHA-256: 23380e77704e46518bf2a469c36d1c65c7169867d0dff9d3e20bcd29fbaf2c43
Validation evidence
- Result: pass
- Path: artifacts/dex-apple-stem-superdex-glm-5.3-flash-max-opencode-go-2026-09-16/validation/judge-replay.public.json
- SHA-256: 6e3f5878bce706915fc0f1bcc2c99dd60841fe82ed67daf6305bfdb1d0392451
- Validator SHA-256: 8d6c638716c7c28701a7369792ab05ddadc72ecc4ad2b6183dfb04b4122b4647
Recorded caveats
- The judge measures task completion only (lift height, hold, stem-only contact, runtime limits); there is no visual quality or blind evaluation for this task.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs and idle gaps; it is not model-only compute.
- The replay and progression clips are rendered by the operator's viewer from the judge's trajectory export with a fixed camera; they are presentation evidence, not a measurement.
- Solver-authored source is published byte for byte as the run left it; it is model output and is not rewritten.
- Two dex-cube tier-2 runs shared the same Mac for the whole window. SuperDex simulation dominates per-iteration cost here, so every wall clock in this cohort carries CPU contention and plausibly cost the slower models real iterations.
- One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
- The 5400 s launcher alarm killed the run mid-request, so the metrics receipt is PARTIAL: the final root request finished as `tool-calls` rather than `stop` and the session ledger records variant `default` where `--variant max` was requested. Every token component reconciles with 0 mismatches; nothing was relabelled.
- The run wrote no submission/NOTES.md, which the prompt requires. The approach field is the archive operator's reading of the solver's own script.
- The run never ran validate.py and suppressed the harness archive with RB_NO_ARCHIVE=1 - a solver-visible off-switch in fixture revision v1.1 that revision v1.2 closes - so it left no attempt archive and no sim-attempt archive of its own; the published verdict is the archive operator's judge run on the submitted log. Its progression is supplied after the fact as six unscored reconstructed attempts.
- This is one of the first four runs on fixture revision v1.1. Verdicts stay comparable with the three v1 runs; only archive availability and the fixture-identity scene_sha256 differ.
- The OpenCode-recorded figure is API-equivalent usage accounting under OpenCode Go, not an itemized marginal cash charge.
- Reconstructed attempts are the solver's own saved arrays judged after the fact by the archive operator. They are unscored evidence of progression: they do not change this run's verdict, stage or cohort standing and are excluded from every comparison.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- the model's own NOTES.md (never written)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.