Measured harness ledgerPublic result
DeepSeek V4.1 Flash

Apple Stem — DeepSeek V4.1 Flash Highest

Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.

Highest reasoningHeadline result
Workflow cost
$0.10
Wall-clock
88m 20.2s wall-clock
Processed tokens
7.82M processed
Record state
judge_fail_independently_replayed
Public summary

DeepSeek V4.1 Flash Highest judge_fail_independently_replayed ledger: 88m 20.2s wall-clock, 7.82M processed, and $0.10 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-apple-stem-superdex-deepseek-v4.1-flash-max-opencode-go-2026-09-16
  • Technical model: DeepSeek V4.1 Flash
  • Provider: OpenCode Go
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenCode Go
  • Stack: OpenCode 1.18.23
  • Stack: Technical model/configuration: DeepSeek V4.1 Flash
  • Stack: Requested variant: max
  • Stack: SuperDex fingertip fixture
  • Stack: Harness apple-stem prompt
  • Stack: Requested tool profile: python-superdex-headless
Primary artifact integrity
  • Kind: open-loop-200hz-joint-target-log
  • Path: artifacts/dex-apple-stem-superdex-deepseek-v4.1-flash-max-opencode-go-2026-09-16/evidence/model-supplied/ctrl.npy
  • SHA-256: 6c7684994be2041004f5ff1c8efcf20cd558f92a724b224e8b320f78ac7309fc
Validation evidence
  • Result: pass
  • Path: artifacts/dex-apple-stem-superdex-deepseek-v4.1-flash-max-opencode-go-2026-09-16/validation/judge-replay.public.json
  • SHA-256: d8fe7589b5ed52172226685e8f7fdde53af72ebb3a83858e37116b0c90d03d76
  • Validator SHA-256: 8d6c638716c7c28701a7369792ab05ddadc72ecc4ad2b6183dfb04b4122b4647
Recorded caveats
  • The judge measures task completion only (lift height, hold, stem-only contact, runtime limits); there is no visual quality or blind evaluation for this task.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs and idle gaps; it is not model-only compute.
  • The replay and progression clips are rendered by the operator's viewer from the judge's trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Solver-authored source is published byte for byte as the run left it. Several scripts contain the workspace's absolute path in a literal string; that is model output, not an operator receipt, and it is not rewritten. Operator provenance, validation and ledger files carry no local paths.
  • Two dex-cube tier-2 runs shared the same Mac for the whole window. SuperDex simulation dominates per-iteration cost here, so every wall clock in this cohort carries CPU contention and plausibly cost the slower models real iterations.
  • One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
  • The 5400 s launcher alarm killed the run mid-request, so the metrics receipt is PARTIAL: the final root request finished as `tool-calls` rather than `stop` and the session ledger records variant `default` where `--variant max` was requested. Every token component reconciles with 0 mismatches; nothing was relabelled.
  • The run wrote no submission/NOTES.md, which the prompt requires. The approach field is the archive operator's reading of the solver's own scripts.
  • The judged log stops at 2.255 s on a runtime violation, so the log-length and hold checks fail as consequences rather than independently; only the first failed check is diagnostic.
  • Both sim-attempt records are the judge's own re-simulations, not solver simulate() calls, because this task's validate.py does not set RB_NO_ARCHIVE=1 for itself.
  • This is one of the first four runs on fixture revision v1.1. Verdicts stay comparable with the three v1 runs; only archive availability and the fixture-identity scene_sha256 differ.
  • The OpenCode-recorded figure is API-equivalent usage accounting under OpenCode Go, not an itemized marginal cash charge.
  • Reconstructed attempts are the solver's own saved arrays judged after the fact by the archive operator. They are unscored evidence of progression: they do not change this run's verdict, stage or cohort standing and are excluded from every comparison.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • the model's own NOTES.md (never written)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console