Measured harness ledgerPublic result
Claude Opus 5

Apple Stem — Claude Opus 5 Max

Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.

Max reasoningHeadline result
Workflow cost
$7.70
Wall-clock
33m 30s wall-clock
Processed tokens
6.98M processed
Record state
judge_pass_independently_replayed
Public summary

Claude Opus 5 Max judge_pass_independently_replayed ledger: 33m 30s wall-clock, 6.98M processed, and $7.70 API-equivalent estimate from session token counts at Anthropic list price (input $5, output $25, cache read $0.50, cache write $6.25 per MTok).

Run identity and stack
  • Result ID: dex-apple-stem-superdex-opus-5-max
  • Technical model: Claude Opus 5
  • Provider: Anthropic Claude Code
  • Client: claude -p --model claude-opus-5 --effort max (headless subagent), one shot, no nudges
  • Stack: Anthropic Claude Code
  • Stack: claude -p --model claude-opus-5 --effort max (headless subagent), one shot, no nudges
  • Stack: Technical model/configuration: Claude Opus 5
  • Stack: SuperDex fingertip fixture
  • Stack: Harness apple-stem prompt
  • Stack: Requested tool profile: python-superdex-headless
Primary artifact integrity
  • Kind: open-loop-200hz-joint-target-log
  • Path: artifacts/dex-apple-stem-superdex-opus-5-max/evidence/model-supplied/ctrl.npy
  • SHA-256: 6e0b26d24fa089d8597044259f7fa2eef895535d605ef7cc95e5ff4f996d86ac
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-apple-stem-superdex-opus-5-max/validation/judge-replay.public.json
  • SHA-256: 6dc6a8919709e8a9f160e192fed031665900c050bb633626a55dbb1296d0e556
  • Validator SHA-256: 8d6c638716c7c28701a7369792ab05ddadc72ecc4ad2b6183dfb04b4122b4647
Recorded caveats
  • The judge measures task completion only (lift height, hold, stem-only contact, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
  • provenance/gold-results.md is the orchestrator's own result file copied verbatim.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Opus cost is an API-list-price equivalent from token counts, not a subscription charge; fresh input was reported as approximately zero.
  • The Opus workspace carried the pre-Y-up-fix export_scene.py, so its archived robot.glb was written by that older exporter; trajectory bins and verdict are unaffected.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console