Measured harness ledgerPublic result
Claude Opus 5.5Apple Stem — Claude Opus 5.5 Max
Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.
Max reasoningHeadline result
- Workflow cost
- $2.11
- Wall-clock
- 16m 39s wall-clock
- Processed tokens
- 2.34M processed
- Record state
- judge_pass_independently_replayed
Public summary
Claude Opus 5.5 Max judge_pass_independently_replayed ledger: 16m 39s wall-clock, 2.34M processed, and $2.11 recorded total_cost_usd from workspace FINAL.json (modelUsage.claude-opus-5-5, costBasis=list, Anthropic list price). Charged to the Claude subscription ($0 charged)..
Run identity and stack
- Result ID: dex-apple-stem-superdex-opus-5.5-max-2026-09-22
- Technical model: Claude Opus 5.5
- Provider: Anthropic Claude Code
- Client: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
- Stack: Anthropic Claude Code
- Stack: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
- Stack: Technical model/configuration: Claude Opus 5.5
- Stack: SuperDex fingertip fixture
- Stack: Harness apple-stem prompt
- Stack: Requested tool profile: python-superdex-headless
Primary artifact integrity
- Kind: open-loop-200hz-actuator-log
- Path: artifacts/dex-apple-stem-superdex-opus-5.5-max-2026-09-22/evidence/model-supplied/ctrl.npy
- SHA-256: 43b13a613b8629063d203634492c2e8734172d61afd7263739e7c5fb909b4452
Validation evidence
- Result: PASS
- Path: artifacts/dex-apple-stem-superdex-opus-5.5-max-2026-09-22/validation/judge-replay.public.json
- SHA-256: d2f841e7906a515347e3f843a135c4ebe9fee547890f15e201288e87bfb3e1d0
- Validator SHA-256: c735d44ca72534f8058a7e1061ccf775f6e9f180cef26200d97a402f3de7403e
Recorded caveats
- The judge measures task completion only; there is no visual quality or blind evaluation for this task.
- The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- Cost is the recorded list-price total_cost_usd from workspace FINAL.json (costBasis=list). Charged to the Claude subscription ($0 charged).
- provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
- The raw Claude CLI receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
- A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EU-CZ-1, 105.6 + 65.4 min, $1.30 + $0.81 = $2.11 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
- Nudge disclosure: none.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.