Measured harness ledgerPublic result
Jev 1.13Apple Carry — Jev 1.13 Unreported
Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.
Unreported reasoningHeadline result
- Workflow cost
- $0.35
- Wall-clock
- Not recorded
- Processed tokens
- Not recorded
- Record state
- judge_verified_decision_loop_sweep_independently_replayed
Public summary
Jev 1.13 Unreported judge_verified_decision_loop_sweep_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $0.35 Not an estimate: OpenRouter's own usage.cost field, reported per call and summed over every call of every episode..
Run identity and stack
- Result ID: apple-carry-xarm7-openroboto-jev-1.13-decision-loop-openrouter-2026-09-20
- Technical model: Jev 1.13
- Provider: TypeSafe via OpenRouter
- Client: OpenRoboto's own reproduce.py --controller jev --seed N --no-render, unmodified, at commit 7a4ed8b72c3c17d7aa790678ed9660df67c10dd3, driven seed by seed by tools/run_seeds.py
- Stack: TypeSafe via OpenRouter
- Stack: OpenRoboto's own reproduce.py --controller jev --seed N --no-render, unmodified, at commit 7a4ed8b72c3c17d7aa790678ed9660df67c10dd3, driven seed by seed by tools/run_seeds.py
- Stack: Technical model/configuration: Jev 1.13
- Stack: Embodied manipulation fixture
- Stack: Harness apple-carry prompt
Validation evidence
- Path: artifacts/apple-carry-xarm7-openroboto-jev-1.13-decision-loop-openrouter-2026-09-20/validation/judge-replay.public.json
- SHA-256: 69ac2a4285078ed460e89bb9e1f6c8ce685176dbc8c1e27aee01286f3738a20b
- Validator SHA-256: 5facf7bb27541fca9a60dbaaded91151a3e4f9962419bfd55877bd0b49db5d81
Recorded caveats
- NOT COMPARABLE WITH THE CODING-AGENT COHORTS. Every coding-agent result in this harness came from an agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason. A coding-agent cohort now exists on apple-carry (GPT-6 Astra max, 5/5 seeds 0-4, apple-carry-fr3-gripper-gpt-6-astra-max-codex-2026-09-21) but is STILL NOT COMPARABLE with this decision-loop row. cohort_eligible is false for exactly this reason.
- TWENTY SEEDS, ONE EPISODE EACH, AND ONE OF THEM IS REUSED. Seed 0 is the 2026-09-20 12:02 trial from OpenRoboto's own checkout, copied in unchanged and judged here; it was not re-run. See seed_zero_provenance.
- THE SEEDS ARE NOT THE OTHER EMBODIMENT'S SEEDS. Here they jitter the apple by up to +/-12 mm and the plate is fixed; in fr3-gripper the apple moves up to +/-4 cm and the plate up to +/-3 cm. The two 20/20 rows are two measurements of the same task shape on two different machines and must not be pooled.
- The executor, the increment ladder, the question text and the answer validation in this embodiment are OPENROBOTO'S, not ours. What is measured is whether a decision model can sequence their primitives from their structured state; it is not a measure of grasping, planning or control skill, and it is not a measurement of code the harness wrote.
- No adapter-ceiling oracle was run in this embodiment. The only in-track reference here is OpenRoboto's own published seed-0 result, registered as apple-carry-xarm7-openroboto-reference-published-recording.
- The seed-0 replay clip is NOT committed; it is pinned by name, sha256, bytes, frames and duration in provenance/clips.public.json. There is no failure clip because no seed failed.
- The recorded run records and journals are not committed for any seed - each record is simultaneously the submission artefact and the complete request/response archive, with per-call provider generation ids - only their per-seed sha256 commitments are, in provenance/source-receipt-hashes.json. Because of that the committed cycle-log.json carries the decisions but NOT the native probability distributions, which is a real asymmetry with the fr3-gripper archive and is stated there too. The API key appears in none of these files and every file in this archive was grepped for it before the archive was assembled.
- The superseded FR3 + Wuji-hand embodiment (Jev 1.13 0/20) is disclosed by hash and is NOT scored.
- The judge's verdict.json and attempt.json echo the `record` path they were given, so for seeds 1-19 those committed files carry the operator's absolute workspace path, and evidence/decision-loop/sweep.json keeps a run_dir per row - the same disclosure the earlier decision-loop archives carry. They are paths, not secrets, and no ledger, README or provenance file in this archive contains one.
Visible evidence gaps
- a coding-agent submission kind for this embodiment (there is none, so there can be no coding-agent cohort here)
- a scripted adapter-ceiling oracle in this embodiment
- a second episode per seed (every row in this track is a single trial on its seed)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.