Measured harness ledgerPublic result
Jev 1.13

Apple Carry — Jev 1.13 Unreported

Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.

Unreported reasoningHeadline result
Workflow cost
$0.24
Wall-clock
Not recorded
Processed tokens
Not recorded
Record state
judge_verified_decision_loop_sweep_independently_replayed
Public summary

Jev 1.13 Unreported judge_verified_decision_loop_sweep_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $0.24 Not an estimate: OpenRouter's own usage.cost field, reported per call and summed over every call of every episode..

Run identity and stack
  • Result ID: apple-carry-fr3-gripper-jev-1.13-decision-loop-openrouter-2026-09-20
  • Technical model: Jev 1.13
  • Provider: TypeSafe via OpenRouter
  • Client: decision-loop executor (harness-authored), scaffold revision B2, DL_INTENT_GATING=0, POST https://openrouter.ai/api/alpha/decisions, driven seed by seed by sweep.py
  • Stack: TypeSafe via OpenRouter
  • Stack: decision-loop executor (harness-authored), scaffold revision B2, DL_INTENT_GATING=0, POST https://openrouter.ai/api/alpha/decisions, driven seed by seed by sweep.py
  • Stack: Technical model/configuration: Jev 1.13
  • Stack: Embodied manipulation fixture
  • Stack: Harness apple-carry prompt
Validation evidence
  • Path: artifacts/apple-carry-fr3-gripper-jev-1.13-decision-loop-openrouter-2026-09-20/validation/judge-replay.public.json
  • SHA-256: 210f7e31b383cb8c2b3a41ed0066936cfda6627702f0819b8e5e2683b20ebdc2
  • Validator SHA-256: e92e63eb3a4c358a423e88d14badd90f1fddbabbf97d0e55857a29d14d939ae9
Recorded caveats
  • NOT COMPARABLE WITH THE CODING-AGENT COHORTS. Every coding-agent result in this harness came from an agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason. A coding-agent cohort now exists on apple-carry (GPT-6 Astra max, 5/5 seeds 0-4, apple-carry-fr3-gripper-gpt-6-astra-max-codex-2026-09-21) but is STILL NOT COMPARABLE with this decision-loop row. cohort_eligible is false for exactly this reason.
  • TWENTY SEEDS, ONE EPISODE EACH. This is a success RATE over twenty scenes, not twenty independent trials of one scene: each row is a single episode on its own seed, so the Wilson interval covers scene-to-scene variation and says nothing about run-to-run variance on a fixed scene.
  • The executor is HARNESS-AUTHORED and owns every physical primitive listed in benchmarks/tracks/decision-loop/README.md - here the Cartesian increment and its magnitude ladder, the jacobian IK, the single force-limited gripper command and the final settle. What is measured is whether a decision model can SEQUENCE those primitives from structured state; it is not a measure of grasping, planning or control skill.
  • Adapter ceiling on the same adapter, budget and frozen judge, measured with the scripted direction oracle: 20/20 PASS stage 3 at mean 67.2 cycles. The model is 4.75 cycles per episode off that ceiling, which is the only in-track comparison these numbers support.
  • The seed-0 replay clip is NOT committed; it is pinned by name, sha256, bytes, frames and duration in provenance/clips.public.json. There is no failure clip because no seed failed.
  • api-archive.jsonl and model_calls.json are not committed for any seed - they carry the complete request/response stream and per-call provider generation ids - only their per-seed sha256 commitments are, in provenance/source-receipt-hashes.json. The API key appears in none of them and every file in this archive was grepped for it before the archive was assembled.
  • The superseded FR3 + Wuji-hand embodiment (Jev 1.13 0/20) is disclosed by hash and is NOT scored.
  • The per-episode run.json files are committed EXACTLY as the executor wrote them, so each one's final.submission field still carries the operator's absolute workspace path, and evidence/decision-loop/sweep.json likewise keeps a run_dir per row - the same disclosure the earlier decision-loop archives carry. They are paths, not secrets, and no ledger, README or provenance file in this archive contains one.
  • The task project's RESULTS.md quotes a max torque ratio of 0.57 for the oracle and 0.61 for Jev; those are seed-0 figures. Over all twenty seeds the maximum is 0.6071 for both, and that is what this ledger records.
Visible evidence gaps
  • a second episode per seed (every row in this track is a single trial on its seed)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console