Measured harness ledgerPublic result
Jev 1.13Claw Unlock — Jev 1.13 Unreported
Find a collision-free rigid motion that unlocks two interlocked wire claws using pure geometry.
Unreported reasoningHeadline result
- Workflow cost
- $0.50
- Wall-clock
- 978.7s wall-clock
- Processed tokens
- Not recorded
- Record state
- judge_verified_decision_loop_run_independently_replayed
Public summary
Jev 1.13 Unreported judge_verified_decision_loop_run_independently_replayed ledger: 978.7s wall-clock, Not recorded, and $0.50 Not an estimate: OpenRouter's own usage.cost field, reported per call and summed over the episode..
Run identity and stack
- Result ID: claw-unlock-geometry-jev-1.13-decision-loop-openrouter-2026-09-20
- Technical model: Jev 1.13
- Provider: TypeSafe via OpenRouter
- Client: decision-loop executor (harness-authored), scaffold revision B2, DL_INTENT_GATING=0, POST https://openrouter.ai/api/alpha/decisions
- Stack: TypeSafe via OpenRouter
- Stack: decision-loop executor (harness-authored), scaffold revision B2, DL_INTENT_GATING=0, POST https://openrouter.ai/api/alpha/decisions
- Stack: Technical model/configuration: Jev 1.13
- Stack: python-fcl geometry fixture
- Stack: Harness claw-unlock prompt
Primary artifact integrity
- Kind: se3-rigid-body-path-polyline
- Path: artifacts/claw-unlock-geometry-jev-1.13-decision-loop-openrouter-2026-09-20/evidence/decision-loop/path.npy
- SHA-256: 6cdb1f454de3e4ad1082cb3c73889406a119de867abc71be11e90dbb2292b162
Validation evidence
- Result: PASS
- Path: artifacts/claw-unlock-geometry-jev-1.13-decision-loop-openrouter-2026-09-20/validation/judge-replay.public.json
- SHA-256: c875dd973cb55ef05a9429f525b09f24fff6512c52f8a22821f8ad69502335b1
- Validator SHA-256: ac0ad531aaa6a4f93213a4445af54d4bcd582fbe105ca369b8253050e49c24cd
Recorded caveats
- NOT COMPARABLE WITH THE CODING-AGENT COHORTS ON THIS TASK. Every other result here came from a coding agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason. This is also ONE SEED, ONE TRIAL, not a success-rate estimate.
- THE EPISODE DID NOT RUN ITS BUDGET. It ended at cycle 2 122 of 3 000 on a provider HTTP 520 that persisted through the bounded retransmits. By cycle 177 the trajectory had already settled into the oscillation described under `approach` and stayed in it for the remaining 1 900 cycles, but the episode is still short of its budget and is reported as a 2 122-cycle run.
- ONE SEED, ONE TRIAL. This is a single episode, not a success-rate estimate, and it must not be quoted as one.
- The executor is HARNESS-AUTHORED and owns every physical primitive listed in benchmarks/tracks/decision-loop/README.md. What is measured here is whether a decision model can sequence those primitives from structured state - not grasping, planning or control skill in general.
- Adapter ceiling, measured with the scripted direction oracle on the same adapter, budget and frozen judge: PASS stage 3 in 1 819 of 3 000 cycles (3 123 rows, 7.666 m, referenceRatio 0.92, interlock 71 -> 0, both milestones, hullGap 6.8 mm; 17 increments rejected = 0.9 %, 102 slid). THIS ORACLE IS PRIVILEGED: it reads the published gold path to pick, each cycle, the six signs that most reduce the SE(3) distance to the next gold waypoint. It is an upper bound on what the ACTION SPACE allows, not an estimate of what a model that does not know the corridor could do.
- The replay clip is NOT committed; it is pinned by name, sha256, bytes, frames and duration in provenance/clips.public.json.
- api-archive.jsonl and model_calls.json are not committed - they carry the complete request/response stream and per-call provider generation ids - only their sha256 commitments are published, in provenance/source-receipt-hashes.json. The API key appears in none of them and was verified absent before this archive was assembled.
- The superseded revision-A episode is disclosed by hash and is NOT scored.
- provenance/orchestrator-results.md is the task project's own decision-loop section copied verbatim with absolute local paths replaced.
- Wall clock is dominated by API waiting, not compute, and is not a measure of model speed.
- evidence/decision-loop/run.json is committed EXACTLY as the executor wrote it, so its `final.submission` field still carries the operator's absolute workspace path - the same disclosure the claw self-check archive carries. It is a path, not a secret, and no other committed file in this archive contains one.
Visible evidence gaps
- a second episode on this task (every row in this track is a single trial)
- an LLM comparator through the identical loop (the scaffold supports one; none is registered here)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.