Measured harness ledgerPublic result
Jev 1.13

Eiffel Tower Drawing — Jev 1.13 Unreported

Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.

Unreported reasoningHeadline result
Workflow cost
$0.03
Wall-clock
139.5s wall-clock
Processed tokens
Not recorded
Record state
judge_verified_decision_loop_run_independently_replayed
Public summary

Jev 1.13 Unreported judge_verified_decision_loop_run_independently_replayed ledger: 139.5s wall-clock, Not recorded, and $0.03 Not an estimate: OpenRouter's own usage.cost field, reported per call and summed over the episode..

Run identity and stack
  • Result ID: dex-draw-mujoco-jev-1.13-decision-loop-openrouter-2026-09-20
  • Technical model: Jev 1.13
  • Provider: TypeSafe via OpenRouter
  • Client: decision-loop executor (harness-authored), scaffold revision B2, DL_INTENT_GATING=0, POST https://openrouter.ai/api/alpha/decisions
  • Stack: TypeSafe via OpenRouter
  • Stack: decision-loop executor (harness-authored), scaffold revision B2, DL_INTENT_GATING=0, POST https://openrouter.ai/api/alpha/decisions
  • Stack: Technical model/configuration: Jev 1.13
  • Stack: MuJoCo pencil-drawing fixture
  • Stack: Harness drawing prompt
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-draw-mujoco-jev-1.13-decision-loop-openrouter-2026-09-20/evidence/decision-loop/ctrl.npy
  • SHA-256: 36cb12db5c01d5fb1660e8f2e14392de1faefb2f25affe9676ef6d43e0f767d8
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-draw-mujoco-jev-1.13-decision-loop-openrouter-2026-09-20/validation/judge-replay.public.json
  • SHA-256: 6e72a09330ea4f3b876dbf48819388ce7df561b3beee799e46d21600abe81391
  • Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
  • NOT COMPARABLE WITH THE CODING-AGENT COHORTS ON THIS TASK. Every other result here came from a coding agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason. This is also ONE SEED, ONE TRIAL, not a success-rate estimate.
  • ONE SEED, ONE TRIAL. This is a single episode, not a success-rate estimate, and it must not be quoted as one.
  • The executor is HARNESS-AUTHORED and owns every physical primitive listed in benchmarks/tracks/decision-loop/README.md. What is measured here is whether a decision model can sequence those primitives from structured state - not grasping, planning or control skill in general.
  • Adapter ceiling, measured with the scripted direction oracle on the same adapter, budget and frozen judge: tier 2 - coverage 0.818, precision 0.703, Chamfer 0.68 mm, 14/14 strokes. TIER 3 IS OUT OF REACH OF THIS ADAPTER AT THIS BUDGET; tier 2 is its ceiling, and Jev is 0.130 of coverage behind the oracle rather than behind the tier-3 bar.
  • The replay clip is NOT committed; it is pinned by name, sha256, bytes, frames and duration in provenance/clips.public.json.
  • api-archive.jsonl and model_calls.json are not committed - they carry the complete request/response stream and per-call provider generation ids - only their sha256 commitments are published, in provenance/source-receipt-hashes.json. The API key appears in none of them and was verified absent before this archive was assembled.
  • The superseded revision-A episode is disclosed by hash and is NOT scored.
  • provenance/orchestrator-results.md is the task project's own decision-loop section copied verbatim with absolute local paths replaced.
  • Wall clock is dominated by API waiting, not compute, and is not a measure of model speed.
  • evidence/decision-loop/run.json is committed EXACTLY as the executor wrote it, so its `final.submission` field still carries the operator's absolute workspace path - the same disclosure the claw self-check archive carries. It is a path, not a secret, and no other committed file in this archive contains one.
Visible evidence gaps
  • a second episode on this task (every row in this track is a single trial)
  • an LLM comparator through the identical loop (the scaffold supports one; none is registered here)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console