Measured harness ledgerPublic result
SemIf + Qwen3.5-4B

Apple Carry — SemIf + Qwen3.5-4B Unreported

Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.

Unreported reasoningHeadline result
Workflow cost
$0.00
Wall-clock
Not recorded
Processed tokens
Not recorded
Record state
judge_verified_decision_loop_sweep_independently_replayed
Public summary

SemIf + Qwen3.5-4B Unreported judge_verified_decision_loop_sweep_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $0.00 Recorded provider cost.

Run identity and stack
  • Result ID: apple-carry-xarm7-openroboto-semif-qwen3.5-4b-decision-loop-remote-4090-2026-09-21
  • Technical model: SemIf + Qwen3.5-4B
  • Provider: RunPod RTX 4090 (SSH tunnel)
  • Client: SemIf mode=direct via POST /score on the pod
  • Stack: RunPod RTX 4090 (SSH tunnel)
  • Stack: SemIf mode=direct via POST /score on the pod
  • Stack: Technical model/configuration: SemIf + Qwen3.5-4B
  • Stack: Embodied manipulation fixture
  • Stack: Harness apple-carry prompt
Validation evidence
  • Path: artifacts/apple-carry-xarm7-openroboto-semif-qwen3.5-4b-decision-loop-remote-4090-2026-09-21/validation/judge-replay.public.json
  • SHA-256: f845d6db49bf039243f6c9bdbbb5fc534f521f37d15e52e33777d72948f4e77d
  • Validator SHA-256: 5facf7bb27541fca9a60dbaaded91151a3e4f9962419bfd55877bd0b49db5d81
Recorded caveats
  • NOT COMPARABLE WITH THE CODING-AGENT COHORTS. Every coding-agent result in this harness came from an agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason.
  • 0/20, 160 cycles every seed, withdraw+hold lock. Pod $2.40 shared with the FR3 sweep.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console