Measured harness ledgerPublic result
SemIf + Qwen3-0.6BApple Carry — SemIf + Qwen3-0.6B Unreported
Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.
Unreported reasoningHeadline result
- Workflow cost
- $0.00
- Wall-clock
- Not recorded
- Processed tokens
- Not recorded
- Record state
- operator_stopped_incomplete_sweep
Public summary
SemIf + Qwen3-0.6B Unreported operator_stopped_incomplete_sweep ledger: Not recorded wall-clock, Not recorded, and $0.00 local inference, $0.
Run identity and stack
- Result ID: apple-carry-fr3-gripper-semif-qwen3-0.6b-decision-loop-local-2026-09-21
- Technical model: SemIf + Qwen3-0.6B
- Provider: local (MPS)
- Client: SemIf mode=direct, batch size 1
- Stack: local (MPS)
- Stack: SemIf mode=direct, batch size 1
- Stack: Technical model/configuration: SemIf + Qwen3-0.6B
- Stack: Embodied manipulation fixture
- Stack: Harness apple-carry prompt
Validation evidence
- Path: artifacts/apple-carry-fr3-gripper-semif-qwen3-0.6b-decision-loop-local-2026-09-21/validation/judge-replay.public.json
- SHA-256: 732f167d9d1e391eea5c7717fa4769cf532c8a7ed0522ec08fa267db2c3ad8f1
- Validator SHA-256: e92e63eb3a4c358a423e88d14badd90f1fddbabbf97d0e55857a29d14d939ae9
Recorded caveats
- NOT COMPARABLE WITH THE CODING-AGENT COHORTS. Every coding-agent result in this harness came from an agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason.
- 10 SEEDS, ONE EPISODE EACH.
- Operator stopped the sweep. On disk: seeds 0-8 judged FAIL, seed 9 ran 160 cycles but was not judged. Files win over the RESULTS.md '0/9' line (seed 9 appeared later) and over the '19 seeds' brief (10 seed directories exist).
- No clip was delivered for this controller.
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.