Demos
Each recording shows the cameras the system sees, what the screener and the judge said at that moment, and a timeline of who was busy: the policy, the screener, or the judge. Each RoboCasa clip runs the same episode twice, side by side: the policy alone on the left, the policy with Spotter on the right.
Clips are cut for viewing: stretches marked ×5 on screen are fast-forwarded, and in RoboCasa the simulator's render and save time is removed. The RoboCasa clock is robot time.
Supervise in parallel, not in series
VLM-led robot agents are serial: the VLM plans a segment, the policy executes it, and the VLM decides what comes next. That forces a trade-off. Short segments mean a VLM call at every step and high latency even on easy tasks. Long segments mean a deviation cannot be caught in time, and the actions that follow keep going wrong.
Spotter removes the trade-off by changing who leads. The embodied model keeps acting. A faster VLM screens each finished chunk in parallel, and the frontier VLM is called only after two flags in a row, to confirm the error and repair it. The VLM's job becomes much simpler, so a locally served, open-source Qwen with no privileged information is enough.
| Aspect | VLM-led agents | Spotter |
|---|---|---|
| Who drives execution | VLM plans every step | Embodied model leads |
| When the VLM is called | Before every step | Only on a confirmed error |
| Policy waits for the VLM | Every step | Only during a repair |
| Needs privileged information | Usually (e.g. the success signal) | No |
| Works with a local Qwen | Only with privileged information | Yes |
Method
Spotter is a robot-agent framework that reverses the usual roles. Instead of a VLM planning every step, the embodied model (a pretrained robot policy such as π₀.₅) stays in charge. Two VLMs run alongside it: a faster one screens every action chunk, and a frontier one steps in to repair an error only after confirming it, then hands control back.
-
Embodied model
1Act
Executes action chunks C₁, C₂, … (16–25 control steps each) back to back, waiting only if it gets K = 2–4 chunks ahead of the latest check.
-
Faster VLM
2Screen
Checks every finished chunk in parallel, one chunk behind: ✓ keep running, ✗ flag.
-
Frontier VLM
3Judge
After two flags in a row (✗C₂, ✗C₃), reviews the flagged chunks and the episode history while the robot keeps moving (C₅).
-
Frontier VLM
4Repair & reflect
If the error is confirmed (✗), takes over the arm with a few motion primitives (Repair, in place of C₆) and hands back (C₇). It predicts what the next check should see and, if that fails, switches to a different kind of repair.
No scene resets at test time
Spotter never resets the scene at test time. Resets are used only beforehand, in simulation, to learn general lessons that the VLM retrieves later, such as grasping a long bottle in the middle rather than at an end.
Return the arm, not the scene
The judge sees an error only after the arm has moved on. Because Spotter records the arm pose at every chunk, a repair can move the arm back to its pose before the failure and correct the next attempt (e.g. shift toward a missed object), with no scene reset, so it also runs on a real robot. This applies when the failed action leaves the scene intact, as in a missed grasp.
Results
Success rate
On RoboCasa (24 tasks, 1,200 episodes), Spotter raises two frozen policies, Cosmos Policy and π₀.₅, by up to 5.6 and 7.5 points. With one learned example (1-shot) it beats Harness VLA, even when Harness VLA can see the task's success signal.
Cosmos Policy
π₀.₅
π₀.₅, RoboTwin 2.0 Hard
| Method | Cosmos Policy | π₀.₅ |
|---|---|---|
| Embodied model 1.8× step limit | 67.9 | 64.3 |
| Harness VLA (Qwen) sees success signal | 71.3+3.4 | 69.1+4.8 |
| Harness VLA (Qwen) no success signal | 65.2−2.7 | 60.5−3.8 |
| Spotter (Qwen) zero-shot | 71.7+3.8 | 68.4+4.1 |
| Spotter (Qwen) 1-shot | 71.8+3.9 | 70.4+6.1 |
| Spotter (GPT) 1-shot | 73.5+5.6 | 71.8+7.5 |
Success rate (%). Change is against the embodied model at 1.8× the step limit, Spotter's budget; Harness VLA gets 5,000 steps. (Qwen)/(GPT): the VLM used.
Latency
Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes just 13 s (Cosmos Policy) and 16 s (π₀.₅) longer than the model alone, and about 70% less time than Harness VLA with the same VLM.
Real robot
On a Franka Research 3 with π₀.₅, Spotter (GPT, 1-shot) raises success from 53% (16/30) to 83% (25/30). The largest gain is on the long-horizon task, 30% → 70%, where without intervention one mistake in either subtask fails the task.
Lesson library and example bank
Spotter never trains the policy. What it learns lives in two small, frozen artifacts that the frontier VLM retrieves at test time.
77 general lessons were distilled from learning episodes on a held-out seed (195), the only place where the judge may reset and retry. They ship with the code in memory_bank/global/.
The example bank holds worked examples for the 1-shot setting: the frames the judge saw, its verdict and the repair that worked, for both Cosmos Policy and π₀.₅. One matching example is added to the judge's prompt per episode.
Download the example bank (151 MB) into the repository:
curl -L https://github.com/zqc3117/Spotter/releases/download/v1.0/fewshot_bank.tar.gz \ | tar -xz -C recovery_explore
A closed gripper is not a grasp; only the object leaving its support proves one
Closing on nothing, pinching an edge, and toppling the object then closing all look identical in the gripper signal alone.
Repair the moment of the missed grasp, not the retreated pose the policy left you in
The error to correct is two or three centimetres at the moment the fingers shut. Every centimetre of travel afterwards is unrelated to that error.
Three failed repairs on one trajectory means stop and let the policy run out
Each intervention leaves the arm somewhere the policy did not choose. The cost compounds while the chance of rescue does not.
Three of the 77 lessons in memory_bank/global, quoted verbatim. Each file records when it applies, the evidence episodes, a confidence level and what would falsify it.
Citation
@article{li2026spotter,
title = {Spotter: Let the Embodied Model Lead, and the VLM Reflect for It},
author = {Li, Long and Zhao, Qichao and Yang, Yue and Xu, Fan and Wang, Zhe and Liew, Alan Wee-Chung and Qu, Chao and Shen, Heng Tao and Pan, Shirui},
journal = {arXiv preprint arXiv:2609.36808},
year = {2026}
}