RoboCasa · microwave → counter  The policy closes on air; the screener flags it, the judge retreats the arm and hands back, and the policy regrasps the squash. Shown at 1.5×, inference pauses cut.

arXiv 2609.36808 Embodied-led robot agents 2026

Spotter

Let the Embodied Model Lead, and the VLM Reflect for It

Long Li1* Qichao Zhao2* Yue Yang3 Fan Xu4 Zhe Wang1 Alan Wee-Chung Liew1 Chao Qu5 Heng Tao Shen6 Shirui Pan1†
1Griffith University 2IIIS, Tsinghua University 3MainCode 4Tencent 5Fudan University 6Tongji University
*Equal contribution  ·  †Corresponding author

In briefA pretrained robot policy leads. A faster VLM screens every action chunk in parallel, and a frontier VLM steps in only to repair a confirmed error, with no scene reset. Spotter lifts π₀.₅ by up to 7.5 points on RoboCasa and real-robot success from 53% to 83%, in about 70% less time than a VLM-led design.

Scroll to explore
See it work

Demos

Each recording shows the cameras the system sees, what the screener and the judge said at that moment, and a timeline of who was busy: the policy, the screener, or the judge. Each RoboCasa clip runs the same episode twice, side by side: the policy alone on the left, the policy with Spotter on the right.

Real robot · up to 5×
Franka Research 3 · “Stack the blue cup onto the pink cup.” The hand stalls 2 cm above the cup with one finger over the rim. On the second alarm in a row, the Qwen screener calls the GPT judge, which measures in the wrist view that the cup sits 2.6 cm to the left of the fingers, nudges the hand sideways and hands back; the policy finishes the stack. 56 s of real time, shown at up to 5×.
RoboCasa · up to 5×
Steak · counter → cabinet. The gripper closes on nothing and the arm backs away with the steak still on the counter. The Qwen screener flags the empty grasp; the judge retreats the arm to the pre-grasp point, sees the steak just below in the wrist view, lowers the hand 3 cm and hands back, and the policy puts the steak in the cabinet. The policy alone fails.
RoboCasa · up to 5×
Jug · counter → cabinet. The gripper closes all the way on nothing and lifts away, leaving the jug on the counter. The judge retreats the arm to the pre-grasp point, lowers the hand 2 cm and hands back; this time the policy closes on the jug, a follow-up check confirms the grasp, and the jug goes into the cabinet. The policy alone fails.
RoboCasa · up to 5×
Potato · plate → pan. The policy closes on nothing and backs off with the potato still on the plate. The screener calls the judge, which retreats the arm to the pre-grasp point, lowers the hand 2 cm and hands back; the policy picks the potato up and puts it in the pan. The policy alone fails.

Clips are cut for viewing: stretches marked ×5 on screen are fast-forwarded, and in RoboCasa the simulator's render and save time is removed. The RoboCasa clock is robot time.

Why embodied-led

Supervise in parallel, not in series

VLM-led robot agents are serial: the VLM plans a segment, the policy executes it, and the VLM decides what comes next. That forces a trade-off. Short segments mean a VLM call at every step and high latency even on easy tasks. Long segments mean a deviation cannot be caught in time, and the actions that follow keep going wrong.

Spotter removes the trade-off by changing who leads. The embodied model keeps acting. A faster VLM screens each finished chunk in parallel, and the frontier VLM is called only after two flags in a row, to confirm the error and repair it. The VLM's job becomes much simpler, so a locally served, open-source Qwen with no privileged information is enough.

AspectVLM-led agentsSpotter
Who drives executionVLM plans every stepEmbodied model leads
When the VLM is calledBefore every stepOnly on a confirmed error
Policy waits for the VLMEvery stepOnly during a repair
Needs privileged informationUsually (e.g. the success signal)No
Works with a local QwenOnly with privileged informationYes
How it works

Method

Spotter is a robot-agent framework that reverses the usual roles. Instead of a VLM planning every step, the embodied model (a pretrained robot policy such as π₀.₅) stays in charge. Two VLMs run alongside it: a faster one screens every action chunk, and a frontier one steps in to repair an error only after confirming it, then hands control back.

(a) Alone, the embodied model does not notice its errors. (b) VLM-led designs such as Harness VLA call the frontier VLM before every step, and the robot idles (grey) while it waits. (c) Spotter: the embodied model keeps acting while the VLMs check in parallel, and gives way only for the repair in place of C₆. Steps 1–4 below walk through this timeline.
  1. Embodied model

    1Act

    Executes action chunks C₁, C₂, … (16–25 control steps each) back to back, waiting only if it gets K = 2–4 chunks ahead of the latest check.

  2. Faster VLM

    2Screen

    Checks every finished chunk in parallel, one chunk behind: ✓ keep running, ✗ flag.

  3. Frontier VLM

    3Judge

    After two flags in a row (✗C₂, ✗C₃), reviews the flagged chunks and the episode history while the robot keeps moving (C₅).

  4. Frontier VLM

    4Repair & reflect

    If the error is confirmed (✗), takes over the arm with a few motion primitives (Repair, in place of C₆) and hands back (C₇). It predicts what the next check should see and, if that fails, switches to a different kind of repair.

Reset-free

No scene resets at test time

Spotter never resets the scene at test time. Resets are used only beforehand, in simulation, to learn general lessons that the VLM retrieves later, such as grasping a long bottle in the middle rather than at an end.

Arm-level repair

Return the arm, not the scene

The judge sees an error only after the arm has moved on. Because Spotter records the arm pose at every chunk, a repair can move the arm back to its pose before the failure and correct the next attempt (e.g. shift toward a missed object), with no scene reset, so it also runs on a real robot. This applies when the failed action leaves the scene intact, as in a missed grasp.

RoboCasa · RoboTwin 2.0 · Franka Research 3

Results

Success rate

On RoboCasa (24 tasks, 1,200 episodes), Spotter raises two frozen policies, Cosmos Policy and π₀.₅, by up to 5.6 and 7.5 points. With one learned example (1-shot) it beats Harness VLA, even when Harness VLA can see the task's success signal.

Cosmos Policy

Embodied model1.8× step limit67.9
Harness VLA (Qwen)sees success signal71.3
Harness VLA (Qwen)no success signal65.2
Spotter (Qwen)zero-shot71.7
Spotter (Qwen)1-shot71.8
Spotter (GPT)1-shot73.5

π₀.₅

Embodied model1.8× step limit64.3
Harness VLA (Qwen)sees success signal69.1
Harness VLA (Qwen)no success signal60.5
Spotter (Qwen)zero-shot68.4
Spotter (Qwen)1-shot70.4
Spotter (GPT)1-shot71.8
SpotterBaselineSuccess rate (%), 24 tasks × 50 episodes, seed 500. Bars start at 0.
MethodCosmos Policyπ₀.₅
Embodied model 1.8× step limit67.964.3
Harness VLA (Qwen) sees success signal71.3+3.469.1+4.8
Harness VLA (Qwen) no success signal65.2−2.760.5−3.8
Spotter (Qwen) zero-shot71.7+3.868.4+4.1
Spotter (Qwen) 1-shot71.8+3.970.4+6.1
Spotter (GPT) 1-shot73.5+5.671.8+7.5

Success rate (%). Change is against the embodied model at 1.8× the step limit, Spotter's budget; Harness VLA gets 5,000 steps. (Qwen)/(GPT): the VLM used.

Versus other embodied models. Ours: Spotter, 1-shot at 1.8× the step limit, on Cosmos Policy for RoboCasa and on π₀.₅ for RoboTwin 2.0 (Hard), where it lifts π₀.₅ from 47.2% to 51.2% with Qwen and 57.0% with GPT. RoboCasa baselines are reported numbers at the official limit.

Latency

Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes just 13 s (Cosmos Policy) and 16 s (π₀.₅) longer than the model alone, and about 70% less time than Harness VLA with the same VLM.

Wall-clock time per episode on RoboCasa. Spotter runs zero-shot; Harness VLA (with the success signal) blocks execution at every round.

Real robot

On a Franka Research 3 with π₀.₅, Spotter (GPT, 1-shot) raises success from 53% (16/30) to 83% (25/30). The largest gain is on the long-horizon task, 30% → 70%, where without intervention one mistake in either subtask fails the task.

Long-horizon task: one cube into the cup, then another onto the plate.
What the VLM learns from

Lesson library and example bank

Spotter never trains the policy. What it learns lives in two small, frozen artifacts that the frontier VLM retrieves at test time.

77 general lessons were distilled from learning episodes on a held-out seed (195), the only place where the judge may reset and retry. They ship with the code in memory_bank/global/.

The example bank holds worked examples for the 1-shot setting: the frames the judge saw, its verdict and the repair that worked, for both Cosmos Policy and π₀.₅. One matching example is added to the judge's prompt per episode.

Download the example bank (151 MB) into the repository:

curl -L https://github.com/zqc3117/Spotter/releases/download/v1.0/fewshot_bank.tar.gz \
  | tar -xz -C recovery_explore
Lesson · perception · verified

A closed gripper is not a grasp; only the object leaving its support proves one

Closing on nothing, pinching an edge, and toppling the object then closing all look identical in the gripper signal alone.

Lesson · strategy · probable

Repair the moment of the missed grasp, not the retreated pose the policy left you in

The error to correct is two or three centimetres at the moment the fingers shut. Every centimetre of travel afterwards is unrelated to that error.

Lesson · strategy · probable

Three failed repairs on one trajectory means stop and let the policy run out

Each intervention leaves the arm somewhere the policy did not choose. The cost compounds while the chance of rescue does not.

Three of the 77 lessons in memory_bank/global, quoted verbatim. Each file records when it applies, the evidence episodes, a confidence level and what would falsify it.

Citation

@article{li2026spotter,
  title   = {Spotter: Let the Embodied Model Lead, and the VLM Reflect for It},
  author  = {Li, Long and Zhao, Qichao and Yang, Yue and Xu, Fan and Wang, Zhe and Liew, Alan Wee-Chung and Qu, Chao and Shen, Heng Tao and Pan, Shirui},
  journal = {arXiv preprint arXiv:2609.36808},
  year    = {2026}
}