When should an imitation-learned policy raise an alarm?

Ongoing · Aug 2026 –

Clean68% on target
Biased from step 3014% on target
Seed 0, 12 s simulated. On the right every commanded position is offset by 40 px from step 30.

The question

A policy learned from human demonstrations has no internal notion of “this is going badly”. When something disturbs the execution, such as a biased actuator or a noisy position estimate, it keeps producing actions until the task fails or times out. I want to know which signals, computed only from what the robot actually has at run time, detect that a run has gone wrong early enough to act on, and at what false-alarm cost.

Setup

The task is Push-T: a round pusher has to push a T-shaped block onto a fixed target pose. It runs in a 2D rigid-body simulation (Pymunk) at 10 Hz. The policy observes positions and angles directly, not camera images, and the pusher is a free-floating disc, not an arm. I chose this deliberately. It keeps contact dynamics while making every run cheap and exactly repeatable.

  • Demonstrations. 206 human demonstrations, 25,650 frames.
  • Policies. A 1-nearest-neighbour baseline, and a partial reproduction of CFDP, a training-free closed-form diffusion policy that generates an action chunk from nearby demonstrations.
  • Disturbances. A constant offset on commanded positions, noise on the observed state, and displacement of the block. Each is injected at a known step.
  • Paired runs. Every disturbed run has a clean twin from the same seed, so the two trajectories are identical up to the injection step.

Three monitors

  • M1: observation novelty. Distance from the current observation to its nearest demonstrations.
  • M2: plan disagreement. How much two consecutive action chunks disagree about the same future steps.
  • M3: accumulated disagreement. A CUSUM over M2, so that persistent small disagreement adds up.

Monitor scores over one paired run

Monitor scores on one paired run. Blue is the clean twin; orange has a 40 px actuator bias from step 60 (red line). Dotted lines are provisional thresholds. The bottom panel is the simulator's reward, which the monitors never see.

The figure shows why the problem is not trivial. After the bias is injected, M2 jumps, but it also spikes in the clean run. M1 rises but stays far below its threshold. M3 only crosses its threshold near the end, long after the task has effectively failed.

Where it stands

The experimental protocol is fixed: development, calibration and test seeds are disjoint, and injection, first-observable and first-check times are logged separately. All three monitors are implemented and have been examined on development seeds. None is calibrated yet.

Two findings from setting it up:

  • With the bandwidth as I first read it from the paper, CFDP behaves almost exactly like 1-NN: the two agree on success or failure on 192 of 200 seeds. Normalising distances per dimension stops the collapse. With 12-step execution it reaches a mean best-coverage score of 0.70 on development seeds, against 0.80 reported in the paper. I stopped chasing the remaining gap.
  • Disturbances don’t reliably cause failures. A 40 px actuator bias clearly lowers the score in about 45% of paired runs, but small disturbances also “lower” it in 15–20%. Any failure label built from outcomes alone is noisy, and the evaluation has to account for that.

Next

Set each monitor’s threshold on held-out clean runs to a common false-alarm rate, then compare detection delay and missed detections on unseen seeds. Replanning after an alarm comes after that, and only against a fixed-rate replanning baseline with the same compute budget.

What this is not

This is not a vision policy, a real robot, or an arm. It is a small, controlled test of run-time failure detection for imitation-learned policies.