A navigation agent can follow a plausible route while drifting away from the actual instruction. Continuing autonomously may compound that mistake. AVERT-VLN uses a separate vision-language monitor to assess execution, request human correction when the controller accepts a lost-state verdict, and turn those corrections into later training data.
Detecting instruction conflicts alongside navigation
The monitor receives the instruction, visual history and current observation. It runs alongside the navigation model rather than replacing that model's action-generation process. Navigation continues while monitoring requests are pending.
The researchers distinguish semantic deviations from distance to a reference path. A valid alternative route may stray from the reference, while a wrong target can remain nearby. Their LostNav dataset introduces controlled region- and object-level target-selection errors, validates the interventions, and rolls them out in a simulator.
Monitor training starts with 40,000 normal trajectories, then combines them with 20,000 counterfactual risk trajectories. Privileged simulator states help construct and verify that data but are not supplied as monitor inputs. This distinction matters when judging whether the monitor could use the observations available during ordinary execution.
Recovery starts from the current state
When the controller commits a Lost verdict, it halts unexecuted actions and invalidates pending monitoring requests for that execution. Previously executed actions are not rolled back. A remote human supplies a pixel goal in the current view or a local turning command; navigation resumes after the correction produces a new observation.
The same interactions support offline preference learning. The method traces a deviation to the nearest preceding boundary judged consistent with the instruction. It pairs a correction-derived preferred decision with the original output under a shared instruction and observation context.
Training targets the relevant decision tokens, such as direction selection or visual grounding. It does not treat an entire unsuccessful trajectory as uniformly wrong. The paper uses a decision-focused objective based on direct preference optimization for DualVLN's higher-level policy.
Human assistance is part of the reported result
On the unseen-environment validation splits of R2R-CE and RxR-CE, the full configuration reaches success rates of 76.2% and 66.3%. Those figures include monitoring and human-assisted recovery, together with the trained navigation system. They are not fully autonomous success rates.
Against DualVLN with the same monitoring and recovery interface, additional training raises success by 0.7 percentage points on each benchmark. The larger improvements against unassisted systems therefore should not all be attributed to preference learning.
The monitor also makes mistakes. On the 2,190-state LostAware evaluation, it records 63.38% recall, 59.52% F1 and 56.10% precision. These figures describe a trade-off between catching deviations and requesting help unnecessarily.
Recovery can change the path as well as task completion. One tested architecture improves success while its RxR-CE trajectory-agreement score falls. AVERT-VLN supplies a concrete recovery-and-learning loop, but deployment would still need to account for missed deviations, monitoring delay and the availability of corrective human input.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!