Vision–Language Navigation for Rapid UAV Deployment in Emergency Response

Carlos Osorio, Jose Martinez-Carranza

Abstract


GPS-denied search-and-rescue missions
require UAVs to navigate quickly through unfamiliar,
cluttered, and dynamic environments while following
high-level human instructions. This work presents a
vision-and-language navigation framework that allows
operators to issue free-form commands, such as “go
to the far corridor, then approach the victim,” without
predefined waypoint maps. Unlike VLM methods that
directly generate text actions, the proposed approach
formulates navigation as a 2D spatial-grounding task.
From a first-person RGB image and a natural-language
instruction, the model iteratively selects image-grounded
waypoint annotations as short-horizon navigation goals.
These 2D waypoints are fused with the estimated travel
distance to generate 3D displacement commands, while
adaptive step-length control balances efficiency and
safety in constrained spaces. The closed-loop system
continuously replans under moving targets, changing
visibility, and emerging obstacles. The framework is
validated through simulation in high-fidelity GPS-denied
environments representative of indoor SAR scenarios,
achieving an online XY RMSE of 0.098 m. The results
show improved performance over strong baselines
across simulated rescue scenarios and ablation studies.

Keywords


Vision–Language Navigation (VLN), GPS-denied UAV autonomy, 2D/3D spatial grounding, Waypoint-based closed-loop control, Search-and-Rescue (SAR).

Full Text: PDF