Plann3r: predicting planning costs grounded in 3D
* Equal contribution.
Predicts pixel-level planning costs as geodesic distances from a query image to an arbitrary subgoal pixel, using a VGGT backbone with a learnable goal token and an MLP cost decoder. The VGGT-Nav pipeline runs the same module twice: offline for mapping and global planning, online for localization and local planning that conditions a learnt controller.