Skip to content

Evaluation

All commands run from the repository root:

export PLANN3R_ROOT=/absolute/path/to/plann3r-release
cd "$PLANN3R_ROOT/plann3r-code"

Reported rows

Oracle localization, all four tasks:

GPU=0 \
ABLATIONS=paper \
TASKS="imitate reverse altgoal shortcut" \
bash baseline/evaluate.sh

The same planner with the controller trained on ground-truth costmaps:

GPU=0 \
ABLATIONS=paper \
CONTROLLER_CONFIG_FILE=configs/controller/gt_trained_costmap.yaml \
TASKS="imitate reverse altgoal shortcut" \
bash baseline/evaluate.sh

MegaLoc retrieval with the oracle 1 m stop:

GPU=0 \
ABLATIONS=megaloc \
TASKS="imitate reverse altgoal shortcut" \
bash baseline/evaluate.sh

MegaLoc retrieval with inferred stopping:

GPU=0 \
ABLATIONS=no-oracle-paper \
TASKS="imitate reverse altgoal shortcut" \
bash baseline/evaluate.sh

Every mode uses the planner models/planner/checkpoint_best.pt, the predicted-costmap controller models/controller/predicted_costmap/latest.pth unless overridden, the Plann3r propagation map costmaps (method.md), 300 steps, a 0.4 m camera height during navigation, and a 1.31 m camera height in the mapping run. The oracle and megaloc modes count an episode as a success when the geodesic distance to the goal is below 1 m. The no-oracle-paper mode lets the agent run until it stops itself, then scores the stop decision against the same 1 m radius.

Launcher options

Variable Default Meaning
GPU 0 CUDA device index
ABLATIONS paper Space-separated modes
TASKS all four Any of imitate reverse altgoal shortcut
CONTROLLER_CONFIG_FILE empty Controller config to load instead of the default
ORACLE_MODE legacy Ground-truth localization rule, legacy (reported) or odometry
MEGALOC_GOAL_LOCK false Keep the MegaLoc submap on the goal frame once it is the top match (not used for reported results)
PROP_MAP_TAG empty Read propagation costmaps built by another planner checkpoint, <name>_<tag>.npy
MAX_STEPS 300 Episode step budget
VISUALIZE false true also writes per-step frames and a video
SUMMARIZE true Write the summary table after all runs
RESULTS_ROOT $PLANN3R_ROOT/runs/ablation_gt Output root
DRY_RUN false Print the resolved commands without running them

setup.md lists the path overrides. The resolved settings of each run are saved in its config.yaml.

Modes

Mode Planner checkpoint Localization Stop
paper checkpoint_best.pt ground-truth pose oracle, 1 m
megaloc checkpoint_best.pt MegaLoc retrieval oracle, 1 m
no-oracle-paper checkpoint_best.pt MegaLoc retrieval inferred
costmap_only ablations/costmap_only.pt ground-truth pose oracle, 1 m
no_pointmap_loss ablations/no_pointmap_loss.pt ground-truth pose oracle, 1 m
no_grad_loss ablations/no_grad_loss.pt ground-truth pose oracle, 1 m
frozen_mlp_goal_token ablations/frozen_mlp_goal_token.pt ground-truth pose oracle, 1 m

Planner ablations

Every ablation navigates with the full planner's released propagation maps, so only the online planner changes:

GPU=0 \
ABLATIONS="paper costmap_only no_pointmap_loss no_grad_loss frozen_mlp_goal_token" \
TASKS="imitate reverse altgoal shortcut" \
bash baseline/evaluate.sh

Summaries are written to:

$RESULTS_ROOT/ablation_summary.csv
$RESULTS_ROOT/ablation_summary.md

Reproducibility

Checkpoint equality does not make two closed-loop runs identical. A small numerical difference in one step changes the next observation, and the change accumulates over up to 300 steps. PyTorch version, CUDA version, GPU architecture, and the Habitat build all cause such differences. The evaluation has no random sampling, so one machine and one software environment give the same result on every run. Compare per-episode outcomes only within one machine and environment, and report the GPU with the numbers.

Planning cost benchmark (MARD)

mard_benchmark/ computes the Mean Absolute Rank Difference between predicted and ground-truth geodesic costmaps (paper Table 1). Ground truth comes from Habitat NavMesh geodesic distances. Each costmap is rank-transformed within the image, the lowest cost getting rank 0 and the highest rank 1. MARD is the mean absolute rank difference over the selected pixels. Pixel selection is either the lowest k percent of cost (cumulative) or the band between two cost percentiles (slab), for k in 5, 15, 30, 50, and 100.

Run the four stages in order. The extractors write under $PLANN3R_ROOT/evaluation/mard. The navmesh stage must finish first because the other two read its anchor pixels.

cd "$PLANN3R_ROOT/plann3r-code/mard_benchmark"
EPISODES="$PLANN3R_ROOT/plann3r-code/episodes_removing_blacklist.txt"
MARD="$PLANN3R_ROOT/evaluation/mard"

# 1. NavMesh geodesic ground truth and anchor pixels
PYTHONNOUSERSITE=1 MAGNUM_LOG=quiet HABITAT_SIM_LOG=quiet \
  pixi run python navmesh_extractor.py \
  --scene-list "$EPISODES" --workers 8 --export-stack --skip-existing \
  --output_dir "$MARD"

# 2. Euclidean distance between VGGT points
PYTHONNOUSERSITE=1 pixi run python euclidean_extractor.py \
  --scene-list "$EPISODES" --navmesh-root "$MARD" --output_dir "$MARD"

# 3. Plann3r predictions
PYTHONNOUSERSITE=1 pixi run python vggtnav_extractor.py \
  --scene-list "$EPISODES" --navmesh-root "$MARD" --output_dir "$MARD"

# 4. Scores
PYTHONNOUSERSITE=1 pixi run python compute_iou_costmaps.py --output_dir "$MARD"

compute_iou_costmaps.py writes iou_overall.csv (cumulative) and iou_overall_slabs.csv (slab), plus per-trajectory and per-pair files. The mean_rank_mae column is MARD on a 0 to 1 scale. Multiply by 10 to compare with the paper, which reports ranks on a 0 to 10 scale. A slab with no pixels in a costmap is skipped for that costmap.

ObjectReact costmaps are not produced by these stages. They come precomputed from ObjectReact's own pipeline, one costmap per frame, and are read from --objectreact-root (default $PLANN3R_ROOT/evaluation/objectreact_costmaps) at the same query frames as the other methods. ObjectReact is skipped if that directory does not exist.

Output structure

$RESULTS_ROOT/<mode>/<task>/
  <timestamp>.console.log
  <task-layout>/<run>/
    config.yaml
    metrics_summary.txt
    results_summary.csv
    <episode>_vggt_nav_topological_pixelwise/
      results.csv
      step_data/

no-oracle-paper runs also write stopping_predictions.csv and stopping_metrics.json in the run folder. A nonzero exit from run_nav.py stops the launcher, and so does a missing controller checkpoint. An episode that fails during setup is recorded as a failure in the run output. After each run the launcher checks that the saved config.yaml has the requested goal method, map directory, and localizer.

Visualization

VISUALIZE=true adds one PNG per navigation step (frames/step_0000.png, frames/step_0001.png, and so on) and one vggt_nav.mp4 to each episode folder. Each frame contains:

  • Current query RGB image.
  • All eight localized submap images.
  • Selected submap frame and goal pixel.
  • Top-down map with mapping trajectory, executed trajectory, current pose, and goal.
  • Plann3r query costmap with its finite range and minimum.
  • Controller waypoint prediction in the robot frame.
  • Step, distance, velocity, collision, localization, anchor, and stopping status.

The costmap white cross marks the minimum predicted value. The selected submap image has a yellow border and the selected pixel is drawn on that image. Rendering adds CPU image encoding and several GB of frames per task, so it is off by default and no reported result used it.