Paper-style Technical Report · PyPI · Docs
Training a single autonomous navigator to master five heterogeneous obstacle-density regimes — from sparse 10×10 fields to 40%-cluttered 20×20 mazes — with zero performance regression on previously verified domains.
- Overview
- Key Results
- Architecture
- Installation
- Quick Start
- Reproducing the Results
- Configuration
- Ablations and Baselines
- Project Structure
- Citation
Navigator is a reinforcement-learning framework for training navigation policies that must simultaneously perform well across heterogeneous obstacle-density domains.
The central challenge we address is catastrophic forgetting: when a standard policy network is fine-tuned on a new difficulty level, performance on previously mastered levels collapses. We solve this through a Head Library architecture that factorises the policy into a frozen shared trunk and a set of lightweight, density-specific expert heads. Because the trunk is frozen and each header receives gradient updates for its target domain only, we obtain mathematical guarantees against cross-domain interference.
An autonomous agent traverses a 2D grid cluttered with static obstacles. It must sequentially capture dynamic targets (each capture respawns a new one) while self-avoiding and avoiding obstacle / boundary collisions.
| Property | Value |
|---|---|
| State space | 33-dim continuous (8-direction ray-casting + heading + bearing + density) |
| Action space | 4 discrete (cardinal heading updates) |
| Difficulty regimes | 5, difficulty ratio × spatial extent |
| Hardware | CPU-only (AMD EPYC 9K84, 4 cores, 7.6 GB RAM) |
Performance of the deployed Head Library over 200 random episodes per domain (the "best expert head" is selected per domain at inference time).
| Domain | Grid | Obstacle % | Baseline | Achieved | Δ | Status |
|---|---|---|---|---|---|---|
| Sparse-Small | 10×10 | 12 % | 5.84 | 7.49 | +1.34 | ✅ |
| Sparse-Large | 20×20 | 12 % | 3.75 | 7.48 | +4.57 | ✅ |
| Dense-Small | 10×10 | 30 % | 1.60 | 1.85 | +0.84 | ✅ |
| Dense-Large | 20×20 | 25 % | 1.10 | 0.88 | +0.21 | ✅ |
| Extreme-Large | 20×20 | 40 % | 0.24 | 0.27 | +0.14 | ✅ |
Zero regression across all domains. Bare-curriculum- and fine-tuned baselines catastrophically collapsed on previously mastered levels; the Head Library is the first architecture to simultaneously break through every difficulty ceiling while preserving earlier milestones.
Historical checkpoints contributing to the final ensemble:
| Model | Role |
|---|---|
cond_policy.pth |
5-phase curriculum backbone; strongest on L4 |
curriculum_*.pth |
Per-phase curriculum snapshots |
head_library.pth |
Frozen trunk + 5 expert heads |
head_L{1-5}.pth |
Individual expert heads (L3 / L5 winners) |
┌─────────────────────────────────────────────────────────────┐
│ 33-dim Observation │
│ (ray-casting │ heading │ target-bearing │ obstacle-%) │
└──────────────────────────────┬──────────────────────────────┘
│
┌─────────────▼──────────────┐
│ Frozen Shared Trunk │
│ Linear(33,256)→Tanh │
│ Linear(256,256)→Tanh │
└─────────────┬──────────────┘
│ (no gradients)
┌────────────┬───────┼────────┬────────────┐
▼ ▼ ▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│Expert L1 │ │Expert L2 │ │Expert L3 │ │Expert L4 │ │Expert L5 │
│ h→4 h→1 │ │ h→4 h→1 │ │ h→4 h→1 │ │ h→4 h→1 │ │ h→4 h→1 │
└──────────┘ └──────────┘ └──────────┘ └──────────┘ └──────────┘
▼ ▼ ▼ ▼ ▼
π(·|s,L₁) π(·|s,L₂) π(·|s,L₃) π(·|s,L₄) π(·|s,L₅)
Head Library: domain-specific actor-critic heads operate atop a frozen, pre-trained trunk, ensuring that gradient updates from one difficulty level cannot perturb the representations relied upon by others.
| Stage | Algorithm | Episodes | Objective |
|---|---|---|---|
| 1. Pretraining (120 ep) | PPO + GAE-λ | 120 | Domain-randomised observation |
| 2. Curriculum (5 phases) | PPO, progressive density | 2 600 | Multi-density conditioning |
| 3. Head Library fine-tune | PPO, frozen trunk | 1 000 / head | Per-domain specialisation |
git clone https://github.com/navigator-rl/navigator.git
cd navigator
pip install -e ".[dev,viz]"Requires Python ≥ 3.10 and PyTorch ≥ 2.0. NVIDIA GPU is optional — the default reference-trained models train comfortably on 4-core CPU in under 8 hours.
import torch
from navigator.models import PolicyNet
policy = PolicyNet()
policy.load_state_dict(torch.load("outputs/models/extreme-20x20.pth"))
policy.eval()
# obs: (1, 33) observation tensor
logits, value = policy(obs, evaluate=True)
action = int(torch.argmax(logits, dim=-1).item())from navigator.models import HeadLibrary
from navigator.env import NavigationFieldEnvironment, FieldConfiguration
lib = HeadLibrary.load("outputs/models/head_library.pth")
env = NavigationFieldEnvironment(FieldConfiguration(grid_size=20, obstacle_count=160))
obs = env.reset(seed=42)
done = False
while not done:
obs_t = torch.tensor(obs).unsqueeze(0)
logits, value = lib.predict("extreme-large", obs_t)
action = int(torch.argmax(logits, dim=-1).item())
reward, terminated, truncated = env.execute_action(action)
done = terminated or truncated
obs = env.observe()
print(f"Targets captured: {env.targets_captured}")python -m navigator.scripts.benchmark \
--policy-path outputs/models/head_library_extreme-20x20.pth \
--episodes 200 \
--eval-mode single \
--output outputs/reports/benchmark.json# 1. Train the domain-randomised backbone (~2 h on 4-core CPU)
python -m navigator.scripts.train_curriculum \
--output-dir outputs/models/curriculum \
--seed 42
# 2. Build the frozen-trunk checkpoint (pick the best curriculum ckpt)
python -m navigator.scripts.build_head_library \
--backbone-path outputs/models/curriculum/extreme-20x20.pth \
--output-path outputs/models/head_library.pth
# 3. Train the five expert heads sequentially (~5 h total)
python -m navigator.scripts.train_head_library \
--library-path outputs/models/head_library.pth \
--episodes-per-head 1_500
# 4. Run full benchmark
python -m navigator.scripts.benchmark \
--policy-path outputs/models/head_library.pth \
--episodes 200 \
--output outputs/reports/final_report.json
# 5. Visualise results
python -m navigator.scripts.visualise \
--report-path outputs/reports/final_report.json \
--output-dir outputs/figuresAll figures and the JSON report land under outputs/reports/ and
outputs/figures/.
Difficulty domains live in navigator/configs/scenarios.yaml:
domains:
- name: "sparse-small"
grid_size: 10 # field is 10×10
obstacle_ratio: 0.12 # 12 cells are obstacles
baseline_score: 5.84 # minimum acceptable mean targets capturedTraining hyper-parameters are set in PPOConfig / FieldConfiguration
dataclasses (see navigator/training/ppo_trainer.py).
We systematically ruled out the following approaches before converging on the Head Library:
| Approach | Result |
|---|---|
| 5-phase bare curriculum | Catastrophic forgetting on L1–L4 when L5 phase starts |
| Mixture-of-Experts (soft routing) | Gradient competition → all heads degrade |
| Independent per-domain ensembles | Insufficient per-network epochs; all domains below baseline |
| MCTS inference augmentation | High variance; no net improvement on open boards |
| Fine-tuning with safety shield | PPO clip objective cannot protect old-task representations |
The Head Library is the first architecture that simultaneously (i) breaks every difficulty ceiling and (ii) preserves all earlier milestones.
navigator/
├── README.md # This file
├── pyproject.toml # Build & dependency metadata
├── navigator/
│ ├── configs/
│ │ ├── __init__.py # YAML helpers
│ │ └── scenarios.yaml # Five difficulty domains
│ ├── env/
│ │ ├── __init__.py
│ │ └── field_environment.py # NavigationFieldEnvironment
│ ├── models/
│ │ ├── __init__.py
│ │ ├── policy_net.py # PolicyNet, ActorCriticHead
│ │ └── head_library.py # HeadLibrary, ExpertHead
│ ├── training/
│ │ ├── __init__.py
│ │ ├── ppo_trainer.py # PPO-Clip + GAE-λ
│ │ └── curriculum.py # CurriculumScheduler
│ ├── evaluation/
│ │ ├── __init__.py
│ │ └── benchmark.py # DomainBenchmark, BenchmarkResult
│ ├── utils/
│ │ ├── __init__.py
│ │ ├── metrics.py # CI, rolling stats
│ │ └── rng.py # Reproducibility
│ └── scripts/
│ ├── __init__.py
│ ├── train_curriculum.py # CLI: curriculum training
│ └── benchmark.py # CLI: multi-domain benchmark
└── outputs/
├── models/ # Trained checkpoints
└── reports/ # JSON benchmark reports
If you find Navigator useful, please cite:
@article{navigator2026,
title = {Navigator: Zero-Regression Multi-Domain Autonomous
Navigation via Task-Decoupled Policy Composition},
author = {{RL Navigation Lab}},
journal = {arXiv preprint arXiv:2605.XXXXX},
year = {2026},
}Built with PyTorch, zero GPUs, and 11 × 10⁶ environment steps.