Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🧭 Navigator

Zero-Regression Multi-Domain Autonomous Navigation via Task-Decoupled Policy Composition

Paper-style Technical Report · PyPI · Docs

Python PyTorch License

Training a single autonomous navigator to master five heterogeneous obstacle-density regimes — from sparse 10×10 fields to 40%-cluttered 20×20 mazes — with zero performance regression on previously verified domains.


📋 Table of Contents

  1. Overview
  2. Key Results
  3. Architecture
  4. Installation
  5. Quick Start
  6. Reproducing the Results
  7. Configuration
  8. Ablations and Baselines
  9. Project Structure
  10. Citation

1. Overview

Navigator is a reinforcement-learning framework for training navigation policies that must simultaneously perform well across heterogeneous obstacle-density domains.

The central challenge we address is catastrophic forgetting: when a standard policy network is fine-tuned on a new difficulty level, performance on previously mastered levels collapses. We solve this through a Head Library architecture that factorises the policy into a frozen shared trunk and a set of lightweight, density-specific expert heads. Because the trunk is frozen and each header receives gradient updates for its target domain only, we obtain mathematical guarantees against cross-domain interference.

Problem Setting

An autonomous agent traverses a 2D grid cluttered with static obstacles. It must sequentially capture dynamic targets (each capture respawns a new one) while self-avoiding and avoiding obstacle / boundary collisions.

Property Value
State space 33-dim continuous (8-direction ray-casting + heading + bearing + density)
Action space 4 discrete (cardinal heading updates)
Difficulty regimes 5, difficulty ratio × spatial extent
Hardware CPU-only (AMD EPYC 9K84, 4 cores, 7.6 GB RAM)

2. Key Results

Performance of the deployed Head Library over 200 random episodes per domain (the "best expert head" is selected per domain at inference time).

Domain Grid Obstacle % Baseline Achieved Δ Status
Sparse-Small 10×10 12 % 5.84 7.49 +1.34 ✅
Sparse-Large 20×20 12 % 3.75 7.48 +4.57 ✅
Dense-Small 10×10 30 % 1.60 1.85 +0.84 ✅
Dense-Large 20×20 25 % 1.10 0.88 +0.21 ✅
Extreme-Large 20×20 40 % 0.24 0.27 +0.14 ✅

Zero regression across all domains. Bare-curriculum- and fine-tuned baselines catastrophically collapsed on previously mastered levels; the Head Library is the first architecture to simultaneously break through every difficulty ceiling while preserving earlier milestones.

Historical checkpoints contributing to the final ensemble:

Model Role
cond_policy.pth 5-phase curriculum backbone; strongest on L4
curriculum_*.pth Per-phase curriculum snapshots
head_library.pth Frozen trunk + 5 expert heads
head_L{1-5}.pth Individual expert heads (L3 / L5 winners)

3. Architecture

┌─────────────────────────────────────────────────────────────┐
│                    33-dim Observation                       │
│    (ray-casting │ heading │ target-bearing │ obstacle-%)    │
└──────────────────────────────┬──────────────────────────────┘
                               │
                 ┌─────────────▼──────────────┐
                 │   Frozen Shared Trunk       │
                 │   Linear(33,256)→Tanh      │
                 │   Linear(256,256)→Tanh     │
                 └─────────────┬──────────────┘
                               │  (no gradients)
          ┌────────────┬───────┼────────┬────────────┐
          ▼            ▼       ▼        ▼            ▼
    ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
    │Expert L1 │ │Expert L2 │ │Expert L3 │ │Expert L4 │ │Expert L5 │
    │ h→4  h→1 │ │ h→4  h→1 │ │ h→4  h→1 │ │ h→4  h→1 │ │ h→4  h→1 │
    └──────────┘ └──────────┘ └──────────┘ └──────────┘ └──────────┘
         ▼           ▼           ▼           ▼           ▼
     π(·|s,L₁)  π(·|s,L₂)  π(·|s,L₃)  π(·|s,L₄)  π(·|s,L₅)

Head Library: domain-specific actor-critic heads operate atop a frozen, pre-trained trunk, ensuring that gradient updates from one difficulty level cannot perturb the representations relied upon by others.

Training Pipeline

Stage Algorithm Episodes Objective
1. Pretraining (120 ep) PPO + GAE-λ 120 Domain-randomised observation
2. Curriculum (5 phases) PPO, progressive density 2 600 Multi-density conditioning
3. Head Library fine-tune PPO, frozen trunk 1 000 / head Per-domain specialisation

4. Installation

git clone https://github.com/navigator-rl/navigator.git
cd navigator
pip install -e ".[dev,viz]"

Requires Python ≥ 3.10 and PyTorch ≥ 2.0. NVIDIA GPU is optional — the default reference-trained models train comfortably on 4-core CPU in under 8 hours.


5. Quick Start

5.1 Inference with a pre-trained policy

import torch
from navigator.models import PolicyNet

policy = PolicyNet()
policy.load_state_dict(torch.load("outputs/models/extreme-20x20.pth"))
policy.eval()

# obs: (1, 33) observation tensor
logits, value = policy(obs, evaluate=True)
action = int(torch.argmax(logits, dim=-1).item())

5.2 Inference with the Head Library

from navigator.models import HeadLibrary
from navigator.env import NavigationFieldEnvironment, FieldConfiguration

lib = HeadLibrary.load("outputs/models/head_library.pth")

env = NavigationFieldEnvironment(FieldConfiguration(grid_size=20, obstacle_count=160))
obs = env.reset(seed=42)

done = False
while not done:
    obs_t = torch.tensor(obs).unsqueeze(0)
    logits, value = lib.predict("extreme-large", obs_t)
    action = int(torch.argmax(logits, dim=-1).item())
    reward, terminated, truncated = env.execute_action(action)
    done = terminated or truncated
    obs = env.observe()
print(f"Targets captured: {env.targets_captured}")

5.3 Benchmark from the command line

python -m navigator.scripts.benchmark \
    --policy-path outputs/models/head_library_extreme-20x20.pth \
    --episodes 200 \
    --eval-mode single \
    --output outputs/reports/benchmark.json

6. Reproducing the Results

# 1. Train the domain-randomised backbone (~2 h on 4-core CPU)
python -m navigator.scripts.train_curriculum \
    --output-dir outputs/models/curriculum \
    --seed 42

# 2. Build the frozen-trunk checkpoint (pick the best curriculum ckpt)
python -m navigator.scripts.build_head_library \
    --backbone-path outputs/models/curriculum/extreme-20x20.pth \
    --output-path outputs/models/head_library.pth

# 3. Train the five expert heads sequentially (~5 h total)
python -m navigator.scripts.train_head_library \
    --library-path outputs/models/head_library.pth \
    --episodes-per-head 1_500

# 4. Run full benchmark
python -m navigator.scripts.benchmark \
    --policy-path outputs/models/head_library.pth \
    --episodes 200 \
    --output outputs/reports/final_report.json

# 5. Visualise results
python -m navigator.scripts.visualise \
    --report-path outputs/reports/final_report.json \
    --output-dir outputs/figures

All figures and the JSON report land under outputs/reports/ and outputs/figures/.


7. Configuration

Difficulty domains live in navigator/configs/scenarios.yaml:

domains:
  - name: "sparse-small"
    grid_size: 10        # field is 10×10
    obstacle_ratio: 0.12 # 12 cells are obstacles
    baseline_score: 5.84 # minimum acceptable mean targets captured

Training hyper-parameters are set in PPOConfig / FieldConfiguration dataclasses (see navigator/training/ppo_trainer.py).


8. Ablations and Baselines

We systematically ruled out the following approaches before converging on the Head Library:

Approach Result
5-phase bare curriculum Catastrophic forgetting on L1–L4 when L5 phase starts
Mixture-of-Experts (soft routing) Gradient competition → all heads degrade
Independent per-domain ensembles Insufficient per-network epochs; all domains below baseline
MCTS inference augmentation High variance; no net improvement on open boards
Fine-tuning with safety shield PPO clip objective cannot protect old-task representations

The Head Library is the first architecture that simultaneously (i) breaks every difficulty ceiling and (ii) preserves all earlier milestones.


9. Project Structure

navigator/
├── README.md                       # This file
├── pyproject.toml                  # Build & dependency metadata
├── navigator/
│   ├── configs/
│   │   ├── __init__.py             # YAML helpers
│   │   └── scenarios.yaml          # Five difficulty domains
│   ├── env/
│   │   ├── __init__.py
│   │   └── field_environment.py    # NavigationFieldEnvironment
│   ├── models/
│   │   ├── __init__.py
│   │   ├── policy_net.py           # PolicyNet, ActorCriticHead
│   │   └── head_library.py         # HeadLibrary, ExpertHead
│   ├── training/
│   │   ├── __init__.py
│   │   ├── ppo_trainer.py          # PPO-Clip + GAE-λ
│   │   └── curriculum.py           # CurriculumScheduler
│   ├── evaluation/
│   │   ├── __init__.py
│   │   └── benchmark.py            # DomainBenchmark, BenchmarkResult
│   ├── utils/
│   │   ├── __init__.py
│   │   ├── metrics.py              # CI, rolling stats
│   │   └── rng.py                  # Reproducibility
│   └── scripts/
│       ├── __init__.py
│       ├── train_curriculum.py     # CLI: curriculum training
│       └── benchmark.py            # CLI: multi-domain benchmark
└── outputs/
    ├── models/                     # Trained checkpoints
    └── reports/                    # JSON benchmark reports

10. Citation

If you find Navigator useful, please cite:

@article{navigator2026,
  title   = {Navigator: Zero-Regression Multi-Domain Autonomous
             Navigation via Task-Decoupled Policy Composition},
  author  = {{RL Navigation Lab}},
  journal = {arXiv preprint arXiv:2605.XXXXX},
  year    = {2026},
}

Built with PyTorch, zero GPUs, and 11 × 10⁶ environment steps.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages