Abstract
Deep reinforcement learning (DRL) for mapless goal navigation often suffers from low sample efficiency in environments with complex geometric structure. We propose Structurally Aligned Prioritized Experience Replay (SAPER), a replay prioritization method for structured mapless navigation. SAPER combines conventional TD-error with an action-deviation-based structural signal derived from a coarse training-time reference, allowing weak structural guidance to act through replay selection rather than reward shaping or policy modification. The learned policy still relies only on onboard sensory inputs at inference time. Experiments on navigation benchmarks show competitive performance against representative replay baselines, with the clearest gains in strongly structured environments. Controlled ablations further show that, for the same structural cue in our setting, replay-level use is more effective than reward-level or policy-level counterparts, and that SAPER remains useful under moderate reference corruption.
| Original language | English |
|---|---|
| Title of host publication | Lecture Notes in Computer Science |
| Subtitle of host publication | Advanced Intelligent Computing Technology and Applications (ICIC 2026) |
| Publisher | Springer Nature |
| Pages | 298 |
| Number of pages | 310 |
| Volume | 16669 |
| ISBN (Electronic) | 978-981-92-3492-9 |
| ISBN (Print) | 978-981-92-3491-2 |
| DOIs | |
| Publication status | Published - 14 Jul 2026 |
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver