TY - GEN
T1 - A Q-learning-based Algorithm for UAV Path Planning under Obstacle Constraints and Wind Disturbances
AU - Zeng, Zhihan
AU - Dong, Qian
AU - Boateng, Gordon Owusu
AU - Yu, Limin
AU - Li, Ji
AU - Xia, Jinbao
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Unmanned aerial vehicles (UAVs) require reliable and adaptive path planning algorithms to operate in complex environments. However, accurately finding an optimal collision-free path remains a significant challenge, particularly in the presence of environmental disturbances and limited prior knowledge. Conventional graph search methods rely on complete environment knowledge and are difficult to extend to unknown or dynamic scenarios. In this paper, we formulate grid-based UAV path planning in obstacle-rich environments with wind disturbances as a markov decision process (MDP). We propose a model-free Q-learning-based framework that enables the UAV to learn the collision-free shortest paths through interaction with the environment. Our method employs a four-direction action space and a reward design that balances path efficiency and obstacle avoidance under wind disturbances. Without assuming prior knowledge of the environment dynamics, the UAV incrementally updates its action-value function using temporal-difference (TD) learning. Simulation results demonstrate that the learned policy reduces the average path length by 47.3% and 43.5% relative to the two benchmark methods of Dijkstra's and A*, respectively. In addition, our proposed approach improves the success rate by 63.3% and 53.0% under stochastic wind conditions compared to the two benchmark methods of Dijkstra's and A∗ algorithms, respectively. These results highlight the effectiveness of reinforcement learning (RL) for UAV navigation in uncertain and dynamic environments.
AB - Unmanned aerial vehicles (UAVs) require reliable and adaptive path planning algorithms to operate in complex environments. However, accurately finding an optimal collision-free path remains a significant challenge, particularly in the presence of environmental disturbances and limited prior knowledge. Conventional graph search methods rely on complete environment knowledge and are difficult to extend to unknown or dynamic scenarios. In this paper, we formulate grid-based UAV path planning in obstacle-rich environments with wind disturbances as a markov decision process (MDP). We propose a model-free Q-learning-based framework that enables the UAV to learn the collision-free shortest paths through interaction with the environment. Our method employs a four-direction action space and a reward design that balances path efficiency and obstacle avoidance under wind disturbances. Without assuming prior knowledge of the environment dynamics, the UAV incrementally updates its action-value function using temporal-difference (TD) learning. Simulation results demonstrate that the learned policy reduces the average path length by 47.3% and 43.5% relative to the two benchmark methods of Dijkstra's and A*, respectively. In addition, our proposed approach improves the success rate by 63.3% and 53.0% under stochastic wind conditions compared to the two benchmark methods of Dijkstra's and A∗ algorithms, respectively. These results highlight the effectiveness of reinforcement learning (RL) for UAV navigation in uncertain and dynamic environments.
KW - obstacle avoidance
KW - Q-learning
KW - reinforcement learning (RL)
KW - unmanned aerial vehicles (UAVs)
KW - wind disturbance
UR - https://www.scopus.com/pages/publications/105044693186
U2 - 10.1109/IWCMC69287.2026.11579965
DO - 10.1109/IWCMC69287.2026.11579965
M3 - Conference Proceeding
AN - SCOPUS:105044693186
T3 - 2026 International Wireless Communications and Mobile Computing Conference, IWCMC 2026
SP - 1047
EP - 1052
BT - 2026 International Wireless Communications and Mobile Computing Conference, IWCMC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 22nd International Wireless Communications and Mobile Computing Conference, IWCMC 2026
Y2 - 1 June 2026 through 6 June 2026
ER -