Bridging the Simulation-to-Reality Gap in Reinforcement Learning-Based Autonomous Robot Navigation | IJCT Volume 13 – Issue 4 | IJCT-V13I4P14

International Journal of Computer Techniques
ISSN 2394-2231
Volume 13, Issue 4  |  Published: July – August 2026

Author

Cecilia Augustine Abi, Joe Essien, Amaku Effiong Amaku

Abstract

Reinforcement learning (RL) is applied to many autonomous robots in simulation before deploying them into the real world; however, the gap between simulation and reality is still a key challenge for many of these applications. RL algorithms have been shown to perform well in simulated environments but have shown poor performance when they are applied to a physical robot platform because of discrepancies in the behavior of the sensors, dynamics of the actuators, changes in environments, and modeling inaccuracies. The current study introduces a novel transfer framework for autonomous robot navigation using RL algorithms that combines the integration of domain randomization, sensor noise modeling, adaptive policy refinement, and multi-sensor fusion. The framework is developed in the Webots simulation environment and tested on a mobile robot built with the Raspberry Pi microcontroller and its various sensors including ultrasonic sensors, an inertial measurement unit, wheel encoders, infrared sensors and a camera module. Soft Actor-Critic (SAC) was used as the reinforcement learning algorithm for training. The Sim-to-Real transfer retention in the simulation experiments was 91% and in the physical-robot experiments was 84%, showing a 92.31% transfer retention between the simulations and reality. The framework was also able to achieve 7% collision rate and 92% generalization in unseen environments, outperforming the Standard SAC and Safety SAC baselines. The results clearly show that adaptive reinforcement learning could significantly improve the application of autonomous robotic systems, especially with the aid of complementary sim-to-real transfer strategies, in dynamic environments.

Keywords

Simulation to reality transfer, Reinforcement learning, Autonomous navigation, Domain randomization, multi-sensor fusion, Soft Actor-Critic.

Conclusion

Using ARL robots in real world applications is still hindered by the ubiquitous Simulation-to-Reality gap. To solve this problem, this study proposed an integrated transfer framework based on the domain randomization, sensor noise modeling, adaptive policy refinement, and multi-sensor fusion to facilitate the transfer of the SAC-based autonomous robot navigation from webots to a physical robot with ultrasonic sensors, IMU, wheel encoders, infrared sensors, and a camera module, which was implemented in Webots and investigated on a physical robot. The reinforcement learning agent converged after about 380 training episodes, which is faster than both Standard SAC (600) and Safety SAC (520) agents and the agent suited to the physical setup with a 84% navigation success rate and a 92.31% Sim-to-Real transfer retention rate after about 370 episodes. The framework also achieved the lowest collision rate of 7%, with the highest generalization performance (92%) on unseen environments, so it is also optimal for handling new scenarios. Ablation results showed that domain randomization was the main factor for generalization, and sensor-noise modeling was the main factor for transfer retention, which proved that the two were not mutually exclusive, neither was one of them more effective than the other, and combination of multiple transfer strategies resulted in a more significant effect than a single one.

References

[1] M. Andrychowicz et al., “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research, vol. 39, no. 1, pp. 3–20, 2020. [2] K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Processing Magazine, vol. 34, no. 6, pp. 26–38, 2017. [3] K. Bousmalis et al., “Using simulation and domain adaptation to improve efficiency of deep robotic grasping,” in Proc. IEEE Int. Conf. on Robotics and Automation, 2018, pp. 4243–4250. [4] X. Chen, Y. Wang, H. Li, and Q. Zhang, “Simulation-to-reality transfer challenges in autonomous robotic systems: A comprehensive review,” Robotics and Autonomous Systems, vol. 184, art. 104821, 2025. [verify] [5] G. Dulac-Arnold, D. Mankowitz, and T. Hester, “Challenges of real-world reinforcement learning,” Communications of the ACM, vol. 64, no. 3, pp. 76–84, 2021. [6] A. Ferreira and F. Barbosa, “Experimental validation of reinforcement learning policies on autonomous mobile robots,” Journal of Intelligent & Robotic Systems, vol. 112, no. 2, pp. 45–63, 2025. [verify] [7] D. Fox, W. Burgard, and S. Thrun, “The dynamic window approach to collision avoidance,” IEEE Robotics & Automation Magazine, vol. 4, no. 1, pp. 23–33, 1997. [8] S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in Proc. Int. Conf. on Machine Learning, 2018, pp. 1587–1596. [9] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016. [10] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proc. Int. Conf. on Machine Learning, 2018, pp. 1861–1870. [11] T. Haarnoja et al., “Soft actor-critic algorithms and applications,” arXiv preprint arXiv:1812.05905, 2018. [12] J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research, vol. 32, no. 11, pp. 1238–1274, 2013. [13] S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research, vol. 17, no. 39, pp. 1–40, 2016. [14] T. P. Lillicrap et al., “Continuous control with deep reinforcement learning,” in Int. Conf. on Learning Representations, 2016. [15] V. Mnih et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015. [16] A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations,” in Proc. Int. Conf. on Machine Learning, 1999, pp. 278–287. [17] OpenAI, “Solving Rubik’s cube with a robot hand,” OpenAI Research Report, 2019. [18] X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in Proc. IEEE Int. Conf. on Robotics and Automation, 2018, pp. 3803–3810. [19] A. S. Polydoros and L. Nalpantidis, “Survey of model-based reinforcement learning,” Journal of Artificial Intelligence Research, vol. 59, pp. 467–555, 2017. [20] F. Sadeghi and S. Levine, “CAD2RL: Real single-image flight without a single real image,” in Proc. Robotics: Science and Systems, 2017. [21] A. Sánchez-González et al., “Learning to simulate complex physics with graph networks,” in Proc. Int. Conf. on Machine Learning, 2020, pp. 8459–8468. [22] J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel, “Trust region policy optimization,” in Int. Conf. on Machine Learning, 2015, pp. 1889–1897. [23] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [24] Y. Shi, L. Zhang, H. Liu, and X. Wang, “Robust simulation-to-reality transfer for autonomous navigation using adaptive reinforcement learning,” IEEE Access, vol. 13, pp. 45678–45695, 2025. [verify] [25] D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proc. Int. Conf. on Machine Learning, 2014, pp. 387–395. [26] R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. Cambridge, MA, USA: MIT Press, 2018. [27] J. Tobin et al., “Domain randomization for transferring deep neural networks from simulation to the real world,” in IEEE/RSJ Int. Conf. on Intelligent Robots and Systems, 2017, pp. 23–30. [28] E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in IEEE/RSJ Int. Conf. on Intelligent Robots and Systems, 2012, pp. 5026–5033. [29] P. Tiwari, S. Khapre, and R. Singh, “Addressing simulation-to-reality challenges in reinforcement learning-based mobile robots,” Robotics and Autonomous Systems, vol. 188, art. 105112, 2026. [verify] [30] S. Thrun, W. Burgard, and D. Fox, Probabilistic Robotics. Cambridge, MA, USA: MIT Press, 2005. [31] H. Wang, Y. Li, X. Chen, and P. Zhao, “Multi-sensor fusion for robust autonomous navigation in dynamic environments,” Sensors, vol. 24, no. 8, art. 2546, 2024. [verify] [32] W. Yu, C. Liu, J. Luo, and Y. Sun, “Adaptive reinforcement learning for autonomous robotic navigation under uncertainty,” IEEE Access, vol. 12, pp. 87654–87670, 2024. [verify] [33] T. Zhang, J. Wang, X. Li, and S. Chen, “Human-aware autonomous navigation using deep reinforcement learning,” Robotics and Autonomous Systems, vol. 175, art. 104625, 2024. [verify] [34] Y. Zhao, M. Xu, K. Liu, and P. Wang, “Domain randomization and transfer learning for simulation-to-reality robotic deployment,” Applied Artificial Intelligence, vol. 39, no. 1, pp. 115–133, 2025. [verify]

How to Cite This Paper

Cecilia Augustine Abi, Joe Essien, Amaku Effiong Amaku (2026). Bridging the Simulation-to-Reality Gap in Reinforcement Learning-Based Autonomous Robot Navigation. International Journal of Computer Techniques, 13(4). ISSN: 2394-2231.

© 2026 International Journal of Computer Techniques (IJCT). All rights reserved.

Submit Your Paper