Reinforcement Learning for Dynamic Spectrum Allocation in 6G Networks

Authors

DOI:

https://doi.org/10.24086/cuesj.v10n2y2026.pp11-28

Keywords:

Cognitive radio networks, Deep reinforcement learning, Dynamic spectrum access, Multi-agent systems, Spectrum sensing

Abstract

In this paper, a multi-agent deep reinforcement learning framework is proposed for opportunistic spectrum access in cognitive radio networks (CRNs). The proposed framework consists of cooperative sensing, adaptive policy optimization, and distributed coordination for dynamic, efficient, and interpretable spectrum utilization. It is trained with diverse datasets from simulated and real environments to enhance generalization across spectral loads and interference levels. Experimental results show that the developed framework achieves a maximum throughput of 31.1 Mbps, an average fairness index of 0.95, and a latency reduction of <5 ms compared with conventional agents. Cross-dataset evaluation verifies accuracy over 95%, throughput retention over 96%, and a loss of <3% when transferring between synthetic and real data domains. Robustness testing revealed an accuracy change of <4.5% under −5 dB signal to noise ratio and a minimal throughput loss of 7.5%. Computational scalability remains stable as the number of agents increases, and interpretability analysis shows consistent policy behavior with entropy between 0.69 and 0.75. These results validate the developed framework as a reliable, scalable, and explainable basis for adaptive and autonomous cognitive radio systems.

Downloads

Download data is not yet available.

Author Biography

Yazen S. Almashhadani, Department of Informatics and Software Engineering, Cihan University-Erbil, Kurdistan Region, Iraq

Yazen Saifuldeen Mahmood is a lecturer in the Department of Informatics and Software Engineering at Cihan University-Erbil, Kurdistan Region, F.R. Iraq. His research interest include Mobile communications and Propagation Models.

References

[1] H. Yigang, F. Ali, C. Weiding, A. Ali, and A. Rahman, “A review on spectrum standardization for wireless networks: Past, present and future advancements,” The Intersection of 6G, AI/Machine Learning, and Embedded Systems, pp. 31–66.

[2] S. Selvarajan, G. Nayak, T. Lalitha, D. Singh, S. Nanda, and R. Narayanswamy, “Dynamic spectrum allocation strategies for mobile broadband efficiency,” Journal of Wireless Mobile Networks, Ubiquitous Computing, and Dependable Applications, vol. 16, no. 2, pp. 204–216, 2025.

[3] R. Priyadarshi, R. R. Kumar, and Z. Ying, “Techniques employed in distributed cognitive radio networks: a survey on routing intelligence,” Multimedia Tools and Applications, vol. 84, no. 9, pp. 5741–5792, 2025.

[4] A. Gbenga-Ilori, A. L. Imoize, K. Noor, and P. O. Adebolu-Ololade, “Artificial intelligence empowering dynamic spectrum access in advanced wireless communications: A comprehensive overview,” AI, vol. 6, no. 6, p. 126, 2025.

[5] F. Golpayegani, N. Chen, N. Afraz, E. Gyamfi, A. Malekjafarian, D. Schäfer, and C. Krupitzer, “Adaptation in edge computing: a review on design principles and research challenges,” ACM Transactions on Autonomous and Adaptive Systems, vol. 19, no. 3, pp. 1–43, 2024.

[6] J. Beckley, “Advanced risk assessment techniques: Merging data-driven analytics with expert insights to navigate uncertain decision-making processes,” Int J Res Publ Rev, vol. 6, no. 3, pp. 1454–1471, 2025.

[7] N. El-haryqy, Z. Madini, and Y. Zouine, “A review of deep learning techniques for enhancing spectrum sensing and prediction in cognitive radio systems: approaches, datasets, and challenges,” International Journal of Computers and Applications, vol. 46, no. 12, pp. 1104–1128, 2024.

[8] X. Bao, J. Zhou, and L. Zhang, “Performance limits of shadowing effects-based passive target localization assisted by active sensor calibration via visible light positioning,” IEEE Transactions on Communications, 2025.

[9] R. Gao, G. Yan, R. Niu, W. Chang, T. Yan, and C. Tang, “A novel spectrum sensing method for multiple unknown signal sources using frequency domain energy detection and dbscan,” IEEE Access, 2025.

[10] Y. Zhang, H. Shan, H. Chen, D. Mi, and Z. Shi, “Perceptive mobile networks for unmanned aerial vehicle surveillance: From the perspective of cooperative sensing,” IEEE Vehicular Technology Magazine, vol. 19, no. 2, pp. 60–69, 2024.

[11] A. G. Olanrewaju, A. O. Ajayi, O. I. Pacheco, A. O. Dada, and A. A. Adeyinka, “Ai-driven adaptive asset allocation: A machine learning approach to dynamic portfolio optimization in volatile financial markets,” Int J Res Finance Manag, vol. 8, no. 1, pp. 320–32, 2025.

[12] A. S. Shethiya, “Adaptive learning machines: A framework for dynamic and real-time ml applications,” Annals of Applied Sciences, vol. 5, no. 1, 2024.

[13] Z. Li, R. Yang, X. Yang, J. Yang, X. Li, H. Lin, F. Qian, Y. Liu, Z. Liao, and D. Hu, “A four-year retrospective of mobile access bandwidth evolution: The inspiring, the frustrating, and the fluctuating,” IEEE Transactions on Mobile Computing, 2025.

[14] S. P. Ghodake, V. R. Malkar, K. Santosh, L. Jabasheela, S. Abdufattokhov, and A. Gopi, “Enhancing supply chain management efficiency: A data-driven approach using predictive analytics and machine learning algorithms.,” International Journal of Advanced Computer Science & Applications, vol. 15, no. 4, 2024.

[15] Y. Liu, Y. Xu, and S. Zhou, “Enhancing user experience through machine learningbased personalized recommendation systems: Behavior data-driven ui design,” Authorea Preprints, 2024.

[16] A. A. Aliyu, J. Liu, and E. Gilliard, “A decentralized and self-adaptive intrusion detection approach using continuous learning and blockchain technology,” Journal of Data Science and Intelligent Systems, 2024.

[17] S. Kuang, J. Zhang, and A. Mohajer, “Reliable information delivery and dynamic link utilization in manet cloud using deep reinforcement learning,” Transactions on Emerging Telecommunications Technologies, vol. 35, no. 9, p. e5028, 2024.

[18] X. Zhang, Z. Chen, Y. Zhang, Y. Liu, M. Jin, and T. Qiu, “Deep-reinforcement-learningbased distributed dynamic spectrum access in multiuser multichannel cognitive radio internet of things networks,” IEEE Internet of Things Journal, vol. 11, no. 10, pp. 17495–17509, 2024.

[19] S. Balhara, N. Gupta, A. Alkhayyat, I. Bharti, R. Q. Malik, S. N. Mahmood, and F. Abedi, “A survey on deep reinforcement learning architectures, applications and emerging trends,” IET Communications, vol. 19, no. 1, p. e12447, 2025.

[20] R. Ali, T. M. Mitcham, T. Brevett, Ò. C. Agudo, C. D. Martinez, C. Li, M. M. Doyley, and N. Duric, “2-d slicewise waveform inversion of sound speed and acoustic attenuation for ring array ultrasound tomography based on a block lu solver,” IEEE transactions on medical imaging, vol. 43, no. 8, pp. 2988–3000, 2024.

[21] Y. Huang, G.-P. Liu, Y. Yu, and W. Hu, “Data-driven distributed predictive tracking control for heterogeneous nonlinear multiagent systems with communication delays,” IEEE Transactions on Automatic Control, vol. 69, no. 7, pp. 4786–4792, 2024.

[22] N. Hussein and P. Musilek, “Enhancing fairness and efficiency in community energy systems: A forecast-driven approach,” Energy, p. 137976, 2025.

[23] C.-f. Chen, R. Napolitano, Y. Hu, B. Kar, and B. Yao, “Addressing machine learning bias to foster energy justice,” Energy Research & Social Science, vol. 116, p. 103653, 2024.

[24] J. Liang, H. Miao, K. Li, J. Tan, X. Wang, R. Luo, and Y. Jiang, “A review of multi-agent reinforcement learning algorithms,” Electronics, vol. 14, no. 4, p. 820, 2025.

[25] H. Lyu, Q. Zhong, D. Jiao, and J. Hua, “Bump feature detection based on spectrum modeling of discrete-sampled, non-homogeneous multi-sensor stream data,” Applied Sciences, vol. 14, no. 15, p. 6744, 2024.

[26] Y.-W. Guo, Y. Liu, P.-C. Huang, M. Rong, W. Wei, Y.-H. Xu, and J.-H. Wei, “Adaptive changes and genetic mechanisms in organisms under controlled conditions: A review,” International Journal of Molecular Sciences, vol. 26, no. 5, p. 2130, 2025.

[27] A. Kantaros, T. Ganetsos, E. Pallis, and M. Papoutsidakis, “From mathematical modeling and simulation to digital twins: Bridging theory and digital realities in industry and emerging technologies,” Applied Sciences, vol. 15, no. 16, p. 9213, 2025.

[28] U. C. Ukpong, O. Idowu-Bismark, E. Adetiba, J. R. Kala, E. Owolabi, O. Oshin, A. Abayomi, and O. E. Dare, “Deep reinforcement learning agents for dynamic spectrum access in television whitespace cognitive radio networks,” Scientific African, vol. 27, p. e02523, 2025.

[29] W. Bai, G. Zheng, W. Xia, Y. Mu, and Y. Xue, “Multi-user opportunistic spectrum access for cognitive radio networks based on multi-head self-attention and multi-agent deep reinforcement learning,” Sensors, vol. 25, no. 7, p. 2025, 2025.

[30] J. Chao and M. Jiao, “Network spectrum resource allocation and optimization based on deep learning and trdm,” Informatica, vol. 49, no. 13, 2025.

[31] M. K. Giri and S. Majumder, “Distributed dynamic spectrum access through multi-agent deep recurrent q-learning in cognitive radio network,” Physical Communication, vol. 58, p. 102054, 2023.

[32] Y. Zhang, X. Han, R. Bai, and M. Jia, “Multi-agent deep reinforcement learning based multiple access for underwater cognitive acoustic sensor networks,” Computers and Electrical Engineering, vol. 120, p. 109819, 2024.

[33] R. Yan, Z. Guo, P. Liu, Q. Lan, X.-P. Zhang, and Y. Dong, “Multi-agent reinforcement learning based channel access optimization for ieee 802.11 bn,” IEEE Transactions on Green Communications and Networking, 2024.

[34] S. Liu, C. Pan, C. Zhang, F. Yang, and J. Song, “Dynamic spectrum sharing based on deep reinforcement learning in mobile communication systems,” Sensors, vol. 23, no. 5, p. 2622, 2023.

[35] S. Ahmad, S. Zain Ul Abideen, M. M. Kamal, M. Al-Khasawneh, G. F. Issa, N. Ullah, O. Alfarraj, A. Tolba, M. Sheraz, and T. C. Chuah, “Resource management for multidrone communications in next-generation noma-enabled wireless networks,” Scientific Reports, vol. 15, no. 1, p. 23585, 2025.

[36] Y. S. Almashhadani and G. A. QasMarrogy, “Dynamic power allocation for downlink noma,” Cihan University-Erbil Scientific Journal, vol. 8, pp. 85–90, June 2024.

[37] G. B. Tarekegn, R.-T. Juang, H.-P. Lin, Y. Y. Munaye, L.-C. Wang, and M. A. Bitew, “Deep-reinforcement-learning-based drone base station deployment for wireless communication services,” IEEE Internet of Things Journal, vol. 9, no. 21, pp. 21899–21915, 2022.

[38] S. Mondal, M. P. Dutta, and S. K. Chakraborty, “A hybrid deep learning based approach for spectrum sensing in cognitive radio,” Physical Communication, vol. 67, p. 102497, 2024.

[39] L. Wang, J. Hu, R. Jiang, and Z. Chen, “A deep long-term joint temporal–spectral network for spectrum prediction,” Sensors, vol. 24, no. 5, p. 1498, 2024.

[40] E. V. Vijay and K. Aparna, “Deep learning-ct based spectrum sensing for cognitive radio for proficient data transmission in wireless sensor networks,” e-Prime-Advances in Electrical Engineering, Electronics and Energy, vol. 9, p. 100659, 2024.

[41] M. Sairam, R. Egala, H. Rajasekhar, and K. Nohith, “Deep learning-based spectrum management to enhance the performance of cognitive radio network using mobilenet,” IRE Journals, vol. 8, no. 6, pp. 274–279, 2024.

[42] G. Narmadha, M. Jeyalakshmi, M. Ponnrajakumari, N. Duraichi, and B. Sakthivel, “Enhancing spectrum prediction in cognitive radio networks using an optimized generative adversarial network,” Results in Engineering, p. 105270, 2025.

[43] S. M. A. Elmorsy, S. M. Osman, and S. A. Gamel, “Enhanced spectrum sensing for 5g and lte signals using advanced deep learning models and hyperparameter tuning,” Scientific Reports, vol. 15, no. 1, p. 24825, 2025.

[44] Y. S. Almashhadani, H. J. A. Alqaysi, and G. A. QasMarrogy, “Performance analysis of OFDM with different cyclic prefix length,” in Proceedings of the 2nd International Conference of Cihan University-Erbil on Communication Engineering and Computer Science (CIC-COCOS’17), (Erbil, Iraq), pp. 66–69, Mar. 2017.

[45] A. Ziya and collaborators, “Cluster-assisted spectrum sensing dataset,” 2022. Accessed: 2025-10-15.

[46] D. Kuester, X. Lu, D. Gu, A. Kord, J. Rezac, K. Carson, M. L. Dowell, E. Eyeson,

A. Feldman, K. Forsyth, et al., “Radio spectrum occupancy measurements amid covid-19 telework and telehealth,” National Institute of Standards and Technology, Tech. Rep.TN-2240, 2022.

[47] A. P. Team, “Aerpaw wireless datasets for public use,” 2023. Accessed: 2025-10-15.

[48] S. Chang, R. Shu, and collaborators, “Csrd2025: A large-scale synthetic radio dataset for spectrum sensing in wireless communications,” 2025. Accessed: 2025-10-15.

[49] U. Rani and C. Prashanth, “Drlnet: a deep reinforcement learning network for hybrid features extraction and spectrum sensing in cognitive radio networks,” Journal of Advances in Information Technology, vol. 14, no. 6, pp. 1321–1330, 2023.

Published

2026-07-01

How to Cite

1.
Almashhadani YS. Reinforcement Learning for Dynamic Spectrum Allocation in 6G Networks. Cihan U Erbil SCI J [Internet]. 2026 Jul. 1 [cited 2026 Jul. 21];10(2):11-28. Available from: https://journals.cihanuniversity.edu.iq/index.php/cuesj/article/view/1753

Issue

Section

Research Article

Similar Articles

<< < 1 2 3 4 5 6 > >> 

You may also start an advanced similarity search for this article.