Reinforcement Learning for Dynamic Spectrum Allocation in 6G Networks
DOI:
https://doi.org/10.24086/cuesj.v10n2y2026.pp11-28Keywords:
Cognitive radio networks, Deep reinforcement learning, Dynamic spectrum access, Multi-agent systems, Spectrum sensingAbstract
In this paper, a multi-agent deep reinforcement learning framework is proposed for opportunistic spectrum access in cognitive radio networks (CRNs). The proposed framework consists of cooperative sensing, adaptive policy optimization, and distributed coordination for dynamic, efficient, and interpretable spectrum utilization. It is trained with diverse datasets from simulated and real environments to enhance generalization across spectral loads and interference levels. Experimental results show that the developed framework achieves a maximum throughput of 31.1 Mbps, an average fairness index of 0.95, and a latency reduction of <5 ms compared with conventional agents. Cross-dataset evaluation verifies accuracy over 95%, throughput retention over 96%, and a loss of <3% when transferring between synthetic and real data domains. Robustness testing revealed an accuracy change of <4.5% under −5 dB signal to noise ratio and a minimal throughput loss of 7.5%. Computational scalability remains stable as the number of agents increases, and interpretability analysis shows consistent policy behavior with entropy between 0.69 and 0.75. These results validate the developed framework as a reliable, scalable, and explainable basis for adaptive and autonomous cognitive radio systems.
Downloads
References
[1] H. Yigang, F. Ali, C. Weiding, A. Ali, and A. Rahman, “A review on spectrum standardization for wireless networks: Past, present and future advancements,” The Intersection of 6G, AI/Machine Learning, and Embedded Systems, pp. 31–66.
[2] S. Selvarajan, G. Nayak, T. Lalitha, D. Singh, S. Nanda, and R. Narayanswamy, “Dynamic spectrum allocation strategies for mobile broadband efficiency,” Journal of Wireless Mobile Networks, Ubiquitous Computing, and Dependable Applications, vol. 16, no. 2, pp. 204–216, 2025.
[3] R. Priyadarshi, R. R. Kumar, and Z. Ying, “Techniques employed in distributed cognitive radio networks: a survey on routing intelligence,” Multimedia Tools and Applications, vol. 84, no. 9, pp. 5741–5792, 2025.
[4] A. Gbenga-Ilori, A. L. Imoize, K. Noor, and P. O. Adebolu-Ololade, “Artificial intelligence empowering dynamic spectrum access in advanced wireless communications: A comprehensive overview,” AI, vol. 6, no. 6, p. 126, 2025.
[5] F. Golpayegani, N. Chen, N. Afraz, E. Gyamfi, A. Malekjafarian, D. Schäfer, and C. Krupitzer, “Adaptation in edge computing: a review on design principles and research challenges,” ACM Transactions on Autonomous and Adaptive Systems, vol. 19, no. 3, pp. 1–43, 2024.
[6] J. Beckley, “Advanced risk assessment techniques: Merging data-driven analytics with expert insights to navigate uncertain decision-making processes,” Int J Res Publ Rev, vol. 6, no. 3, pp. 1454–1471, 2025.
[7] N. El-haryqy, Z. Madini, and Y. Zouine, “A review of deep learning techniques for enhancing spectrum sensing and prediction in cognitive radio systems: approaches, datasets, and challenges,” International Journal of Computers and Applications, vol. 46, no. 12, pp. 1104–1128, 2024.
[8] X. Bao, J. Zhou, and L. Zhang, “Performance limits of shadowing effects-based passive target localization assisted by active sensor calibration via visible light positioning,” IEEE Transactions on Communications, 2025.
[9] R. Gao, G. Yan, R. Niu, W. Chang, T. Yan, and C. Tang, “A novel spectrum sensing method for multiple unknown signal sources using frequency domain energy detection and dbscan,” IEEE Access, 2025.
[10] Y. Zhang, H. Shan, H. Chen, D. Mi, and Z. Shi, “Perceptive mobile networks for unmanned aerial vehicle surveillance: From the perspective of cooperative sensing,” IEEE Vehicular Technology Magazine, vol. 19, no. 2, pp. 60–69, 2024.
[11] A. G. Olanrewaju, A. O. Ajayi, O. I. Pacheco, A. O. Dada, and A. A. Adeyinka, “Ai-driven adaptive asset allocation: A machine learning approach to dynamic portfolio optimization in volatile financial markets,” Int J Res Finance Manag, vol. 8, no. 1, pp. 320–32, 2025.
[12] A. S. Shethiya, “Adaptive learning machines: A framework for dynamic and real-time ml applications,” Annals of Applied Sciences, vol. 5, no. 1, 2024.
[13] Z. Li, R. Yang, X. Yang, J. Yang, X. Li, H. Lin, F. Qian, Y. Liu, Z. Liao, and D. Hu, “A four-year retrospective of mobile access bandwidth evolution: The inspiring, the frustrating, and the fluctuating,” IEEE Transactions on Mobile Computing, 2025.
[14] S. P. Ghodake, V. R. Malkar, K. Santosh, L. Jabasheela, S. Abdufattokhov, and A. Gopi, “Enhancing supply chain management efficiency: A data-driven approach using predictive analytics and machine learning algorithms.,” International Journal of Advanced Computer Science & Applications, vol. 15, no. 4, 2024.
[15] Y. Liu, Y. Xu, and S. Zhou, “Enhancing user experience through machine learningbased personalized recommendation systems: Behavior data-driven ui design,” Authorea Preprints, 2024.
[16] A. A. Aliyu, J. Liu, and E. Gilliard, “A decentralized and self-adaptive intrusion detection approach using continuous learning and blockchain technology,” Journal of Data Science and Intelligent Systems, 2024.
[17] S. Kuang, J. Zhang, and A. Mohajer, “Reliable information delivery and dynamic link utilization in manet cloud using deep reinforcement learning,” Transactions on Emerging Telecommunications Technologies, vol. 35, no. 9, p. e5028, 2024.
[18] X. Zhang, Z. Chen, Y. Zhang, Y. Liu, M. Jin, and T. Qiu, “Deep-reinforcement-learningbased distributed dynamic spectrum access in multiuser multichannel cognitive radio internet of things networks,” IEEE Internet of Things Journal, vol. 11, no. 10, pp. 17495–17509, 2024.
[19] S. Balhara, N. Gupta, A. Alkhayyat, I. Bharti, R. Q. Malik, S. N. Mahmood, and F. Abedi, “A survey on deep reinforcement learning architectures, applications and emerging trends,” IET Communications, vol. 19, no. 1, p. e12447, 2025.
[20] R. Ali, T. M. Mitcham, T. Brevett, Ò. C. Agudo, C. D. Martinez, C. Li, M. M. Doyley, and N. Duric, “2-d slicewise waveform inversion of sound speed and acoustic attenuation for ring array ultrasound tomography based on a block lu solver,” IEEE transactions on medical imaging, vol. 43, no. 8, pp. 2988–3000, 2024.
[21] Y. Huang, G.-P. Liu, Y. Yu, and W. Hu, “Data-driven distributed predictive tracking control for heterogeneous nonlinear multiagent systems with communication delays,” IEEE Transactions on Automatic Control, vol. 69, no. 7, pp. 4786–4792, 2024.
[22] N. Hussein and P. Musilek, “Enhancing fairness and efficiency in community energy systems: A forecast-driven approach,” Energy, p. 137976, 2025.
[23] C.-f. Chen, R. Napolitano, Y. Hu, B. Kar, and B. Yao, “Addressing machine learning bias to foster energy justice,” Energy Research & Social Science, vol. 116, p. 103653, 2024.
[24] J. Liang, H. Miao, K. Li, J. Tan, X. Wang, R. Luo, and Y. Jiang, “A review of multi-agent reinforcement learning algorithms,” Electronics, vol. 14, no. 4, p. 820, 2025.
[25] H. Lyu, Q. Zhong, D. Jiao, and J. Hua, “Bump feature detection based on spectrum modeling of discrete-sampled, non-homogeneous multi-sensor stream data,” Applied Sciences, vol. 14, no. 15, p. 6744, 2024.
[26] Y.-W. Guo, Y. Liu, P.-C. Huang, M. Rong, W. Wei, Y.-H. Xu, and J.-H. Wei, “Adaptive changes and genetic mechanisms in organisms under controlled conditions: A review,” International Journal of Molecular Sciences, vol. 26, no. 5, p. 2130, 2025.
[27] A. Kantaros, T. Ganetsos, E. Pallis, and M. Papoutsidakis, “From mathematical modeling and simulation to digital twins: Bridging theory and digital realities in industry and emerging technologies,” Applied Sciences, vol. 15, no. 16, p. 9213, 2025.
[28] U. C. Ukpong, O. Idowu-Bismark, E. Adetiba, J. R. Kala, E. Owolabi, O. Oshin, A. Abayomi, and O. E. Dare, “Deep reinforcement learning agents for dynamic spectrum access in television whitespace cognitive radio networks,” Scientific African, vol. 27, p. e02523, 2025.
[29] W. Bai, G. Zheng, W. Xia, Y. Mu, and Y. Xue, “Multi-user opportunistic spectrum access for cognitive radio networks based on multi-head self-attention and multi-agent deep reinforcement learning,” Sensors, vol. 25, no. 7, p. 2025, 2025.
[30] J. Chao and M. Jiao, “Network spectrum resource allocation and optimization based on deep learning and trdm,” Informatica, vol. 49, no. 13, 2025.
[31] M. K. Giri and S. Majumder, “Distributed dynamic spectrum access through multi-agent deep recurrent q-learning in cognitive radio network,” Physical Communication, vol. 58, p. 102054, 2023.
[32] Y. Zhang, X. Han, R. Bai, and M. Jia, “Multi-agent deep reinforcement learning based multiple access for underwater cognitive acoustic sensor networks,” Computers and Electrical Engineering, vol. 120, p. 109819, 2024.
[33] R. Yan, Z. Guo, P. Liu, Q. Lan, X.-P. Zhang, and Y. Dong, “Multi-agent reinforcement learning based channel access optimization for ieee 802.11 bn,” IEEE Transactions on Green Communications and Networking, 2024.
[34] S. Liu, C. Pan, C. Zhang, F. Yang, and J. Song, “Dynamic spectrum sharing based on deep reinforcement learning in mobile communication systems,” Sensors, vol. 23, no. 5, p. 2622, 2023.
[35] S. Ahmad, S. Zain Ul Abideen, M. M. Kamal, M. Al-Khasawneh, G. F. Issa, N. Ullah, O. Alfarraj, A. Tolba, M. Sheraz, and T. C. Chuah, “Resource management for multidrone communications in next-generation noma-enabled wireless networks,” Scientific Reports, vol. 15, no. 1, p. 23585, 2025.
[36] Y. S. Almashhadani and G. A. QasMarrogy, “Dynamic power allocation for downlink noma,” Cihan University-Erbil Scientific Journal, vol. 8, pp. 85–90, June 2024.
[37] G. B. Tarekegn, R.-T. Juang, H.-P. Lin, Y. Y. Munaye, L.-C. Wang, and M. A. Bitew, “Deep-reinforcement-learning-based drone base station deployment for wireless communication services,” IEEE Internet of Things Journal, vol. 9, no. 21, pp. 21899–21915, 2022.
[38] S. Mondal, M. P. Dutta, and S. K. Chakraborty, “A hybrid deep learning based approach for spectrum sensing in cognitive radio,” Physical Communication, vol. 67, p. 102497, 2024.
[39] L. Wang, J. Hu, R. Jiang, and Z. Chen, “A deep long-term joint temporal–spectral network for spectrum prediction,” Sensors, vol. 24, no. 5, p. 1498, 2024.
[40] E. V. Vijay and K. Aparna, “Deep learning-ct based spectrum sensing for cognitive radio for proficient data transmission in wireless sensor networks,” e-Prime-Advances in Electrical Engineering, Electronics and Energy, vol. 9, p. 100659, 2024.
[41] M. Sairam, R. Egala, H. Rajasekhar, and K. Nohith, “Deep learning-based spectrum management to enhance the performance of cognitive radio network using mobilenet,” IRE Journals, vol. 8, no. 6, pp. 274–279, 2024.
[42] G. Narmadha, M. Jeyalakshmi, M. Ponnrajakumari, N. Duraichi, and B. Sakthivel, “Enhancing spectrum prediction in cognitive radio networks using an optimized generative adversarial network,” Results in Engineering, p. 105270, 2025.
[43] S. M. A. Elmorsy, S. M. Osman, and S. A. Gamel, “Enhanced spectrum sensing for 5g and lte signals using advanced deep learning models and hyperparameter tuning,” Scientific Reports, vol. 15, no. 1, p. 24825, 2025.
[44] Y. S. Almashhadani, H. J. A. Alqaysi, and G. A. QasMarrogy, “Performance analysis of OFDM with different cyclic prefix length,” in Proceedings of the 2nd International Conference of Cihan University-Erbil on Communication Engineering and Computer Science (CIC-COCOS’17), (Erbil, Iraq), pp. 66–69, Mar. 2017.
[45] A. Ziya and collaborators, “Cluster-assisted spectrum sensing dataset,” 2022. Accessed: 2025-10-15.
[46] D. Kuester, X. Lu, D. Gu, A. Kord, J. Rezac, K. Carson, M. L. Dowell, E. Eyeson,
A. Feldman, K. Forsyth, et al., “Radio spectrum occupancy measurements amid covid-19 telework and telehealth,” National Institute of Standards and Technology, Tech. Rep.TN-2240, 2022.
[47] A. P. Team, “Aerpaw wireless datasets for public use,” 2023. Accessed: 2025-10-15.
[48] S. Chang, R. Shu, and collaborators, “Csrd2025: A large-scale synthetic radio dataset for spectrum sensing in wireless communications,” 2025. Accessed: 2025-10-15.
[49] U. Rani and C. Prashanth, “Drlnet: a deep reinforcement learning network for hybrid features extraction and spectrum sensing in cognitive radio networks,” Journal of Advances in Information Technology, vol. 14, no. 6, pp. 1321–1330, 2023.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Yazen S. Almashhadani

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License [CC BY-NC-ND 4.0] that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).



