Webpage Recommendation System for Healthcare Issue Using Reinforcement Learning Algorithm
Published 2026-09-18
Keywords
- Recommender system,
- Q-learning,
- Contrastive learning,
- Self-supervised Learning
Copyright (c) 2026

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Abstract
World Wide Web is an immense repository of information, the sheer volume of data makes it challenging to locate specific information easily. Recommendation Systems (RS) are designed to alleviate this issue by identifying similar items and customers based on their behaviour, subsequently suggesting items tailored to individual preferences. The RS involves RL (Reinforcement Learning)-based RS approaches used to overcome existing limitations by fusing RL with self-supervised sequential learning. However, these approaches often suffer from biases in estimating Q-values. It exclusively depends To rectify this, a novel strategy incorporating Supervised Negative Q-learning and Supervised Advantage Actor-Critic has been introduced. This work proposes a solution to enhance stability by merging RL with sequential modelling, contrastive-based objectives, and negative sampling methods, alongwith contrastive learning and conservative Q-learning. Together, these components improve performance and stability. The proposed method validates its effectiveness through empirical results obtained from real-world datasets.
References
- Chen, M. (2021). Exploration in recommender systems. In Proceedings of the 15th ACM Conference on Recommender Systems, pages 551–553.
- Emilio Parisotto, H Francis Song, Jack W Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant M Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, et al. 2019. Stabilizing Transformers for Reinforcement Learning. arXiv preprint arXiv:1910.06764 (2019).
- Fajie Yuan, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M Jose, and Xiangnan He. 2019. A simple convolutional generative network for next item recommendation. In WSDM. 582–590
- Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019. BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In CIKM. 1441–1450
- Gao, R.; Xia, H.; Li, J.; Liu, D.; Chen, S.; Chun, G. DRCGR: Deep Reinforcement Learning Framework Incorporating CNN and GAN-Based for Interactive Recommendation. In Proceedings of the 2019 IEEE ICDM, Beijing, China, 8–11 November 2019; pp. 1048–1053.
- H. F. Germany, “Diabetes data set,” 2023, https://www.kaggle. com/johndasilva/diabetes.
- Hado V Hasselt. 2010. Double Q-learning. In Advances in Neural Information Processing Systems. 2613–2621
- Jiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He, Liang Chen, Jianxun Lian, and Xing Xie. 2021. Self-supervised graph learning for recommendation. In SIGIR. 726–735.
- Jiaxi Tang and Ke Wang. 2018. Personalized top-n sequential recommendation via convolutional sequence embedding. In WSDM. 565–573.
- Jin Huang, Zhaochun Ren, Wayne Xin Zhao, Gaole He, Ji-Rong Wen, and Daxiang Dong. 2019. Taxonomy-aware multi-hop reasoning networks for sequential recommendation. In WSDM. 573–581.
- Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020). Conservative q-learning for offline reinforcement learning. Advances in Neural Information Processing Systems, volume 33, pages 1179–1191. Curran Associates, Inc.
- Lei, Y.; Li, W. Interactive Recommendation with User-Specific Deep Reinforcement Learning. ACM Trans. Knowl. Discov. Data 2019, 13, 1–15.
- Li, J., Zhang, W., Wang, T., Xiong, G., Lu, A., and Medioni, G. (2023). Gpt4rec: A generative framework for personalized recommendation and user interests interpretation.
- Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H Chi. 2019. Top-k off-policy correction for a REINFORCE recommender system. In WSDM. 456–464.
- OpenAI (2023). Gpt-4 technical report.
- Pima Indians Diabetes Database. https://kaggle.com/uciml/pima-indians-diabetes-database (accessed August. , 2023).
- Ruihong Qiu, Zi Huang, Hongzhi Yin, and Zijian Wang. 2022. Contrastive learning for representation degeneration problem in sequential recommendation. In WSDM. 813–823.
- Ruining He, Chen Fang, Zhaowen Wang, and Julian McAuley. 2016. Vista: A visually, socially, and temporally-aware model for artistic recommendation. In RecSys. 309–316.
- Van Den Oord, A., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. CoRR, abs/1807.03748. 11
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017). Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R., editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
- Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In ICDM. 197–206.
- Wu, C.-Y., Ahmed, A., Beutel, A., Smola, A. J., and Jing, H. (2017).Recurrent recommender networks. In Proceedings of the tenth ACM international conference on web search and data mining, pages 495–503.
- Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M Jose. 2020. Self Supervised Reinforcement Learning for Recommender Systems. SIGIR (2020).
- Xin, X., Karatzoglou, A., Arapakis, I., and Jose, J. (2020). Self-supervised reinforcement learning for recommender systems. In Proceedings of the 43th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20).
- Xin, X., Karatzoglou, A., Arapakis, I., and Jose, J. M. (2022). Supervised advantage actor-critic for recommender systems. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, WSDM ’22, page 1186–1196, New York, NY, USA. Association for Computing Machinery.
- Xu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu, Jinyang Gao, Jiandong Zhang, Bolin Ding, and Bin Cui. 2022. Contrastive learning for sequential recommendation. In ICDE. 1259–1273.
- Zhao, X.; Xia, L.; Zhang, L.; Tang, J.; Ding, Z.; Yin, D. Recommendations with Negative Feedback via Pairwise Deep Reinforcement Learning. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK, 19–23 August 2018; pp. 1040–1048.
- Zhou, F.; Luo, B.; Hu, T.; Chen, Z.; Wen, Y. A Combinatorial Recommendation System Framework Based on Deep Reinforcement Learning. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data), Orlando, FL, USA, 15–18 December 2021; pp. 5733–5740.
- Zhou, K., Wang, H., Zhao, W. X., Zhu, Y., Wang, S., Zhang, F., Wang, Z., and Wen, J.-R. (2020). S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization. page 1893–1902, New York, NY, USA. Association for Computing Machinery.