All systems operational
Q1 2026

Adaptive reinforcement learning for recommendation via large language models and knowledge graphs

Qiang Fan · Hongfeng Han · Jingqi Xing · Zhiyong Zhang
10.1007/s44443-026-00745-z 398 Views 0 Citations
0
Citations
398
Views
Abstract

Abstract

Interactive recommendation systems (IRS) have become a prominent research topic as they dynamically optimize user experience through real-time feedback loops. To model the evolving dynamics of user preferences and maximize long-term rewards, reinforcement learning (RL) has been incorporated into IRS by formulating the recommendation process as a Markov decision process (MDP). However, RL policies trained on static offline data still face two major challenges: (1)
distribution shift
, where the mismatch between offline logs and dynamic online environments often leads to suboptimal long-term decision-making; and (2)
sample efficiency
, as the large action space in recommendation tasks requires substantial interaction before achieving optimal performance. To address these issues, we propose
ARLK
(Adaptive Reinforcement Learning with Large Language Models and Knowledge Graphs), a novel adaptive framework that combines large language model (LLM)-guided offline pretraining and knowledge graph (KG)-enhanced online learning via an adaptive policy fusion mechanism that smoothly transitions from offline initialization to online adaptation. LLMs provide strong semantic understanding that can capture user preferences and simulate interaction feedback, thereby improving policy pretraining and ensuring high-quality initial recommendations in simulation-based online evaluation. Meanwhile, the structured information in KGs is utilized during policy learning to guide candidate generation and significantly reduce exploration cost. Experiments on three benchmark datasets demonstrate that ARLK achieves substantial improvements in both initial recommendation quality and long-term performance compared with state-of-the-art baselines, with average reward improvements of 5.15%, 3.40%, and 1.80% on LFM, Industry, and Coat datasets, respectively, and up to 12.73% gain in Recall@10 on the Coat dataset.

Cite this Article (APA)
Qiang, F., Hongfeng, H., Jingqi, X., Zhiyong, Z. (2026). Adaptive reinforcement learning for recommendation via large language models and knowledge graphs. Journal of King Saud University - Computer and Information Sciences. https://doi.org/10.1007/s44443-026-00745-z
Related Papers
A lightweight model for indoor object detection in unstructured scenes based on joint attention and …
Zhizhong Xing; Leping Li; Ying Yang; Wei Zhou; Guolan Ma; Shaochun Chen; Lechun · 2026
13
cites
425
DDM-YOLO: A lightweight oriented detection model for mature daylily fruits in complex environments
Minqiu Kuang; Xuejie Zou; Fangping Xie; Xiaojian Li; Shang Chen; Dawei Liu; Yuxu · 2026
8
cites
430
Information guided Levy flight for robot search in unknown environments
Weitao Zhao; Zati Hakim Azizul; Xin Lyu; Weijie Kuang · 2026
4
cites
411
3
cites
504
Bridging the gap: A comprehensive survey on AI-driven digital twin networks for future wireless syst…
Yousef Sanjalawe; Salam Fraihat; Salam Al-E’mari; Sharif Naser Makhadmeh · 2026
3
cites
418
Access
View Full Text via DOI
Published in
ISSN 1319-1578
Quartile Q1
AMS Score 100
Field Computer Science & AI
Publisher Elsevier / King Saud University
Country 🇸🇦 Saudi Arabia
View Journal Profile →
Authors
Publication Details
Year 2026
Language English
Added 06 Jul 2026