TY - GEN
T1 - Cache Policy Design via Reinforcement Learning for Cellular Networks in Non-Stationary Environment
AU - Srinivasan, Ashvin
AU - Amidzadeh, Mohsen
AU - Zhang, Junshan
AU - Tirkkonen, Olav
N1 - Publisher Copyright: © 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - We consider wireless caching both at the network edge and at User Equipment (UE) to alleviate traffic congestion, aiming to find a joint cache placement and delivery policy by maximizing the Quality of Service (QoS) while minimizing backhaul load and User Equipment (UE) power consumption. We assume unknown and time-variant file popularities which are affected by the UE cache content, leading to a non-stationary Partial Observable Markov Decision Process (POMDP). We address this problem in a deep reinforcement learning framework, employing Feed Forward Neural Network (FFNN) and Long Short Term Memory (LSTM) networks in conjunction with Advantageous Actor Critic (A2C) algorithm. LSTM exploits the correlation of the file popularity distribution across time slots to learn information of the dynamics of the environment and A2C algorithm is used due to its ability of handling continuous and high dimensional spaces. We leverage LSTM and A2C tools based on its virtue to find an optimal solution for the POMDP environment. Simulation results show that using LSTM-based A2C outperforms a FFNN-based A2C in terms of sample efficiency and optimality. An LSTM-based A2C gives a superior performance under the non-stationary POMDP paradigm.
AB - We consider wireless caching both at the network edge and at User Equipment (UE) to alleviate traffic congestion, aiming to find a joint cache placement and delivery policy by maximizing the Quality of Service (QoS) while minimizing backhaul load and User Equipment (UE) power consumption. We assume unknown and time-variant file popularities which are affected by the UE cache content, leading to a non-stationary Partial Observable Markov Decision Process (POMDP). We address this problem in a deep reinforcement learning framework, employing Feed Forward Neural Network (FFNN) and Long Short Term Memory (LSTM) networks in conjunction with Advantageous Actor Critic (A2C) algorithm. LSTM exploits the correlation of the file popularity distribution across time slots to learn information of the dynamics of the environment and A2C algorithm is used due to its ability of handling continuous and high dimensional spaces. We leverage LSTM and A2C tools based on its virtue to find an optimal solution for the POMDP environment. Simulation results show that using LSTM-based A2C outperforms a FFNN-based A2C in terms of sample efficiency and optimality. An LSTM-based A2C gives a superior performance under the non-stationary POMDP paradigm.
KW - Advantageous Actor Critic
KW - Deep Reinforcement Learning
KW - Long Short Term Memory
KW - Non-Stationary POMDP
KW - Wireless caching
UR - https://www.scopus.com/pages/publications/85177834257
UR - https://www.scopus.com/pages/publications/85177834257#tab=citedBy
U2 - 10.1109/ICCWorkshops57953.2023.10283680
DO - 10.1109/ICCWorkshops57953.2023.10283680
M3 - Conference contribution
T3 - 2023 IEEE International Conference on Communications Workshops: Sustainable Communications for Renaissance, ICC Workshops 2023
SP - 764
EP - 769
BT - 2023 IEEE International Conference on Communications Workshops
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2023 IEEE International Conference on Communications Workshops, ICC Workshops 2023
Y2 - 28 May 2023 through 1 June 2023
ER -