Abstract
Reinforcement learning (RL) is increasingly emplo-yed for real-time control in Open Radio Access Networks (O-RANs). However, the asynchronous nature of O-RAN feedback introduces stochastic delays between actions and their observable rewards, leading to mis-aligned credit assignment and unstable training. This letter formulates the O-RAN control problem as a Markov decision process (MDP) with delayed and stochastic rewards, where the true reward corresponding to an action is received after an unknown delay. To address this, we propose a delay-aware deep RL framework that integrates a long short-term memory (LSTM)-based delay predictor and a temporal reward re-alignment module within an off-policy learning pipeline. Using empirical delay traces collected from the Open AI Cellular (OAIC) testbed, we demonstrate that the proposed framework achieves faster convergence and significantly reduces latency violations and packet drops compared to conventional RL methods that ignore feedback delay. The results confirm that reward-delay compensation is critical for reliable AI-driven O-RAN control.
| Original language | English |
|---|---|
| Pages (from-to) | 263-267 |
| Number of pages | 5 |
| Journal | IEEE Networking Letters |
| Volume | 8 |
| DOIs | |
| State | Published - 2026 |
Keywords
- delay compensation
- delayed reward
- O-RAN
- radio resource control
- reinforcement learning
- stochastic feedback
Fingerprint
Dive into the research topics of 'Delay-Aware Reinforcement Learning for O-RAN Control Under Stochastic Reward Feedback'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver