NeurIPS · 2025

Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach

Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman

Advances in Neural Information Processing Systems

How can reinforcement learning use imperfect forecasts beyond one step?

A Bayesian value function and Bellman Jensen gap quantify the value of imperfect transition forecasts. BOLA separates offline value learning from online adaptation, with sample efficiency analysis and experiments in synthetic environments and wind energy storage control.

Research themes

Connections

Algorithms guide learned decision makers

Q value advice and imperfect transition forecasts provide structured information for reinforcement learning. LEAD applies a related principle by supporting language agents with dueling bandit algorithms.

Open questions

Prediction quality as a decision resource

Which prediction errors actually matter for the decision being made?

Explore open directions

Cite this paper

Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman. Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach. Advances in Neural Information Processing Systems, 2025. https://papers.nips.cc/paper_files/paper/2025/hash/940a7634dab556b67af15bacd337f7db-Abstract-Conference.html

BibTeX
@inproceedings{tongxin-bellman-jensen,
  title = {{Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach}},
  author = {Chenbei Lu and Zaiwei Chen and Tongxin Li and Chenye Wu and Adam Wierman},
  year = {2025},
  booktitle = {Advances in Neural Information Processing Systems},
  url = {https://papers.nips.cc/paper\_files/paper/2025/hash/940a7634dab556b67af15bacd337f7db-Abstract-Conference.html}
}