NeurIPS · 2023

Anytime-Competitive Reinforcement Learning with Policy Prior

Jianyi Yang, Pengfei Li, Tongxin Li, Adam Wierman, Shaolei Ren

Advances in Neural Information Processing Systems

Can reinforcement learning enforce cost protection throughout an episode?

ACRL constrains cumulative costs relative to a policy prior at each round, rather than only in expectation over episodes. The analysis establishes cost guarantees and regret relative to the constrained optimum, with experiments in carbon aware computing.

Research themes

Cite this paper

Jianyi Yang, Pengfei Li, Tongxin Li, Adam Wierman, Shaolei Ren. Anytime-Competitive Reinforcement Learning with Policy Prior. Advances in Neural Information Processing Systems, 2023. https://papers.nips.cc/paper_files/paper/2023/hash/f53437debdd397c42929d929614bc705-Abstract-Conference.html

BibTeX
@inproceedings{tongxin-anytime-rl,
  title = {{Anytime-Competitive Reinforcement Learning with Policy Prior}},
  author = {Jianyi Yang and Pengfei Li and Tongxin Li and Adam Wierman and Shaolei Ren},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  url = {https://papers.nips.cc/paper\_files/paper/2023/hash/f53437debdd397c42929d929614bc705-Abstract-Conference.html}
}
2026Applied Energy

Counterfactual load forecasting with LLM-structured events and representation learning

Yujie Chen, Yifei Gao, Runyao Yu, Yuhe Wu, Guangyu Wang, Yue Chen, Tongxin Li

NACF turns news into structured treatments and estimates load trajectories under alternative event conditions. Reweighting and representation balancing address observed confounding. Experiments examine factual accuracy and interpretable demand perturbations without claiming that unobserved counterfactual outcomes can be directly validated.

How might electricity demand change under a different news event? AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization large language models LLM agents contextual control world models reinforcement learning dueling bandits
2026AISTATS

Leveraging Machine-Learned Advice in Strategic Interactions with No-Regret Learners

Tinashe Handina, Tongxin Li, Kishan Panaganti, Eric Mazumdar, Adam Wierman

A measure of advice quality connects simulators and payoff predictions to strategic performance. The paper establishes benefits of reliable advice for approximate Stackelberg play and limitations on simultaneously exploiting accurate advice and protecting against inaccurate advice.

How useful is imperfect advice against an adaptive opponent? game theory learning in games Bayesian games Stackelberg strategies equity public models learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization
2026IEEE PES International Meeting

PEARL: A Physics-Enhanced Adaptive Residual Learning Framework for PV Modeling

Yikai Lu, Yujie Chen, Tongxin Li

PEARL separates a physical photovoltaic simulator from a learned residual correction. A controlled benchmark compares several machine learning architectures and a coupled modeling baseline, examining accuracy while retaining the physical model as an interpretable reference.

Can learned residuals improve physical models of solar generation? AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization