NeurIPS · 2023
Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions
Advances in Neural Information Processing Systems
Does the structure of advice improve robust decision making?
Q value predictions expose more information than an opaque policy recommendation. For a single trajectory MDP, the analysis characterizes consistency and robustness tradeoffs and shows how structured advice improves the guarantees available from a robust baseline.
Research themes
Connections
Algorithms guide learned decision makers
Q value advice and imperfect transition forecasts provide structured information for reinforcement learning. LEAD applies a related principle by supporting language agents with dueling bandit algorithms.
Open questions
Prediction quality as a decision resource
Which prediction errors actually matter for the decision being made?
Explore open directionsCite this paper
Tongxin Li, Yiheng Lin, Shaolei Ren, Adam Wierman. Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions. Advances in Neural Information Processing Systems, 2023. https://papers.nips.cc/paper_files/paper/2023/hash/8e806d3c56ed5f1dab85d601e13cbe38-Abstract-Conference.html
BibTeX
@inproceedings{tongxin-q-value-advice,
title = {{Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions}},
author = {Tongxin Li and Yiheng Lin and Shaolei Ren and Adam Wierman},
year = {2023},
booktitle = {Advances in Neural Information Processing Systems},
url = {https://papers.nips.cc/paper\_files/paper/2023/hash/8e806d3c56ed5f1dab85d601e13cbe38-Abstract-Conference.html}
}
Related papers
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
The study identifies a gap between quick preference discovery and sustained exploitation by language agents. LEAD combines LLM reasoning with dueling bandit algorithms to obtain weak and strong regret guarantees, with evaluations under noisy and adversarial prompts.
Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach
A Bayesian value function and Bellman Jensen gap quantify the value of imperfect transition forecasts. BOLA separates offline value learning from online adaptation, with sample efficiency analysis and experiments in synthetic environments and wind energy storage control.