Research theme

Language and decision agents

How can language and world models support dependable decisions?

Language models provide context, preferences, and predictions for sequential decisions. Algorithmic guidance and structural certification connect these capabilities to explicit performance guarantees.

Explore the research map

Papers

2026Applied Energy

Counterfactual load forecasting with LLM-structured events and representation learning

Yujie Chen, Yifei Gao, Runyao Yu, Yuhe Wu, Guangyu Wang, Yue Chen, Tongxin Li

NACF turns news into structured treatments and estimates load trajectories under alternative event conditions. Reweighting and representation balancing address observed confounding. Experiments examine factual accuracy and interpretable demand perturbations without claiming that unobserved counterfactual outcomes can be directly validated.

How might electricity demand change under a different news event? AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization large language models LLM agents contextual control world models reinforcement learning dueling bandits
2026ICML

World Models in Pieces: Structural Certification for General Agents

Yikai Lu, Yifei Wu, Xinyu Lu, Tongxin Li

Structural certification links performance on compositional goals to local guarantees on world model transitions. The results identify reliable pieces of a model without requiring universal competence, and establish limits on the precision of such guarantees.

Which parts of an agent's world model can support reliable planning? large language models LLM agents contextual control world models reinforcement learning dueling bandits information theory graph learning sample complexity compressed sensing graph neural networks Riemannian geometry
2025ACL Findings

Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents

Fanzeng Xia, Hao Liu, Yisong Yue, Tongxin Li

The study identifies a gap between quick preference discovery and sustained exploitation by language agents. LEAD combines LLM reasoning with dueling bandit algorithms to obtain weak and strong regret guarantees, with evaluations under noisy and adversarial prompts.

Can language agents learn reliably from pairwise preferences? large language models LLM agents contextual control world models reinforcement learning dueling bandits learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization
2025CDC

INSTRUCT MPC: A Human-LLM-in-the-Loop Framework for Context-Aware Control

Ruixiang Wu, Jiahao Ai, Tongxin Li

A language to distribution module converts contextual instructions into disturbance predictions for MPC. The framework closes the loop between instructions, predictions, and control, with a regret analysis for linear dynamics under the stated training assumptions.

How can human instructions improve predictive control? large language models LLM agents contextual control world models reinforcement learning dueling bandits model predictive control MPC stability recursive feasibility distributed control system level synthesis AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization
2025ACM e-Energy demo

Open In-Context Energy Management Platform

Yikai Lu, Tinko Sebastian Bartels, Ruixiang Wu, Fanzeng Xia, Xudong Wang, Yifei Wu, Haoxiang Yang, Tongxin Li

OpenCEM presents a platform design connecting energy time series, events, human context, and simulation. An on campus solar and battery installation motivates evaluation of context sensitive control. The paper describes planned data and API capabilities.

What would a shared benchmark for contextual energy management provide? large language models LLM agents contextual control world models reinforcement learning dueling bandits AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization model predictive control MPC stability recursive feasibility distributed control system level synthesis
2025NeurIPS

Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach

Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman

A Bayesian value function and Bellman Jensen gap quantify the value of imperfect transition forecasts. BOLA separates offline value learning from online adaptation, with sample efficiency analysis and experiments in synthetic environments and wind energy storage control.

How can reinforcement learning use imperfect forecasts beyond one step? learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization large language models LLM agents contextual control world models reinforcement learning dueling bandits AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization
2023NeurIPS

Anytime-Competitive Reinforcement Learning with Policy Prior

Jianyi Yang, Pengfei Li, Tongxin Li, Adam Wierman, Shaolei Ren

ACRL constrains cumulative costs relative to a policy prior at each round, rather than only in expectation over episodes. The analysis establishes cost guarantees and regret relative to the constrained optimum, with experiments in carbon aware computing.

Can reinforcement learning enforce cost protection throughout an episode? large language models LLM agents contextual control world models reinforcement learning dueling bandits learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization model predictive control MPC stability recursive feasibility distributed control system level synthesis
2023NeurIPS

Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions

Tongxin Li, Yiheng Lin, Shaolei Ren, Adam Wierman

Q value predictions expose more information than an opaque policy recommendation. For a single trajectory MDP, the analysis characterizes consistency and robustness tradeoffs and shows how structured advice improves the guarantees available from a robust baseline.

Does the structure of advice improve robust decision making? learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization large language models LLM agents contextual control world models reinforcement learning dueling bandits
2021IEEE Transactions on Smart Grid

Learning-Based Predictive Control via Real-Time Aggregate Flexibility

Tongxin Li, Bo Sun, Yue Chen, Zixin Ye, Steven H. Low, Adam Wierman

Maximum entropy feedback summarizes the feasible actions of controllable loads. Reinforcement learning approximates this feedback for Penalized Predictive Control, reducing information and computation requirements. Charging data demonstrates the coordination approach.

How can an aggregator communicate flexibility in real time? model predictive control MPC stability recursive feasibility distributed control system level synthesis AI for energy electric vehicle charging demand response renewable energy load forecasting decarbonization large language models LLM agents contextual control world models reinforcement learning dueling bandits

Connected themes

From language context to constrained actions

INSTRUCT MPC translates human instructions into disturbance predictions. World model certification and VigilMPC study complementary ways of deciding which learned models can support reliable planning and safe updates.

Algorithms guide learned decision makers

Q value advice and imperfect transition forecasts provide structured information for reinforcement learning. LEAD applies a related principle by supporting language agents with dueling bandit algorithms.

Open directions

Certified contextual control

Can contextual and world model updates be certified before they affect a physical system?

Connect language conditioned predictions with local world model guarantees and feasibility checks. An open challenge is preserving safety when both the context and the learned dynamics change during operation.

Prediction quality as a decision resource

Which prediction errors actually matter for the decision being made?

Move from a single forecast accuracy score to guarantees that reflect the prediction, the task, and the affected transitions. This suggests adaptive confidence allocation across sources, horizons, and latent disturbances.

Accountable learning across agents

How should shared models balance efficiency, strategic risk, and equity?

Connect the downstream objectives of public models to robust strategic decision making. A key challenge is evaluating advice when its benefits and failures are distributed unevenly across participating agents.