ACL Findings · 2025

Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents

Fanzeng Xia, Hao Liu, Yisong Yue, Tongxin Li

Findings of the Association for Computational Linguistics

Can language agents learn reliably from pairwise preferences?

The study identifies a gap between quick preference discovery and sustained exploitation by language agents. LEAD combines LLM reasoning with dueling bandit algorithms to obtain weak and strong regret guarantees, with evaluations under noisy and adversarial prompts.

Research themes

Connections

Algorithms guide learned decision makers

Q value advice and imperfect transition forecasts provide structured information for reinforcement learning. LEAD applies a related principle by supporting language agents with dueling bandit algorithms.

Open questions

Accountable learning across agents

How should shared models balance efficiency, strategic risk, and equity?

Explore open directions

Cite this paper

Fanzeng Xia, Hao Liu, Yisong Yue, Tongxin Li. Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of the Association for Computational Linguistics, 2025. https://doi.org/10.18653/v1/2025.findings-acl.519

BibTeX
@inproceedings{tongxin-lead,
  title = {{Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents}},
  author = {Fanzeng Xia and Hao Liu and Yisong Yue and Tongxin Li},
  year = {2025},
  booktitle = {Findings of the Association for Computational Linguistics},
  url = {https://doi.org/10.18653/v1/2025.findings-acl.519},
  doi = {10.18653/v1/2025.findings-acl.519}
}