IEEE Open Journal of Control Systems · 2023
Certifying Black-Box Policies With Stability for Nonlinear Control
IEEE Open Journal of Control Systems
How can model based advice stabilize a learned nonlinear policy?
Simply blending two stabilizing policies can cause instability. An adaptive confidence policy instead uses approximate model information to certify stability, with competitive guarantees under bounded nonlinearity and evaluations on Cart Pole and charging under distribution shift.
Research themes
Connections
From prediction error to control performance
Trust adaptation in linear quadratic control and perturbation bounds for MPC connect the quality of a forecast to the cost of acting on it. Nonlinear certification adds stability to this picture.
Cite this paper
Tongxin Li, Ruixiao Yang, Guannan Qu, Yiheng Lin, Adam Wierman, Steven H. Low. Certifying Black-Box Policies With Stability for Nonlinear Control. IEEE Open Journal of Control Systems, 2023. https://doi.org/10.1109/OJCSYS.2023.3241486
BibTeX
@article{tongxin-certified-policies,
title = {{Certifying Black-Box Policies With Stability for Nonlinear Control}},
author = {Tongxin Li and Ruixiao Yang and Guannan Qu and Yiheng Lin and Adam Wierman and Steven H. Low},
year = {2023},
journal = {IEEE Open Journal of Control Systems},
url = {https://doi.org/10.1109/OJCSYS.2023.3241486},
doi = {10.1109/OJCSYS.2023.3241486}
}
Related papers
Disentangling Linear Quadratic Control with Untrusted ML Predictions
DISC learns confidence in predictions of latent disturbance components. Competitive analysis covers linear and more general mixing functions, showing how online confidence adaptation can exploit accurate forecasts while maintaining protection against large errors.
Can a controller learn which components of a forecast to trust? learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization model predictive control MPC stability recursive feasibility distributed control system level synthesisBounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and Nonlinearity
A general analysis pipeline converts perturbation bounds for finite horizon control into dynamic regret bounds for MPC. It handles prediction errors in costs, dynamics, and disturbances, and extends the analysis to constrained and nonlinear settings under suitable regularity conditions.
How do forecast errors translate into regret for MPC? model predictive control MPC stability recursive feasibility distributed control system level synthesis learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimizationRobustness and Consistency in Linear Quadratic Control with Untrusted Predictions
A confidence parameter governs how much a linear quadratic controller trusts disturbance predictions. Competitive bounds describe the tradeoff between accurate and inaccurate advice, and a self tuning policy adapts confidence online using observed prediction quality.
Can predictive control balance consistency and robustness automatically? learning augmented algorithms algorithms with predictions competitive analysis robustness consistency online optimization model predictive control MPC stability recursive feasibility distributed control system level synthesis