Second-Order Policy Gradient Method for the LQR

Published in Engineering Applications of Artificial Intelligence, 2025

Most policy-gradient methods lean on first-order updates, which converge slowly on ill-conditioned control problems. Working with Arash Bahari Kordabad and Prof. Sadegh Soudjani at the Max Planck Institute for Software Systems, we derived an algorithm for exact second-order updates in deterministic policy-gradient methods applied to the Linear Quadratic Regulator (LQR), where the Hessian of the cost can be computed in closed form. The resulting method converges faster than first-order policy gradient while remaining computationally tractable.

Read the paper on arXiv.

Recommended citation: Amirreza Velae, Arash Bahari Kordabad, and Sadegh Soudjani. "Second-Order Policy Gradient Method for the LQR." Engineering Applications of Artificial Intelligence.
Download Paper