CAIDAS Paper Receives Outstanding Paper Award at RLC 2026
17.08.2026We are happy to announce that the paper "Gradient Iterated Temporal-Difference Learning" by Théo Vincent, Kevin Gerhardt, Yogesh Tripathi, Habib Maraqten, Adam White, Martha White, Jan Peters, and Carlo D'Eramo has received the Outstanding Paper Award on Empirical Reinforcement Learning Research at the Reinforcement Learning Conference (RLC) 2026 in Montreal, Canada.
The Reinforcement Learning Conference is one of the leading peer-reviewed venues dedicated specifically to reinforcement learning research. Its Outstanding Paper Awards recognize a small number of papers each year across several categories, with the Empirical Reinforcement Learning Research award honoring work that makes significant contributions to the empirical practice of the field, including new methodologies, benchmarks, and evaluation techniques carried out with a high standard of experimental rigor.
The awarded paper addresses a long-standing tension in temporal-difference (TD) learning between stability and speed. Most TD methods use semi-gradient updates that learn quickly but can diverge, as illustrated by the classical Baird's counterexample. Gradient TD methods fix this stability issue but have historically been slower and therefore less widely adopted. Building on the recently introduced idea of iterated TD learning, which learns a sequence of action-value functions in parallel to speed up training, the authors propose Gradient Iterated TD learning: a method that computes gradients over the moving targets in this sequence, combining the stability of gradient TD methods with competitive learning speed. Evaluations across a range of benchmarks, including Atari games, show that the approach matches the speed of semi-gradient methods while retaining the theoretical guarantees of gradient TD — a result no prior gradient TD method has demonstrated.
Carlo D'Eramo leads the Reinforcement Learning and Computational Decision-Making group at CAIDAS, University of Würzburg.
