D2 Actor Critic: Diffusion Actor Meets Distributional Critic

Lunjun Zhang; Shuo Han; Hanrui Lyu; Bradly C. Stadie

D2 Actor Critic: Diffusion Actor Meets Distributional Critic

Lunjun Zhang, Shuo Han, Hanrui Lyu, Bradly C. Stadie

Published: 06 Oct 2025, Last Modified: 06 Oct 2025Accepted by TMLREveryoneRevisionsBibTeXCC BY 4.0

Abstract: We introduce D2AC, a new model-free reinforcement learning (RL) algorithm designed to train expressive diffusion policies online effectively. At its core is a policy improvement objective that avoids the high variance of typical policy gradients and the complexity of backpropagation through time. This stable learning process is critically enabled by our second contribution: a robust distributional critic, which we design through a fusion of distributional RL and clipped double Q-learning. The resulting algorithm is highly effective, achieving state-of-the-art performance on a benchmark of eighteen hard RL tasks, including Humanoid, Dog, and Shadow Hand domains, spanning both dense-reward and goal-conditioned RL scenarios. Beyond standard benchmarks, we also evaluate a biologically motivated predator-prey task to examine the behavioral robustness and generalization capacity of our approach.

Submission Length: Regular submission (no more than 12 pages of main content)

Video: https://d2ac-actor-critic.github.io/

Assigned Action Editor: ~Manuel_Haussmann1

Submission Number: 5221

Loading