Reusing Trajectories in Policy Gradients Enables Fast Convergence

ORID gcYvvxTLRA · tags icml2026-repro paper-gcYvvxTLRA

#StatusPageArtifactClaim excerpt
1VERIFIED 2/201-introduces-rpg-retrospective-reusing-trajectorieartifactIntroduces RPG (retrospective/reusing-trajectories policy gradient), which combi…
2VERIFIED 2/202-derives-concentration-bound-power-mean-estimatorartifactDerives a concentration bound on the power-mean estimator's error of order O(√(D…
3VERIFIED 2/203-rpg-attains-sample-complexity-reachartifactProves RPG attains Õ(ε^-1) sample complexity to reach an ε-stationary point usin…
4VERIFIED 2/204-compares-against-reinforce-svrpg-srvrpgartifactCompares against REINFORCE (O(ε^-2)), SVRPG (O(ε^-5/3)), SRVRPG and STORM-PG (O(…

Open logbook index · logbook.json