连续延迟记忆随机梯度下降与连续时间强化学习
原标题:Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies
AI 摘要
该论文综述了随机微分方程(SDE)与神经网络参数化在天体物理类星体光变曲线建模中的历史应用,并在此基础上提出连续延迟记忆随机梯度下降(Continuous-Delayed-Memory SGD),在二维景观模拟中相比普通 SGD 展现出更广的探索性和更精确的收敛行为。作者还提出一种无需求解 HJB 偏微分方程的连续时间策略梯度强化学习结构,并证明其最优性条件可恢复 Gibbs 策略。
正文节选
Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies Debartha Paul Juncheng Yi debartha@iastate.edu ccyi3@iastate.edu May 8th 2026 Abstract Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves