ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

PACT:LLM强化学习token级credit,被三个条件唯一确定

PACT:LLM强化学习token级credit,被三个条件唯一确定 PACT: From Credit Assignment to Critic Alignment作者Jiayan Fu, Hang Xu, Yong Zhang, Zhaokai Luo, Yao Hu, Dongyan Zhao, Mu Chuan核心发表机构AllSpark Research论文链接arXiv:2609.26355v1发布于arXiv 预印本cs.LG|—————| Base Model | 46.67 | 51.25 | 28.63 | 37.50 | 41.01 || GRPO (ϵ h i g h 0.28 \epsilon_{\mathrm{high}}0.28ϵhigh​0.28) | 76.50 | 74.78 | 41.81 | 63.19 | 64.07 || PPO (λ 0.95 \lambda0.95λ0.95) | 32.33 | 29.38 | 17.86 | 25.67 | 26.31 || PPO (λ 1.0 \lambda1.0λ1.0) | 66.04 | 73.96 | 40.31 | 58.54 | 59.71 || SAO | 51.25 | 63.33 | 36.63 | 53.33 | 51.14 || PACT w/o IS | 76.04 | 82.50 | 51.38 | 61.04 | 67.74 || PACT | 83.12 | 85.21 | 51.69 | 71.46 | 72.87 |Agentic coding 主结果如下。在 SWE-bench Verified 上PACT 达到最高 pass rate 67.4%比 GRPO 高 2.0 个百分点比 PPOλ 1.0 \lambda1.0λ1.0高 2.4 个百分点比 SAO 高 3.8 个百分点。MethodBase ModelGRPOPPO (λ 1.0 \lambda1.0λ1.0)SAOPACTPass1 Rate60.865.465.063.667.4下图给出基础模型以及 PPO、GRPO、SAO、PACT 训练模型在数学推理与编码基准上的整体比较。)这些结果从三个层面支持论文主张。第一唯一 credit 表示指出的 critic 对齐问题并非纯理论细节而是影响最终性能。第二λ 1 \lambda1λ1在 PPO 内部已优于λ 0.95 \lambda0.95λ0.95而 PACT 在此基础上进一步通过 Actor-then-Critic 与重要性校正改进了 critic 对齐。第三PACT 在数学推理与 coding 两类差异较大的 agentic 任务上都取得提升说明方法并不局限于单一环境或单一奖励类型。4.3 消融实验 / Ablation Study消融主要围绕两个设计critic 目标选择与重要性采样校正。第一critic objective 消融。论文进行固定策略 critic 预训练从相同初始化、相同 on-policy rollout 数据训练 BCE 与 MSE critics并在 Qwen3.5-4B 数学推理与 Qwen3.6-35B-A3B agentic coding 两种设置下比较。结果是在两个模型规模上BCE-trained critics 都取得更低的 BCE 和 MSE losses以及更大的 value separation。Value separation 定义为 positive samples 上的平均预测值减去 negative samples 上的平均预测值记为Δ ± \Delta_{\pm}Δ±​。更大的 separation 说明 BCE-trained critics 对成功与失败轨迹赋予更不同的值。这一消融支持了 PACT 使用概率 BCE critic 而非传统 MSE critic 的选择。第二importance sampling correction 消融。PACT w/o IS 保留 Actor-then-Critic 更新顺序、BCE critic objective、actor-side 优化设置但移除 critic-target importance sampling correction。加入 importance correction 后平均准确率从 67.74% 提升到 72.87%增益 5.13 个百分点且在四个 benchmark 上均有提升。训练也更稳定见训练奖励曲线。该比较支持 PACT 中 importance sampling correction 的有效性。此外论文对三个正则条件进行了理论层面的“消融”附录证明 Completeness、Prefix Consistency、Neutrality 均不可由另外两个推出。这一必要性分析说明唯一表示并非依赖冗余条件而是需要三个条件共同排除三类病理行为。因此消融不仅体现在算法组件上也体现在理论公理的选择上。4.4 训练动态与稀疏性证据训练动态方面论文报告了响应长度与平均∣ V ^ t − V ^ t − 1 ∣ |\widehat V_t-\widehat V_{t-1}|∣Vt​−Vt−1​∣在 PACT 训练中的变化。在 PACT 训练中响应逐渐变长同时平均∣ V ^ t − V ^ t − 1 ∣ |\widehat V_t-\widehat V_{t-1}|∣Vt​−Vt−1​∣下降。这一趋势与近似 credit 稀疏性定理的直觉一致随着序列变长单个 token 上的局部 credit 增量趋于变小而结果奖励方差被分散到更多 token 上。下图展示该动态。训练奖励方面论文给出 Qwen3.5-4B 在 agentic mathematical reasoning 上的训练奖励曲线以及 Qwen3.6-35B-A3B 在 agentic coding 上的训练奖励曲线。图中浅色线为原始平均轨迹奖励粗线为指数移动平均系数α 0.3 \alpha0.3α0.3。数学推理图中每个 rollout round 对应四个 training steps。这些曲线用于说明加入 importance correction 后训练更稳定并与主结果一致。))从实验整体看论文的证明链条是理论给出唯一 credit 表示该表示揭示 critic 对齐的重要性GAE 分析说明λ 1 \lambda1λ1可消除中间 critic 误差PACT 通过 Actor-then-Critic 与重要性校正改进 critic 对齐消融显示 BCE 与重要性校正各自有效主结果在数学推理与 coding 上均取得提升。需要指出的是给定片段中未给出完整消融表格、逐基准的 PACT w/o critic IS 数值也未展开 actor 更新内部损失的完整形式因此本文不补造这些细节。五、相关工作 / Related Work在 LLM 强化学习方面RL 已成为 LLM post-training 的重要部分支持 preference alignment 与复杂推理能力。早期应用主要关注 preference alignment代表性 RLHF 从人类偏好训练 scalar reward models再用 PPO 优化语言模型策略其 PPO 实现使用 learned value functions 构造 token-level advantage estimates。近期 RL 越来越多用于通过 outcome-level 或 verifiable rewards 改进复杂推理。Critic-free 方法包括 ReMax用 greedily decoded response 的 reward 作为 baselineRLOO用同一 prompt 下其他 responses 的 rewards 构造 leave-one-out baselineGRPO用 group-relative advantage estimates 替代 learned value modelREINFORCEcritic-free 优化加 global advantage normalizationDAPO改进 clipping、sampling、loss aggregationDr. GRPOnormalization-bias correctionGSPOsequence-level importance weighting and clipping。Value-based 方法包括 VC-PPO将 PPO 在 long-chain-of-thought 训练中的失败归因于 value-initialization bias 和 GAE 中 terminal reward signals 的衰减引入 value pretraining 与 decoupled GAEVAPO在这些技术基础上引入 length-adaptive GAE异步训练系统包括 AReaL解耦 rollout generation 与 policy optimization显式处理 stale training samplesSAO结合 single-rollout asynchronous training、learned value model、token-level GAE estimator。与这些研究不同本文工作发展了 long-sequence LLM RL 中 token-level credit 的一般刻画并研究其对 response-level baselines 与 learned critics 的启示。在 credit assignment 方面credit assignment 关注早期决策如何影响后续结果是学习系统的基本困难。经典方法包括 Monte Carlo用 sampled returns 作为 value estimation 目标REINFORCE-style policy gradient用 observed rewards/returns 构造梯度估计TD learning从时间上相继的预测中 bootstrappingEligibility traces把后续 TD errors 分配给先前访问过的状态和动作Actor-critic用 learned value functions 估计 policy gradientsGAE组合 TD residuals 控制 advantage estimation 的 bias-variance trade-off。长期延迟场景中RUDDER 通过 return predictions 的 contribution analysis 学习 return-equivalent reward redistributionsTemporal Value Transport 用 attentional memory retrieval 将价值估计从后续事件传输到远处相关事件Hindsight Credit Assignment 根据过去决策导致观测结果的 likelihood 分配 creditCounterfactual Credit Assignment 构造 future-conditioned baselines解耦动作影响与外部因素及后续动作COCOA 通过询问在替代动作下后续 reward 是否仍会获得估计动作贡献。Pignatelli 等人的 survey 将这些发展组织为 temporal contiguity、return decomposition、future conditioning 等家族。这些方法通过不同目标量和估计过程处理延迟奖励传播、轨迹 return 分解或动作影响估计。在 LLM 长时程 RL 中该问题尤其突出许多 outcome-supervised 方法用单个 response-level 或 terminal reward 训练大量 token-level 决策。本文研究三个 regularity conditions 下 credit 的数学刻画并用所得唯一表示分析现有算法、指导 critic training。与具体方法的区别也值得强调。与 PPO 相比PACT 指出 PPO-style 训练存在 actor–critic 策略滞后critic 值来自上一策略且更新后不重算当前 actor advantage。PACT 采用 Actor-then-Critic 更新顺序对 critic 训练进行重要性采样校正使 critic 与更新后策略对齐并用 BCE 替代 MSE使用λ 1 \lambda1λ1。与 GRPO/RLOO 相比RLOO 的响应级基线在期望上与唯一 token-level credit 的梯度贡献一致但仅是期望等价不保证统计效率PACT 则追求更细粒度、更准确的价值估计与 credit 对齐。与 OPD 相比已有工作表明 OPD 目标可代数重写为密集 KL 约束 RLteacher–student 对数密度比扮演 token-level advantage本文进一步建立 credit 级连接理想 teacher 下的 OPD 更新在期望上与唯一 token-level credit 诱导的策略梯度成正比。与 SAO 相比SAO 采用长度自适应 GAE 系数随响应长度增加趋近 1PACT 基于误差分解直接采用λ 1 \lambda1λ1因为λ 1 \lambda1λ1可消除中间评论家误差项仅保留前缀价值误差。与 GAE 相比论文证明在近似 credit 稀疏性下λ 1 \lambda1λ1时中间 critic 误差可能变得与真实 credit 信号相当甚至主导λ 1 \lambda1λ1时中间误差项消失。六、局限性与展望 / Limitations Future Work论文自述的局限性有三点。第一唯一性结果条件于所提出的正则条件。本文对 credit 的刻画并不意味着所有可能的 credit 概念都必须采取本文推导的形式。唯一性结果条件于 Completeness、Prefix Consistency、Neutrality。正如替换欧几里得平行公设会产生不同但内部一致的非欧几何采用不同的 credit 要求也可能导致不同表示。第二本文研究的 credit 是统计的而非因果的。它通过条件期望定义不刻画替换单个 token 或 action 的反事实因果效应。第三论文未解决实际 LLM RL 中的精确估计问题。本文刻画了在所提条件下 token-level credit 应该是什么但没有解决在实际 LLM 强化学习中精确估计它的问题。开发准确且高效的 token-level credit 估计器仍是重要未来方向。从理论假设看唯一表示定理假设R RR可积且F τ \mathcal F_\tauFτ​-可测τ ≤ L \tau\le Lτ≤L为停止时间。理想教师定义假设R RR有界且p t p_tpt​在词表V \mathcal VV上全支撑q t ⋆ q_t^\starqt⋆​相对于当前策略定义并在相应 OPD 更新期间固定。GAE 误差分析假设奖励可归一化到[ 0 , 1 ] [0,1][0,1]取γ 1 \gamma1γ1终端预测V ^ τ R \widehat V_\tauRVτ​R终端价值误差为零。RLOO 等价性仅在期望意义上成立不意味着统计效率相同。PACT 推导假设环境动态不变、continuation distributions 绝对连续精确 continuation ratio 在长回复中可能有高方差因此实践采用 detached current-token importance ratio 并进行 mask。这些假设并不自动在所有真实 LLM RL 设置中成立尤其当环境动态变化、奖励无界或 continuation 分布不绝对连续时理论保证需要重新审视。展望方面可以沿多条路径推进。其一开发更准确且高效的 token-level credit 估计器避免长视界下 importance ratio 方差过大同时保持 critic 与更新后策略同步。其二研究三正则条件之外的替代公理体系探索不同但内部一致的 credit 概念并将其与因果 credit assignment 结合以刻画反事实 token 贡献。其三将 PACT 的 Actor-then-Critic 与重要性校正思想扩展到异步、分布式或超长上下文训练处理 stale samples 与推理-训练不匹配。其四在更多任务、更多模型规模与更多奖励类型上验证 BCE critic、λ 1 \lambda1λ1与重要性掩码的鲁棒性。其五进一步研究 credit 稀疏性对采样效率、优势估计方差与训练稳定性的影响从而设计自适应λ \lambdaλ、自适应 critic 更新频率或分层 credit 估计方案。七、总结 / Conclusion论文从 token-level credit assignment 缺乏通用数学定义这一问题出发提出 Completeness、Prefix Consistency 与 Neutrality 三个正则条件并证明它们唯一确定 token-level credit 为条件奖励预测的增量C i V i − V i − 1 C_iV_i-V_{i-1}Ci​Vi​−Vi−1​且该序列是鞅差序列。这一表示提供了一个统一视角OPD 中的理想教师可视为隐式 critic其期望策略梯度与唯一 credit 诱导的梯度成正比RLOO 的响应级基线尽管粒度更粗其期望策略梯度贡献仍与 token-level credit 匹配在有界结果奖励下credit 具有近似稀疏性而 GAE 中λ 1 \lambda1λ1时的中间 critic 误差可能变得与真实 credit 相当甚至主导λ 1 \lambda1λ1则可消除这些中间误差。基于这些分析论文提出 PACT。PACT 采用 Actor-then-Critic 更新顺序在 actor 更新后对同一 rollout batch 做额外前向用当前 token 重要性比构造校正后的 critic 目标Y ρ ⊙ R Y\rho\odot RYρ⊙R并仅在接受范围内计算归一化 BCE 损失。概率 critic 与 BCE 目标在不改变最优预测的前提下改善了价值估计经验表现而重要性校正使 critic 更好地对齐更新后的策略。实验上PACT 在 agentic mathematical reasoning 四个基准上平均准确率 72.87%超过 GRPO 与 PPO 分别 8.80 与 13.16 个百分点在 SWE-bench Verified 上通过率 67.4%超过 PPO、GRPO、SAO 分别 2.4、2.0、3.8 个百分点。消融显示 BCE critic 与重要性采样校正均有贡献训练动态也与近似 credit 稀疏性一致。总体而言论文的核心价值在于它既给出了 token-level credit 的数学刻画又用这一刻画诊断了现有 actor-critic 训练中的 critic 对齐问题并据此设计出在数学推理与 coding 任务上均有效的 PACT。原文摘要:Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit. This characterization provides a unified basis for explaining phenomena across existing algorithms and guides the development of an improved actor-critic training procedure. Through this lens, an ideal teacher in On-Policy Distillation (OPD) acts as an implicit critic, yielding an expected policy gradient proportional to that induced by token-level credit. Response-level REINFORCE Leave-One-Out (RLOO) signals match the expected policy-gradient contribution of token-level credit despite their coarser granularity. We further establish approximate credit sparsity under bounded outcome rewards and show how intermediate critic errors in Generalized Advantage Estimation (GAE) can become comparable to the underlying credit. These motivate Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy. In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively. On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.PDF链接:https://arxiv.org/pdf/2609.26355v1部分平台可能图片显示异常请以我的博客内容为准
返回列表