<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>信用分配 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E4%BF%A1%E7%94%A8%E5%88%86%E9%85%8D/</link><description>Recent content in 信用分配 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 23 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E4%BF%A1%E7%94%A8%E5%88%86%E9%85%8D/index.xml" rel="self" type="application/rss+xml"/><item><title>Critical-State RL：为多轮工具调用诊断「可训练的模型调用」 —— 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-critical-state-rl-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-critical-state-rl-paper-reading/</guid><description>深度精读 Salesforce AI Research 的 Critical-State RL。它指出多轮工具调用失败往往卡在单个模型调用上，但「奖励有差异」并不等于「该调用值得训练」。方法先用训练前三闸门诊断（动作充分性 / 提升空间 / 可训练性），再用嵌套同前缀采样分离「动作依赖奖励方差」与「后续噪声」，最后只对选中的关键调用做 occurrence-local RL（上下文赌博机式训练）。BFCL 上对缺失函数任务 miss_func 恢复率 0.14→0.283（+14.3pp），错位训练反而 −4.5pp；记忆子任务 34.54%→50.54%；Nemotron 重复调用一致性 37%→75%。</description></item><item><title>FLARE：用生成式奖励模型为长程编码智能体提供全生命周期稠密监督 —— 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-flare-grm-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-flare-grm-paper-reading/</guid><description>深度精读北京大学、南京大学、北京邮电大学与独立研究者联合提出的 FLARE。它通过 RADAR 双轨数据合成训练一个 4B 生成式奖励模型（GRM），输出结构化的步级风险诊断（风险等级+错误类+修复建议），在推理时做断点再生、训练时做 SFT 筛选与 RL 稠密奖励，首次把测试时干预与训练时对齐闭环到同一套诊断信号。F2P Pass@1 14.10% 近乎翻倍于全局重采样 7.80%，token 省 5×；ROC-AUC 75.74%；SFT 相对提升 19.13%，RL 平均提升 9.19%。</description></item><item><title>DRACO 精读：没有验证器时，如何给长程 Agent 训练信号分步定责</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-06-draco-outcome-blind-credit-assignment-paper-reading/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-06-draco-outcome-blind-credit-assignment-paper-reading/</guid><description>DRACO（IBM×CMU）形式化&amp;rsquo;outcome-blind&amp;rsquo;训练设定——长程 Agent 任务往往没有程序化验证器可依赖。方法用训练中动态生成的 rubric 逐轨迹打一次分，再按&amp;rsquo;步骤涉及哪些标准&amp;rsquo;闭式分摊到每步 GRPO advantage，不引入任何可学习归因模块。AppWorld TN 上 Qwen3.6-27B TGC/SGC 69.4/41.1→85.3/70.6，反超偷看真值奖励的 GRPO +5.3/+11.3，τ-bench 零样本迁移 SR 15.8→20.4。</description></item><item><title>HarnessEvo 精读：Harness 自进化的价值藏在控制槽位里</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-harnessevo-value-localization-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-harnessevo-value-localization-paper-reading/</guid><description>HarnessEvo 把 Agent harness 分解为 role/strategy/format/control 四个可独立进化的槽位，用 leave-one-in/out 协议做价值归因：整体指标&amp;rsquo;看似无效&amp;rsquo;（0.657 vs 0.642），但收益完全 localized 于 reflection/control 槽位（+0.119, p=0.0046）；等预算下多槽位同进反而互相稀释——预算分摊陷阱。本文基于全文阅读拆解其归因协议与对自进化领域的方法论警示。</description></item></channel></rss>