<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>奖励模型 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E5%A5%96%E5%8A%B1%E6%A8%A1%E5%9E%8B/</link><description>Recent content in 奖励模型 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E5%A5%96%E5%8A%B1%E6%A8%A1%E5%9E%8B/index.xml" rel="self" type="application/rss+xml"/><item><title>训练前沿四重奏：蒸馏外推、规模解除、奖励治理与具身评测 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-02-training-frontier-quartet-paper-reading/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-02-training-frontier-quartet-paper-reading/</guid><description>本篇合读 2026 年 9 月底集中出现的四篇训练侧前沿论文：RIDE 把「教师是方向不是终点」操作化为表示空间的 RL 残差外推，四组基座全部追平或超越教师；OASIS 诊断出在线自蒸馏的规模瓶颈源于「在教师不可模仿处监督」，用验证脚手架解除；ProRubric 把 Rubric-RL 奖励攻击的根因从准则文本转移到聚合方式，合取式协议级评分回收三分之一损失；GPT-6 Astra 具身评测则以六域系统实验界定前沿模型「任务决策强、物理控制弱」的能力边界。四篇合起来构成一条主线：监督信号放对了空间、放对了状态、用对了聚合、并诚实面对能力边界。</description></item><item><title>Diffusion Reward Models × SOLO：成熟技术的首次规模化双案例 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-paradigm-first-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-paradigm-first-duet-paper-reading/</guid><description>同日两篇论文在各自领域做了同一件事：把一项被搁置多年的成熟技术，用一个机制创新首次推到现代规模。清华 thunlp+港中文+UIUC 的 Diffusion Reward Models（DRM，当日 Hugging Face 日榜 up=22）把扩散模型首次引入奖励建模——奖励从点估计变为条件密度估计 p(r|x,y)，轻量 DiT 头无参数族假设地表示人类偏好的多峰结构（多峰率随标注分歧 37.6%→63.2%，Wasserstein 距离全表最低），同骨干同数据受控对比平均 66.2 vs ArmoRM 62.3（+3.9）；中科院自动化所的 SOLO 把局部学习（local learning）首次推到十亿参数 LLM 预训练——用一个所有模块共享的终端 readout 只读副本打破 update locking，梯度对齐 cos 从 0.52 升至 0.70，流水线激活内存降近 p 倍、吞吐达 1F1B 的 1.44×。二者方法论同构：识别被搁置的旧技术、诊断其规模化障碍、用一个机制创新解锁，为「旧思想在新规模下复活」提供了可复用的模板。</description></item><item><title>Critical-State RL：为多轮工具调用诊断「可训练的模型调用」 —— 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-critical-state-rl-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-critical-state-rl-paper-reading/</guid><description>深度精读 Salesforce AI Research 的 Critical-State RL。它指出多轮工具调用失败往往卡在单个模型调用上，但「奖励有差异」并不等于「该调用值得训练」。方法先用训练前三闸门诊断（动作充分性 / 提升空间 / 可训练性），再用嵌套同前缀采样分离「动作依赖奖励方差」与「后续噪声」，最后只对选中的关键调用做 occurrence-local RL（上下文赌博机式训练）。BFCL 上对缺失函数任务 miss_func 恢复率 0.14→0.283（+14.3pp），错位训练反而 −4.5pp；记忆子任务 34.54%→50.54%；Nemotron 重复调用一致性 37%→75%。</description></item><item><title>FLARE：用生成式奖励模型为长程编码智能体提供全生命周期稠密监督 —— 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-flare-grm-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-flare-grm-paper-reading/</guid><description>深度精读北京大学、南京大学、北京邮电大学与独立研究者联合提出的 FLARE。它通过 RADAR 双轨数据合成训练一个 4B 生成式奖励模型（GRM），输出结构化的步级风险诊断（风险等级+错误类+修复建议），在推理时做断点再生、训练时做 SFT 筛选与 RL 稠密奖励，首次把测试时干预与训练时对齐闭环到同一套诊断信号。F2P Pass@1 14.10% 近乎翻倍于全局重采样 7.80%，token 省 5×；ROC-AUC 75.74%；SFT 相对提升 19.13%，RL 平均提升 9.19%。</description></item><item><title>StudentSim: Training LLM-based Student Simulators 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-21-studentsim-paper-reading/</link><pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-21-studentsim-paper-reading/</guid><description>微软研究院与 UIUC 的 StudentSim 把&amp;rsquo;AI 学生模拟器&amp;rsquo;形式化为可优化的双目标问题：行为保真度 F（复现特定学生的真实行为）与指导响应度 R（被导师教会的能力）。两阶段训练——跨学生池化预训练学共享模式 + 每生 LoRA 特化——让 Qwen3-4B 在国际象棋、二语写作、数学三个领域 F/R 双指标全面超过 prompted GPT-5.4（chess 0.51/0.91 vs 0.23/0.72），用其做奖励的导师 RL 经专家盲评三轴全胜（准确率 90.5% vs GPT-5.4 奖励的 71.6%）。本精读覆盖 F×R 分解的问题化、&amp;lsquo;池化贡献多样性而非更新量&amp;rsquo;的消融证据、4B 特化胜过前沿 API 的机制根源，以及&amp;rsquo;模拟器即基础设施&amp;rsquo;的通用性灵感。</description></item><item><title>EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-evors-self-evolving-reward-systems-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-evors-self-evolving-reward-systems-paper-reading/</guid><description>复旦大学针对开放 RL 的奖励系统自进化框架 EvoRS：rubric 奖励与 policy 构成动态反馈回路——policy 优化当前奖励时，初始有用的奖励系统会因 reward hacking 或区分度退化而失效。EvoRS 把奖励系统表示为可执行 Reward-DAG，agentic designer 从 on-policy rollout 与奖励轨迹更新它。写作/角色扮演任务上三种 judge 下质量最佳，超固定奖励 policy 2.107/4.767 分，reward hacking 与覆盖失败双降。</description></item><item><title>GenV 精读：把 Z3 等价性判定蒸馏成语言模型的“第六感”</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-genv-generative-reward-autoformalization-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-genv-generative-reward-autoformalization-paper-reading/</guid><description>CWRU×AWS（实习合作）提出 GenV：把离线 Z3 等价 oracle 蒸馏为 reference-free 的连续等价分，专治 autoformalization 中“表面合法但语义错位”的欺骗性轨迹。GenV+HN 在 VPU 检测上 F1 0.832（process RM 仅 0.246），logit lens/SAE 机制分析证明信号真实存在于残差流而非捷径。</description></item><item><title>HackProbe: 自进化语言模型的奖励黑客检测与免疫 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-hackprobe-reward-hacking-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-hackprobe-reward-hacking-paper-reading/</guid><description>Fullive-AI×北大×京东×NTU×武大的 HackProbe 是一个通过两个黑盒钩子挂载到任意自进化回路的监控器：秘密固定分布对比核心保证跨代可比，轮换新鲜层抗共适应；四项检验+Šidak 校正输出族校准 p 值，风险感知免疫层从候选池重选诚实更新。本文精读&amp;rsquo;诊断之外还能恢复&amp;rsquo;的奖励黑客治理闭环。</description></item><item><title>Cliff: Learning Process Rewards from the First Mistake 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-cliff-first-mistake-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-cliff-first-mistake-paper-reading/</guid><description>RLVR 用最终答案的对错给整条推理链打分，粒度太粗；PRM 要训练专用奖励模型，OPD 要求师生推理同构。Amazon Web Services 联合 UIUC 的 Cliff 给出过程监督的最小充分形式：只用现成 LLM 定位每个错误 rollout 的&amp;rsquo;第一个错误步&amp;rsquo;（Pitfall Step），把轨迹切成正确前缀与错误后缀，前缀正优势、后缀负优势。12 个场景上一致超越：比 On-Policy Distillation 高约 15%、比标准 GRPO 高约 7%，且弱教师（27B）下依然有效。本精读拆解&amp;rsquo;错误前缀之后无信息&amp;rsquo;这一核心洞察、token 级优势的构造细节，以及教师定位能力与人类标注 80% 一致率的验证实验。</description></item><item><title>Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-ns-prm-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-ns-prm-paper-reading/</guid><description>工具增强 LLM 已经能消灭绝大多数计算与语法错误，但一个隐蔽的失败模式仍在：公式用对了变量却用错了——推理步语法完好、数学可执行、量纲一致，却完全脱离题意。本文提出 NS-PRM，把『推理正确』解耦为符号有效性 V 与语义接地性 G 两个维度：确定性符号验证器（JSON Schema 语法检查+数学等价+Pint 量纲分析，覆盖 128 个计算原语与 15 个逻辑原语）作为硬过滤器保证 V；PRM 只在验证器接受的空间上对 G 条件评分。训练侧引入反事实符号扰动 CSP，算法化生成『完美通过验证器但逻辑错误』的硬负例；推理侧采用验证器优先的约束 beam 搜索。最终 ProcessBench 平均 F1 达 74.2、PRMBench 77.3，均超 9 个 PRM 基线；同等算力下 beam 宽度近翻倍，MATH 零额外成本提升 4.5 个百分点。</description></item><item><title>Privacy Without Regret: Differentially Private Inference-Time Alignment 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-privbon-dp-alignment-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-privbon-dp-alignment-paper-reading/</guid><description>印度理工坎普尔提出推理时对齐的差分隐私方法 PrivBoN 与 PrivITP。核心洞察是：差分隐私与抗 reward hacking 本质上是同一个干预——把 Best-of-N 的硬 argmax 软化。PrivBoN 在奖励分数上加尺度 σ=2Δr/ε 的 Gumbel 噪声，等价于指数机制实现 ε-DP，同时等价于 KL 正则化对齐；当隐私预算超过阈值 ε* 时，隐私要求的噪声恰好就是对齐最优的正则，隐私零成本。PrivITP 进一步用 χ² 正则化拒绝采样加两阶段高斯机制，把正则参数与隐私参数解耦，隐私代价只随实际停时增长。实验中弱奖励模型下 BoN 出现负提升（GSM8K 上 −5.18%），而 PrivITP 反而 +0.81%，并在固定隐私预算下靠 FSRC 组合多答约 3 倍查询。</description></item><item><title>J-Zero: Unified Challenger-Solver-Judge Co-Evolution from Zero Data 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-jzero-judge-coevolution-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-jzero-judge-coevolution-paper-reading/</guid><description>自进化大模型在可验证领域已有成熟方案，但在没有标准答案的开放域，学习信号只能来自打分模型，而固定的Judge只能把Solver推到自己内化偏好的上限，饱和后奖励失去区分度。J-Zero让Challenger、Solver、Judge三者零数据共同进化：出题者与解题者通过GRPO对抗博弈，Judge依靠角色不对称与子任务放大两类结构性偏好对做Bradley-Terry更新，标签由构造方式先验决定而非Judge打分，避免自我强化偏差。Qwen3-4B上可验证域提升9.47分、不可验证域提升11.23分，超越R-Zero 4.74分，并持续改进十个迭代而baseline两迭代后退化。本文按九部分结构精读其动机、机制、实验证据与可迁移灵感。</description></item><item><title>Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-thermodpo-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-thermodpo-paper-reading/</guid><description>本文精读西湖大学、浙江大学与快手可灵团队合著的 arXiv 2608.20011。论文形式化了流匹配生成模型偏好优化中的流形漂移现象：偏好更新会把轨迹终端样本推离预训练数据流形，产生非零法向位移，这是 reward hacking 的结构性根因。理论上，最优流匹配能精确恢复数据分布，而偏好更新一旦含非零法向分量必然离流形。方法上提出 THERMODPO，用带温度的三态能量 softmax 在胜者侧加流形锚，统一了 RFT 与 FlowDPO，并对流形距离给出逐点上界。实验显示玩具基准 StrictScore 达 0.899，SD3.5-M 四指标宏平均提升 16.0%、OCR 提升 47.5%，FLUX.2 上复现同趋势。</description></item><item><title>OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-08-osreward-paper-reading/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-08-osreward-paper-reading/</guid><description>深度精读港大与腾讯联合出品的 OSReward——首个系统检验计算机使用 Agent（CUA）轨迹评判器可靠性的基准与开源奖励模型工作。论文构建了覆盖 Web/Windows/macOS/Ubuntu/Mobile 五大平台的 1,019 条人工标注轨迹基准，揭示了所有主流 VLM 评判器存在的系统性「宽容偏差」，并训练出成本降低 30-60 倍的开源奖励模型 OS-Shepherd。本精读从背景补全、研究脉络定位、问题抽象、方法机制、评估证据、效果根源、必要知识反推到可推广灵感，九部分完整拆解这项为 CUA 评判建立标准化评估体系的开创性研究。</description></item><item><title>Verbalizable Representations Form a Global Workspace in Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-07-global-workspace-paper-reading/</link><pubDate>Tue, 07 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-07-global-workspace-paper-reading/</guid><description>Anthropic团队在Claude模型内部发现了一个类似人脑全局工作空间的特权表示子空间——J-space。它由一小撮不断演变的&amp;rsquo;未说出的词语&amp;rsquo;组成，仅占激活方差不到10%，却承担着言语报告、内部推理、灵活泛化和自我监控的核心功能。本文深度解析Jacobian Lens技术、J-space的五个功能属性、以及反事实反思训练这一全新对齐范式。</description></item><item><title>QUBRIC: 协同设计查询与评分标准，突破可验证奖励的强化学习瓶颈 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-10-qubric-paper-reading/</link><pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-10-qubric-paper-reading/</guid><description>深度精读 Amazon 与佐治亚理工学院联合发表的 QUBRIC 论文——一个协同设计查询（Query）和评分标准（Rubric）的框架，通过关键点锚定的查询重写、对比式评分标准生成和可学习性过滤，突破传统 Rubric-based RL 中&amp;rsquo;评分标准质量受限于查询结构&amp;rsquo;的根本瓶颈。在 ArenaHard 上取得 +5.5 的增益，并零样本迁移到法律、道德、叙事推理三个未见基准上，平均提升 +6.3 分。本文揭示了查询与评分标准的结构耦合关系，为将 RL 扩展到不可验证领域提供了实用路径。</description></item><item><title>Skill-RM: 通过 Agent Skill 统一异构奖励评估标准 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-10-skill-rm-paper-reading/</link><pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-10-skill-rm-paper-reading/</guid><description>深度精读中山大学、香港中文大学、北京大学、ETH苏黎世与阿里巴巴通义千问团队联合发表的 Skill-RM 论文——将奖励建模重新定义为可复用的&amp;rsquo;奖励评估技能&amp;rsquo;（Reward-Evaluation Skill），通过结构化的 Agent 技能编排异构评估资源（评分准则、验证器、检查清单、聚合规则），在 RewardBench2、RM-Bench、JudgeBench 三大基准上全面超越传统 LLM-as-a-Judge 和专用奖励模型，为 LLM 后训练提供统一、可解释、可扩展的奖励信号框架。</description></item></channel></rss>