<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>RLHF on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/rlhf/</link><description>Recent content in RLHF on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 23 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/rlhf/index.xml" rel="self" type="application/rss+xml"/><item><title>onPanda: 通过 Token 级纠错高效标注 LLM 与 Agent 的同策略对齐数据 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-onpanda-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-onpanda-paper-reading/</guid><description>onPanda（阶跃星辰 StepFun + 厦门大学）提出以 token 级纠错为核心的新型标注范式：标注者只需定位第一个不合适的 token，从模型候选集中点选或自由改写，系统随即截断后续内容并由模型从修正后的前缀继续生成（locate-correct-continue 循环）。该范式让绝大多数 token 由 rollout 模型原生生成，从而在低成本标注的同时高度保留同策略（on-policy）保真度，并自动产出可精细到位置的监督信号。本文按九部分结构精读其动机、系统设计、实验证据，并通过外部检索交叉验证标注效率、同策略数据价值与 RLHF 标注工具（Argilla/POTATO/Reptile）等相关工作。</description></item><item><title>Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-proof-carrying-cognition-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-proof-carrying-cognition-paper-reading/</guid><description>这篇论文提出一个普适论题：生成能力已规模化而验证没有——验证瓶颈是 AI 能力提升的普适约束，数据、算力、对齐等其它瓶颈最终都归约为它。理论上证明 verifier correlation 是&amp;rsquo;算力-能力汇率&amp;rsquo;（ρ=0.5 的验证器只实现可靠验证器 50.2% 的增益，与命题精确吻合）；实证上展示 asymmetric verification 攻击使不可靠验证器在优化压力下失效（hacking gap 0.27→~0），并提出 reality-settled reward——标签由现实执行结算而非 proxy 判分：GRPO 训练中冻结 RM 的 proxy 攀升而真实奖励崩溃 90%，RM 以 10% 结算流重训后执行奖励保持为冻结组 6×、on-policy 结算比随机标注标签效率高 10×。</description></item></channel></rss>