<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>训练范式 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E8%AE%AD%E7%BB%83%E8%8C%83%E5%BC%8F/</link><description>Recent content in 训练范式 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E8%AE%AD%E7%BB%83%E8%8C%83%E5%BC%8F/index.xml" rel="self" type="application/rss+xml"/><item><title>NSD 精读：教推理模型“别这么错”，比教它“该怎么对”更有效</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-nsd-negative-self-distillation-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-nsd-negative-self-distillation-paper-reading/</guid><description>UVA+Stanford 提出负面自蒸馏（NSD）：针对 on-policy 自蒸馏（OPSD）模仿“带标准答案的伪自信轨迹”导致难推理任务退化的失败模式，构造“负条件”（注入已知缺陷的解）让学生显式规避。token 级自适应门控+gated unlikelihood 在七基准上 1.7B/4B/8B 平均 +2.3%/+7.5%/+6.0%，且保留自纠错行为——模型越大增益越高。</description></item></channel></rss>