<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>数据质量 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%95%B0%E6%8D%AE%E8%B4%A8%E9%87%8F/</link><description>Recent content in 数据质量 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%95%B0%E6%8D%AE%E8%B4%A8%E9%87%8F/index.xml" rel="self" type="application/rss+xml"/><item><title>When Synthetic Data Hurts 精读：Agent 技能检索器的合成数据灾难遗忘</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-synthetic-data-hurts-skill-retrieval-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-synthetic-data-hurts-skill-retrieval-paper-reading/</guid><description>Manulife（加拿大金融集团）实证研究：Agent 技能检索器用合成任务微调后，最激进配置下 OOD recall 从 0.850 跌至 0.650（−20pp）；合成数据使 Hit@10 持平但 Recall@10 −0.021——部分重排把额外正确技能挤出 top-10。同时证明 0.6B 紧凑检索器可追平更大混合系统：监督质量&amp;gt;模型规模。skill 数据飞轮假设的第一份系统性反例。</description></item></channel></rss>