<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>测量学 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%B5%8B%E9%87%8F%E5%AD%A6/</link><description>Recent content in 测量学 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%B5%8B%E9%87%8F%E5%AD%A6/index.xml" rel="self" type="application/rss+xml"/><item><title>MCP 注册表随机抽样审计精读：48.8% 握手率背后的工具生态幸存者偏差</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-mcp-registry-random-draw-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-mcp-registry-random-draw-paper-reading/</guid><description>独立研究者 Haseeb Mohammed Afsar 对 MCP 注册表做首个未修复概率样本审计：24,135 服务器普查中概率抽 400 个 npm/stdio 服务器在线探测，仅 48.8% 完成 initialize 握手（手工精选框架 66.7%），37.5% 根本无法启动；能跑的 195 个硬一致性 100%，但安全注记缺失率 58.8% vs 精选 41.5%——整个领域的采样偏差第一次被量化。</description></item><item><title>The Double Measurement Confound in Agent Benchmarks 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-double-measurement-confound-benchmarks-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-double-measurement-confound-benchmarks-paper-reading/</guid><description>这篇来自西班牙团队的论文给 agent benchmark 的分数有效性下了诊断书：执行关键决策由固定 scaffold 而非模型做出（第一重测量混杂），scorer 用与任务正确性脱节的标准打分（第二重），两者合谋让排行榜测的是&amp;rsquo;评测管线属性&amp;rsquo;而非&amp;rsquo;模型能力&amp;rsquo;。论文提出测量论框架+审计修复协议三步——把执行决策移交给模型（de-scaffolding）、种子化金标评分替代形状匹配、用最差情形/尾部风险报告超越均值的可靠性。在 ComtradeBench 上：无 LLM 的规则基线得 96.8 分 vs Kimi/Claude 的 97.5，联合干预把平坦排行榜变成&amp;rsquo;平均性能×种子鲁棒性&amp;rsquo;的可靠性谱。</description></item></channel></rss>