<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>基准 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E5%9F%BA%E5%87%86/</link><description>Recent content in 基准 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E5%9F%BA%E5%87%86/index.xml" rel="self" type="application/rss+xml"/><item><title>IdeaAMBIG 精读：从论文想法到能跑的代码之间，隔着 660 个“没人写的细节”</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-ideaambig-research-idea-specs-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-ideaambig-research-idea-specs-paper-reading/</guid><description>Yale×TUM×腾讯发布 IdeaAMBIG：660 个证据落地的“实现关键缺口”基准（163 个真实缺口来自可复现性报告与 GitHub issue + 497 个受控合成缺口），评测 LLM 能否发现科研想法规格中的缺失决策。13 个 LLM 最好者 Macro Defect Recovery 仅 9.6%——想法到实现的鸿沟被首次量化，且当前模型几乎看不见它。</description></item></channel></rss>