<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>代码评审 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E4%BB%A3%E7%A0%81%E8%AF%84%E5%AE%A1/</link><description>Recent content in 代码评审 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E4%BB%A3%E7%A0%81%E8%AF%84%E5%AE%A1/index.xml" rel="self" type="application/rss+xml"/><item><title>SPDF × Silent Failures 精读：LLM 代码安全评测的双警报日</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-spdf-silent-failures-eval-crisis-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-spdf-silent-failures-eval-crisis-paper-reading/</guid><description>同日两篇论文从两端夹击“静态/测试通过=安全”的假设：Toronto Metropolitan 的 SPDF 度量静态过-动态败缺口（654 个静态干净样本中 14.53% 被运行时利用验证击穿）；Tampere 大学的静默失败实证（1,030 条 Agent 修复轨迹中 170 例确认静默失败，Omission 占 48.2%）建立四维分类学。与 SWE-Gate、PatchBench 共同固化“测试通过≠安全”证据链。</description></item><item><title>RISE 之外的第二条线：When Models Edit Too Much — 代码过编辑与最小编辑保真度 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-over-editing-fidelity-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-over-editing-fidelity-paper-reading/</guid><description>NUS 团队构造 400 个带已知最小补丁的修复任务（BigCodeBench 注入 AST 级损坏），首次系统量化 LLM 代码修复的&amp;rsquo;过编辑&amp;rsquo;：GPT-5.5 等前沿模型普遍重写过度。保持性指令使超额 Levenshtein 距离 0.195→0.131、认知复杂度 -26.6%、Pass@1 +2.3；SFT 过拟合损坏模式而 RL 取得最佳 OOD 保真。本文精读&amp;rsquo;最小性&amp;rsquo;作为修复一等目标的评测与训练路径。</description></item><item><title>RealSWE 精读：真实用户请求正在让编码 Agent 榜单失真</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-06-realswe-realistic-user-requests-paper-reading/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-06-realswe-realistic-user-requests-paper-reading/</guid><description>RealSWE（成均馆大学）用六类信息分类学×四维语言风格对照 SWE-chat 真实用户 prompt 与 SWE-bench 任务，发现 88% 真实请求只带问题描述而基准任务仅 7%；据此构建 381 个多变体任务族，测得 7 个主流模型在真实输入下平均掉 6.4pp 且排行榜改写——MiMo V2.5 Pro 反超更贵模型升到第 2。控制变量消融进一步证明：Desired Behavior 字段值 8pp，复现步骤与环境信息几乎一文不值。</description></item><item><title>SWE-Gate 精读：通过功能测试对软件工程 Agent 并不够</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-swe-gate-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-swe-gate-paper-reading/</guid><description>SWE-Gate（中山大学/浙大/重大）从真实 PR 评审评论中提取约束并构造 303 个仓库级修复实例，发现 644 个通过功能测试的补丁中 221 个（34.3%）违反评审约束——SWE-bench 式功能唯一评测系统性高估了 Agent 的真实修复能力。本文基于全文逐页阅读，拆解其约束提取管线、双测试设计与 221 个隐藏失败的分布规律。</description></item></channel></rss>