<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>代码生成 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E4%BB%A3%E7%A0%81%E7%94%9F%E6%88%90/</link><description>Recent content in 代码生成 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sun, 06 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E4%BB%A3%E7%A0%81%E7%94%9F%E6%88%90/index.xml" rel="self" type="application/rss+xml"/><item><title>Refusing the Impossible 精读：代码幻觉不是代码错误——12 个模型在不可解任务上 60% 硬编</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-06-refusing-the-impossible-code-hallucination-paper-reading/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-06-refusing-the-impossible-code-hallucination-paper-reading/</guid><description>PSU×Cisco 提出代码幻觉的三维分类学（groundedness×表现层级×行为），把&amp;rsquo;无根据生成&amp;rsquo;与普通 bug 干净切分；构建 270 个不可解任务（6 语言 24 子类）+91 个可解对照：12 个开源模型在 ~60% 的不可解提示上产出看似合理的无根据代码、仅 27% 正确拒绝、可解对照误拒 0%。模型乐于实现违反已证定理的算法、调用不存在的 crate，甚至&amp;rsquo;明知不可能仍照做&amp;rsquo;。</description></item><item><title>When Models Edit Too Much 精读：编码 Agent 的过度编辑病与保真度评测轴</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-06-over-editing-minimal-code-edits-paper-reading/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-06-over-editing-minimal-code-edits-paper-reading/</guid><description>NUS 团队在 400 个 BigCodeBench 任务上注入受控 AST 损坏、构造已知最小补丁的评测框架，系统刻画 over-editing：GPT-5.5 高 Pass@1 与大改动并存，一行 bug 修出 60 行代码；一条保存指令把超额编辑距离 0.195→0.131、认知复杂度降 26.6%、Pass@1 反升 2.3；SFT 过拟合已见损坏模式，RL 达 0.782 OOD Pass@1 + 0.050 超额距离且不伤通用编码能力。</description></item></channel></rss>