<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>LLM安全 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/llm%E5%AE%89%E5%85%A8/</link><description>Recent content in LLM安全 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 09 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/llm%E5%AE%89%E5%85%A8/index.xml" rel="self" type="application/rss+xml"/><item><title>Safety for Whom：把安全边界从话题级细化到边界级的数据配方 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-09-boundary-aware-safety-refusal-paper-reading/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-09-boundary-aware-safety-refusal-paper-reading/</guid><description>Multiverse Computing 揭示话题级安全对齐的粒度错配：部署需要窄边界（拒操纵、答选举事实）而非话题级一刀切。提出窄边界形式化+离线自生成数据框架（受控话题生成/覆盖修复/补偿数据/边界对），Qwen3-8B 目标域拒绝 9.47%→84.75%、不安全率 26.26%→0.14%，并量化数据成分对安全-可用权衡的决定作用（过度拒绝 74%→5.2%）。</description></item><item><title>Refusing the Impossible 精读：代码幻觉不是代码错误——12 个模型在不可解任务上 60% 硬编</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-06-refusing-the-impossible-code-hallucination-paper-reading/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-06-refusing-the-impossible-code-hallucination-paper-reading/</guid><description>PSU×Cisco 提出代码幻觉的三维分类学（groundedness×表现层级×行为），把&amp;rsquo;无根据生成&amp;rsquo;与普通 bug 干净切分；构建 270 个不可解任务（6 语言 24 子类）+91 个可解对照：12 个开源模型在 ~60% 的不可解提示上产出看似合理的无根据代码、仅 27% 正确拒绝、可解对照误拒 0%。模型乐于实现违反已证定理的算法、调用不存在的 crate，甚至&amp;rsquo;明知不可能仍照做&amp;rsquo;。</description></item></channel></rss>