<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>安全 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E5%AE%89%E5%85%A8/</link><description>Recent content in 安全 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E5%AE%89%E5%85%A8/index.xml" rel="self" type="application/rss+xml"/><item><title>CoordPoison × Pretext × TrustProbe × ActionGuard：Skill 生态的信任危机——攻防测四面体 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-02-skill-security-quartet-paper-reading/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-02-skill-security-quartet-paper-reading/</guid><description>本精读合读四篇 Skill 安全新工作，构成「攻×测×防」完整对抗格局：北航+百度 CoordPoison 将恶意执行与情境借口解耦到两个 skill（ASR 76.19%、跨模型迁移 96.88%、跨生命周期 cASR 100%），证伪孤立 skill 审计；华为苏黎世 Pretext 用白盒 LLM 攻击者击穿 NVIDIA SkillSpector（冻结检测器 ASR 至 96.7%），证明「静态规则+LLM 语义 judge」类检测器设计性缺陷；中科院信工所 TrustProbe 以污点分析+定向模糊在 11 个 agent 中挖出 104 个已验证漏洞（成本仅 $2.19），揭示 skill 递送机制使攻击面放大 3 倍；高丽大学 ActionGuard 用上下文分离+fail-closed 授权将 ASR 从 29.05% 压至 8.65%。四篇共同宣判：孤立 skill 审计的防御假设已被系统性证伪，安全边界必须移到运行时执行点。</description></item><item><title>CoDeL × ReproBench：智能体安全的攻防共进化与漏洞复现评估 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-security-training-eval-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-security-training-eval-duet-paper-reading/</guid><description>本精读合读两篇智能体安全新工作：北航 CoDeL 将间接提示注入防御形式化为攻防共进化训练，首创「攻击潜伏期」可度量信号，在 AgentDojo 上将攻击成功率从 0.364 压到 0.042（−88.5%）同时受攻效用提升 38%；中科院软件所 ReproBench 则把 LLM 漏洞复现评估推入 pre-environment 设定，用真目标门控暴露出 45.3% 的「仿真替代」失败模式。一篇教智能体抵御攻击、一篇测智能体发起攻击的能力，恰构成安全能力的攻守双向标尺。</description></item><item><title>Failure-Transparent Agents × FCD × CoSec：智能体安全的三个新失效面 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-security-trio-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-security-trio-paper-reading/</guid><description>本精读覆盖三篇 2026 年 9 月底的 Agent 安全论文：FTA 把「工具失败后模型谎报成功」从端到端评估中剥离出来，发现六模型平均 22.8% 的假成功率，而一个四字段证据契约把它压到 0.8%；FCD 命名并防御「schema 没变但 handler 语义变了」的版本漂移——GitHub MCP v1.4→v1.3 让同一省略参数的建仓调用从私有变公开；CoSec 则把授权边界放进多用户社区，证明同一模型换一个 harness 隐私违规率差 24 个百分点。三者共同把 Agent 安全从「注入攻击」扩展到汇报失真、版本漂移、社区边界三个系统性失效面，与产业界 NVIDIA Open Agent Safety Platform 和白宫超级智能协定的「安全在模型之外的层」思路同频。</description></item><item><title>Scan the Skill, Govern the Action 精读：agent 技能的「许可 ≠ 恶意」，66,192 个技能全语料测量出的运行时治理缺口</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-14-scan-skill-govern-action-oats-paper-reading/</link><pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-14-scan-skill-govern-action-oats-paper-reading/</guid><description>Pheo 团队对 ClawHub 全部 66,192 个 agent 技能版本做了测量：705 个被所有扫描器和 LLM 判官共同判为「清白」的技能，仍在指示 agent 执行 CIS/NIST 明令禁止的操作；活体实验中 agent 对 43.4% 的此类技能真的伸手，运行时门控 23/23 全部拦截。论文提出 OATS——无模型决策路径的确定性解析器 + 按资源×类别键控的信任账本 + 从操作者风险容忍度统计推导的晋升阈值，把 agent 安全从「发布时扫描」的单层世界重构为分层组合的世界。</description></item><item><title>A2ABreak 精读：把 A2A 协议规范编译成状态机之后，11 个新漏洞自己浮出水面</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-a2abreak-protocol-security-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-a2abreak-protocol-security-paper-reading/</guid><description>Purdue+UT Dallas（Elisa Bertino 组）对 Linux 基金会 A2A 协议做首个系统性安全分析：NL 规范→验证 FSM→受限 LLM 推理+对抗验证的三阶段框架。FSM 构建在 TCP ground-truth 上恢复 11/11 状态、19/20 转移（F1 0.84）；在“攻击者完全合规”假设下发现 11 个新漏洞——跨客户端上下文注入、委托链多跳身份丢失凭证收割等，全部无需实现缺陷。</description></item><item><title>EvoSafeHarness 精读：Agent 安全没有万能线束，那就让线束自己进化</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-evosafeharness-agent-security-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-evosafeharness-agent-security-paper-reading/</guid><description>JHU/UC Berkeley/NVIDIA/UIUC/UW-Madison 五机构发布 EvoSafeHarness：为冻结 LLM Agent 自动搜索&amp;rsquo;模型×领域&amp;rsquo;专用安全 harness，DecodingTrust-Agent 上 ASR 45.6%→10.0%（utility 仅损 3.3 分），AgentDojo 82.8% utility @ 0 ASR。核心洞察：模型变体决定 enforcement 强度、领域变体决定谓词与状态——universal 安全 harness 在结构上就不存在。</description></item><item><title>SPDF × Silent Failures 精读：LLM 代码安全评测的双警报日</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-spdf-silent-failures-eval-crisis-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-spdf-silent-failures-eval-crisis-paper-reading/</guid><description>同日两篇论文从两端夹击“静态/测试通过=安全”的假设：Toronto Metropolitan 的 SPDF 度量静态过-动态败缺口（654 个静态干净样本中 14.53% 被运行时利用验证击穿）；Tampere 大学的静默失败实证（1,030 条 Agent 修复轨迹中 170 例确认静默失败，Omission 占 48.2%）建立四维分类学。与 SWE-Gate、PatchBench 共同固化“测试通过≠安全”证据链。</description></item><item><title>HookPry 精读：Agent Harness 的 hook 更新通道是全新的供应链攻击面</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-hookpry-agent-harness-security-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-hookpry-agent-harness-security-paper-reading/</guid><description>HookPry（北邮/网信办数据中心/北航/浙大）首次系统揭示 AI Agent Harness 生命周期 hook 的更新通道攻击面：良性插件上架获取信任后，一次携带 hook 的恶意更新即可在 LLM 完全不可见的路径上以宿主权限执行任意命令。1000 次端到端攻击攻破全部 7 个 harness（最高 92.5%），Microsoft Defender 召回率 0%。本文基于全文阅读拆解其 AMO/TD/LCI 三组件与防御失灵的机制根源。</description></item><item><title>PatchBench 精读：AI 漏洞修复的解决率被高估了 1.83 倍</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-patchbench-vuln-patching-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-patchbench-vuln-patching-paper-reading/</guid><description>马里兰大学的 PatchBench 揭示 AI 漏洞修复评测的两大效度威胁：25% 的 Agent 补丁与历史开发者补丁高度相似（补丁记忆），PoC-only 验证使 11 个 SOTA Agent 的解决率平均虚增 1.83×。其解法是只选 ground-truth 修复在 crash stack 之外的漏洞 + 漏洞移植 + 安全/语义双重验证。本文基于全文阅读拆解 DiffBLEU 记忆检测与验证协议设计。</description></item><item><title>从DeepSeek到Kimi K3，中国开源模型如何逼出黄仁勋的'开源联盟'</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-13-ai-open-source-deepseek-kimi-huangrenxun/</link><pubDate>Thu, 13 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-13-ai-open-source-deepseek-kimi-huangrenxun/</guid><description>DeepSeek V4 Pro登场，中国开源模型连续逼近前沿能力，迫使黄仁勋牵头组建美国&amp;quot;开放安全AI联盟&amp;quot;，Sam Altman、Sundar Pichai等闭源掌门人罕见支持。这期硅谷101系统拆解了AI&amp;quot;开源&amp;quot;到底开的是什么——从七步训练流程到Open Weights与Open Source的本质区别，以及许可证之争、开源公司如何赚钱、闭源阵营的安全担忧与商业焦虑。</description></item><item><title>「模型能力已经够了，要卷就卷 Infra」｜对话戴冠兰：从 Cloudflare 到 Runta，为十亿个 Agent 造执行底座</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-10-runta-agent-infra-daiguanlan/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-10-runta-agent-infra-daiguanlan/</guid><description>Runta 创始人戴冠兰（前 Cloudflare/Kong 核心）在十字路口播客中提出核心判断：模型能力爬坡已放缓，真正制约 Agent 落地的是执行层基础设施。Runta 刚完成 2000 万美元种子轮（a16z 领投，Jeff Dean、李飞飞天使），定位是为 Agent 打造确定性执行底座——在概率性大模型之上加入隔离、权限、审计和热迁移等系统能力，让企业敢于把生产权限交给智能体。文章梳理了 Token Maximizing 到 Minimizing 的反转、Agent 安全必然爆发的逻辑、以及公有云和基模厂商为何难以抢占这一赛道。</description></item><item><title>何谓蒸馏？硅谷如何看中国开放模型逼近前沿</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-01-%E4%BD%95%E8%B0%93%E8%92%B8%E9%A6%8F%E7%A1%85%E8%B0%B7%E5%A6%82%E4%BD%95%E7%9C%8B%E4%B8%AD%E5%9B%BD%E5%BC%80%E6%94%BE%E6%A8%A1%E5%9E%8B%E9%80%BC%E8%BF%91%E5%89%8D%E6%B2%BF/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-01-%E4%BD%95%E8%B0%93%E8%92%B8%E9%A6%8F%E7%A1%85%E8%B0%B7%E5%A6%82%E4%BD%95%E7%9C%8B%E4%B8%AD%E5%9B%BD%E5%BC%80%E6%94%BE%E6%A8%A1%E5%9E%8B%E9%80%BC%E8%BF%91%E5%89%8D%E6%B2%BF/</guid><description>月之暗面K3开源权重发布震动硅谷，开源模型首次在多项能力上追平甚至超越最强闭源前沿模型。两位嘉宾——前Hugging Face开源生态负责人王铁镇和TinyFace联合创始人TJ——深度拆解了&amp;quot;蒸馏&amp;quot;争议的技术真相、中国开源模型为何成本更低、Kimi License商业模式对闭源实验室估值体系的冲击，以及开源模型安全之争的真正焦点。核心判断：没有开源，才是这个时代最不安全的事情。</description></item></channel></rss>