<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>推理效率 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%8E%A8%E7%90%86%E6%95%88%E7%8E%87/</link><description>Recent content in 推理效率 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 05 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%8E%A8%E7%90%86%E6%95%88%E7%8E%87/index.xml" rel="self" type="application/rss+xml"/><item><title>Random Attention 精读：KV 缓存驱逐的选择信号几乎买不到任何东西</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-random-attention-kv-eviction-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-random-attention-kv-eviction-paper-reading/</guid><description>Salesforce AI Research×UIUC 的 Random Attention 证明：长推理场景下 KV 缓存驱逐的&amp;rsquo;重要性打分&amp;rsquo;几乎无用——在每个注意力头内均匀随机驱逐、完全不计算分数，即可在 4 个模型×6 个推理任务上匹配最强选择器，且在 vLLM 分页 serving 下因省去打分 pass 快 32–43%。本文基于全文阅读拆解其三层机制解释与实验设计。</description></item></channel></rss>