<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>推理加速 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%8E%A8%E7%90%86%E5%8A%A0%E9%80%9F/</link><description>Recent content in 推理加速 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 09 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%8E%A8%E7%90%86%E5%8A%A0%E9%80%9F/index.xml" rel="self" type="application/rss+xml"/><item><title>Uno：扩散增强 LLM 的无损加速范式——AR 与扩散在同一架构内的参数解耦 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-09-uno-lossless-diffusion-speedup-paper-reading/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-09-uno-lossless-diffusion-speedup-paper-reading/</guid><description>MBZUAI×Cerebras 等推出 Uno：在同一 Transformer 内解耦 AR 权重（质量）与 rank-128 LoRA 扩散适配器（速度），冻结 AR 后仅用 7B token 做块级单步扩散蒸馏，配合 Ψ-Spec 采样器做 AR 验证的拒绝采样，实现严格无损、全 batch 区间保持的至高 3× 加速。8B 模型 SWE-bench Verified 68.4%，系统吞吐 5733 toks/s 全面超越 EAGLE-3/DFlash 与闭源 Mercury 2。</description></item><item><title>ReCache: 工具增强Agent的组合不变KV缓存复用 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-recache-kv-reuse-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-recache-kv-reuse-paper-reading/</guid><description>工具增强Agent每个请求都要重新编码一遍以不同组合、不同顺序出现的工具与技能schema，标准前缀缓存对此无能为力。ReCache提出resource-wise attention，切断资源间注意力并重置资源内位置索引，使每个资源的KV块具有组合不变性、可独立缓存复用；再叠加贡献选择的层-KV头组路由与字段感知的语义剪枝，把KV张量内存降低92.43%、注意力加速1.423倍，同时Inv-F1基本不降。本精读覆盖其动机、机制、七数据集基准与效果根源分析。</description></item></channel></rss>