<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI Infra on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/ai-infra/</link><description>Recent content in AI Infra on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/ai-infra/index.xml" rel="self" type="application/rss+xml"/><item><title>Diffusion Reward Models × SOLO：成熟技术的首次规模化双案例 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-paradigm-first-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-paradigm-first-duet-paper-reading/</guid><description>同日两篇论文在各自领域做了同一件事：把一项被搁置多年的成熟技术，用一个机制创新首次推到现代规模。清华 thunlp+港中文+UIUC 的 Diffusion Reward Models（DRM，当日 Hugging Face 日榜 up=22）把扩散模型首次引入奖励建模——奖励从点估计变为条件密度估计 p(r|x,y)，轻量 DiT 头无参数族假设地表示人类偏好的多峰结构（多峰率随标注分歧 37.6%→63.2%，Wasserstein 距离全表最低），同骨干同数据受控对比平均 66.2 vs ArmoRM 62.3（+3.9）；中科院自动化所的 SOLO 把局部学习（local learning）首次推到十亿参数 LLM 预训练——用一个所有模块共享的终端 readout 只读副本打破 update locking，梯度对齐 cos 从 0.52 升至 0.70，流水线激活内存降近 p 倍、吞吐达 1F1B 的 1.44×。二者方法论同构：识别被搁置的旧技术、诊断其规模化障碍、用一个机制创新解锁，为「旧思想在新规模下复活」提供了可复用的模板。</description></item><item><title>MassAlloc Attention × DaRoPE：注意力计算分配与位置编码的双子星重构 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-architecture-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-architecture-duet-paper-reading/</guid><description>本精读合并解读两篇重构 Transformer 核心算子的论文：HKUST(GZ)×BAAI×巴黎西岱的 MassAlloc Attention（MALA）把稀疏注意力改写为「分布条件化的运行时计算分配」——保留全部 QK 打分，用 online-softmax 演化归一化因子做逐 tile 贡献比测试，跳过低贡献 tile 的 post-score 计算，同一容差 τ=1 统一前向/反向/prefill/解码，训练 FLOPs -23.1% 而 PPL 无损、关联回忆 89.67% 贴近 FullAttn（NSA 仅 22.61%）；Meta AI 巴黎×Inria 的 DaRoPE（ICLR 2027）诊断出 RoPE 慢频带这一具体弱点，快带保序、慢带换成 sigmoid 有界的逐头内容坐标，外推免配置，K=256 键值回忆 30.5% vs 其他方法 ≤7.25%。二者共享同一哲学：不做一刀切裁剪，让算子自己知道「哪里重要」。</description></item><item><title>Skill2Env × QwenGyre × AgentPerfBench：智能体强化学习的数据、系统与推理三层基建 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-rl-infra-trio-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-rl-infra-trio-paper-reading/</guid><description>同日挂出的三篇 arXiv 论文恰好拼出 Agent 强化学习的三层基础设施：数据层（AllSpark 的 Skill2Env 把「环境合成」从扩展覆盖升级为能力参数化——100 个带控制旋钮的难度模式+基于 solver 执行证据的迭代加难，1.5K 轨迹 SFT 换 7 基准 +8.4pt）；系统层（阿里 Token Hub 牵头的 QwenGyre 用 cell 级弹性调度+轨迹树处理让 2.4T 旗舰模型的百万 token rollout 在线 RL 提速 1.78×，pass rate 52.48%→58.54%）；推理层（帝国理工+剑桥+牛津的 AgentPerfBench 用 22 个负载画像+饱和扫描+NCU roofline 证明 chat→coding 的 TTFT 差 4.8×、操作强度差 46×，agentic 负载在并发爬升时吞吐崩塌 43–76% 而 chat 仍在扩张）。本文按九部分结构合并精读，并给出「Agent RL 全栈工程」的公共图景。</description></item><item><title>WeEnv: The Environment for Agentic Reinforcement Learning at WeChat 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-weenv-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-weenv-paper-reading/</guid><description>腾讯微信 AI 提出 agentic RL 的「环境税」：环境初始化独占迭代时间 53.4%，比训练本身还慢。WeEnv 对环境做打包-初始化-供给的全生命周期管理——layer group 活页夹式组合把 39,471 个镜像缩到 133 个、按需拉取让环境 10.6 秒启动（E2B 要 150.6 秒）、弹性配额化解秒级 47 倍的资源突发。本精读覆盖动机量化、三大设计的机制因果链、与 E2B/AgentENV/SkyRL 的外部交叉验证，以及企业生产部署视角的通用启发。</description></item><item><title>Agent 越能干，越等不起：云栖2026「AI实时数据智能」论坛复盘——卡住生产级 Agent 的不是模型，是数据的实时性、语义与断点</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-realtime-data-intelligence/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-realtime-data-intelligence/</guid><description>2026 云栖「AI 实时数据智能」论坛复盘：六场演讲共答一个矛盾——Agent 任务从秒级问答变成长程执行后，卡住生产的不是模型，而是数据新鲜度、业务语义与任务断点。Kafka 流算湖一体收敛实时链路，SLS 沉淀轨迹资产，RocketMQ LiteTopic 支撑手脑分离，AgentBridge 以逻辑统一取代数据集中，震坤行与 Qoder 给出客户证词。</description></item><item><title>当互联网的主体不再是人：云栖2026云网络专场，网络为什么重新变成稀缺品</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-cloud-network-ai-innovation/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-cloud-network-ai-innovation/</guid><description>2026云栖大会云网络专场（9月24日）年度发布梳理：当AI相关流量超过公网流量一半，网络从&amp;quot;尽力而为的管道&amp;quot;重新变成确定性交付的稀缺品。本文从尾延时与推理性价比、用通信换计算、推理下沉边缘、Agent成为网络新主体、数据语言统一五条机制链，拆解英特尔、阿里云、网易互娱、小鹏汽车、Maxinsights、Megaport六方的一线说法，并给出可迁移的判断框架。</description></item><item><title>模型每2.8天更新一次、企业采购却要等半年：云栖2026百炼专场复盘——从 Model 到 Token，卡住企业的不再是模型本身</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-bailian-model-to-token/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-bailian-model-to-token/</guid><description>2026 云栖百炼专场复盘：信通院姜春宇给出「模型 2.8 天一更、任务时长每 7 个月翻倍」对撞「企业一采购就落后」，提出 AI 原生方法论；安克商渭清展示日均四千亿 token 的 DOM×Launch 实践；百炼于文渊拆解「今年 90% token 来自 Agent」；圆桌把价格 K 型分化、智能路由与算力卡点收敛为「从有模型到用好模型」。</description></item><item><title>真武是系统，不只是芯片：平头哥算力峰会上的超节点方法论</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-yunmo-zhenwu-ai-chip/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-yunmo-zhenwu-ai-chip/</guid><description>整理2026云栖大会平头哥算力峰会「云模之芯，共筑非凡」全场内容。平头哥副总裁李伟良梳理真武810E→M890→V900→J900的年度迭代节奏；阿里云王超给出「系统语义不在物理边界断裂」的真超节点判据；Kimi许欣然拆解100毫秒decode里的访存经济学；元戎启行曹通易与无界动力夏中谱讲物理AI的数据量级与快慢脑；圆桌与倚天CPU、磐脉920网卡两场把Agent时代的关键路径补完整。芯片竞争已从单卡指标转向系统协同。</description></item><item><title>Memory Compression for High-Fanout Agent Sandboxes (AgentZip) 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-20-agentzip-paper-reading/</link><pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-20-agentzip-paper-reading/</guid><description>Agent 平台把每个动作都关进沙箱，而 RL 训练和并行推理让一个任务扇出几十个沙箱——内存（而非算力）成为并发上限。HKUST 的 AgentZip 是首个专为 Agent 沙箱设计的内存压缩系统：用模板增量、同类群字典、页内 RLE 三种编解码器榨取&amp;rsquo;近似相同&amp;rsquo;页面的冗余，用恢复期预取取代保守选页，把昂贵压缩搬进 LLM 思考的空闲窗口。实测沙箱内存最高降 8.7 倍（Linux 配置仅 2.1 倍），激进压缩的减速从 3.1 倍压到 1.40 倍。本精读逐页拆解其 How/What/When 三问重构与全部消融，并对照 DeltaBox、DroidSpeak 等同期系统工作交叉验证。</description></item><item><title>DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-19-deepseek-v4-1-flash-paper-reading/</link><pubDate>Sat, 19 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-19-deepseek-v4-1-flash-paper-reading/</guid><description>DeepSeek-V4.1-Flash 用三件武器把长上下文智能体的部署成本打下来：CED 非对称架构让 prefill 只激活 8B 参数（decode 16B），CSA2 跨层 KV 复用 + FP4 量化把全局 KV 缓存压到 890 字节/token（较 V1 降 437 倍），SWA Bounded Replay 把持久化缓存再压到 1/8。在 Codeforces 3348→3471、DeepSWE v1.1 达 74.2% 的同时，45T token 多模态预训练完全开源。本文拆解其架构因果链与&amp;rsquo;智能体负载第一性&amp;rsquo;的设计哲学。</description></item><item><title>ComPO 零阶偏好对齐 与 SpectralShift 线性注意力长上下文扩展 精读（二重奏）</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-18-compo-spectralshift-paper-reading/</link><pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-18-compo-spectralshift-paper-reading/</guid><description>两篇训练方法学论文合并精读。UC Berkeley×NYU×阿里达摩院的 ComPO 提出 LLM 偏好对齐的零阶范式：不在偏好对上直接优化可微损失，而是用 comparison oracle 提取方向信息——规避 DPO 类方法在低似然边际对上的 likelihood displacement 失效，五个模型家族上改进含长度控制胜率，并给出收敛与性能保证。人大高瓴×MSRA 的 SpectralShift 从转移矩阵谱视角重新审视 Gated DeltaNet 的长上下文扩展：慢谱带宽度决定长程检索、快衰减模式负责状态清理，重参数化 alpha 投影初始化+学习率缩放即可让 10B 模型 8K→128K 课程扩展持续增益（RULER 64K +4.2）。</description></item><item><title>SSD-LLaMA 精读：单张 RTX 5090 跑万亿参数 MoE 的 SSD 原生推理系统</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-18-ssd-llama-paper-reading/</link><pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-18-ssd-llama-paper-reading/</guid><description>港科大联合中科院深圳先进院、南科大发布 SSD-LLaMA：面向万亿参数 MoE 的 SSD 原生本地推理系统——SSD I/O 流水线优化专家投递、SSD-RAM-VRAM 三层存储动态驻留、CPU-GPU 均衡混合执行，保证每个选中专家无剪枝无替换。三大前沿 MoE 家族上 prefill 提速 1.52–4.19×、decode 提速 2.10–15.58×，单张 RTX 5090+32GB RAM 实现万亿模型 &amp;gt;1 token/s。本精读拆解『把带宽受限问题转化为层次调度问题』的系统设计。</description></item><item><title>Mo' Models, Mo' Problems × Co-Skill 精读：多智能体选型与边云技能演化双视角</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-mo-models-coskill-mas-pool-evolution-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-mo-models-coskill-mas-pool-evolution-paper-reading/</guid><description>两篇互补的 multi-agent 工程研究：NVIDIA×哥本哈根的 Mo&amp;rsquo; Models 用 23 模型×3 科学基准证明 MAS 模型池&amp;rsquo;加模型常降性能&amp;rsquo;——oracle 潜力与实际达成存在鸿沟、同族池是唯一稳定正收益、准确模型解集高度嵌套（rM=0.931）；哈工大的 Co-Skill 诊断边云 skill 演化的&amp;rsquo;盲通信&amp;rsquo;根因（上传 token 25-42% 是重复前缀），用前缀合并轨迹 trie+渐进 skill 树双向解盲，token 省 15.6-41.9%、成功率提升 25.8-76.4%。合读视角：多 agent 系统的两个新瓶颈——选谁进队（异构组合的聚合噪声）与怎么通信（协作双方的信息结构）。</description></item><item><title>OPEN-1B: A Fully Auditable Training Run 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-open1b-auditable-training-run-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-open1b-auditable-training-run-paper-reading/</guid><description>开源 LLM 即使放出全部数据与配方，也因浮点非结合性无法逐位复现——你无法验证发布的 checkpoint 真是声明的配方训出来的。Gensyn 的 OPEN-1B 定义第四层透明度&amp;rsquo;完全可审计&amp;rsquo;：RepOps 跨硬件逐位复现算子（固定规约序、统一 FMA/次正规数约定、计数器式 RNG）、拓扑不变数据流（token 流=种子的纯函数）、确定性 butterfly all-reduce，多审计者各验若干步拼出全程、整跑收敛为单一哈希。代价是 MFU 从一个数量级掉到 5%——可验证性与速度的明码标价，以及首个该层级的开源 LLM 全套产物。</description></item><item><title>推理芯片之战：带宽成为新算力，Groq、Cerebras 与 OpenAI 的三条路线与 Bill Dally 的 location 哲学</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-inference-chip-wars-groq-cerebras-openai/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-inference-chip-wars-groq-cerebras-openai/</guid><description>推理的瓶颈正在从算力转向带宽：每生成一个 token，都要把整个模型从存储介质读一遍。硅谷101本期请到两位正在做推理芯片创业的嘉宾——Bill Dally 的博士生 Mark 与前亚马逊 Annapurna 芯片软件栈科学家子阳，拆解 Groq 的静态调度、Cerebras 的晶圆级集成与 OpenAI Jalapeño 的 HBM4 通用路线为何是不同约束条件下的各自最优解，为什么中国关心 tokens per dollar 而美国关心 tokens per watt，以及推理速度如何决定模型智能的上限。</description></item><item><title>AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-amdkernelvault-amd-gpu-kernel-corpus-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-amdkernelvault-amd-gpu-kernel-corpus-paper-reading/</guid><description>AMD 开源 HIP/Triton 内核优化语料与 agentic 训练框架：HIPKernelGen/TritonKernelGen 管线把 PyTorch 参考实现转为 HIP/Triton 内核、在 ROCm 下编译验证、上硬件延迟剖析——产出 62,153 个执行验证 HIP 内核 + 2,377 条 ROCm 库 QA + 39,893 个 Triton 内核。演示价值：Qwen3-8B 经 SFT+执行感知 RL 后在 PyTorch→HIP 达 34.0% Pass@1、TritonBench-G 33.2% Corr@3、ROCmBench 41.94% Corr@3——打破 CUDA/NVIDIA 中心主义的开放生态基建。</description></item><item><title>GLIE 精读：几何先验驱动的检索压缩——100 万页 258GB 到 1GB 的流形参数化</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-13-glie-generative-late-interaction-paper-reading/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-13-glie-generative-late-interaction-paper-reading/</guid><description>KAUST×Edge Hill 发现视觉文档检索的页面向量恰在单位球面上且集中在本征维度 5-6 的低维流形附近（三个编码器一致验证），据此提出 GLIE：k≪N 个向量既作轻量索引又作全页嵌入的再生基底——归一化质心免费 +0.093 nDCG@5，k=4 时 1040 字节/页 vs 未压缩 257.8KB，保留未压缩系统近 80% 性能（先前最佳 70%）；415K 参数网络 3 GPU 分钟千页训练，骨干全程冻结。</description></item><item><title>X-AuT 精读：渐进剪枝+跨尺度蒸馏的语音编码器压缩，18→16 层错误率不降反升</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-13-x-aut-audio-encoder-compression-paper-reading/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-13-x-aut-audio-encoder-compression-paper-reading/</guid><description>小鹏汽车提出 X-AuT：短行为探针选择可恢复的层组合，渐进式剪枝 Qwen3-ASR-0.6B 音频塔 18→16→14 层，表征对齐+跨尺度蒸馏（教师强制+计划学生策略）+LoRA 恢复，解码器骨干全程冻结。18→16 层宏平均错误率 5.61%→5.27%（不降反升）；14 层 5.75%、参数 -20.7%、车载加速器延迟 -21.4%；渐进剪枝 5.75% vs 直接剪枝 6.73%。</description></item><item><title>BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-beaconkv-thought-revisiting-kv-compression-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-beaconkv-thought-revisiting-kv-compression-paper-reading/</guid><description>KV cache 压缩方法都用&amp;rsquo;最近的查询&amp;rsquo;预测未来注意力——本文发现长程推理中这个假设根本不成立：解码会不定期产生&amp;rsquo;思维重访令牌&amp;rsquo;（TRT），重新关注数千 token 之前的推理计划，而近期查询无法预知这次重访。关键观察是 TRT 对应的全局查询在嵌入空间聚成少数簇——只需为每簇维护一个&amp;rsquo;信标查询&amp;rsquo;（Continual FPS 在线选取），就能预判哪些 KV 将被重访。训练自由、无需改动架构：四个开源推理模型上内存最高压缩 5.8×、精度近全量、吞吐 +4.3×，对 RPC/R-KV 最高领先 31.7 个百分点。</description></item><item><title>Iris: Climbing to the Search Frontier — 开源搜索智能体的数据反构造与 SFT-RL 攀爬配方 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-iris-search-agent-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-iris-search-agent-paper-reading/</guid><description>AllSpark 团队发布 Iris-mini/pro 两个开源搜索智能体（35B-A3B 与 397B-A17B）：从网页超链接实体图反向构造多跳任务，把非答案实体改写为描述性引用以杜绝字符串匹配作弊；提出 SFT-RL climbing 交替训练——每轮 RL 把最难与最高效轨迹回流进下一轮 SFT。BrowseComp 88.6 / HLE 56.4，同参数段开源最强。本文精读其&amp;rsquo;不可作弊任务合成&amp;rsquo;与&amp;rsquo;爬坡式两阶段循环&amp;rsquo;的完整配方。</description></item><item><title>Compile by Training: Turning Natural-Language Specifications into Local Neural Functions 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-compile-by-training-neural-functions-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-compile-by-training-neural-functions-paper-reading/</guid><description>滑铁卢大学×哈佛的 Compile by Training 把&amp;rsquo;编译&amp;rsquo;概念引入神经函数：教师模型从自然语言规格合成监督数据，训练 LoRA 适配器特化冻结的 Qwen3-0.6B 解释器，产出可存储、可版本化、可组合的 .paw 程序。在 PAW 快速编译器零精确匹配的 FuzzyBench-Hard 上语义准确率从 0.224 提升至 0.836，编译仅需约 50 秒。本文精读其&amp;rsquo;训练即编译&amp;rsquo;范式、分钟级编译服务工程与速度-精度新权衡点。</description></item><item><title>LatentPress: Context Compression Beyond Text and Vision 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentpress-soft-token-context-compression-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentpress-soft-token-context-compression-paper-reading/</guid><description>压缩后的上下文通常仍以文本或图像这两种&amp;rsquo;给人看&amp;rsquo;的形式存在。LatentPress 提出第三种表示：小型 writer 把对话/文档直接写成连续记忆 token，冻结解码器经输入嵌入接口读取，推理时零文本重建。LongMemEval 上 7.7× 压缩反而比未压缩证据更准（0.504 vs 0.490），写入 43ms 比摘要快一个量级——机器原生的记忆表示从此有了实证立足点。</description></item><item><title>Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-minima-gdn-4bit-quantization-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-minima-gdn-4bit-quantization-paper-reading/</guid><description>社区量化混合架构 LLM 时一致保留循环半区（Gated DeltaNet）的高精度，理由是&amp;rsquo;循环误差会累积&amp;rsquo;。Minima 直接把 NVFP4 W4A4 打满全部 496 个线性层：五任务平均仅 -0.52（种子噪声内），显存 17.5 GiB 最小、prefill 提速 14-19%。四重机制研究（块缩放局域化离群值→门控非线性压缩误差→delta-rule 主动遗忘→逐 token 代价被冲刷）解释了为什么直觉是错的。</description></item><item><title>EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-earlyeval-agent-eval-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-earlyeval-agent-eval-paper-reading/</guid><description>跑一遍前沿模型在 SWE-bench Verified 上要花数百到数千美元，而 agent 开发需要反复评测。上海交大联合新加坡管理大学等提出 EarlyEval：agent 的最终成败往往在轨迹中段就已注定——训练一对 LightGBM 成功/失败分类器，一旦置信度过阈值就提前终止运行。三个基准上砍掉 13%–26% 步数、最高省 44.1% 输入 token，预测精度 89%–97%，排行榜排序保真度 Spearman ρ 高达 0.99。本精读拆解&amp;rsquo;轨迹内降本&amp;rsquo;与&amp;rsquo;基准蒸馏降任务数&amp;rsquo;的正交关系、行为特征为何比参考解更有用，以及阈值-保真度的可调权衡。</description></item><item><title>Language Models Can Control Their Own Attention 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-declarative-attention-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-declarative-attention-paper-reading/</guid><description>长上下文解码时，模型每生成一个 token 都要把整个 KV cache 读一遍——1M token 上下文意味着每步约 15GB 的内存搬运，而注意力其实高度集中。KAIST AI 联合 Google DeepMind 提出 Declarative Attention：让模型在思维链里用 &lt;global&gt;/&lt;focus&gt;/&lt;local&gt; 三种标签自己声明&amp;rsquo;现在需要看哪里&amp;rsquo;，推理引擎像解析工具调用一样解析声明并跳过绝大部分 KV 读取。零训练、零外部打分器，15 个长上下文任务上 Gemma-4-31B 注意 token 降 52.0%、精度仅降 1.27pp。本精读拆解三模式协议、与代理打分式稀疏注意力的机制差异，以及&amp;rsquo;模型自己最知道该看哪里&amp;rsquo;的第一性原理。</description></item><item><title>Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-repo-to-skill-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-repo-to-skill-paper-reading/</guid><description>自主 ML 研究 agent 缺的不是更强的模型或更聪明的流程，而是&amp;rsquo;怎么把方法跑通&amp;rsquo;的操作知识层。BAAI 联合中科大、人大、港理工提出 DisCo 蒸馏框架，把 1000 个 GitHub 仓库蒸馏成 5353 个经过验证的技能，构建 AREX-Skill Library。在固定 GPT-5.5+Codex 的对照实验下，技能让 MLE-bench 相对提升 134.3%、PaperBench 提升 34.4%、FrontierCS 提升 9.2%、PassNet 提升 14.0%，并以更低 token 消耗帕累托支配 Claude Code。本精读拆解技能图三层结构、四阶段蒸馏流水线、对照实验设计，以及&amp;rsquo;试错成本越高、操作知识价值越大&amp;rsquo;的机制根源。</description></item><item><title>On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-02-qwen38-next-architecture-paper-reading/</link><pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-02-qwen38-next-architecture-paper-reading/</guid><description>Qwen Team 的 Qwen3.8-Flash-Next 架构报告展示了一次教科书级的&amp;rsquo;联合设计&amp;rsquo;实践：GDN 线性注意力混合 + 压缩索引稀疏注意力（QSA）+ 四分支门控残差 + 主机内存 n-gram 嵌入，125B 总参 6B 激活的模型以约 1/9 训练 FLOPs 在 14 个基准上 8 个超越 397B 旗舰，1M 上下文预填充加速 7.6 倍，且 4 倍学习率压力测试下全程零 loss spike。报告最宝贵的不是单个组件，而是&amp;rsquo;loss、基准、效率、稳定性必须当一个问题解&amp;rsquo;的方法论，以及大量&amp;rsquo;loss 与下游精度背离&amp;rsquo;的诚实披露。</description></item><item><title>Affix Cache for Diffusion Large Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-acache-dllm-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-acache-dllm-paper-reading/</guid><description>扩散语言模型（DLLM）以双向注意力实现并行解码，但也因此丧失了自回归模型前缀缓存的红利——共享文本段的 KV 状态与不断演化的生成 token 相互耦合，直接跨请求复用会立刻过期。本文提出 ACache：在请求开始时用一次跨注意力探针找出少数对生成 token 影响最大的『锚点令牌』，之后只重算锚点与请求特有位置的 KV，其余共享段缓存跨请求复用。在 LLaDA-8B 与 Dream-7B 上，约 20% 锚点率即可恢复大部分精度损失；基于 Nano-vLLM 的原型将重算延迟最高降低 55.7%，端到端吞吐最高提升 1.68 倍，峰值 KV 显存最高下降 43.3%。这项工作首次把『跨请求缓存复用』引入 DLLM 推理系统，揭示了双向注意力下共享上下文管理的新范式。</description></item><item><title>Fast Weight Attention for Continual Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-falcon-fast-weight-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-falcon-fast-weight-paper-reading/</guid><description>深度精读 ByteDance Seed、Princeton、清华、UCLA 与 Hyperbolic Labs 合作的 Falcon 论文：把线性注意力与状态空间模型的状态转移显式化为在线学习规则，发现读后写语义下快记忆的正确训练对应是前缀配对 φ(k(t-1))→v(t) 而非常见的同对配对，并从平方误差回归与内积两个局部目标统一推导出 NLMS 归一化的六个变体族（Falcon-1/2/3 与 Falcon-1A/2A/3A），全部兼容 SSD 式 chunk 并行训练。语言建模上与最强递归基线互有胜负，变长加法外推上 Falcon-3A.3 以 87.2 平均精度大幅领先 Transformer 的 65.8。</description></item><item><title>LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-lowrankarena-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-lowrankarena-paper-reading/</guid><description>SVD 低秩压缩号称能省显存又保精度，但各家论文的评测基准、压缩率定义、推理后端各不相同，「进步」到底是方法变强还是协议不同造成的？杜克大学与维克森林大学团队发布 LowRankArena：统一任务版本（LM-Eval-Harness v0.4.11）、统一精度下的 keep ratio 定义、统一 vLLM 0.18.1 端到端推理测量，并开放 3TiB 以上压缩 checkpoint 动物园。对 5 种代表方法（ASVD、SVD-LLM、DoBi-SVD、Basis Sharing、MoDeGPT）×3 个骨干模型的对齐审计发现：排名随骨干剧烈变动、MCQ 准确率可掩盖困惑度坍塌（ASVD 0.353 恰好贴着 0.357 随机地板）、低秩省的是 FLOPs 而非时间——prefill 提速最高 4.20 倍，decode 场景 SVD-LLM 吞吐反而跌到 0.80 倍，且 5 个方法在单张 H200 上压缩 70B 全部因工程问题失败。</description></item><item><title>Muon with Finite Newton-Schulz 精读：有限迭代不是误差，而是收敛的功臣</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-muon-finite-ns-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-muon-finite-ns-paper-reading/</guid><description>Muon 已被用于 Kimi K2、GLM-4.5 等前沿模型训练，其核心是用几步 Newton-Schulz 迭代近似动量的正交化。此前的理论要么把这一迭代换成精确极分解，要么把有限深度当作需要控制的逼近误差。本文反转视角：通过折扣在线到非凸转换框架证明，有限 Newton-Schulz 恰恰是让 Muon 在非光滑非凸目标上收敛的平滑机制。深度 q 调节惩罚项与稳定性项的权衡，取 q=O(log(1/ε)) 即可得到匹配最优界的平稳点复杂度，而精确极分解版 Muon 反而可能不收敛。</description></item><item><title>Sliding-window beats linear attention 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-swa-beats-linear-attention-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-swa-beats-linear-attention-paper-reading/</guid><description>深度精读微软 Applied Sciences Group 的评测批判型论文：后训练线性注意力（LoLCATs、MOHAWK、QRWKV 等）一直以『低成本解决 KV 缓存膨胀』为卖点，但从未与真正的强基线——免训练的带注意力 sink 滑窗注意力 SWA(64,4) 公平比较。本文在 11 个基础模型（1.3B 到 70B）上补上这场缺失的对比：短上下文 SWA 恢复基线平均性能 99.0%，长上下文 S-NIAH 与 BABILong 上领先 2 到 10 倍，且零后训练 token、解码最快、内存最低。一篇动摇整个『线性化改造』研究方向价值主张的论文。</description></item><item><title>Survival-Guided Length Control for Efficient Diffusion Language Models 精读：用生存分析一次前向测出生成长度</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-survival-dllm-length-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-survival-dllm-length-paper-reading/</guid><description>扩散语言模型（DLM）迭代去噪生成文本，但标准 AOAR 解码对全体样本使用统一的保守长度预算 Lmax=1024，大量去噪步浪费在精修真实结束位置之后的冗余画布上。本文把生成长度选择重新表述为 [EOS] token 上的离散时间生存问题：对长 mask 画布做一次前向传播，把各位置的 [EOS] 概率解释为危险率，经均场近似连乘得生存曲线，再用期望等于生存概率求和的标准恒等式闭式算出期望长度，作为该样本的解码预算。方法免训练、即插即用，在 LLaDA-8B-Base 与 Dream-7B-Base 上取得 3.1-6.6 倍解码加速（instruct 模型最高 7.3 倍），任务精度变化全部落在标准差之内；固定均值长度消融与逐步漂移分析进一步证明逐样本长度预测的必要性。</description></item><item><title>TreeGraft 精读：多起草器共嫁接一棵草稿树，让投机解码跨越快与好的单选题</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-treegraft-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-treegraft-paper-reading/</guid><description>树形投机解码用一个起草器建草稿树，陷入『小模型快但树差、大模型树好但慢』的两难。TreeGraft 让免训练 n-gram 小起草器与预训练中间起草器共同构建一棵共享草稿树：中间起草器可回看全树历史节点重新打分并复活被低估的路径（Where），嫁接时不覆盖既有子树（How），再由离线价值系统蒸馏出的 3265 参数轻量调度器逐步决定是否值得调用中间起草器（When）。在 10 个模型对 × 6 个基准上平均加速 1.60 倍，超两个单起草器端点中较优者 15.1%，且在留出模型对与完全留出的 MT-Bench 任务上仍保持 1.48 倍与 1.60 倍，证明调度器学到的是可迁移的成本-质量权衡。</description></item><item><title>Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-redwood-ai-accelerator-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-redwood-ai-accelerator-paper-reading/</guid><description>Architect Labs 的 AI 系统在两名人类架构师只写高层规范的前提下，用不到两周自主生成性能模型、RTL、UVM 验证环境、形式证明、固件与计算内核，所有模块达成 95% 代码与功能覆盖率，首次上 FPGA 零 bug，架构变更 48 小时内完成再验证再部署；投影至 Samsung 8nm 工艺后，Redwood 以 49 &lt;a href="mailto:tokens/s@1.335W"&gt;tokens/s@1.335W&lt;/a&gt; 对比 Jetson Orin Nano 同模型实测 28 &lt;a href="mailto:tokens/s@2.59W"&gt;tokens/s@2.59W&lt;/a&gt;，能效提升 3.4 倍。本精读面向软件背景读者，补全 RTL、UVM、形式验证、FPGA、roofline 等概念，剖析单一规范源全栈协同生成的因果链，并提炼七条可迁移的系统设计灵感。</description></item><item><title>SKILL.state: Scalable Long-Horizon Agent Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-skill-state-execution-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-skill-state-execution-paper-reading/</guid><description>让 Agent 执行长任务时，主流运行时把所有推理、动作、观察不断追加进对话历史，prompt 随步数二次膨胀，token 烧钱、噪声污染、过期事实还会诱发幻觉。Google 与 Purdue 合作的 SKILL.state 干脆废除这个 append-only 历史：每一步模型只看到技能规范、结构化执行状态和最新观察，推理轨迹用完即弃，状态以 JSON 补丁形式确定性合并。prompt 尺寸从 O(T) 降为 O(1)，百步任务 token 缩减 16.2 倍，CTF pass@1 提升 7.8 个点，外部篡改状态后零步恢复而基线幻觉 5 至 8 步。本文精读其运行时设计、预算匹配对照实验与效果根源。</description></item><item><title>The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-reasoning-tax-token-economics-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-reasoning-tax-token-economics-paper-reading/</guid><description>推理模型在回答前会生成大量思考 token，但这些思考是否值得其成本？联想基础设施方案集团的这篇论文提出边际效率指标 Token Economy Score（TES），用准确率增益除以生成 token 倍数来度量推理模式的性价比，并在 151 个模型-基准组合上给出三个部署结论：任务结构比名义难度更能预测推理收益，序列推理任务（数学竞赛、代码）高 TES 而知识回忆型任务（MMLU-Pro、GPQA）低 TES；提高推理努力档位呈尖锐边际递减甚至负增益；思考链占总推理成本中位数高达 94.7%，而自建 MoE 集群可把云端成本再降 2.0 至 25.6 倍。本精读覆盖指标设计、实验证据、因果解释与可迁移的通用灵感。</description></item><item><title>Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-agent-mesh-reliability-primitives-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-agent-mesh-reliability-primitives-paper-reading/</guid><description>当Agent编排系统直接搬用服务网格的重试、超时、错误率熔断这三件套时会发生什么？这篇论文对一台生产级Agent交付平台（66,185行代码、59个模块）的147起编号事故做了回顾性失效研究，量化展示了三大可靠性假设在Agent场景全部失效：54次连续成功调用让错误率熔断全程失明、21个事件跨6次调用累积让完全正确的幂等组件永远无法通过测试、12起执法层阻断正确工作的事故。论文提炼出贯穿5个子系统的横切根因——身份充分性，并推导出以delegation为执法单元的7个新可靠性原语。对构建Agent基础设施的工程师来说，这是一份罕见的、带测量成本的真实现场失效档案。</description></item><item><title>AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-asymspec-decode-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-asymspec-decode-paper-reading/</guid><description>Agent 流水线的推理成本随上下文累积而飙升，压缩输入省钱却伤精度，而投机解码（SD）受制于「drafter 与 verifier 必须读同一份上下文」的对称性约束，无法破解这个两难。华为与中国科大的 AsymSpec 打破对称：小 drafter 读全文、大 verifier 只读压缩视图，通过同模型跨上下文的对比 δ-fusion 把被压缩丢弃的信息在 logit 空间回注，再用免调参的 JSD 散度门保持接受率稳定。结果：以 0.23 倍计算拿到 LongBench 59.7 F1（恢复压缩损失差距的 72%）、1.45 倍吞吐；恢复量随压缩严重度单调增长，近无损压缩处自动失活；还能让读原图的视觉 drafter 操纵只读文字的 verifier。</description></item><item><title>Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-kope-npu-kernel-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-kope-npu-kernel-paper-reading/</guid><description>深度精读 KOPE 论文——香港城市大学与华为联合提出的硬件内核优化自进化 Agent 框架。在公共语料极度稀缺的昇腾 NPU 场景下，KOPE 用经验图记忆保留「决策-结果」证据链，配合预算化三层上下文注入，使模型参数完全冻结的前提下通过率达 84.6%（最强基线 57.8%），token 消耗反而下降 93%。本文从内核优化领域背景、经验记忆机制、主动上下文管理、双消融实验到 RISC-V 跨硬件迁移，完整拆解「经验复用为何在语料稀缺场景碾压模型能力」的因果链。</description></item><item><title>Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-cheaper-agent-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-cheaper-agent-paper-reading/</guid><description>斯坦福单作者实证研究：固定模型、系统性变化任务描述本身，量化 prompt 信息量对编码 agent token 开销的影响。2,700 次受控运行显示——把完整规格砍到裸 user story 使成本 +29.7%、轮数 +16.4%（五个任务全部同向）；prompt 只动均值不动方差（重复运行几何标准差恒为 ×1.34）；输出 token 仅占 2.7% 却占 51.1% 花费；单次 $0.11 探测可把未知任务成本预测误差从 161% 降到 36%。「具体性而非要求的存在」才是省轮数关键。</description></item><item><title>Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in MoE LLMs 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-groundhog-bitflip-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-groundhog-bitflip-paper-reading/</guid><description>MoE（混合专家）架构靠稀疏激活省钱省算力，但 Louisiana State、UCLA 等五校联合团队发现：终止 token（EOS/EOT）的生成权竟集中在极少数「终止相关专家」手里，路由层因此成了一个局部化的攻击面。Groundhog Bit-Flip Attack（GBFA）——首个针对 MoE LLM 的 bit-flip 型 Denial-of-Wallet 可用性攻击——只需翻转路由器权重中平均不到 4 个专家对应的少量 bit，就让平均输出膨胀 5912%（最严重 87 倍）、多数样本顶满 token 上限，而模型语义基本无损、PPL 几乎不变。攻击波及对话、推理（思考永不停）、Agent 规划（10 个沙盒全部顶满步数）三种模式。本精读拆解「专家-终止 token 特化」的发现、免推理的脆弱 bit 搜索，以及为什么防御如此棘手。</description></item><item><title>Parason: Revealing Subtask- and Trial Parallelism in LLM Reasoning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-parason-trial-parallelism-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-parason-trial-parallelism-paper-reading/</guid><description>清华、MIT与NVIDIA合作的Parason瞄准推理模型的延迟瓶颈：自回归解码把整条推理链串行执行，难题要等数小时甚至数天。论文首先提出推理并行性的语义分类——Subtask并行（AND分支，分而治之）与Trial并行（OR分支，多路试探），并测量发现Trial并行占了可并行推理计算的多数（DeepSeek-V4在HLE上65.5%），而此前系统几乎只利用了前者。Parason用上下文无关文法把串行推理轨迹改写成引擎可解析的并行结构，配合PA-GRPO多目标奖励（正确性+关键路径延迟+两类并行比例）训练，经SGLang真实执行。AIME24/25等基准上平均加速约1.7×，8k token延迟预算下用25%预算匹配全额性能。</description></item><item><title>Prefix Sliding for efficient test-time scaling 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-prefix-sliding-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-prefix-sliding-paper-reading/</guid><description>斯坦福、华盛顿大学联合 Prime Intellect 等机构提出 Prefix Sliding：推理时只保留前缀（系统指令+prompt，充当注意力锚点）+ 最近数千 token 滑动窗口，丢弃中间推理 token。基于「中间 token 完成子任务后即失去重要性」的注意力概率观察，免训练即可让现有模型提速 3 倍而性能不降；配合截断反向传播，RL 可训练 10 万 token 级长推理轨迹。方法极简但解决了「无限测试时扩展必须有界每 token 成本」这一根本问题。</description></item><item><title>Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-reflection-steering-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-reflection-steering-paper-reading/</guid><description>香港四校合作提出 Reflection Steering：一个免训练的激活空间干预框架，用于抑制大推理模型中冗余的反思行为。方法通过 PCA 去噪、对共享推理方向正交化、逐层校准和有界投影删除四阶段，把「反思方向」从「通用推理方向」中解缠出来，在 6 个模型×基准设置中平均节省 16.9% 的思考 token，MATH-500 上精度统计等效。本精读完整拆解其方向净化机制、层校准协议与统计验证方法。</description></item><item><title>Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-samuon-spectral-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-samuon-spectral-paper-reading/</guid><description>剑桥大学与清华大学合作，用 out-of-sample 谱探测回答优化器领域的核心开放问题：Muon 为何快于 Adam？沿真实训练轨迹把动量缓冲做 SVD，在留出批次上估计每个奇异方向的最优步长，发现稳定且各向异性的谱剖面——单一「易变头」处于稳定性边缘（允许步长最小、恰等于 Muon 实际步长），「宽容主体」允许数倍步长。据此提出 SAMuon：头钉住、主体放大，比调优 Muon 少 13.3-24.0% token 达到同等验证损失。SGD→Adam→Muon→SAMuon 在统一谱分配视角下排成一条线。</description></item><item><title>领读Kimi K3技术报告：一个清华架构博士眼中的注意力谱系与「有效scaling」</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-26-kimi-k3-tech-report-architecture-lead-read/</link><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-26-kimi-k3-tech-report-architecture-lead-read/</guid><description>一集面向技术读者的Kimi K3技术报告领读播客，嘉宾孙宇涛（清华计算机系博士生、上海创智学院pre-doc，研究方向LLM架构与预训练）从K3出发串联起十多篇前作，把KDA线性注意力的每一项公式还原成RetNet→Mamba→DeltaNet→Gated DeltaNet的历史叠加，讲清channel-wise衰减、low-rank dk与BF16 tile的kernel co-design，MLA+QK-norm式门控的稳定性逻辑，Latent MoE对通信开销的削减，以及Quantile Balancing如何用线性规划一步求出负载均衡bias。预训练侧K3反潮流回归cosine decay、在混合注意力里用NoPE让长上下文免调参外推；后训练侧on-policy蒸馏成为多teacher多reward的「多模型合板」方案。嘉宾的暴论：大模型架构没有本质创新了，K3最核心的变量是size——2.8T总参、百B激活、K2的2.5倍scaling效率，而把size做work才是真创新。</description></item><item><title>AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-asmevo-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-asmevo-paper-reading/</guid><description>深度精读 AMD 与南方科技大学合作的 arXiv 2026 论文 AsmEvo：当 GPU 内核源码不可得、部署二进制是唯一行为基准时，用智能体直接在汇编级优化已编译的 AMDGPU code object。恢复可重汇编表示、profiling 定位热窗编辑、ABI 保持重建、差分验证门控接受，在 MI308X 上让 30 个 KernelBench 内核中的 29 个提速，几何均值 1.35 倍、最大 3.88 倍；MI300X 生产负载（AITer、vLLM、SGLang）全部提升且全程保持功能等价。</description></item><item><title>Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-optimal-skill-selection-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-optimal-skill-selection-paper-reading/</guid><description>当 Agent 技能库膨胀到成千上万份文档，往上下文装哪几份技能直接决定任务成败与 token 账单。清华交叉信息研究院 Longbo Huang 组首次把「技能选择」形式化为硬 token 预算下最大化「单调次模收益减线性上下文惩罚」，并提出多项式算法 BPS，证明该问题首个双准则(1−1/e, 1)近似保证，收益系数多项式时间最优。目标函数从执行记录拟合，拟合误差可证转移到有界选择regret。在污染受控 BigCodeBench 变体上，BPS 达 0.73 实测成功率，对已发布路由器、检索器与执行器自选的 0.20–0.52 全面胜出，且比最强路由器省 28% token。本精读拆解其形式化、BPS 算法、预算对齐插值证明，以及「上下文价值是集合级而非单体可打分」的核心洞察。</description></item><item><title>ReCache: 工具增强Agent的组合不变KV缓存复用 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-recache-kv-reuse-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-recache-kv-reuse-paper-reading/</guid><description>工具增强Agent每个请求都要重新编码一遍以不同组合、不同顺序出现的工具与技能schema，标准前缀缓存对此无能为力。ReCache提出resource-wise attention，切断资源间注意力并重置资源内位置索引，使每个资源的KV块具有组合不变性、可独立缓存复用；再叠加贡献选择的层-KV头组路由与字段感知的语义剪枝，把KV张量内存降低92.43%、注意力加速1.423倍，同时Inv-F1基本不降。本精读覆盖其动机、机制、七数据集基准与效果根源分析。</description></item><item><title>FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-flashprefill-v2-paper-reading/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-flashprefill-v2-paper-reading/</guid><description>深度精读中科院自动化所、国科大与腾讯微信联合发表的 FlashPrefill V2——面向长上下文 LLM 服务的训练免费块稀疏 prefill 注意力方案。针对前作「精度失控、内核落后、不兼容生产引擎」的三重差距，V2 以块均值校正补偿被剪枝块的注意力贡献、以对齐 FA3/4 的 warp 特化稀疏内核兑现稀疏收益、以原生 paged KV 与连续批处理无缝接入 SGLang。128K 上下文下 H20 算子级较 FA3/4 对齐 dense 基线加速 17.5 倍（FP8 达 30 倍），端到端 TTFT 最高 4.83 倍，而 RULER/LongBench 精度损失不足 1.1 分，是「算法-内核-系统」三层协同落地的教科书级范例。</description></item><item><title>Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-agentic-esopt-evolution-strategies-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-agentic-esopt-evolution-strategies-paper-reading/</guid><description>新加坡国立大学、南方科技大学与牛津大学团队论证：在长程agent微调场景中，进化策略（ES）不只是更便宜的RL替代品，而是结构性更优的选择。Agentic ESOpt以参数空间扰动+奖励加权更新实现免反传全参优化，GPU显存与推理持平（8.41GB，比GRPO低85.7%），在可控Sudoku实验中呈现horizon-dependent crossover——H*=15时超最强GRPO 12.5个百分点；WebArena-Lite上完成27B模型全参适配（29.47%→36.16%），并支持prompt-参数协同进化。</description></item><item><title>LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-lego-rl-harness-native-coding-rl-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-lego-rl-harness-native-coding-rl-paper-reading/</guid><description>华为LegoX团队联合港中文发布LEGO-RL——在不修改原生编码agent harness内部控制流的前提下接入可扩展策略梯度训练。三大支柱：进程内LLM代理捕获原始生成流实现token级对齐与训练端logprob重算（即使harness压缩/重序列化上下文）、Nydus镜像缓存+分级防御抑制reward hacking、插件化校验监控+Live UI轨迹诊断。训练Qwen3.5-35B-A3B（GSPO）在三大harness上全面提升：OpenHands SDK 64.0%→70.4%、Claude Code 62.4%→68.2%、OpenCode 57.2%→66.6%，rollout-训练概率相关性保持0.99以上。</description></item><item><title>MoNe: Modular Neural Memory for Efficient Long Context Inference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-mone-modular-neural-memory-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-mone-modular-neural-memory-paper-reading/</guid><description>三星研究院提出MoNe——附加到任意冻结预训练Transformer上的模块化神经记忆，实现O(N)预处理+O(1)查询的长上下文推理。两阶段设计：测试时以固定分段读入上下文、快权重神经记忆网络做层局部梯度更新；推理时仅从查询token生成KV，不再重读任何上下文token。128K token时计算与峰值显存较ICL均降约80%（参数开销仅6.4%），RULER上S-NIAH 128K达0.96（ICL仅0.28、MK-NIAH上ICL完全归零而MoNe保持），且可泛化到骨干原生窗口远之外的长度。</description></item><item><title>Agentic Kernel Optimization: Generating State-of-the-Art GPU Kernels Without Hand-Written CUDA 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-agentic-kernel-optimization-paper-reading/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-agentic-kernel-optimization-paper-reading/</guid><description>Intellifusion（云天励飞）技术报告：验证通用代码Agent能否在无任何手写CUDA的前提下产出SOTA GPU内核。在Houmao多Agent编排框架中构建“正确性门控”的内核优化工作流——人类仅做编排（定义流程、强制正确性与反作弊约束、提供关键参考、卡住时重定向搜索），完全不审阅内核代码；起点仅为PyTorch参考实现+工作负载定义+基准命令+紧凑CUDA优化技能集。约19亿agent token在NVIDIA B200上产出：Fused MoE加速92.68×（FlashInfer库为47.08×）、DSA TopK 1101.02×（FlashInfer 52.03×）、DSA Sparse Attention 181.35×（FlashInfer 10.33×）；MLSys 2026 FlashInfer竞赛官方评测中Fused MoE内核1.71×超FlashInfer基线并超过agent-assisted赛道第一名（1.68×）。</description></item><item><title>Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-ventor-qtest-api-audit-paper-reading/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-ventor-qtest-api-audit-paper-reading/</guid><description>腾讯朱雀实验室提出第三方托管LLM API的威胁模型驱动黑盒审计方法：把“服务商是否真的在跑你购买的模型”形式化为随机路由过程估计问题——模型名只是声明而非密码学证明。Ventor-QTest无需目标API任何概率信息：重复请求组件对冻结约束上下文多次重发、从返回文本计数重构类别输出分布，联合报告平均保真损失（AFL）与最坏情况期望保真损失（EFL）双指标，覆盖“平均诚实但最坏降级”的攻击面。论文论证两指标须联合报告（尤其对长时程agentic任务的最坏行为敏感），工具已开源并入腾讯AI-Infra-Guard。</description></item><item><title>Massive Activations in Hybrid Linear Attention Large Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-massive-activations-hla-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-massive-activations-hla-paper-reading/</guid><description>混合线性注意力（HLA）LLM 用少量全注意力层搭配大量线性注意力层来兼顾长上下文与效率，但巨量激活（Massive Activations）在其中如何表现此前几乎无人研究。本精读覆盖首个系统性分析工作：论文发现两种与架构严格对齐的激活形态——注意力前尖峰（PAS）与尖峰间平台（ISP），在 5 种线性架构、6 种混合配置、5 个数据域、1.2B–397B 的 12 个开源检查点上高度复现，Sink-spike 对齐率高达 99.4%–100%；并提出统一的「写入-汇聚-抵消」生命周期机制：MA 的抵消时机决定形态——快速抵消形成 PAS，延迟抵消形成 ISP，全注意力极限下恢复传统 LLM 的稳定 MA。门控实验的不对称效应进一步印证全注意力层是组织 MA 动力学的核心节点。</description></item><item><title>Thought-Level Beam Search for Reasoning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-thought-beam-search-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-thought-beam-search-paper-reading/</guid><description>测试时算力扩展是大推理模型性能的主引擎，但并行采样极度浪费、减法剪枝又让GPU挨饿，核心问题已从“花多少算力”变成“把算力花在哪”。本文提出思维级束搜索 Gambit：用轻量评分器探测步骤边界隐状态，周期性剪除劣迹并立即从高分前缀分支，以零和交换维持固定容量活跃池，同时保持硬件满载。在 5 个基准 × 3 个模型上严格支配 SC、Slim-SC、DeepConf、STEP 四大基线：HMMT-24 最高 +6.7%，token 消耗较并行采样最高降 68.5%，吞吐超 2 倍，管理开销仅 0.97%。本精读覆盖其问题形式化、方法细节、实验证据与因果链根源分析。</description></item><item><title>从DeepSeek到Kimi K3，中国开源模型如何逼出黄仁勋的'开源联盟'</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-13-ai-open-source-deepseek-kimi-huangrenxun/</link><pubDate>Thu, 13 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-13-ai-open-source-deepseek-kimi-huangrenxun/</guid><description>DeepSeek V4 Pro登场，中国开源模型连续逼近前沿能力，迫使黄仁勋牵头组建美国&amp;quot;开放安全AI联盟&amp;quot;，Sam Altman、Sundar Pichai等闭源掌门人罕见支持。这期硅谷101系统拆解了AI&amp;quot;开源&amp;quot;到底开的是什么——从七步训练流程到Open Weights与Open Source的本质区别，以及许可证之争、开源公司如何赚钱、闭源阵营的安全担忧与商业焦虑。</description></item><item><title>Addressable Memory for Video World Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-worldtrace-video-memory-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-worldtrace-video-memory-paper-reading/</guid><description>交互式视频世界模型在长时程生成中会遇到一个隐蔽的「记忆失效」问题：KV cache 里明明存着过去的画面，模型却读不出来。本文精读 NVIDIA/Princeton/ Toronto 联合提出的 WorldTrace 框架，它精准定位了 RoPE 旋转位置编码超出训练范围导致的「内容不可寻址」根因，并用一套无需训练的虚拟槽位机制，在时间一致性上提升 15.5%、在 LoopBench 情景回忆上提升 19.5%。本精读将从世界模型的记忆机制讲起，逐层揭开位置编码、相位抵消、虚拟槽位、规范 key 平均等关键技术。</description></item><item><title>DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-dcas-scaffold-decoupling-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-dcas-scaffold-decoupling-paper-reading/</guid><description>深度精读华为加拿大软件卓越中心与女王大学的 DCAS 论文——首个系统揭示开源 CLI Agent 存在&amp;rsquo;scaffold 锁定&amp;rsquo;现象的工作。论文发现：在 OpenHands 单一 scaffold 下微调的模型，迁移到其他 scaffold 时性能可从 52.6% 暴跌至 8.4%。通过提出 DCAS 后端替换拦截层和区分显式/隐式规划两种形式，论文给出了一条从 scaffold 制品到模型能力的可行迁移路径，仅用 576 条规划感知轨迹即可让模型在非训练 scaffold 上一致提升。</description></item><item><title>「模型能力已经够了，要卷就卷 Infra」｜对话戴冠兰：从 Cloudflare 到 Runta，为十亿个 Agent 造执行底座</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-10-runta-agent-infra-daiguanlan/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-10-runta-agent-infra-daiguanlan/</guid><description>Runta 创始人戴冠兰（前 Cloudflare/Kong 核心）在十字路口播客中提出核心判断：模型能力爬坡已放缓，真正制约 Agent 落地的是执行层基础设施。Runta 刚完成 2000 万美元种子轮（a16z 领投，Jeff Dean、李飞飞天使），定位是为 Agent 打造确定性执行底座——在概率性大模型之上加入隔离、权限、审计和热迁移等系统能力，让企业敢于把生产权限交给智能体。文章梳理了 Token Maximizing 到 Minimizing 的反转、Agent 安全必然爆发的逻辑、以及公有云和基模厂商为何难以抢占这一赛道。</description></item><item><title>一片晶圆的赌局：Cerebras如何从 Scaling Law 的实验室走向千亿推理市场</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-07-cerebras-wafer-scale-inference/</link><pubDate>Fri, 07 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-07-cerebras-wafer-scale-inference/</guid><description>硅谷101专访Cerebras早期投资人周南、前百度研究员Greg Diamos及产品高管Angela，复盘这家&amp;quot;晶圆级引擎&amp;quot;公司十年史。从百度三亿参数模型发现Scaling Law、2016年赌注式投资、良率工程突破，到G42订单、OpenAI两百亿美元推理合同、千亿市值IPO，文章梳理了Cerebras如何从一个物理学的疯狂赌注，变成推理时代绕开HBM与CoWoS瓶颈的差异化基础设施，以及上市后毛利率、供应链和客户自研芯片带来的真实风险。</description></item><item><title>特修斯之船：Kimi K3如何把Transformer的零件全部换掉，还能逼近前沿</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-06-kimi-k3-tech-report-deep-dive/</link><pubDate>Thu, 06 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-06-kimi-k3-tech-report-deep-dive/</guid><description>月之暗面K3是首个达到3T级别的开放权重模型，其47页技术报告揭示了一种&amp;quot;特修斯之船&amp;quot;式的架构哲学：注意力改成了线性+全局混合（KDA+MLA），残差变成了深度方向的Attention，FFN变成压缩空间的稀疏专家，甚至连位置编码都几乎被删掉。RadixArc创始成员赵晨阳和华盛顿大学博士生曾志远分别从Infer和算法两条线拆解K3：线性注意力在2.8T规模上实现了6.3倍解码加速，Quantile Balancing路由是3T稳定训练的关键之一，MOPD让九个领域专家模型高效合板，而KDA（Kernel Development Agent）证明RSI已在kernel优化领域高速运转。核心判断：权重只是一次训练的产物，环境才是能反复产出下一代权重的护城河。</description></item><item><title>Infra 的浪漫与 AI 平权：盛颖从 SGLang 到 RadixArk</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-03-shengying-sglang-radixark-infra-romance/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-03-shengying-sglang-radixark-infra-romance/</guid><description>SGLang 发起人、前 xAI 推理负责人盛颖在 101 视频播客中讲述了从斯坦福形式化验证到 xAI 推理系统、再到创立 RadixArk（1 亿美元种子轮）的完整路径。她提出 Infra 不应是 support 角色，而应成为产品本身；RadixArk 的使命不止于推理引擎，而是让所有人拥有制造 AI 的能力。文章梳理了推理引擎的竞争格局、开源社区的商业化困境，以及一位女性研究者在技术圈中的真实体验。</description></item><item><title>何谓蒸馏？硅谷如何看中国开放模型逼近前沿</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-01-%E4%BD%95%E8%B0%93%E8%92%B8%E9%A6%8F%E7%A1%85%E8%B0%B7%E5%A6%82%E4%BD%95%E7%9C%8B%E4%B8%AD%E5%9B%BD%E5%BC%80%E6%94%BE%E6%A8%A1%E5%9E%8B%E9%80%BC%E8%BF%91%E5%89%8D%E6%B2%BF/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-01-%E4%BD%95%E8%B0%93%E8%92%B8%E9%A6%8F%E7%A1%85%E8%B0%B7%E5%A6%82%E4%BD%95%E7%9C%8B%E4%B8%AD%E5%9B%BD%E5%BC%80%E6%94%BE%E6%A8%A1%E5%9E%8B%E9%80%BC%E8%BF%91%E5%89%8D%E6%B2%BF/</guid><description>月之暗面K3开源权重发布震动硅谷，开源模型首次在多项能力上追平甚至超越最强闭源前沿模型。两位嘉宾——前Hugging Face开源生态负责人王铁镇和TinyFace联合创始人TJ——深度拆解了&amp;quot;蒸馏&amp;quot;争议的技术真相、中国开源模型为何成本更低、Kimi License商业模式对闭源实验室估值体系的冲击，以及开源模型安全之争的真正焦点。核心判断：没有开源，才是这个时代最不安全的事情。</description></item><item><title>GPU其实很闲：AI Infra四层架构与榨干硅极限的效率革命</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-31-ai-infra-gpu-utilization/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-31-ai-infra-gpu-utilization/</guid><description>当AI行业的重心从训练转向推理，一个被忽视的事实浮出水面：GPU大多数时间其实很&amp;quot;闲&amp;quot;。Azure推理负载高达65%的能耗消耗在空转等待上，OpenAI的Chat类请求也达到52%。本文基于硅谷101播客，系统梳理AI Infra四层架构，拆解SGLang/vLLM等开源推理引擎如何通过KV Cache复用、连续批处理、PD分离、投机采样、强化学习训练框架MegaScale等技术，把GPU利用率从50%推向90%+——软件层的每一次优化都变成直接的商业问题。</description></item><item><title>清华程序员很聪明：清程极智如何把Token成本砍掉75%——AI Infra创业的降本逻辑</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-31-qingcheng-jizhi-ai-infra-token/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-31-qingcheng-jizhi-ai-infra-token/</guid><description>清华系AI Infra创业公司清程极智（八卦炉+赤兔推理引擎+AI Ping）联合创始人师天麾深度访谈。高二信息学奥赛金牌保送清华、博士师从翟季冬做高性能计算，2023年底创立公司，一年融资过亿。本篇梳理其核心观点：为什么推理引擎是AI的操作系统、赤兔如何通过FP8/FP4让四台服务器变一台、Token经济爆发后AI Infra被投资人追着投、以及Token服务市场为何是个&amp;quot;黑盒&amp;quot;。</description></item></channel></rss>