<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>模型量化 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%A8%A1%E5%9E%8B%E9%87%8F%E5%8C%96/</link><description>Recent content in 模型量化 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Mon, 07 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%A8%A1%E5%9E%8B%E9%87%8F%E5%8C%96/index.xml" rel="self" type="application/rss+xml"/><item><title>Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-minima-gdn-4bit-quantization-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-minima-gdn-4bit-quantization-paper-reading/</guid><description>社区量化混合架构 LLM 时一致保留循环半区（Gated DeltaNet）的高精度，理由是&amp;rsquo;循环误差会累积&amp;rsquo;。Minima 直接把 NVFP4 W4A4 打满全部 496 个线性层：五任务平均仅 -0.52（种子噪声内），显存 17.5 GiB 最小、prefill 提速 14-19%。四重机制研究（块缩放局域化离群值→门控非线性压缩误差→delta-rule 主动遗忘→逐 token 代价被冲刷）解释了为什么直觉是错的。</description></item></channel></rss>