<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>模型合并 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%A8%A1%E5%9E%8B%E5%90%88%E5%B9%B6/</link><description>Recent content in 模型合并 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Tue, 29 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%A8%A1%E5%9E%8B%E5%90%88%E5%B9%B6/index.xml" rel="self" type="application/rss+xml"/><item><title>New LoRA Skills Should Read but Never Write 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-read-lora-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-read-lora-paper-reading/</guid><description>LoRA 让大模型微调变得便宜，但把多个独立训练的适配器合并成一个模型始终是难题：直接在权重空间相加会互相干扰，全量重训昂贵且伤害旧技能，路由方案则放弃了「单一模型」的目标。本文追溯其困难根源，指出每个组合方法都在隐式做两个选择——因子坐标（规范自由度）与耦合方向（读写不对称），并提出 READ 方法：规范化因子坐标、只训练新技能的「读行」、折叠回基权重。在 32 个测试谱系中，READ 以平均 +0.073（95% CI +0.047 至 +0.101）战胜 14 种可折叠基线的逐谱系最强者，且全部 8 次失败均集中于 BBH 的新技能习得环节。本精读覆盖其背景、定位、方法、实验证据与优势根源的因果解释，并交叉验证相关谱系的研究结论。</description></item></channel></rss>