<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>模型路由 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%A8%A1%E5%9E%8B%E8%B7%AF%E7%94%B1/</link><description>Recent content in 模型路由 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 16 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%A8%A1%E5%9E%8B%E8%B7%AF%E7%94%B1/index.xml" rel="self" type="application/rss+xml"/><item><title>The Router Within: Eliciting Native Skill Routing from a Frozen LLM（Gavel）精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-gavel-native-skill-routing-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-gavel-native-skill-routing-paper-reading/</guid><description>部署的 harness 把所有 skill 元数据预加载进上下文（注意力稀释+库规模受限），检索管线把选择移出上下文但也移出了模型能力。Gavel 证明冻结 LLM 的前向传播已携带路由信号——两个线性映射（唯一被训练的参数）读出任务与各 skill 的 mid-layer 状态，对紧凑 per-skill bank 打分完成全库路由，skill 文本不进上下文。本精读覆盖&amp;rsquo;模型已隐式知道该用什么&amp;rsquo;的探针证据、线性读出的参数效率与路由内部化对库规模扩展的意义。</description></item><item><title>Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-dsr-diverse-skill-routing-dpp-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-dsr-diverse-skill-routing-dpp-paper-reading/</guid><description>Virginia Tech 提出 DSR：当技能注册表达到数万级，top-k 独立排序会返回功能冗余的技能集合浪费上下文预算。DSR 用行列式点过程（DPP）把路由从&amp;rsquo;排序问题&amp;rsquo;升维为&amp;rsquo;集合选择问题&amp;rsquo;，核心创新 query-residual 多样性核先扣除技能表示中与查询对齐的成分再算冗余——在 80K 技能池的 SkillRouter benchmark 上 recall 与 full coverage 双升，多技能查询增益最大。</description></item><item><title>NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-neohorse-1-routing-harness-rsi-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-neohorse-1-routing-harness-rsi-paper-reading/</guid><description>NeoHorse-1 把部署中的模型路由 harness 变成递归自改进（RSI）的数据飞轮：路由层天然记录每次交互的&amp;rsquo;能力需求预测-实际执行-结果&amp;rsquo;三元组，这些记录被转化为保留交错推理与工具调用的 user-turn 训练样本，路由分数进一步组织成三阶段课程 SFT 与路由引导的在线策略蒸馏。4B/9B 模型十项基准宏平均分别从 58.94/65.60 提升至 64.87/69.04，路由 harness 数据比公开 Agent 数据平均高 6.26 分。本文精读拆解其数据管线、课程设计、OPD 机制与&amp;rsquo;评估-选择-更新&amp;rsquo;闭环为何能成立。</description></item><item><title>TACIT-SWITCH: Cost-Aware Model Escalation for LLM Agents from Censored Supervision 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-tacit-switch-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-tacit-switch-paper-reading/</guid><description>深度精读北师大统计学院与香港理工大学合作的 TACIT-SWITCH。论文研究 LLM Agent 运行中何时把控制权从便宜小模型永久移交给强模型这一停时问题，把医学统计的生存分析工具箱搬进 agent 路由：配对 Cheap-Strong 双 rollout 结局加教师标注的粗移交窗口构成区间删失监督，混合治愈模型拆开两个不确定性——强模型能否救回与累积风险何时越过阈值，部署时无需教师。机制仿真 73.52% 对三类基线提升 7.39-11.12 个百分点；ALFWorld 4B→27B 上 48.5% vs 22.4%，DABench 73.1% 且成本最低。统计学家跨界的范例之作。</description></item><item><title>The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-handoff-tax-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-handoff-tax-paper-reading/</guid><description>AWS Agentic AI团队用58,000次agent运行、200万次API调用、360亿token的系统实验，测量了coding agent中途切换模型的隐性代价——Handoff Tax。核心发现呈方向二重性：升级（便宜模型→贵模型）时Raw全轨迹移交只恢复不到一半质量差距且成本可达LC的4-6倍，Claude家族下甚至被&amp;rsquo;弃用重启&amp;rsquo;严格支配；降级（贵模型→便宜模型）却是甜点区，保住大部分质量同时省下大头成本。最有工程价值的是接口反转现象：升级时应丢弃前模型的轨迹只留代码改动，降级时恰恰相反。本文从实验设计讲到成本机制分解，给出模型切换策略的实操建议。</description></item><item><title>Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (Risa) 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-risa-routing-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-risa-routing-paper-reading/</guid><description>Risa（复旦大学）首次把稀疏 MoE 模型的原生路由轨迹用作软件 Agent 测试时扩展的&amp;rsquo;行为坐标系&amp;rsquo;：把每层每 token 的专家路由权重积分成路由指纹，探索阶段选与历史最不相似的候选（disagree to explore），写补丁阶段在同伴收敛处提交，跨尝试仲裁取&amp;rsquo;决策 token&amp;rsquo;上一致性最高者（agree to commit）。SWE-bench Verified 宏平均 44.9%→48.2%，跨家族迁移到 Qwen3.6 仍 +3.5pp——全程无需外部 judge、无需测试执行。</description></item><item><title>RGAO 精读：多智能体代码生成的检索条件化拓扑选择与可证明预算守恒</title><link>https://inkeast.github.io/MessageDaily/posts/2026-05-27-rgao-paper-reading/</link><pubDate>Wed, 27 May 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-05-27-rgao-paper-reading/</guid><description>深度精读 arxiv:2605.05657——提出 RGAO 架构，通过检索代码仓库提取结构复杂度向量来动态选择编排拓扑，配合预算代数系统实现静态预算守恒。误路由率从 30.1% 降至 8.2%（p&amp;lt;10⁻⁶），预算验证在任何 LLM 调用前完成。首次将检索条件化路由与形式化预算代数组合，产生两者单独都不具备的可证明安全性。</description></item><item><title>Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-05-27-router-r1-paper-reading/</link><pubDate>Wed, 27 May 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-05-27-router-r1-paper-reading/</guid><description>深度精读 UIUC NeurIPS 2025 论文——首次将 LLM 模型路由从单轮一对一映射升级为多轮序贯决策过程。Router-R1 将路由器本身实例化为 LLM，通过强化学习训练其交替执行&amp;rsquo;思考&amp;rsquo;和&amp;rsquo;路由&amp;rsquo;动作，动态整合多个 LLM 的互补优势。在 7 个 QA 基准上平均 EM 达到 0.416，超越 RouteLLM、GraphRouter、FrugalGPT 等 14 种基线方法，同时通过成本奖励实现性能-成本 Pareto 优化。</description></item></channel></rss>