<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>上下文召回 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E4%B8%8A%E4%B8%8B%E6%96%87%E5%8F%AC%E5%9B%9E/</link><description>Recent content in 上下文召回 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E4%B8%8A%E4%B8%8B%E6%96%87%E5%8F%AC%E5%9B%9E/index.xml" rel="self" type="application/rss+xml"/><item><title>MassAlloc Attention × DaRoPE：注意力计算分配与位置编码的双子星重构 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-architecture-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-architecture-duet-paper-reading/</guid><description>本精读合并解读两篇重构 Transformer 核心算子的论文：HKUST(GZ)×BAAI×巴黎西岱的 MassAlloc Attention（MALA）把稀疏注意力改写为「分布条件化的运行时计算分配」——保留全部 QK 打分，用 online-softmax 演化归一化因子做逐 tile 贡献比测试，跳过低贡献 tile 的 post-score 计算，同一容差 τ=1 统一前向/反向/prefill/解码，训练 FLOPs -23.1% 而 PPL 无损、关联回忆 89.67% 贴近 FullAttn（NSA 仅 22.61%）；Meta AI 巴黎×Inria 的 DaRoPE（ICLR 2027）诊断出 RoPE 慢频带这一具体弱点，快带保序、慢带换成 sigmoid 有界的逐头内容坐标，外推免配置，K=256 键值回忆 30.5% vs 其他方法 ≤7.25%。二者共享同一哲学：不做一刀切裁剪，让算子自己知道「哪里重要」。</description></item><item><title>编码智能体经济学三重奏精读：成本行为、紧凑文档与上下文蒸馏</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-coding-agent-economics-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-coding-agent-economics-paper-reading/</guid><description>一次读完三篇 2026 年 9 月 25 日同日发布的编码智能体经济学论文：Purdue 的成本低效行为实证研究（三种行为覆盖 79%–98% 任务、最高吃掉 22.75% 成本，7 条开发者原则降本 41.73% 反超检索工具与智能体自合成技能）、孟加拉 DIU 与夏威夷马诺阿分校的紧凑文档基准（源码扣留时 0.08→0.71 的大提升 vs 源码在场时 33 vs 29/30 的大规模 null 结果）、北大的 LOHA+ACD 上下文蒸馏（压缩所见而非所言，上下文降 43%–57%，32K 限制下解决率反升至 21.1%，吞吐 1.9 倍）。本精读逐篇覆盖九部分结构，并在合并结语中回答同一个问题：什么信息值得放进上下文。</description></item><item><title>证据时效性二重奏精读：过期文档投毒与分层协作记忆</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-validity-duet-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-validity-duet-paper-reading/</guid><description>本精读合并解读两篇互补论文：南丹麦大学的「过期文档投毒」证明一条真实的过期检索证据就能推翻模型本来正确的答案（中性检索下 Llama 30%、Qwen 37% 被投毒，GPT-5.5 也有 74/83 被推翻），且模型「会读日期但不会推断适用性」，date-only 仅 6/50 转换而显式失效边界达 50/50；NTU 等机构的 HiCoMER 则从记忆侧给出解法——把冲突消解前移到写入时（SFT+GRPO 训练的分层维护器，Conflict F1 从 46.13 提至 87.88），再叠加有效性感知检索，ORR@5 从 28.63 降至 14.18。一篇证明「过期证据有害且模型不会用日期」，一篇证明「写入时维护+有效性感知检索可救」，问题与解法构成证据时效性的完整二重奏。</description></item><item><title>Just-in-Time Memory×EnSIMem：智能体记忆的读时策展与实体索引 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-25-agent-memory-duet-paper-reading/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-25-agent-memory-duet-paper-reading/</guid><description>本篇合并精读两篇同期 Agent 记忆论文：Salesforce 的 Just-in-Time Memory（JITMEM）把记忆策展从写时推迟到读时，读取时刻按当前任务即时从原始轨迹合成上下文，在 ALFWorld/WebShop/τ2-bench 上较最强基线提升 16.2/16.3/3.9 个成功率点；UIUC+Amazon 的 EnSIMem 用实体-属性索引重组长期记忆，检索沿实体关系而非时间线推进，在 LoCoMo/LongMemEval 上取得 90.6%/92.8% 的新 SOTA。两篇论文恰好回答了记忆系统两个正交的问题——何时策展（读时 vs 写时）与如何组织（实体索引 vs 时间线），合并精读可以拼出智能体记忆设计的完整坐标图。</description></item><item><title>StateComp×PaMER：长程智能体的历史压缩时机与记忆控制信号 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-25-statecomp-pamer-paper-reading/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-25-statecomp-pamer-paper-reading/</guid><description>深度精读同一团队（TierFlow+中国人民大学+清华大学）的两篇姐妹论文：StateComp 首次将「何时压缩历史」建模为状态条件判定问题，构建 KEEP/READY 显式监督数据集与冻结模型隐状态路由器，token 减少 52.27% 且 reward 持平；PaMER 进一步发现压缩与召回控制信号在动作发生前已可从隐状态线性读出（AUROC 0.831/0.765），证明记忆操作是提前计划的，据此构建的门控压缩系统 token 减少 71.7% 且 reward 反升 1.7。</description></item><item><title>EvoOntology: A Self-Evolving Ontology Layer for Data Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-22-evoontology-paper-reading/</link><pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-22-evoontology-paper-reading/</guid><description>EvoOntology（中国人民大学 ruc-datalab）提出一个面向数据智能体的「自进化本体层」：把数据库的领域概念、字段映射与约束封装成可被 agent 在运行时按需查询的 MCP 服务，并用 builder agent 自动构建初版本体、用「诊断—归因—修补—门控」四步环从失败轨迹中持续进化。本文按九部分结构精读，重点拆解三层架构、四步进化环，以及为何「静态语义层全量注入反而掉分」是全篇最有证明力的实验设计，并从因果链上解释其优势根源。</description></item><item><title>BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-beaconkv-thought-revisiting-kv-compression-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-beaconkv-thought-revisiting-kv-compression-paper-reading/</guid><description>KV cache 压缩方法都用&amp;rsquo;最近的查询&amp;rsquo;预测未来注意力——本文发现长程推理中这个假设根本不成立：解码会不定期产生&amp;rsquo;思维重访令牌&amp;rsquo;（TRT），重新关注数千 token 之前的推理计划，而近期查询无法预知这次重访。关键观察是 TRT 对应的全局查询在嵌入空间聚成少数簇——只需为每簇维护一个&amp;rsquo;信标查询&amp;rsquo;（Continual FPS 在线选取），就能预判哪些 KV 将被重访。训练自由、无需改动架构：四个开源推理模型上内存最高压缩 5.8×、精度近全量、吞吐 +4.3×，对 RPC/R-KV 最高领先 31.7 个百分点。</description></item><item><title>Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentstream-progressive-latent-memory-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentstream-progressive-latent-memory-paper-reading/</guid><description>流式视频理解的主流范式是&amp;rsquo;存历史、按需检索&amp;rsquo;，但外部证据永远只是临时上下文。南京理工×蚂蚁×NUS×港中文的 LatentStream 把范式翻转为&amp;rsquo;检索并内化&amp;rsquo;：分层流记忆 + 渐进扩张感受野的 latent token 把历史证据固化进固定长度潜记忆，用熵构造的渐进置信奖励在测试时联合优化。OVO-Bench 64.2%、StreamingBench 76.9%、MLVU +6.1，全部 SOTA。</description></item><item><title>LatentPress: Context Compression Beyond Text and Vision 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentpress-soft-token-context-compression-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentpress-soft-token-context-compression-paper-reading/</guid><description>压缩后的上下文通常仍以文本或图像这两种&amp;rsquo;给人看&amp;rsquo;的形式存在。LatentPress 提出第三种表示：小型 writer 把对话/文档直接写成连续记忆 token，冻结解码器经输入嵌入接口读取，推理时零文本重建。LongMemEval 上 7.7× 压缩反而比未压缩证据更准（0.504 vs 0.490），写入 43ms 比摘要快一个量级——机器原生的记忆表示从此有了实证立足点。</description></item><item><title>Language Models Can Control Their Own Attention 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-declarative-attention-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-declarative-attention-paper-reading/</guid><description>长上下文解码时，模型每生成一个 token 都要把整个 KV cache 读一遍——1M token 上下文意味着每步约 15GB 的内存搬运，而注意力其实高度集中。KAIST AI 联合 Google DeepMind 提出 Declarative Attention：让模型在思维链里用 &lt;global&gt;/&lt;focus&gt;/&lt;local&gt; 三种标签自己声明&amp;rsquo;现在需要看哪里&amp;rsquo;，推理引擎像解析工具调用一样解析声明并跳过绝大部分 KV 读取。零训练、零外部打分器，15 个长上下文任务上 Gemma-4-31B 注意 token 降 52.0%、精度仅降 1.27pp。本精读拆解三模式协议、与代理打分式稀疏注意力的机制差异，以及&amp;rsquo;模型自己最知道该看哪里&amp;rsquo;的第一性原理。</description></item><item><title>Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-03-skill-following-rae-paper-reading/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-03-skill-following-rae-paper-reading/</guid><description>崇实大学提出 Skill Following（SF）评测框架：现有&amp;rsquo;检索 vs 不检索任务的聚合分差&amp;rsquo;衡量技能库价值存在严重选择偏差。论文形式化 RAE（Retrieval-Invoked Actual-Use Effect）指标——仅在 agent 主动发生检索的任务上，比较同一任务开/关技能的配对执行差。17 个 LLM × 编码（MBPP+）/数学（Math500）的实测揭示&amp;rsquo;评测悖论&amp;rsquo;：多个模型聚合检索提升为正、RAE 却为负——系统层面看似受益，恰恰在真正调用了技能的任务上反而有害。诊断分析证明&amp;rsquo;上下文出现技能内容&amp;rsquo;远不等于&amp;rsquo;模型遵循技能&amp;rsquo;，当前工具使用能力被系统性高估。</description></item><item><title>CorporateBench: Large-Scale Q&amp;A Benchmarking with Temporal Knowledge Bases 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-corporatebench-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-corporatebench-paper-reading/</guid><description>深度精读 Epiq AI Labs 与康奈尔大学联合发布的企业级问答基准 CorporateBench。论文用程序化生成的时序知识库（KB）构建四家虚拟公司（12 到 10210 名员工、共 26.3 万封邮件），并从 KB 用人工验证的 SPARQL 查询确定性导出标准答案，保证任意规模下的跨文档逻辑一致性。五个前沿模型测试显示：实体抽取基本不随规模衰减，但关系抽取与时序关系严重崩坏；KB 直连（SQL 工具）显著优于 RAG，且两者差距随规模从 0.24 扩大到 0.37，揭示大规模企业通信网络仍是当前 LLM 的重大短板。</description></item><item><title>Lost in Compression 精读：抽取式提示压缩器的跨语言审计</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-lost-in-compression-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-lost-in-compression-paper-reading/</guid><description>提示压缩号称能砍掉 LLM 推理成本，但主流学习型压缩器几乎全部用英语训练和评估。这篇来自 Hostinger 与考纳斯理工大学的论文做了一次严格的受控跨语言审计：在十种语言、五种文字系统的完全平行数据上，用目标模型分词器做预算匹配对照，超过 25.8 万次评估调用后发现迁移差距真实存在且随压缩强度急剧放大——保持率 0.33 时英语保留 57 至 62 个百分点的上下文价值，中文几乎归零。差距由监督语言而非模型架构驱动：三种英语训练的压缩器全部复现差距，多语训练的 XProvence v1 完全没有差距，确定性基线也无差距。非英语的安全压缩预算只有英语的一半左右。</description></item><item><title>MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-memuse-natural-integration-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-memuse-natural-integration-paper-reading/</guid><description>对话式 LLM 的记忆系统一直用“直接问答“评测：问模型能否回忆先前对话里的事实 X。京都大学团队做了 4 个月真实部署（40 用户、1872 会话、7 种记忆条件）检验这个假设——Direct QA 准确率随容量从 19.7% 涨到 70.1%，用户满意度却纹丝不动。他们从中检出 72 个用户主动引用记忆的真实时刻，构建 MEMUSE 基准，用“自然整合“（回复是否真正织入被引用的记忆）替代召回评测：同一模型同一上下文，Direct QA 78.8% vs 自然整合仅 7.9%，71 分鸿沟；且只有自然整合与满意度相关（ρ=+0.29），Direct QA 完全不相关。Two-step 消融把瓶颈定位于对话生成层而非检索层——即使提取步骤已给出正确细节，生成仍有 77% 不使用。</description></item><item><title>Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-scaleqa-episode-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-scaleqa-episode-paper-reading/</guid><description>深度精读 UC San Diego 的 SCALE-QA 与 TSIM 论文——针对真实助手使用形态「单线程多主题混杂的长对话」定义并测量 episode integrity failure：决定性证据明明在对话里，系统却检索到貌似合理的碎片而非让局部约束生效的完整 episode。SCALE-QA 用 3000 题反事实基准（防预训练泄漏）+ 确定性运行时打包（16k-1M）；TSIM 以语义漂移在线分段 + 三视图 episode 索引，在三个后端全部第一（比最强基线高 5.6-17.6 点），1M 诊断中用约 1.3k token 达 96.5% 而 Full Context 用 1.05M token 只有 87.2%。本文拆解「找回 episode 而非答对问题」才是瓶颈的完整证据链。</description></item><item><title>Recuris: Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-recuris-memory-evolution-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-recuris-memory-evolution-paper-reading/</guid><description>Recuris（NUS × Stanford × Oxford × Princeton）把递归自我改进从『改模型/改智能体』收缩到『只演化外置记忆控制层』：工作记忆维护经检查器验证的任务状态并按需调用技能，跨任务的固定 Meta-Agent 读结构化轨迹、把失败归因到 E/W/ρ/C 四组件之一并只修补被归因组件，经修复源任务且不回退开发集的验证门才准入。在 4 个长程基准 × 10 个模型上 35/37 完成的模型-基准对成功率提升，GPT-5.6 Sol +17.8、Claude Opus 5 +15.6、最长任务 +32.2 分，六类长程失败模式下降 20–86%；机制上证明长程失败是执行问题而非检索问题，技能价值是『调用条件性』的，结构化轨迹使故障定位从 13.0% 提升到 64.8%。</description></item><item><title>Active Inference as Context Acquisition for AI Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-active-inference-context-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-active-inference-context-paper-reading/</guid><description>深度精读 arXiv 2608.19202——把 Agent 的上下文获取（澄清提问、检索、工具调用、提示试验）形式化为主动推理：内层更新对潜在任务状态的信念，外层在上下文动作、任务动作与停止动作之间做选择，最小化包含风险、认知价值与成本的期望自由能。论文推出 OQA 基准，把提问变成属性表上的二十问游戏，用动态规划给出最优 oracle；七个前沿模型全部落后于 oracle。在提示补全实验中，定向澄清把验证器合规率从 0.0417 提升到 0.375，最佳设定 ε*=0.01、K*max=2。核心主张：主动推理是模型无关的上下文获取层设计原则。</description></item><item><title>The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search (Ascp) 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-ascp-context-allocation-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-ascp-context-allocation-paper-reading/</guid><description>Ascp（北京大学 × 腾讯）为生成式搜索建立了&amp;rsquo;上下文分配定律&amp;rsquo;：同预算下窄窗多轮（k=2 检索×T=12 轮生成）比宽窗单轮（k=24×T=1）的 portfolio recall 高 0.144，T:1→12 带来 16.8-20.5pp 提升，且验证到 32B 规模。其测量工具是因果留一（LOO）探针——teacher-forced 反事实消融直接测量每篇文档对生成文本的因果利用率，在 same-query 硬负例下 AUC 0.876，而 embedding 相似度坍缩到 0.484（随机水平）。相关性代理测的是&amp;rsquo;话题相关&amp;rsquo;，不是&amp;rsquo;真被用了&amp;rsquo;。</description></item><item><title>MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-memtrapbench-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-memtrapbench-paper-reading/</guid><description>浙江大学联合新加坡国立大学、东北大学、Heriot-Watt 大学与腾讯提出 MemTrapBench，首次系统评测记忆诱发的认知陷阱：忠实记录且语义相关的记忆，仍可能扭曲大模型的推理与信念，使表现跌破无记忆水平。基准按植入陷阱、噪声掩埋、触发陷阱三阶段生成 1050 个多轮对话实例，覆盖认知偏差、任务边界、创伤、安全四类场景；五个主流记忆框架在 Gemini 与 Qwen 上全部落后无记忆基线逾 10 个百分点，创伤场景去陷阱对照的正确性从 66.40% 回升至 91.07%，证明退化源于陷阱语义而非上下文长度。论文进一步提出推理时提示技能 AdaptiveMem，把是否该用记忆显式化为决策前的静默校验，最高提升 14.9 个百分点且不损害常规记忆基准表现。</description></item><item><title>Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-qcr-trajectory-reuse-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-qcr-trajectory-reuse-paper-reading/</guid><description>深度精读 arxiv:2608.12847——西安交大团队提出 QCR（Query-Conditioned Reuse），指出轨迹记忆的真正瓶颈不在检索而在检索之后的“复用”环节：历史轨迹里的用户名、路径、日期等绑定值会随时间过期，直接注入会诱导模型照抄旧值。QCR 在检索与执行之间插入一步改写，把选中轨迹转化为“工作流不变量/需重取绑定/适用条件/验证护栏”四字段笔记。在 WebArena/WorkArena/AppWorld 共 2391 个目标上，平均 Success 62.3%（比注入完整轨迹高 10.7 点），在线 token 省 48.9%；大绑定偏移下过期绑定错误率从 46.9% 降至 10.9%。</description></item><item><title>Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-08-activity-frames-paper-reading/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-08-activity-frames-paper-reading/</guid><description>深度精读 arXiv:2608.05784——独立研究者 Nossa Iyamu 提出的 Activity Frames，一个零模型、确定性的屏幕活动编译管道。它将屏幕捕获流分割为携带应用、站点、时序、输入量和证据指针的『活动帧』，在 128,756 帧、51 活跃天的真实语料上把单日上下文从 126,812 token 压缩到 1,469 token（86×），编译延迟仅 68ms，下游问答准确率 98.4%（LLM 摘要仅 66-80%），幻觉率 0%。同一编译器还首次测量了代理成本模型假设但从未实测的两个参数：Routine Overhead Ratio R=60-343x 和可委托复发率 h=7.7%（样本外）。核心洞察：把『解释』从『测量』中剥离，用最无趣的确定性代码填补捕获与记忆之间的缝隙。</description></item><item><title>LLMs Get Lost in Evolving User Intent 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-01-evolving-intent-paper-reading/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-01-evolving-intent-paper-reading/</guid><description>本文精读 Microsoft Research 团队发表于 2026 年 7 月的论文《LLMs Get Lost in Evolving User Intent》。论文提出一个将任意静态单轮基准测试转化为动态多轮对话的框架，通过三种意图转移（论点揭示、论点修正、函数切换）模拟用户意图的真实演化过程，同时保留原始评估协议实现免标注的自动验证。跨数学、Text-to-SQL、搜索、编程四个领域的实验揭示了一个一致现象：在单轮设置下表现优异的模型，一旦用户意图动态演化，性能便大幅下降，最严重时直接归零。这一发现暴露了静态评估的盲区，对协作式 Agent 的未来发展具有关键启示。</description></item><item><title>FastContext: Training Efficient Repository Explorer for Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-17-fastcontext-paper-reading/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-17-fastcontext-paper-reading/</guid><description>深度精读微软 FastContext 论文——把代码仓库探索从主 Agent 中剥离、交给一个专门训练的小模型来干。用 4B-30B 的专用探索模型替代昂贵的探索过程，端到端解决率最高提升 5.5%，主模型 token 消耗最高降低 60%。从编码 Agent 瓶颈分析、模型分工解构、SFT+RL 训练配方，到必要知识反推与通用性灵感，全面拆解这项将仓库探索&amp;rsquo;模块化、可训练化&amp;rsquo;的工作。</description></item><item><title>GPS: Graph-Guided Proactive Information Seeking in Large Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-05-26-gps-paper-reading/</link><pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-05-26-gps-paper-reading/</guid><description>深度精读北京大学 ICLR 2026 论文 GPS，提出用 DAG 显式建模文档中的条件规则结构，通过图遍历引导 LLM 在 RAG 系统中高效主动追问，成功率超最强基线 7.5%，追问效率提升 4.2%。</description></item></channel></rss>