<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>持续学习 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%8C%81%E7%BB%AD%E5%AD%A6%E4%B9%A0/</link><description>Recent content in 持续学习 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Tue, 29 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%8C%81%E7%BB%AD%E5%AD%A6%E4%B9%A0/index.xml" rel="self" type="application/rss+xml"/><item><title>New LoRA Skills Should Read but Never Write 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-read-lora-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-read-lora-paper-reading/</guid><description>LoRA 让大模型微调变得便宜，但把多个独立训练的适配器合并成一个模型始终是难题：直接在权重空间相加会互相干扰，全量重训昂贵且伤害旧技能，路由方案则放弃了「单一模型」的目标。本文追溯其困难根源，指出每个组合方法都在隐式做两个选择——因子坐标（规范自由度）与耦合方向（读写不对称），并提出 READ 方法：规范化因子坐标、只训练新技能的「读行」、折叠回基权重。在 32 个测试谱系中，READ 以平均 +0.073（95% CI +0.047 至 +0.101）战胜 14 种可折叠基线的逐谱系最强者，且全部 8 次失败均集中于 BBH 的新技能习得环节。本精读覆盖其背景、定位、方法、实验证据与优势根源的因果解释，并交叉验证相关谱系的研究结论。</description></item><item><title>Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-27-jaz-invoke-paper-reading/</link><pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-27-jaz-invoke-paper-reading/</guid><description>MIT CSAIL 团队提出 JAZ：一个只比 agent loop 多一点点的极简智能体框架。它仅暴露一个 LLM 原语 invoke——一个函数体由 LLM 在运行时生成的函数——加上动态作用域与 hooks，就涌现出传统上需要专门 harness 才能实现的长程记忆与持续自改进能力：在 StuLife 远程回忆子集上以约 43% 的成本超越 Letta（MemGPT）8 个百分点，在 AppWorld 上以更低成本胜过专门的自改进框架 ACE。本文从语言原语的第一性原理出发，拆解 invoke 的两条定义性质、tail-recursive delegation 如何统一各类上下文管理为特例，并用 MemGPT、RLM、ACE、context rot 研究等外部文献交叉验证其效果优势的根源。</description></item><item><title>Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-27-jit-memory-paper-reading/</link><pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-27-jit-memory-paper-reading/</guid><description>Salesforce AI Research 提出 JITMEM，把智能体记忆的「塑形」时机从写时推迟到读时：写时零损耗保存原始轨迹，读时由 curator LLM 联合当前任务与检索轨迹合成任务自适应 payload，用后即焚。因为 payload 在当前任务上被立即消费，curator 可用 GRPO 直接以任务原生得分训练，信用分配零步延迟。在 ALFWorld、WebShop、τ²-bench 上分别领先最强基线 16.2、16.3、3.9 个成功率点，输入 token 省约一半。本精读覆盖其问题定义、方法拆解、实验证据、根因分析与可迁移灵感。</description></item><item><title>内核证据检测与记忆家族隔离：Agent 基础设施二重奏精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-26-duet-kernel-memory-paper-reading/</link><pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-26-duet-kernel-memory-paper-reading/</guid><description>「二重奏」精读两篇 Agent 基础设施论文。第一篇《On the Effectiveness of Kernel-Level Evidence for Agent Security》构建 ACE 配对语料库（4047 会话×17 威胁模型），首次系统测量内核 syscall 证据对 Agent 安全检测的增益：Kernel-only 最高 OOD AUROC 0.922，跨层拼接普遍优于任一单层，Falco 默认规则近乎随机，证明内核证据应成为 Agent 检测的一等输入。第二篇《Scope Before You Persist》针对持久技能记忆的跨家族干扰，提出「认证范围=部署范围」原则与 Scoped-ORC：仅改变检索范围即把有害接受从 6/12 降到 0/63，27 流效用 +0.063，证明范围匹配而非更强的验证器才是持续适应的关键。</description></item><item><title>Continual Learning Mechanisms Compose for Long-Horizon Memorization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-cl-mechanisms-compose-long-horizon-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-cl-mechanisms-compose-long-horizon-paper-reading/</guid><description>让模型依次学 100 个知识任务且不留旧例、不给任务 ID——&amp;lsquo;长程记忆化&amp;rsquo;设定下，任何单一持续学习机制都崩盘（保留率普遍个位数）。Johns Hopkins 的系统学研究证明机制要&amp;rsquo;组合&amp;rsquo;：锚点类型（data 复演/function 蒸馏/权重正则——保留什么）×低秩分配（merged LoRA——保留在哪）两维设计，任务级逐次减半搜索组合空间+因子实验量化交互。最优组合（三锚点+merged LoRA）把最终保留率从 1.2% 拉到 34.9%（28 倍），且 data 锚点×merged LoRA 在三个数据集上一致超可加——遗忘来源互补，组合解决结构问题。HF 日榜 281 赞当日第一。</description></item><item><title>Dream-RSI: Recursive Self-Improvement through Evolving Worlds 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-dream-rsi-replay-simulator-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-dream-rsi-replay-simulator-paper-reading/</guid><description>RSI（递归自我改进）的核心瓶颈是探索策略管理：固定策略无法适应搜索空间扩张，在线策略优化又受困于长程 rollout 的延迟昂贵反馈。Dream-RSI 的关键洞察是——积累的发现历史本身就是已实现搜索空间上的重放模拟器，把探索策略的改进从昂贵的真实环境 rollout 搬到廉价的历史重放（做梦即训练）。Lasso 求解器发现任务上 agent 调用较 SimpleTES 削减 162×，held-out 运行时 3587→2931ms。本精读覆盖三环循环机制、重放模拟器的信息学根基与发现求解器的算法细节。</description></item><item><title>RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-rsiagent-causal-memory-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-rsiagent-causal-memory-paper-reading/</guid><description>数字 Agent 进入新环境（接口/工具/失败模式预训练未覆盖）时如何无监督适应？RSIAgent 给出 training-free 答案：curriculum/actor/verifier 三类 Agent 协同自主探索，把&amp;rsquo;动作-条件-后果&amp;rsquo;因果关系沉淀为可冻结复用的记忆；广度+深度双探索消融显示完整 RSI 74.54% 显著优于单策略（65.52%/56.50%），并让 Kimi-K3、GLM-5.3 在 OSWorld-v2 与 Agent&amp;rsquo;s Last Exam 上反超 GPT-6 Astra。本精读覆盖因果记忆与轨迹记忆的本质差异、广深互补的机制解释与开源反超闭源的信号意义。</description></item><item><title>When Agents Slow Down: Elo-per-token 分析与 Agent 测试时策略 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-elo-per-token-agents-slow-down-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-elo-per-token-agents-slow-down-paper-reading/</guid><description>Agent 在测试时的&amp;rsquo;减速&amp;rsquo;行为——更多 token 换来多少真实能力提升？本文提出 Elo-per-token 度量：以独立采样为理论参照（Elo 随 log compute 线性增长），定义 scaling inflection point（边际 Elo 增益跌至参照线的每会话预算）。4 个通用 Agent × 4 个开放基准、单会话最高 1 亿 token 的实验给出反直觉发现：AtCoder Heuristic Contest 上 Agent 超越历史最强人类选手的超线性提升是持续学习的证据——减速之后仍有巨大 headroom。本精读覆盖测试时 scaling 的度量学、独立采样参照的设计逻辑与&amp;rsquo;持续学习 vs 收益递减&amp;rsquo;的分界证据。</description></item><item><title>LifeMem: Enabling Lifelong Experience Reuse for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-lifemem-lifelong-experience-reuse-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-lifemem-lifelong-experience-reuse-paper-reading/</guid><description>北理工 BITHLP 实验室的 agent 记忆工作 LifeMem：针对跨环境经验迁移与灾难性遗忘两大难题，按底层 workflow 聚类交互轨迹提取可复用技能（结构级抽象而非表层相似），推理时召回技能+轨迹引导动作。在 5 场景 10 环境 13k+ 任务上验证（其中 4 环境新标注 2k+ 轨迹），遗忘降低与跨任务迁移双优；并发现任务流顺序影响学习、结构相似巩固有增益。数据集代码全开源。</description></item><item><title>Memory as Plans 精读：把记忆从执行期条件重构为规划期证据，机器人非马尔可夫任务 SOTA</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-13-memory-as-plans-map-wam-paper-reading/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-13-memory-as-plans-map-wam-paper-reading/</guid><description>哈工大×NTU×山大提出 MaP-WAM：将记忆依赖的世界-动作建模拆解为记忆锚定规划与计划条件执行两层——情景区段记忆作为规划期证据，因果世界模型（WAN-2.2-5B 微调）生成视觉计划，World-Action-Progress 模型把任务进度升级为一等模态。RMBench 83.3% SOTA、真机 78.0%，执行器延迟随历史增长恒定；Swap T/Press Button 达 96%。</description></item><item><title>Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-coevolving-harness-model-imitation-fit-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-coevolving-harness-model-imitation-fit-paper-reading/</guid><description>harness 进化后，让弱模型模仿更强专家的轨迹——这个&amp;rsquo;显然正确&amp;rsquo;的配方在七个企业任务上全部翻车（平均 -14.9 分），而同样的做法在未进化 harness 下却有增益。论文定位出根源：模仿让弱模型学会了专家的知识，却也继承了专家的规划风格，破坏了它与&amp;rsquo;围绕自身原生风格进化出来的 harness&amp;rsquo;的拟合。解法是 on-policy 专家修正：meta-MLE agent 定位失败 turn、专家只重写那一轮，平均 +1.7 分且规划失败桶保持地板水平。本文精读拆解&amp;rsquo;模型-harness 拟合&amp;rsquo;这一新概念与其共进化配方。</description></item><item><title>Experience Funnel: A State–Policy Alternating Loop for Self-Evolving Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-experience-funnel-state-policy-loop-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-experience-funnel-state-policy-loop-paper-reading/</guid><description>自进化 Agent 面临双时间尺度困境：文本状态（技能/记忆）快而外部依赖重，参数策略持久而更新慢。华为与港理工的 Experience Funnel 用交替环打通两者：轨迹先蒸馏为显式状态快速适配，再选择性把&amp;rsquo;跨状态修订仍有效&amp;rsquo;的行为经 transition-aware distillation 固化入策略——经验像漏斗一样从原始轨迹逐级过滤为可复用能力。多基准上一致超越 state-only 进化与 policy-internalization 两条单路线。</description></item><item><title>Procedural Graphs: Self-Evolving Execution Structures for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-procedural-graphs-self-evolving-agents-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-procedural-graphs-self-evolving-agents-paper-reading/</guid><description>知识图把事实组织成 (实体, 关系, 实体) 三元组来回答&amp;rsquo;是什么&amp;rsquo;；Google 团队的 Procedural Graph 用 (过程, 关系, 过程) 三元组回答&amp;rsquo;怎么做&amp;rsquo;。Agent 每步决策时定位活跃节点，引导模型把邻域子图翻译成步级情境引导；离线自进化循环对比成败轨迹编辑图拓扑与属性，验证门通过才采纳、拒绝项存为负约束。六个基准三个 LLM 全面超越 ReAct/ExpeL/AWM 等记忆基线——Gemini 3.1 Pro 上 τ-bench 72.17→80.00、GDPval 56.39→78.78、ALFWorld 满分，零骨架自进化图匹配乃至超越手工设计。本文精读拆解过程性知识的表示设计与&amp;rsquo;验证门+拒绝记忆&amp;rsquo;的进化机制。</description></item><item><title>EvoHarnessBench: 智能体能跟上不断进化的 Harness 吗 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-evoharnessbench-evolving-harness-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-evoharnessbench-evolving-harness-paper-reading/</guid><description>Salesforce Research×UNC Chapel Hill×UW–Madison 的 EvoHarnessBench 把非平稳性从任务流转移到 harness 本身：17 条受控 harness 进化流（802 任务、520 工具、42 技能、62 智能体），分部署评估（能力保持）与自进化适应两设定。基准回答一个此前无人系统提问的问题：当工具、技能、子智能体持续增加时，已部署 agent 的既有能力何去何从。</description></item><item><title>Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentstream-progressive-latent-memory-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentstream-progressive-latent-memory-paper-reading/</guid><description>流式视频理解的主流范式是&amp;rsquo;存历史、按需检索&amp;rsquo;，但外部证据永远只是临时上下文。南京理工×蚂蚁×NUS×港中文的 LatentStream 把范式翻转为&amp;rsquo;检索并内化&amp;rsquo;：分层流记忆 + 渐进扩张感受野的 latent token 把历史证据固化进固定长度潜记忆，用熵构造的渐进置信奖励在测试时联合优化。OVO-Bench 64.2%、StreamingBench 76.9%、MLVU +6.1，全部 SOTA。</description></item><item><title>Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-bcit-conditional-experience-transfer-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-bcit-conditional-experience-transfer-paper-reading/</guid><description>自主 LLM 后训练系统不断积累&amp;rsquo;过去什么更新有效&amp;rsquo;的经验，但父模型一旦变化，旧经验就可能是毒药。本文把这一困境形式化为条件经验迁移问题，提出 BCIT：把效果绑定到源上下文、更新前检查适用性、具名硬冲突否决、必要时小预算试验取证。等预算对比中 BCIT 更少授权有害更新、最终模型质量更高，为自进化 Agent 补上&amp;rsquo;免疫排异&amp;rsquo;机制。</description></item><item><title>Aspire: Can Models Self-Evolve from Vague Goals? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-aspire-vague-goals-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-aspire-vague-goals-paper-reading/</guid><description>现有 LLM 自进化研究都从人类定义好的显式任务出发，agent 只搜索&amp;rsquo;怎么优化&amp;rsquo;；但人类学习往往始于&amp;rsquo;成为更好的物理学家&amp;rsquo;这样的模糊目标。ByteDance Seed 联合 SUTD、M-A-P 等发布 Aspire 基准：只给一句自然语言能力目标，评测集对 agent 完全隐藏，agent 必须自己决定优化什么、怎么训练、如何验证。实验给出罕见的机制级阴性结果——24 次 final-only 运行仅 1 次超过基线分，最佳进化 harness 仍低于人工 Qwen-Agent。本精读拆解隐藏评测设计、三条研究问题（RQ1-RQ3）的实验逻辑，以及&amp;rsquo;代理增益不迁移&amp;rsquo;这一失败模式的根源。</description></item><item><title>S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-s3gym-self-improvement-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-s3gym-self-improvement-paper-reading/</guid><description>Agent 每天与环境交互积累海量轨迹，但经验真的变成了能力吗？ByteDance Seed 姊篇基准 S3Gym 把&amp;rsquo;自改进&amp;rsquo;拆成自测试、自判断、自改进三个可测环节，在 7 个可执行验证的文本游戏上比较三种经验注入通路：原始历史 ICL、摘要记忆、参数训练。7 个前沿模型的核心发现：自改进既不自动也不均匀——GPT-5.5 在 PvZ 上 History ICL 的 AUC⁺ 高达 548.5，换摘要记忆暴跌到 33.2；同一模型同一环境换个通路结果天差地别。本精读拆解宽松探索/严格评测的分离设计、自评分与环境真值的对照记录，以及&amp;rsquo;经验压缩可行性决定通路优劣&amp;rsquo;的机制规律。</description></item><item><title>Fast Weight Attention for Continual Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-falcon-fast-weight-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-falcon-fast-weight-paper-reading/</guid><description>深度精读 ByteDance Seed、Princeton、清华、UCLA 与 Hyperbolic Labs 合作的 Falcon 论文：把线性注意力与状态空间模型的状态转移显式化为在线学习规则，发现读后写语义下快记忆的正确训练对应是前缀配对 φ(k(t-1))→v(t) 而非常见的同对配对，并从平方误差回归与内积两个局部目标统一推导出 NLMS 归一化的六个变体族（Falcon-1/2/3 与 Falcon-1A/2A/3A），全部兼容 SSD 式 chunk 并行训练。语言建模上与最强递归基线互有胜负，变长加法外推上 Falcon-3A.3 以 87.2 平均精度大幅领先 Transformer 的 65.8。</description></item><item><title>CAFE: Self-Improving Search Agents Need Co-Evolving Feedback 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-cafe-coevolving-feedback-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-cafe-coevolving-feedback-paper-reading/</guid><description>CAFE（复旦大学 × 腾讯 LLM 部门）把纠正性反馈做成搜索智能体轨迹内可请求的干预：共享参数模型分饰 agent/critic 两角色，冷启动 SFT 用『保留自身错误前缀+教师插入反馈+成功续跑』的恢复演示；在线 RL 用比较反馈估计（同提示请求组 vs 跳过组的成功率差）塑造请求回报，反馈感知优势塑形按请求边界分离『跑偏前缀』与『修复续段』的信用；离线 RDPO 从前缀匹配的成功/失败对学反馈生成；100 步在线×1 次 RDPO 交替 5 轮。7 个 SearchQA 基准上 Qwen2.5-7B 平均 EM 52.5/F1 60.7 为最强 RL 搜索方法，6 个 OOD 基准全保持，答案级幻觉率 17.6%→12.6%；单向消融证明只改 agent 或只改 critic 都会平台化，交替优化持续上升——末代 critic 配末代 agent 84.0 vs 配 SFT critic 80.6。</description></item><item><title>Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-frag-unlearning-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-frag-unlearning-paper-reading/</guid><description>KAIST 与东京大学合作研究机器遗忘的「复活」问题：遗忘后的 LLM 经短暂微调即可恢复已删知识。论文反驳「权重移动距离决定鲁棒性」的流行假说，提出免训练预测器 FRAG——度量遗忘更新是否集中在 forget 关键权重而避开 retain 关键权重，Spearman 相关达 −0.78（全局 L2 仅 −0.36）；同原理 instantiated 为剪枝方法 FRP，在三种重学习攻击下 post-attack 遗忘分数全面最优。「哪些权重动了」比「动了多远」更能解释鲁棒性。</description></item><item><title>Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-unfolding-papers-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-unfolding-papers-paper-reading/</guid><description>字节跳动 Seed、南京大学与 Evolvent AI 合作提出「论文展开」管线：保持论文原文逐字不动，用教师模型反向重建写作请求、全局规划与逐节写作前的思考，把整篇论文展开为多轮生成轨迹用于持续预训练。1.8M 篇 arXiv 论文的 30B token 原文被展开为 57–60B token 轨迹，中位文档长度从 11.2K 提升到 28–29K token。同一反向构造还产出 200K 样本 SFT 数据集与 2,940 题的 PAW-Bench 学术写作基准。受控实验（同预算纯论文文本对照）证明增益来自「构造」而非论文内容：写作四项平均 54.34 对纯文本 52.12/无 CPT 51.90，推理不降反稳，长文档理解提升；且 4B 小生成器的语料最难拟合（loss 1.453）却下游最好——生成器越弱、数据越难拟合，收益越大。</description></item><item><title>领读Kimi K3技术报告：一个清华架构博士眼中的注意力谱系与「有效scaling」</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-26-kimi-k3-tech-report-architecture-lead-read/</link><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-26-kimi-k3-tech-report-architecture-lead-read/</guid><description>一集面向技术读者的Kimi K3技术报告领读播客，嘉宾孙宇涛（清华计算机系博士生、上海创智学院pre-doc，研究方向LLM架构与预训练）从K3出发串联起十多篇前作，把KDA线性注意力的每一项公式还原成RetNet→Mamba→DeltaNet→Gated DeltaNet的历史叠加，讲清channel-wise衰减、low-rank dk与BF16 tile的kernel co-design，MLA+QK-norm式门控的稳定性逻辑，Latent MoE对通信开销的削减，以及Quantile Balancing如何用线性规划一步求出负载均衡bias。预训练侧K3反潮流回归cosine decay、在混合注意力里用NoPE让长上下文免调参外推；后训练侧on-policy蒸馏成为多teacher多reward的「多模型合板」方案。嘉宾的暴论：大模型架构没有本质创新了，K3最核心的变量是size——2.8T总参、百B激活、K2的2.5倍scaling效率，而把size做work才是真创新。</description></item><item><title>ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-continualskillbench-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-continualskillbench-paper-reading/</guid><description>北大×BIGAI 提出 ContinualSkillBench，首次系统回答「Agent 技能库能否自主进化」：五个领域各 100 个按难度与技能依赖排序的关联子任务，三回合协议让 Codex CLI 与 Claude Code 在执行-反馈-反思中自建技能。15 组模型-领域设置中 14 组顺序执行提升归一化奖励（整体相对 +16.9%），但关键对照实验揭示：纯 ICL（不维护显式技能）平均 0.605 vs 显式技能 0.602，几乎无差异——顺序收益大部分来自保留上下文与反馈适应而非可复用技能抽象；且弱模型（GPT-4o）堆积 384 个碎片化技能，远多于强模型（GPT-5.3-Codex）的 205 个。</description></item><item><title>EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-evoharness-rl-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-evoharness-rl-paper-reading/</guid><description>深度精读 UIUC×Meta AI 合作论文 EvoHarness-RL（已被 LLA@COLM 2026 接收）：把长程智能体对外部 harness（记忆、工具、状态跟踪）的访问从提示词硬编码变成可学习的策略决策。通过 BPE 三态抽象（Belief/Progress/Experience）与四个元动作（track/commit/recall/note），配合教师轨迹 SFT 与代价感知 GRPO 两阶段训练，Qwen3-8B 在 ALFWorld 上从 ReAct 的 47.9% 跃升至 96.9%，逼近 Claude Opus 4.5；训练中还揭示“harness 退火”与“harness 演化”两个动力学过程。</description></item><item><title>LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-longwof-bench-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-longwof-bench-paper-reading/</guid><description>LongWoF-Bench（EvoMap × 清华大学，778 个机器可验证长工作流任务）回答了一个技能资产化的核心问题：什么样的&amp;rsquo;经验&amp;rsquo;才值得复用？对照实验给出干净答案——&amp;lsquo;验证器确认的执行经验&amp;rsquo;（Gene）在 7 个消费模型上稳定超越静态技能文档 8.7~15.5pp 且 token 更省；而没有经过验证器确认的&amp;rsquo;参考蒸馏&amp;rsquo;经验反而全面落后。经验的有效性来自&amp;rsquo;经过端到端验证的失败与修正信息&amp;rsquo;，而非表示形式。</description></item><item><title>MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-merchantbench-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-merchantbench-paper-reading/</guid><description>深度精读阿里巴巴×浙大×北大×复旦合作论文 MerchantBench——首个通过 365 天订单级电商仿真评测 LLM Agent 长期连贯性的基准。基于 1688 平台 98,843 条真实商品记录与 26 个工具，8 个主流大模型在 48 次全年运营中无一接近人类：最佳配置（Qwen3.7-Max + Hermes）最终净资产仅为人类参与者的 27.3%。论文提出操作连贯性与战略连贯性双维分析框架，揭示『活动衰减』与『战略漂移』两类渐进性失败，并提出 SWR（持续窗口率）这一可诊断『高分掩盖停摆』的过程指标。</description></item><item><title>Prime Agent: A Self-Improving RLM Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-prime-agent-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-prime-agent-paper-reading/</guid><description>Prime Agent（Prime Intellect × Princeton × MIT）用一个持久 IPython REPL + 递归子 Agent 的抽象，证明同一模型仅更换 harness 即可把 ARC-AGI-3 成绩从 30.2% 推到 95.5%、超过人类专家基线 95.4%。本精读拆解其两层核心抽象——Recursive Language Model（把上下文当变量、子 Agent 委派当函数调用）与 Continual Harness（把 harness 自身状态变成可 CRUD、可在线自我改进的数据），并解释为什么&amp;rsquo;harness 表达力&amp;rsquo;是被严重低估的能力放大器。</description></item><item><title>Repo2Skill-Evo: Repository Skills Go Stale in Silence 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-repo2skill-evo-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-repo2skill-evo-paper-reading/</guid><description>Repo2Skill-Evo（字节跳动 × 北京大学 × 北京交通大学）提出并评测了一个此前无人命名的问题：&amp;lsquo;仓库技能静默失效&amp;rsquo;——仓库版本升级后，从旧版蒸馏的 Agent 技能不报任何错、继续被加载检索，但内容已全面过时。基准要求 Agent 依据官方 release patch 维护技能集（删掉过时内容），用人工逐行验证的 12,217 行&amp;rsquo;黄金过时行集&amp;rsquo;做删除式指标：6 个前沿模型全部不及格，最强的 Claude-opus-4.6 也只有 69.7% F1，85/105 个版本转换低于 0.65 的 Easy 阈值。</description></item><item><title>FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-flowevo-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-flowevo-paper-reading/</guid><description>FlowEvo 提出一个免训练框架，让工作流与可执行技能在推理期共进化：它把验证通过的成功轨迹在线编译成带接口与回放测试的技能存入持久库，经直接执行、技能条件化生成与动态生成三路分层路由加以复用，并用对比效用机制抑制持续负迁移的技能。在 GPT-4o-mini 骨干上，FlowEvo 于 ALFWorld、HumanEval、MBPP、GSM8K、MATH-500 五个全量基准全面超越八个基线，ALFWorld 达 85.6%（超最强基线 26.4 点）且每任务 token 约为基线的三分之一，跨十种骨干模型 49/50 项对比胜出。</description></item><item><title>PRAXIS: Graph-Grounded Tacit Knowledge for Domain Code Generation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-praxis-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-praxis-paper-reading/</guid><description>为什么最强的编码 Agent 一进专业仓库就失灵？北京大学团队把根因锁定在「隐性知识」——那些只存在于开发者脑中、从不写进文档的业务规则、接口契约与操作约定。它们潜伏在开发实践之下、沿代码依赖图分散传播、且 Agent 根本不知道自己缺什么，三重性质让一切检索式方案天然失效。PRAXIS 给出四阶段闭环：让 Agent 在目标仓库里真实写代码暴露行为差异、蒸馏为带触发条件的结构化四元组、锚定到依赖图上双向传播与去重仲裁、并在任务初始化与工具交互时主动注入，支持在线演化。KoCo-Bench 四域平均 Pass@1 达 32.06%，较次优基线相对提升 16.7%，且随实践积累持续上涨。</description></item><item><title>MidTool: 面向Agent工具使用的中期训练数据合成 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-midtool-midtraining-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-midtool-midtraining-paper-reading/</guid><description>工具使用是 LLM Agent 的核心能力，但此前几乎全靠后训练习得。MidTool（华盛顿大学 + Snowflake + UNC，工作完成于 Snowflake 实习）提出首个面向通用工具使用的开放中期训练语料管线：从网页、PDF、代码、真实 API 与 MCP 技能四类源出发，经「上下文接地增广」与「原生 Agent 轨迹合成」两条分支构建 20.3B token 的 MidTool-Mix，中期训练 Qwen3-4B/8B-Base 后再统一 SFT+RL。在 BFCL、τ²-Bench、MCP Universe 三基准上，两种后训练配方下均一致超过 SFT-only 基线，RL 通常进一步放大增益，MCP-Universe 上 4B/8B 全面超过 Qwen3 官方同尺寸模型。本精读覆盖背景、定位、问题抽象、管线解法、实验证据、机制因果链、必要知识反推与可迁移灵感。</description></item><item><title>Inducing Task Models from Computer-Use Traces 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-task-model-induction-paper-reading/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-task-model-induction-paper-reading/</guid><description>计算机使用 agent 要真正进入真实工作，必须先搞清楚&amp;rsquo;这项工作实际是怎么做的&amp;rsquo;。Stanford 与 CMU 的这篇论文提出 TMI（TaskModelInduction），从被动录制的自然计算机使用轨迹中诱导结构化任务模型：先把多条交织的并发任务解缠成独立潜任务，再为每个任务构建&amp;rsquo;层次目标模型（做什么）+ 过程模型（怎么做）&amp;lsquo;双模型。在受控轨迹上任务分组与 ground-truth 一致性达 0.974，重建 74.9% 的观测执行步骤；由任务模型派生的技能使 held-out 任务准确率提升 30.0%。本精读覆盖其问题定义、双模型解法、内外双层评估设计、优势根源与可推广的通用性灵感。</description></item><item><title>MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-memtrapbench-paper-reading/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-memtrapbench-paper-reading/</guid><description>深度精读浙大 ZJUNLP 联合 NUS、东北大学、赫瑞瓦特大学与腾讯的 MemTrapBench——首个系统评估&amp;rsquo;记忆诱导认知陷阱&amp;rsquo;的基准。论文发现：忠实记录、语义相关的记忆仍可能扭曲模型推理与信念，1050 个对抗实例上所有记忆框架全面低于无记忆基线，最好的 EverMemOS 也落后 13.99 个百分点。文章拆解两类四情景陷阱分类、三段式对抗构建流水线、四组归因消融实验，以及仅靠推理时提示就挽回 14.9 个百分点的 AdaptiveMem 修复方案。</description></item><item><title>SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-skillforge-self-distilling-skills-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-skillforge-self-distilling-skills-paper-reading/</guid><description>上海交通大学顾晓东组提出SkillForge——面向项目特定issue解决的自蒸馏框架。核心洞察是冷启动问题：agent在特定仓库上缺乏项目知识，历史驱动方法依赖过往issue信号、在线方法每题付出昂贵探索成本。SkillForge反其道行之：主动重新实现仓库中带测试覆盖的核心功能来合成项目特定issue，解决后把经验蒸馏为实体锚定技能。SWE-bench Verified上DeepSeek-V3.2达72.2%（+5.8超基线，超最强对手+3.0），GPT-5-mini 60.6%（+5.6），SWE-bench Pro上同样领先，单issue成本仅$0.069-0.087。</description></item><item><title>Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-qcr-trajectory-reuse-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-qcr-trajectory-reuse-paper-reading/</guid><description>深度精读 arxiv:2608.12847——西安交大团队提出 QCR（Query-Conditioned Reuse），指出轨迹记忆的真正瓶颈不在检索而在检索之后的“复用”环节：历史轨迹里的用户名、路径、日期等绑定值会随时间过期，直接注入会诱导模型照抄旧值。QCR 在检索与执行之间插入一步改写，把选中轨迹转化为“工作流不变量/需重取绑定/适用条件/验证护栏”四字段笔记。在 WebArena/WorkArena/AppWorld 共 2391 个目标上，平均 Success 62.3%（比注入完整轨迹高 10.7 点），在线 token 省 48.9%；大绑定偏移下过期绑定错误率从 46.9% 降至 10.9%。</description></item><item><title>DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-dive-skill-evolution-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-dive-skill-evolution-paper-reading/</guid><description>深度精读佐治亚理工与 Cisco Research 合作的 DIVE 论文——让冻结大模型通过多样性驱动的技能进化实现自我改进。文章从冻结模型为何记不住经验讲起，拆解其核心设计：K=10 个独立技能种群、四种异构进化算子加 UCB 预算分配、联合选择至多 M=10 个互补技能、推理时候选排序。在六个数学与逻辑推理基准上，GPT-5-nano 借助 DIVE 从 52.3 跃升至 81.5，反超 GPT-5 few-shot，推理成本还降低 42.5%。本文逐节还原问题形式化、机制细节、完整实验数据与效果根源解释，并给出可迁移的方法论灵感。</description></item><item><title>SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-skillevo-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-skillevo-paper-reading/</guid><description>深度精读腾讯云 Andon 与浙江大学合作的 SkillEvo 论文。核心论断：技能自进化的瓶颈不在编辑能力也不在迭代次数，而在评估反馈能否持续供给可信的进化梯度。框架用两根支柱支撑这一命题——把多轮用户模拟从评估终点反转为反馈生成器（意图状态机、双侧正交评估、集体归因），再用独立治理层主动修复事实退化与结构膨胀（双锚点硬约束、图结构诊断软约束）。在腾讯云 6 类云服务、9 个生产 Skill、2000 张升级工单上，TSR 从 30.0 提升到 81.8，较单轮 QA 进化高 15.4 个点，膨胀率仅 2.8%，并已部署于生产环境。</description></item><item><title>DIVE: 多样性驱动的冻结模型技能进化 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-16-dive-diversity-skill-evolution-paper-reading/</link><pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-16-dive-diversity-skill-evolution-paper-reading/</guid><description>深度精读佐治亚理工与 Cisco Research 合作的 DIVE 论文——让只能通过 API 访问的冻结大模型，把任务经验进化成可持续复用的自然语言技能。从三大挑战（自修订噪声、经验超上下文、进化路径依赖）出发，拆解其核心机制：多种群独立进化维持假设多样性、异构算子组合 + UCB 自适应分配进化预算、算子本身也能进化、验证集上联合选择互补技能集。GPT-5-nano 借此平均 81.5 分反超 GPT-5 + ICL 且推理成本降低 42.5%，小模型逆袭大模型的路径首次如此清晰。</description></item><item><title>ERSkill: 检索技能与路由器共进化 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-16-erskill-retrieval-skill-evolution-paper-reading/</link><pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-16-erskill-retrieval-skill-evolution-paper-reading/</guid><description>Agent 记忆系统的进化大多发生在“写入侧”——怎么抽取、压缩、组织记忆。深圳国际工业与应用数学中心等机构的 ERSkill 把目光转向被忽视的“读取侧”：检索机制本身。它把检索行为表示为由固定原语（实体搜索/BM25/稠密检索 + 三种扩展 + LLM 处理）组合成的可执行技能，用训练好的 router 按查询的信息需求派发技能；进化时用经验 trie 记录所有探索过的原语路径以避免重复提议，用 Pareto 式双前沿把“能力探索”与“router 面向的部署”解耦。三大记忆基准上整体平均提升 31.3%（Qwen3-Next-80B-A3B 骨干），LongMemEval 零训练迁移仍居首。本精读逐部分拆解其机制，并建立“查询异构性→技能化→证据密集型任务受益”的因果链。</description></item><item><title>RippleMem: 从孤立检索到联想式回忆 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-16-ripplemem-associative-recollection-paper-reading/</link><pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-16-ripplemem-associative-recollection-paper-reading/</guid><description>问一个 Agent&amp;rsquo;该不该听 Sam 的推荐带 Maya 去 Harbor Grill 吃饭&amp;rsquo;，正确的回答需要三条分散在不同会话里的证据：晚餐计划、Maya 的海鲜过敏、这家店是海鲜餐厅——直接检索只命中第一条，无向图扩展可能带出无关的订座偏好却恰恰漏掉过敏这条安全约束。本文精读中国传媒大学联合智联英才科技的 RippleMem：它把记忆访问从&amp;rsquo;一次性查找&amp;rsquo;重构为&amp;rsquo;证据条件化的联想式回忆&amp;rsquo;——已召回的记忆不是检索的终点，而是寻找缺失支持的线索。系统把交互历史写成线索丰富的情景记忆单元，组织成事件中心图，查询时从初始锚点沿语义与结构双通道局部扩散，定向找回缺失证据。在 LoCoMo 上 F1 52.49、LLM 裁判准确率 87.14 均为最佳，temporal 类超 SimpleMem 9.66 分，multi-session 从 60.92 提到 78.20；建图成本约为 Mem0g/Zep 的 1/30。本精读重点拆解其写读两阶段设计，并用因果链解释&amp;rsquo;evidence-distributed 题型为何增益最大&amp;rsquo;与'30 倍成本下降为何来自延迟建图&amp;rsquo;。</description></item><item><title>AI for Science 爆发前夜：曹原谈验证瓶颈、概念抽象与 AGI 的最后一公里</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-yuancao-unreasonable-labs-ai-for-science/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-yuancao-unreasonable-labs-ai-for-science/</guid><description>Jeff Dean 带三位谷歌元老出走创立 Discovery Loop 当周，前 DeepMind 科学家、Unreasonable Labs 联合创始人曹原在硅谷101拆解 AI for Science 为何在此时爆发：代码与数学能力到位后科研成为下一个智能爆发点，但真正的瓶颈从建模移到了物理世界的验证环节；LLM 无法凭训练数据产生真正新的科学概念，概念抽象可能是 AGI 的&amp;quot;最后一公里&amp;quot;甚至不可计算；他主张 AI and Science 而非 AI for Science——科学难题是驱动 AI 本身进化的催化剂，并给出&amp;quot;AI 拿诺贝尔奖至少还需二三十年&amp;quot;的长期判断。</description></item><item><title>ComBodied Agents: a New Paradigm of Human-Centric Agentic AI 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-12-combodied-agents-paper-reading/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-12-combodied-agents-paper-reading/</guid><description>Bang Liu 团队 38 页范式论文，提出继 Digital Agents（数字状态）和 Embodied Agents（物理状态）之后的第三种 Agentic AI 行动基底——Combodied Agents（以人的演化状态为核心）。文章构建了一个以事件级多模态感知、可纠正纵向记忆、Personal World Models、可接受干预策略四模块组成的闭环框架，并把&amp;rsquo;保留并增强人类 agency&amp;rsquo;首次系统化为可评估的指标体系。本文从背景、定位、问题抽象、解法机制、评估体系、优势根源、必要知识反推、通用性灵感八个维度逐层精读。</description></item><item><title>SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-12-skillzip-paper-reading/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-12-skillzip-paper-reading/</guid><description>深度精读阿里+浙大+杜克联合提出的 SkillZip——首个无需任务回放（evaluation-free）的 Agent 技能压缩方法。它把&amp;rsquo;自进化积累的技能&amp;rsquo;视为一份带类型签名的契约，用&amp;rsquo;解释一次，引用多次&amp;rsquo;的直觉统一了规则共享、作用域提升、工作流复用与例外编码，形式化为一个带硬覆盖约束的类型化最小描述长度（MDL）目标。实验显示：平均压缩 31.2%，性能甚至略超未压缩技能，压缩速度比最强基线 SkillReducer 快 3.5 倍，且零次任务 rollout。</description></item><item><title>DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-dcas-scaffold-decoupling-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-dcas-scaffold-decoupling-paper-reading/</guid><description>深度精读华为加拿大软件卓越中心与女王大学的 DCAS 论文——首个系统揭示开源 CLI Agent 存在&amp;rsquo;scaffold 锁定&amp;rsquo;现象的工作。论文发现：在 OpenHands 单一 scaffold 下微调的模型，迁移到其他 scaffold 时性能可从 52.6% 暴跌至 8.4%。通过提出 DCAS 后端替换拦截层和区分显式/隐式规划两种形式，论文给出了一条从 scaffold 制品到模型能力的可行迁移路径，仅用 576 条规划感知轨迹即可让模型在非训练 scaffold 上一致提升。</description></item><item><title>SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-sft-rl-multitask-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-sft-rl-multitask-paper-reading/</guid><description>当大语言模型需要同时掌握数学、代码、科学、逻辑多种推理能力时，SFT（监督微调）和 RL（强化学习）会表现出截然相反的行为：SFT 在多阶段训练中因梯度方向冲突而性能崩溃，RL 却因为「优势归一化 + on-policy 采样」产生的近似正交更新而稳定共存。本文通过参数级几何分析和高维浓度不等式，首次从理论上揭示了「SFT 干扰是范数受限的、RL 干扰是方差受限的」这一本质差异，并提出 Parallel-RL 范式——各任务独立 RL 后合并参数，在 DeepSeek-R1-Distill-Qwen-1.5B 上实现 ΔBase +10.7%、Retention 103.2%。本精读将从零讲清 SFT/RL/GRPO 的机制差异，建立「方法差异→参数更新几何→理论边界→指标提升」的完整因果链。</description></item><item><title>SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-skillprox-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-skillprox-paper-reading/</guid><description>港科大的 SkillProx 把 LLM Agent 的「技能自进化」重新拆解为「近端梯度下降」的前向-后向两阶段：前向用闭环重执行拦截退化的诊断补丁，后向用冻结的留一效用审计配合验证门控选择性整合/降级/删除知识单元。相比最强梯度基线 SkillGrad 平均提升 3.0pp，且消融清晰地揭示了「闭环诊断 -1.5、近端收缩 -2.5」的因果分工。本精读以九部分结构，详解这套自进化框架的方法机制、实验证据、效果根源与可迁移灵感。</description></item><item><title>TEPA: Revoking Stale Memories for Conflict-Robust Language Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-tepa-memory-revocation-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-tepa-memory-revocation-paper-reading/</guid><description>深度精读 TEPA 论文——首次将 Agent 长期记忆的&amp;rsquo;记忆污染&amp;rsquo;形式化为可证伪性问题，提出可撤销的证据-记忆机制，让有效性成为记忆的显式状态。在完全反转场景下，append-only 跌至 0.210（甚至低于无记忆基线 0.309），而 TEPA 保持 0.950。真实文件执行场景同样再现这一模式，MemoryAgentBench SH-6k 上匹配强 last-write-wins 缓存（0.890）。边界测试揭示多跳和超长上下文是下一阶段架构挑战。</description></item><item><title>特修斯之船：Kimi K3如何把Transformer的零件全部换掉，还能逼近前沿</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-06-kimi-k3-tech-report-deep-dive/</link><pubDate>Thu, 06 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-06-kimi-k3-tech-report-deep-dive/</guid><description>月之暗面K3是首个达到3T级别的开放权重模型，其47页技术报告揭示了一种&amp;quot;特修斯之船&amp;quot;式的架构哲学：注意力改成了线性+全局混合（KDA+MLA），残差变成了深度方向的Attention，FFN变成压缩空间的稀疏专家，甚至连位置编码都几乎被删掉。RadixArc创始成员赵晨阳和华盛顿大学博士生曾志远分别从Infer和算法两条线拆解K3：线性注意力在2.8T规模上实现了6.3倍解码加速，Quantile Balancing路由是3T稳定训练的关键之一，MOPD让九个领域专家模型高效合板，而KDA（Kernel Development Agent）证明RSI已在kernel优化领域高速运转。核心判断：权重只是一次训练的产物，环境才是能反复产出下一代权重的护城河。</description></item><item><title>TARL：面向长期Agent的可执行记忆管理精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-05-tarl-paper-reading/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-05-tarl-paper-reading/</guid><description>长期Agent的持久记忆里，一次错误的更新会像多米诺骨牌一样反复扭曲未来的检索与推理。现有系统把记忆更新简化为二元Write/Hold决策，无法区分&amp;rsquo;新增/忽略/修订/拒绝/延迟验证&amp;rsquo;这五种本质不同的处置。TARL把每条语句映射到五种可执行操作，通过Accepted/Pending/History三账本管理记忆生命周期，并用反事实执行监督——在训练时执行所有候选动作、比较产生的记忆状态质量——来训练模型选择导致正确结果的操作。5-way Macro F1 0.8286、Next State Accuracy 0.6621、Memory Pollution Rate改善10.1%，且完美二元标签仅能恢复28.6%的状态、五动作Oracle可达100%——这篇论文从机制因果上证明了为什么二元监督从根本上不足。</description></item><item><title>Recursive Harness Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-23-rhi-recursive-harness-self-improvement-paper-reading/</link><pubDate>Thu, 23 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-23-rhi-recursive-harness-self-improvement-paper-reading/</guid><description>Sakana AI 与 UC Berkeley 提出 RHI（递归式框架自改进）：把多智能体框架当作提示词级对象，仅用当前与上一版本的自我比较来迭代优化，少数几轮就能让低推理强度的 Agent 超越同族最高推理强度设置，同时把推理成本降低最高 60%。本文从 Harness 是什么、模型-框架协同进化讲起，拆解 RHI 的轨迹局部目标、算法流程、信息论隐式目标，并提炼可推广的通用性灵感。</description></item><item><title>AI时代什么值得学？知识、代码都贬值了，经验和技能才是硬通货</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-22-waic2026-ai-learning/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-22-waic2026-ai-learning/</guid><description>WAIC 2026期间，科大讯飞AI大学堂发布AI热点和AI Vault两大新功能。围绕&amp;quot;AI时代什么值得学&amp;quot;，极客时间业务负责人王一鹏、科大讯飞开放平台总经理李佳琪、野生AI Hacker许恒在围炉夜话中展开了深度讨论：AI的角色正从&amp;quot;知识百科&amp;quot;转向&amp;quot;私人教练&amp;quot;和&amp;quot;军师谋士&amp;quot;，个人知识库被AI记忆系统取代，系统化学习回归经典课程，人才供需的gap在持续拉大——这个时代真正奖励的是热爱、执行力和深度判断力。</description></item><item><title>2026-06 arXiv 智能体记忆系统（Agent Memory）领域综述：113 篇全文精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-17-agent-memory-survey-2026-06/</link><pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-17-agent-memory-survey-2026-06/</guid><description>智能体记忆子领域纵深综述。从 11930 篇 6 月 arXiv 预印本中筛定 113 篇核心集，全部下载 PDF 抽全文逐篇精读，提炼 7 大共识性问题（检索不等于使用、固化的保留与遗忘决策、上下文成本爆炸、一致性/矛盾解决、评测混淆变量、记忆即新攻击面、遗忘治理）与多项原创问题定义，关键新框架已联网交叉验证。</description></item><item><title>智能体技能演化（Skill Evolution 与 Self-Evolving Agents）综述：53 篇核心论文精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-17-skill-evolution-survey-2026-06/</link><pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-17-skill-evolution-survey-2026-06/</guid><description>技能演化与自演化智能体综述。从 116 篇候选中筛定 53 篇 CORE 论文下载全文精读，提炼技能库选择退化、技能创建与部署脱节、自演化缺乏可靠接受准则、上下文无界膨胀等共性问题，以及 16 个范式级新转变（PACE、Bayesian-Agent、Red Queen Godel、Trellis、MMG2Skill 等 5 个已联网验证）。</description></item><item><title>Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-13-incomplete-learning-sft-paper-reading/</link><pubDate>Mon, 13 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-13-incomplete-learning-sft-paper-reading/</guid><description>这篇 ACL 2026 Main 论文首次系统性地揭示了「不完全学习现象」（ILP）：即使训练损失收敛，LLM 仍有约15%的训练样本无法被正确复现。论文将这一现象归因为五个可诊断的来源，并提出了一个「先诊断、再对症下药」的框架，证明了SFT失败的异质性——不同原因需要不同解法。</description></item><item><title>Harness Engineering for Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-07-harness-engineering-paper-reading/</link><pubDate>Tue, 07 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-07-harness-engineering-paper-reading/</guid><description>Lilian Weng（Thinking Machines Lab 联合创始人、前 OpenAI 研究副总裁）在这篇万字综述中系统梳理了「Harness 工程」——围绕基础模型的运行时系统——作为通往递归自我改进（RSI）现实路径的核心命题。文章从 RSI 的思想起源讲起，把 Harness 定义为决定模型如何思考、规划、调用工具、管理上下文、评估结果的系统层，并梳理了三大设计模式（工作流自动化、文件系统持久记忆、子代理并行）、四大优化方向（上下文工程、工作流设计、自我改进、进化搜索）以及与模型权重的联合优化，最后坦诚列出七大瓶颈。本精读将这篇综述放在 RSI→Harness 的研究脉络中定位，提炼其方法论骨架与可迁移的普适灵感。</description></item><item><title>SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-29-skill-disco-paper-reading/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-29-skill-disco-paper-reading/</guid><description>深度精读微软研究院与北京外国语大学合著的 SKILL-DISCO 论文——将 Agent 成功执行轨迹蒸馏为可重用的参数化控制流子图（PFSM），再编译为可调用、可执行、可验证的过程技能。在 ALFWorld 和 WebArena 上，仅用 5 个技能（对比 ASI 的 110 个）就将成功率推高至 99.3%，且技能可跨模型迁移——GPT-4o 归纳的技能让 Qwen3.5-9B 在 ALFWorld 上达到 98.5%。</description></item><item><title>Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-15-language-models-need-sleep-paper-reading/</link><pubDate>Mon, 15 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-15-language-models-need-sleep-paper-reading/</guid><description>深度精读马里兰大学与卡内基梅隆大学合著的「语言模型需要睡眠吗？」论文——受人类睡眠记忆巩固机制启发，提出离线循环（Offline Recurrence）机制：模型在&amp;rsquo;睡眠&amp;rsquo;阶段对累积上下文执行 N 轮离线反复遍历，将信息蒸馏为持久化快速权重（Fast Weights），然后清空 KV 缓存。在不增加在线推理延迟的前提下，成功解决常规 Transformer 和 SSM-Attention 混合模型均失败的多跳推理和数学推理任务。增加睡眠轮数 N 可以持续提升性能，且推理越深的样本获益越大（相关系数 r=0.92）。</description></item><item><title>Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-rho-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-rho-paper-reading/</guid><description>深度精读香港城市大学 × 微软亚洲研究院 RHO 论文——首个仅利用无标签历史轨迹实现 Agent Harness 全链路自监督优化的工作。从 Harness 工程概念补全、六项前序工作定位、自偏好估计的问题抽象、三阶段核心方法详解、必要知识反推到七条通用性灵感，全面拆解这项在 SWE-Bench Pro 上将通过率从 59% 提升至 78% 的开创性研究。</description></item><item><title>Role-Agent: 通过双角色自举实现 LLM 智能体-环境协同进化 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-role-agent-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-role-agent-paper-reading/</guid><description>深度精读中科大 × 阿里AMAP团队 Role-Agent 论文——首个利用单一 LLM 同时扮演智能体与环境双角色，实现自举式协同进化的工作。从 Agent 强化学习背景补全、智能体-环境协同优化的研究脉络定位、双角色自举的问题抽象、WIA（世界内化于智能体）+ AIW（智能体内化于世界）双模块方法详解、必要知识反推到六条通用性灵感，全面拆解这项在 ALFWorld、WebShop、搜索增强QA 三大场景平均提升超 4% 的开创性研究。</description></item><item><title>Self-Harness: Harnesses That Improve Themselves 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-self-harness-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-self-harness-paper-reading/</guid><description>深度精读上海人工智能实验室 Self-Harness 论文——首个让 LLM Agent 自主改进自身操作套件（Harness）的范式。从 Harness 概念补全、三大范式对比定位、三阶段闭环机制详解、模型特异性验证到通用性灵感提取，全面拆解这项在 Terminal-Bench-2.0 上取得高达 21.4% 绝对提升的开创性研究。</description></item><item><title>HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-09-harnessforge-paper-reading/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-09-harnessforge-paper-reading/</guid><description>深度精读北京航空航天大学与清华大学合著的 HarnessForge 论文——一个元自适应框架，将 LLM Agent 系统形式化为 harness-policy 对，通过故障引导的 harness 裁剪和 harness 条件化的策略对齐实现协同演化。在 5 个跨领域基准上超越所有 harness-only 和 policy-only 基线，最高增益达 12.0%，揭示了一个关键洞察：harness 和 policy 之间的可执行兼容性是 Agent 系统适应的核心。</description></item><item><title>Rethinking Continual Experience Internalization for Self-Evolving LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-09-rethinking-continual-experience-internalization-paper-reading/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-09-rethinking-continual-experience-internalization-paper-reading/</guid><description>深度精读中国人民大学高瓴人工智能学院与美团联合发表的 Rethinking Continual Experience Internalization 论文——系统性地揭示了 LLM 智能体在多轮经验内化中出现的&amp;rsquo;能力崩塌&amp;rsquo;现象，从经验粒度、注入模式和内化机制三个维度诊断根因，提出&amp;rsquo;原则级经验 + 逐步注入 + Off-policy 蒸馏&amp;rsquo;的稳定自进化配方，使模型在连续迭代中实现可持续的性能提升而非渐进退化。</description></item><item><title>From Context to Skills: Can Language Models Learn from Context Skillfully? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-06-ctx2skill-paper-reading/</link><pubDate>Sat, 06 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-06-ctx2skill-paper-reading/</guid><description>深度精读清华大学、DeepLang AI、UIUC 等机构联合发表的 Ctx2Skill 论文——一个无需人工标注和外部反馈的自进化技能发现框架。通过多智能体自博弈循环让 Challenger 和 Reasoner 共同进化技能集，配合 Cross-Time Replay 机制防止对抗性坍塌，在 CL-bench 四个上下文学习任务上跨模型一致提升解决率。</description></item><item><title>SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-03-skilladaptor-paper-reading/</link><pubDate>Wed, 03 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-03-skilladaptor-paper-reading/</guid><description>深度精读浙江大学 ZJUNLP 团队 SkillAdaptor 论文——一个免训练的步骤级技能自适应框架。从 Skill/Harness 概念补全、步骤级归因 vs 轨迹级反思的核心区别、三阶段适应流水线（归因-修改-资格验证）、必要知识反推到通用性灵感提取，全面拆解这项在 WebShop/PinchBench/Claw-Eval 三个基准上均优于基线的创新工作。</description></item><item><title>Yunjue Agent: 零起点原位自进化智能体系统精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-02-yunjue-agent-paper-reading/</link><pubDate>Tue, 02 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-02-yunjue-agent-paper-reading/</guid><description>深度精读 Yunjue Agent 技术报告——首个完全可复现的零起点原位自进化 Agent 系统。从自进化 Agent 三大支柱（工具/上下文/工作流）背景补全、四类自进化方法定位、工具进化为关键路径的问题抽象、多 Agent 协作+并行批进化+进化泛化损失指标详解、必要知识反推到通用性灵感，全面拆解这项在5个基准上零起点超越专有系统的开创性工作。</description></item><item><title>δ-mem: Efficient Online Memory for Large Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-02-delta-mem-paper-reading/</link><pubDate>Tue, 02 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-02-delta-mem-paper-reading/</guid><description>深度精读南洋理工大学 DeCLaRe Lab 的 δ-mem 论文——在冻结的大语言模型上添加仅 8×8 的在线联想记忆状态矩阵，通过 delta-rule 学习实现高效动态记忆。从 LLM 记忆困境背景补全、三大记忆路线定位、核心问题抽象、四步计算流程与三种写入策略详解、必要知识反推到六条通用性灵感，全面拆解这项在 MemoryAgentBench 上取得 1.31× 提升的轻量级记忆工作。</description></item></channel></rss>