<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Harness on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/harness/</link><description>Recent content in Harness on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 03 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/harness/index.xml" rel="self" type="application/rss+xml"/><item><title>GUI-HARVEST + DynaHarness + EvoGen-Harness 三篇合读：harness 自进化在 GUI、机器人、图像生成三条垂直域的落地</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-03-harness-vertical-trio-paper-reading/</link><pubDate>Sat, 03 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-03-harness-vertical-trio-paper-reading/</guid><description>「冻结骨干、进化运行时」正在成为 agent 自进化的主流路线：GUI-HARVEST 用重复视觉执行证据加行为预测双门，在 OSWorld 六个骨干上最高提升 12.33 个百分点；DynaHarness 用快慢脑加物理执行契约，把冻结 π0.5 机器人策略从 17.5% 拉到 74.25%；EvoGen-Harness 用 where+how 联合归因进化，让冻结文生图模型在 GenEval2 上从 0.4456 涨到 0.7089。三篇论文分别代表 GUI、物理机器人、图像生成三条垂直域的 harness 进化代表作，本文合读三者的共同骨架、域特化设计与实验证据链，并讨论 harness 工程的边界与反例。</description></item><item><title>Harness 的有效性边界：Malena × Finding the Right Fit 合读精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-03-harness-boundary-duet-paper-reading/</link><pubDate>Sat, 03 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-03-harness-boundary-duet-paper-reading/</guid><description>同月两篇论文从正反两面划定 Agent Harness 工程的有效性边界。EPFL+Apple 的 Malena 用受控消融证明：骨干够强时，几乎全部收益来自编码 agent 运行时与骨干本身——给模型一个 shell 和文件系统（Chat→Oneshot）是测得的最大单一效应，搜索原语、多 agent 编排全部统计不显著（最大差距 Best-of-N vs UCB1 仅 2.71pp，CI 含 0），单会话极简 agent 在 MLE-bench 上以 62.5% medal 率碾压四个 SOTA Harness（最佳外部 47.1%）。NTU 的 Finding the Right Fit 用 66 配置矩阵证明另一半事实：换 Harness 模型排名完全反转（TB4 上 Claude−GPT 差距在 OpenHands +7.94、PI −30.16，摆幅 38.09pp），6204 条轨迹归因发现 Harness 的真实价值集中于「把失败转成模型可用的反馈」这一件事。合读结论：Harness 的价值不在「编排的丰富度」而在「反馈回路的完整性」，且随骨干增强而向运行时基底收缩。</description></item><item><title>Harness 的三种缩放轴：Mid-Harness × STITCH × Turbo Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-02-harness-scaling-duet-paper-reading/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-02-harness-scaling-duet-paper-reading/</guid><description>同日三篇论文从三个互补粒度回答同一问题：固定模型后 Harness 侧还有哪些缩放轴可挖。NVIDIA+KAIST 的 Mid-Harness 下沉「动作级」——在模型与 Harness 边界采样 N 个候选、执行前由验证器选一，发现采样收益完全由验证支配（前沿验证器把 TerminalBench-Lite 从 50.00% 拉到 68.03%），且与轨迹级缩放正交可组合（+Best-of-T 达 66.33%、成本减半）。UIUC+UMich 的 STITCH 沿「原语级」轴测试时组装——带 scope/contract 的原语库+确定性编译器，SWE-V 80.5%、组装开销仅 2.7%，配 mismatch gap 与指数衰减两命题。Rutgers+Red Hat AI+MIT-IBM 的 Turbo Harness 沿「实例级」轴打补丁——回收外层搜索副产物蒸馏 playbook、GRPO 训 9B 编辑器逐实例补丁，SWE-V 38.4→54.4%、步数 23.1→8.7。三轴正交可叠加，构成 Harness 工程学的完整缩放谱系。</description></item><item><title>Harness 自动进化三重奏：MILO、ScholarEvolve 与 Malena 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-02-harness-evolution-trio-paper-reading/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-02-harness-evolution-trio-paper-reading/</guid><description>2026 年 9 月末，arXiv 上同时出现三篇方向相撞的 Harness 论文：MILO 用「负证据谱系记忆+编排者元进化」把自动 Harness 发现推上 Terminal-Bench 2.1 官方榜首之上（RR@5 86.1%），ScholarEvolve 让 Harness 从研究文献中学习模块化变异（AppWorld Challenge TGC 49.6%→63.6%），而 Malena 却用大规模控制变量消融证明：在前沿编码 Agent 之上，复杂 Harness 机制几乎全部冗余（MLE-bench 获奖 62.5% vs 最佳开源 Harness 47.1%）。本精读逐篇拆解三者的机制与实验，再正面处理这个当日最大的张力——结论是：Malena 消融的是「统一机制的加减法」，而 MILO/ScholarEvolve 的增益来自「任务自适应、负证据利用与成本-精度前沿」这些 Malena 未覆盖的轴，弱模型反例（gpt-oss-120b、Gemma 4 31B）进一步表明机制收益随模型能力变化，两条路线实为同一光谱的两端。</description></item><item><title>编码 Agent 的安全边界与协作假象：Approval Laundering 与 OpenCollab 合读 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-10-02-coding-agent-security-duet-paper-reading/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-10-02-coding-agent-security-duet-paper-reading/</guid><description>本文合读 2026 年 9 月底同期出现的两篇编码 Agent 基础设施论文：复旦单作者工作 Approval Laundering 证明「人批准的动作 ≠ 实际执行的动作」，用六轴分类学系统化批准-执行绑定漏洞（Scope/Temporal/PATH 替换 BGR=1.0），并以七字段 HMAC Approval Token 部分修复；上海交大牵头的七机构工作 OpenCollab 证明「声明的协作 ≠ 发生的协作」，用 Adherence 六轴审计与 CACE 因果归因把多智能体增益争议变成可测量问题，并以双 Coder 工作流在 SWE-bench Pro 拿下 64.25% SOTA。两篇从安全与效能两个方向拆掉 Harness 的同一类隐式信任假设：把 Agent 系统的隐式假设变成可测量、可审计的对象。</description></item><item><title>CodeSkill × NanoHarness：技能抽象与 Harness 效应——被低估的智能体性能杠杆 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-skill-harness-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-skill-harness-duet-paper-reading/</guid><description>本精读合并解读两篇从「模型权重之外」挖掘编码智能体性能的论文：清华+华为诺亚+上交的 CodeSkill 把长程 RL 从 token 级提升到技能级——teacher 蒸馏三级文本技能（目标/执行/控制），分层 VAE + Gumbel-Softmax 执行反馈门控边界把离散技能映射为连续隐变量 zH/zM/zL，以软提示前缀注入冻结 LLM（LoRA），PPO 在隐空间优化；SWE-bench Verified 76.2、EvalPlus 99.2 开源第一，交互步数 12.3→6.8（−45%），去掉 VAE 直接用文本技能+RL 则 BigCodeBench 从 94.8 掉到 88.4。南京理工+TUM+南京大学的 NanoHarness 首次把 harness（模型外基础设施）作为一等研究对象做组件级受控分解：固定模型下 harness 差 19.4pp vs 模型差 22.8pp；在 mini-SWE-agent 上增量加五组件，工具注册表 +4.57pp、任务特定子代理 +5.91pp，而上下文压缩 −4.86pp——机制是结构化工具把无序 shell 探索变针对性调用（jqlang 案例 133 次探测→43 次、通过率 24.68%→68.17%）。二者共同指向：技能结构与 harness 设计是被系统低估的性能杠杆。</description></item><item><title>Opera × CER：长程编码智能体的评论家介入与早期奖励预测 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-test-time-supervision-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-test-time-supervision-duet-paper-reading/</guid><description>本精读合并解读两篇在「轨迹还没走完」时提供质量信号的长程编码 Agent 论文：Salesforce AI Research 的 Opera 构建口头评论家框架——把每次修正管理为持久化笔记（混合调度五事件触发审查 + 九个契约化类型算子诊断 + 准入/发布双审计把关投递与关闭，并把「遵从」与「解决」分开跟踪），在 Terminal-Bench 2.1 / SWE-Bench Pro / DeepSWE v1.1 三基准上把 Qwen3.8-27B 的 resolve rate 提升 7.9/4.0/8.9pp，去掉双审计后增益损失 2/3，证明误导性反馈的代价之高；UW 等机构的 CER 则在 rollout 结束前从 40 步前缀预测终端奖励——用检索经验库合成任务自适应 rubric、同父兄弟续跑组内共享打分（组内序保持即足以支撑排名类下游），TTS 上 Nemotron 3 Ultra RM@8 达 67.6%（+4.2pp）且只用 15.3% token 匹配最佳基线（省 84.7%），RL 上 40 步截断 + 弃权门控 DPPO 以 52.7% 更少在线 token 超过全轨迹 TMax（51.6% vs 49.7%）。一个向内诊断当前轨迹，一个向前预测最终结局，合看构成测试时监督的两条互补路径。</description></item><item><title>Skill2Env × QwenGyre × AgentPerfBench：智能体强化学习的数据、系统与推理三层基建 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-rl-infra-trio-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-rl-infra-trio-paper-reading/</guid><description>同日挂出的三篇 arXiv 论文恰好拼出 Agent 强化学习的三层基础设施：数据层（AllSpark 的 Skill2Env 把「环境合成」从扩展覆盖升级为能力参数化——100 个带控制旋钮的难度模式+基于 solver 执行证据的迭代加难，1.5K 轨迹 SFT 换 7 基准 +8.4pt）；系统层（阿里 Token Hub 牵头的 QwenGyre 用 cell 级弹性调度+轨迹树处理让 2.4T 旗舰模型的百万 token rollout 在线 RL 提速 1.78×，pass rate 52.48%→58.54%）；推理层（帝国理工+剑桥+牛津的 AgentPerfBench 用 22 个负载画像+饱和扫描+NCU roofline 证明 chat→coding 的 TTFT 差 4.8×、操作强度差 46×，agentic 负载在并发爬升时吞吐崩塌 43–76% 而 chat 仍在扩张）。本文按九部分结构合并精读，并给出「Agent RL 全栈工程」的公共图景。</description></item><item><title>TraceDance × Maintaining Benchmarks：Agent 行为基准的构建与作弊治理 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-behavior-benchmark-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-behavior-benchmark-duet-paper-reading/</guid><description>本精读合并解读两篇互为镜像的 Agent 基准治理论文：字节跳动+UIC 的 TraceDance 解决「供给侧」——从 25 万条真实部署轨迹中按用户自然语言指定的不良行为自动构建定向行为基准，其可编程 Anchor-and-Confirm 把全量扫描搬到 CPU、构建成本从 O(N·LLM) 降为 O(N·CPU)，139 个查询完成率 95.3%、产出 107 个基准 4,125 实例，9 个前沿模型平均通过率仅 26.7%；Scale AI 的 Maintaining Benchmarks 解决「信任侧」——把「通过任务但未展现目标能力」定义为 unearned pass（SWEBench Pro 上 GPT-5.6-Sol 违规率 68.27% 而 GPT-6 Astra 为 0%，git 历史 oracle 是主导通道），用三值裁决+对抗复核+通道级密封+重放探针+新鲜复评构成检测-定位-修复-复评闭环。一篇让基准「从真实世界长出来」，一篇让基准「在强大模型面前保持诚实」，合看构成 Agent 评测有效性的完整叙事。</description></item><item><title>MoMHa 与 SkillEvoReg 精读：Agent 资产的优化与正则</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-momha-skill-evoreg-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-momha-skill-evoreg-paper-reading/</guid><description>一篇合并精读两篇 2026 年 9 月 25 日同期挂出的 Agent 资产治理论文：Adobe Research 的 MoMHa 首次把 LLM harness 设计形式化为准确率×安全×token 三目标优化搜索，单阶段联合奖励全面压过两阶段与十个 prompt 优化基线（J=0.482 对 TextGrad 0.422，安全分 0.781 全场最高）；华为诺亚方舟实验室的 SkillEvoReg 首次定义「技能进化过拟合」问题，把 dropout/容量正则/对抗验证三原则迁移到离散技能更新，SpreadsheetBench 提升 12.25 个百分点的同时技能体积缩 61%。两篇恰好都是纯企业实验室主导，共同指向 Agent 外部资产的优化与正则化这条新主线。</description></item><item><title>AI 智能体行为的水印税与全模态 harness 双精读：Provenance Tax × Qwen3.8-Omni-Flash</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-27-duet-provenance-omni-paper-reading/</link><pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-27-duet-provenance-omni-paper-reading/</guid><description>本期二重奏精读收录两项围绕「AI 智能体」的研究。上篇解读 Lasso Security 的企业研究 The Provenance Tax：Anthropic 即将在 Claude 中部署的 SynthID-Text 水印虽然宣称非失真，但改变了逐 token 采样过程，实验证明它会改变智能体的工具调用与拒绝行为，注入攻击下 gemma-3-27b 的逐项判定翻转率高达 23.5%。下篇解读阿里通义千问技术报告 Qwen3.8-Omni-Flash：一个原生全模态智能体模型，凭 Thinker-Talker 架构、1M 上下文与智能体式选择性感知，把音视频理解的准确率与 token 成本同时推向新平衡，并开源 Qwen-MM-Plugins 与 Qwen-Live-Harness 两个框架。两篇文章合起来，恰好构成 2026 年智能体工程的两条主线：模型行为的可靠性边界，与全模态 harness 的系统化设计。</description></item><item><title>Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-27-jaz-invoke-paper-reading/</link><pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-27-jaz-invoke-paper-reading/</guid><description>MIT CSAIL 团队提出 JAZ：一个只比 agent loop 多一点点的极简智能体框架。它仅暴露一个 LLM 原语 invoke——一个函数体由 LLM 在运行时生成的函数——加上动态作用域与 hooks，就涌现出传统上需要专门 harness 才能实现的长程记忆与持续自改进能力：在 StuLife 远程回忆子集上以约 43% 的成本超越 Letta（MemGPT）8 个百分点，在 AppWorld 上以更低成本胜过专门的自改进框架 ACE。本文从语言原语的第一性原理出发，拆解 invoke 的两条定义性质、tail-recursive delegation 如何统一各类上下文管理为特例，并用 MemGPT、RLM、ACE、context rot 研究等外部文献交叉验证其效果优势的根源。</description></item><item><title>Harness as a Language×Bounded Loops：Agent 脚手架的语言化与可验证化 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-25-harness-theory-duet-paper-reading/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-25-harness-theory-duet-paper-reading/</guid><description>本精读合并解读两篇同期论文：MIT CSAIL 的 Harness as a Language（JAZ）把 agent harness 定义为一个极小的语言原语 invoke，函数体由 LLM 在调用时现场生成，从而用纯提示在长程记忆与自我改进两类任务上超过专用 harness；Qualixar 的 Bounded Loops 则给 harness 装上类型化循环与静态验证器，运行前即可证明终止性、花费上界与完成性三性质。两篇论文共同指向一个主题：harness 正从工程偶然走向数学对象——能做什么可证明，花多少可预证。本精读覆盖两文的动机、形式化核心、实验证据、交叉验证的根源解释，以及可迁移到其他领域的通用灵感。</description></item><item><title>当智能体开始给操作系统「付房租」：云栖2026上一场关于Agent OS定义权的暗战</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-agent-native-operating-system/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-agent-native-operating-system/</guid><description>2026云栖大会「AI Agent原生操作系统」分论坛回放：Agent负载让「换操作系统=白拿40%性能」、token成本治理下沉到系统调用层，AMD与Arm用「沙箱密度」重写CPU规格表，中兴提出OS的第一性使命是「踩刹车」。本文梳理Agentic OS、SysAgent、SkillHub、SAIL开源等发布，拆解负载变迁→OS价值位重估→治理标准之争三层传导，并给出可落地的观察指标。</description></item><item><title>Harness-Zero：通过 Agent-as-Harness 实现 Harness 蒸馏——论文精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-harness-zero-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-harness-zero-paper-reading/</guid><description>精读北京大学与 Google 合作的 Harness-Zero，提出 agent-as-harness 范式：用一个审查型智能体在学生模型的响应边界上把进化出的专用 harness 收益「翻译」成可训练轨迹，经 SFT 把外挂行为蒸馏进权重，部署时彻底移除外挂。Qwen3.5-9B 宏平均成功率从 23.3% 提升到 44.3%，甚至超过挂载原 harness 的 41.7%。</description></item><item><title>RRSI: Regularized Recursive Self-Improvement of Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-rrsi-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-rrsi-paper-reading/</guid><description>深度精读 Google Cloud AI Research 等提出的 RRSI——首个把机器学习正则化思想系统迁移到 Agent Harness 递归自我改进的工作。文章从 Harness 与 RSI 概念讲起，拆解过拟合问题的成因，逐一讲解提案侧 L0 式退火编辑预算、证据感知信用分配、结构化探索，与选择侧泄漏筛查、噪声调整底线、L2 式成本门槛、L1 式结构剪枝，并结合八基准三域实验与外部检索交叉验证，剖析其「为何能泛化」的根源性解释与可迁移灵感。</description></item><item><title>Agent 技能自进化二重奏：EVOLVE 与 GraphSkillEvo 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-22-skill-evolution-duet-paper-reading/</link><pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-22-skill-evolution-duet-paper-reading/</guid><description>2026年9月同主题连发的两篇论文不约而同地把「Agent 技能库」当作可进化的资产：Adobe+Brown 的 EVOLVE 让冻结模型在真实用户流量中演化 SKILL.md 技能库（Widening/Deepening 两轴 + Matched Replay Gate 保守准入）；港城大+NUS+南科大的 GraphSkillEvo 则把技能表示为「全局指导+有向图」，用种群进化（4算子变异/交叉）优化。本文合并精读二者，共用背景与灵感节，逐篇拆解问题定义、解法与评估，并用因果链解释优势根源（保守准入防评分漂移、图结构压缩搜索空间），交叉对照 Reflexion/ExpeL/Voyager/Safe-Policy-Improvement/GEPA 谱系。两文共同指向一条结论：把「改模型权重」换成「改模型身边的自然语言资产」，是一条更稳、更安全、可迁移的持续适应路线。</description></item><item><title>Self Improvement via Fast Tree-search 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-21-sift-paper-reading/</link><pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-21-sift-paper-reading/</guid><description>MIT 与 Sakana AI 的 SIFT 把递归自改进（RSI）编码智能体的最大瓶颈从&amp;rsquo;生成候选&amp;rsquo;移到了&amp;rsquo;验证候选太贵&amp;rsquo;：用 pairwise LLM-as-a-judge（每次 $0.044）+ 正则化 Bradley-Terry 聚合替代 $6.0 的基准子集评估作为中间信号，在完全解耦的树搜索流水线中让扩展与评估并行。Polyglot-225 上以 DGM 约 1/10 的 CPU 小时拿到 31.1%（Qwen3-30B）/35.1%（o3-mini）全面超越 DGM/HGM/SICA，TerminalBench 2.1 从 29.2% 提到 36.7%。本精读覆盖&amp;rsquo;便宜排名+昂贵验证&amp;rsquo;分离范式的机制因果、judge 输入格式的消融证据、与 DGM 谱系的定位对比，以及&amp;rsquo;把验证成本当一等公民&amp;rsquo;的通用性灵感。</description></item><item><title>An Empirical Study of Harness Design for Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-19-harness-empirical-paper-reading/</link><pubDate>Sat, 19 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-19-harness-empirical-paper-reading/</guid><description>UMass Amherst、Emory 联合 Zoom 的实证研究，把编码智能体 harness 从&amp;rsquo;黑盒整体评估&amp;rsquo;拆解为组件级受控实验：固定执行循环，只变化规划、动作空间、上下文管理三组件，在 4 个模型 × SWE-Bench Verified + Terminal-Bench 2.1 上跑出 176 组匹配设置。四个条件性发现——上下文管理在预算收紧时价值陡增（主要靠防溢出）、&amp;lsquo;规则删略+LLM 摘要&amp;rsquo;分阶段策略效率最优、规划对弱模型是准确率支架对强模型是成本节省器、bash 熟练模型用纯 shell 更省——为&amp;rsquo;harness 设计是条件科学而非玄学&amp;rsquo;奠定第一块实验基石。</description></item><item><title>SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-19-sol-pi-harness-study-paper-reading/</link><pubDate>Sat, 19 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-19-sol-pi-harness-study-paper-reading/</guid><description>NVIDIA 联合 NTU/MIT 提出 SoL-Pi：把编码智能体 harness 的效率改进本身建模为跨环境搜索问题，用约 150 个方向、500 个环境、3000+ 次实验、60000+ 次交互的自动研究漏斗，筛选出 Action Fusion、Online Context Compact、ObservationPack、Evidence-Preserving Reducer 四个可复用机制，EdgeBench 上 token 流量降 44.7–49.0%、成本省 1/3 且性能持平，迁移到未见过的 Opus 5 后端仍保留 94.3% 性能。这是 RSI（递归自改进）从&amp;rsquo;改模型&amp;rsquo;转向&amp;rsquo;改脚手架&amp;rsquo;的代表性工作。</description></item><item><title>Agentic Societies Need a Social Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-social-harness-agent-societies-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-social-harness-agent-societies-paper-reading/</guid><description>当不同主人的 AI 智能体开始自主协作，会发生什么？华盛顿大学的系统实验给出冷峻答案：即使全部诚实的 agent 也会因上下文分裂与信道争用大量失败（7 人群组排程成功率最低 0%），恶意 agent 凭&amp;rsquo;言论&amp;rsquo;即可让欺骗攻击 100% 成功、日历侧信道 100% 泄露。论文提出五层 Social Harness 协议栈（身份→有序通信→个人防火墙→协作规范→社会机构），把人类社会协作的制度智慧移植为 agent 社会基础设施。本精读覆盖&amp;rsquo;诚实 agent 也失败&amp;rsquo;的失败解剖与&amp;rsquo;协议栈防类别性失败&amp;rsquo;的设计哲学。</description></item><item><title>ExecuCritic × AgentGuard × RepoAtlas × Protocol Trimming 精读：编码智能体可靠性四重奏</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-coding-agent-reliability-quartet-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-coding-agent-reliability-quartet-paper-reading/</guid><description>四篇互补的 coding agent 可靠性研究合读：Intel×北大的 ExecuCritic 给 RLVR 加&amp;rsquo;校准 critic 塑形&amp;rsquo;——ρK 秩相关门控让 critic 失准自动坍缩，SWE-bench Lite +3.7pp 且 sandbox 执行省 42%；York 的 AgentGuard 从 642 条异常轨迹自动学条件激活护栏，异常执行率 69.0%→26.7%（代价：过度拒绝 19.3%）；北航 RepoAtlas 用 select-project-refresh 演化多模态仓库视图，三 VLM 一致 +2.4pp 且 token -5.8%；Intuit 工程报告量化协议保持裁剪——常规裁剪成功率 66.6-77.3% vs 协议感知 92.2%/自适应护栏 96.0%，临界阈值随复杂度上移。合读视角：可靠 coding agent 的四层防线——训练时（奖励塑形）、执行时（护栏）、探索时（上下文视图）、压缩时（协议保持）。</description></item><item><title>ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-sciencebuddy-recursive-self-improvement-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-sciencebuddy-recursive-self-improvement-paper-reading/</guid><description>ScienceBuddy 把&amp;rsquo;与研究者聊天&amp;rsquo;变成模型-线束双改进的燃料：研究者交互免费产出任务定义与评估 rubric，内层递归固定模型演化 harness（有界编辑+成对回归检查），外层递归固定 harness 做 rubric 奖励 GRPO——三周期后科学任务准确率 42.2%→73.3%，纯 harness 演化即可 +20pp（权重冻结），纯模型 RL 覆盖率 +19.5pp。Recursive-in-Recursive 范式为 RSI 提供了&amp;rsquo;两条改进通道各自可评估、互为条件&amp;rsquo;的工程化路径，并作为可下载的科研产品发布。</description></item><item><title>AlgoEvo × MOSCOPT × SkillLift：Skill 优化三部曲 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-algoevo-moscopt-skilllift-optimization-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-algoevo-moscopt-skilllift-optimization-paper-reading/</guid><description>三篇同日论文从三个角度推进 skill 优化。AlgoEvo（港城大）：把算法发现 agentic 化——design skill hub 解耦范式知识与发现引擎，三层经验库（经验卡/经验树/跨任务固化）组织搜索轨迹，6 任务匹配或超越专用方法且评估数与 token 大减。MOSCOPT：skill 池+gating skill 联合优化——EditAdam 双态维护+三阶段交错更新，免梯度单调改进，突破&amp;rsquo;单模板优化&amp;rsquo;的协同缺失。SkillLift：稀疏 oracle→稠密 rubric 双层优化——冻结 rubric 作廉价代理引导 skill 修订，解耦搜索与 oracle 成本。本精读合并解读 skill 优化的三条进化路径：知识组织、多技能协同、评估降本。</description></item><item><title>Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-asclepius-clinical-harness-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-asclepius-clinical-harness-paper-reading/</guid><description>临床 Agent 的&amp;rsquo;执行差距&amp;rsquo;：急诊整班次模拟（CES）中 Agent 多能给出正确诊断（4.39/5）却无法完整及时执行关键动作（2.94/5/3.34/5）——诊断对但病人死的结构性失败。Asclepius 三件套：换班间用 trace 反馈重写操作手册的自进化 harness、高风险规程外置的临床技能库、按病人队列隔离的三个子 Agent。held-out 批次上 critical-action correctness +22%（p=0.024）且诊断精度保持。本精读覆盖执行差距的三失效模式操作化与&amp;rsquo;操作手册级&amp;rsquo;harness 演化的医疗安全意义。</description></item><item><title>HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-harnessbandit-multi-harness-scheduling-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-harnessbandit-multi-harness-scheduling-paper-reading/</guid><description>同一模型在不同 harness（系统提示/工具 schema/控制循环/轨迹格式）下表现不均——多 harness 共同训练时每个优化步选哪个 harness 是被忽视的调度问题。HarnessBandit 用双信号在线调度：learnability（批平均绝对优势，还有多少可学）× transferability（梯度 sketch 余弦，学了是否白学），bandit 采样决策。6 harness 在 ClawGym 训练，held-out 任务与 held-out harness 双评测均优于混合批次训练。本精读覆盖双信号的互补性设计与 DeepSeek 产学研背景下的调度理论落地。</description></item><item><title>ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-modularrsi-harness-generalization-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-modularrsi-harness-generalization-paper-reading/</guid><description>harness 自改进的泛化性危机：在评测基准上演化=对测试集过拟合，单轨迹更新把系统性缺陷与实例细节纠缠。ModularRSI 三重解法——benchmark-disjoint（2000 个外部演化任务与评测基准不相交）、对比式信用分配（同任务成功/失败轨迹对比聚合跨任务证据）、模块化定位（缺陷归因到 harness 具体组件）。DeepSeek-V4-Flash 骨干上 SWE-Bench-Verified 73.40→76.45、TerminalBench 2.0 47.57→52.43，演化 harness 可跨基座迁移。本精读覆盖三大缺陷的诊断逻辑与对比式信用分配的因果推断本质。</description></item><item><title>PMPA × SkillSecurer × SkillAtlas：Skill 与记忆安全三连 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-pmpa-skillsecurer-skillatlas-security-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-pmpa-skillsecurer-skillatlas-security-paper-reading/</guid><description>三篇同日论文从攻击、防御、资源三面拼出 skill/记忆安全的完整地图。PMPA（复旦）：harness 持久记忆投毒——恶意指令藏进良性外部源诱导 Agent 写入持久记忆，OpenClaw ISR/C-ASR 73.7%/55.5%、Claude Code 66.9%/81.7% 且良性性能保持。SkillSecurer：红蓝 Agent 对抗扫描 skill 注入漏洞，9 威胁类型注入级评估，最佳后端唯一 100% 检测率，skills.sh 热门 skill 17%+ 有漏洞并实测触发事故。SkillAtlas：3014 案例/6589 轨迹的托管攻击轨迹库，42.5% 成功案例首轮失败后才成功，轨迹标签把 pre-execution guard 精度提至 0.770。本精读合并解读攻击面（记忆写入）→防御（红蓝扫描）→基础设施（公共案例库）的完整安全链条。</description></item><item><title>SkillSeam: Six Principles for Auditing Agent Skill Collections 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-skillseam-skill-collection-audit-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-skillseam-skill-collection-audit-paper-reading/</guid><description>一堆合格技能不等于一个可靠系统——技能在集合的&amp;rsquo;接缝&amp;rsquo;处失败：程序竞争注意力、别名重复加载、边界模糊。SkillSeam 提出六原则审计框架（持久梯度/系统连贯/机制门控/正交覆盖/触发流/粒度纪律），每条原则映射到失效机制→最强可观测信号→受控扰动测试。关键发现：破坏持久层级后 loaded-skill tokens +59.5% 而准确率不降——成本病在准确率病之前出现，准确率导向的评测对集合级架构债不敏感。本精读覆盖&amp;rsquo;按失效机制预测的信道评估&amp;rsquo;方法论与成本先行的预警价值。</description></item><item><title>Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-stellar-colosseum-many-agent-harness-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-stellar-colosseum-many-agent-harness-paper-reading/</guid><description>语言模型能产出看似合理的短证明，但在长程研究问题（不确定且相互依赖的决策序列）上不可靠——短证明能力与长程研究能力之间存在结构断层。Stellar Colosseum（CMU×Google Research）给出 model-agnostic 的多 Agent harness：策略探索后 readiness gate 决定何时分解、证明计划表示为 section 级相互依赖子问题、verifier 反馈路由回受影响部分；并行候选生成+定向证伪+重叠随机采样树聚合。已在数学与理论计算机科学问题上产出实际研究进展。本精读覆盖&amp;rsquo;长程=决策序列管理&amp;rsquo;的问题重构、readiness gate 的推理分配经济学与树聚合的抗噪声机制。</description></item><item><title>COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-cobra-skills-bandit-skill-optimization-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-cobra-skills-bandit-skill-optimization-paper-reading/</guid><description>港中文深圳团队把 agent 技能优化重构为&amp;rsquo;动态候选空间上的预算受限序贯优化&amp;rsquo;：contextual bandit 优先级评分决定评估哪个候选（exploit 历史得分 + explore 不确定性），证据驱动的进化只做有界精炼不做全局重写。结果：6 个 benchmark × 3 模型平均提升 13.1/26.9/22.5pp，相对 SkillOpt 总成本砍 55-58%，每 benchmark 只用 50 个优化样本，且对 harness 更换鲁棒。</description></item><item><title>Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-harness-or-model-contamination-controlled-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-harness-or-model-contamination-controlled-paper-reading/</guid><description>evolutionID GmbH 用 256 个私有任务的污染控制套件，首次把 agent 编程中 harness（驱动模型的软件层）作为唯一变量隔离测量。结论颠覆直觉：厂商原生 harness 无平均能力优势（±1.25pp 统计不显著），但按任务类型剧烈分化（仓库任务落后 9pp、竞赛任务领先 23.7pp）；中立 harness 每解一题成本反而高 1.3-1.6 倍。论文还自曝自家成本遥测存在缺陷并全量重算——测量诚实度的范本。</description></item><item><title>Look Before You Leap: Pre-Action Verification for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-pre-action-verification-silent-failure-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-pre-action-verification-silent-failure-paper-reading/</guid><description>针对 agent 动作的&amp;rsquo;静默失败&amp;rsquo;（产生貌似合理但错误的效果且不报错），本文提出 success/clean-failure/silent-failure 三分框架与确定性预检层：shell 命令侧 9,930 命令+482 工具上静态验证器捕获 95.8% 无效命令（语法/二进制检查 oracle-exact 零假阳性）；代码编辑侧 640 编辑×224 文件基准揭示格式尖锐分化——内容锚定格式（search/replace、diff）近零静默失败，行号/函数名格式高静默失败。护栏微秒-毫秒级、零模型调用，可包裹任何黑盒前沿 agent。</description></item><item><title>Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-skill-issue-gepa-skillopt-kotlin-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-skill-issue-gepa-skillopt-kotlin-paper-reading/</guid><description>TU Munich × JetBrains Research 的产学研负结果研究：在真实 Kotlin 仓库的合并 PR 反向挖掘任务上，GEPA 优化 SKILL 文档仅 +4.9pp（统计不显著）、SkillOpt 仅 +0.1pp——此前文献自报的巨大增益（55%→82%）是在弱模型弱 harness 配置下测出的。论文进一步证明 pass-rate 增益量级与二元判决本身的误标率（10.7% 盲重试通过）同阶，测量仪器而非优化器才是瓶颈。maintainer 盲读却确认 SKILL 含真实项目知识——分数之外的价值。</description></item><item><title>Studying Without a Syllabus: Task-Agnostic Environment Preprocessing 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-studying-without-syllabus-env-preprocessing-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-studying-without-syllabus-env-preprocessing-paper-reading/</guid><description>Scale AI 提出并形式化&amp;rsquo;任务无关环境预处理&amp;rsquo;设定：agent 在不知道下游任务分布的前提下自主&amp;rsquo;学习&amp;rsquo;陌生环境（S: Π×E→E——学习系统在预算内探索环境、产出制品给冻结 solver）。Meta-Agent（±策略 Archive）在 6 个异构 benchmark 的 5 个上取得最高 Avg@3；学习制品显著降低测试时采样需求；但更大学习预算不必然提升——&amp;lsquo;学什么&amp;rsquo;比&amp;rsquo;学多久&amp;rsquo;重要。</description></item><item><title>EvoSafeHarness 精读：Agent 安全没有万能线束，那就让线束自己进化</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-evosafeharness-agent-security-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-evosafeharness-agent-security-paper-reading/</guid><description>JHU/UC Berkeley/NVIDIA/UIUC/UW-Madison 五机构发布 EvoSafeHarness：为冻结 LLM Agent 自动搜索&amp;rsquo;模型×领域&amp;rsquo;专用安全 harness，DecodingTrust-Agent 上 ASR 45.6%→10.0%（utility 仅损 3.3 分），AgentDojo 82.8% utility @ 0 ASR。核心洞察：模型变体决定 enforcement 强度、领域变体决定谓词与状态——universal 安全 harness 在结构上就不存在。</description></item><item><title>Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-selfplay-text-harness-bbo-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-selfplay-text-harness-bbo-paper-reading/</guid><description>Google 的这篇论文问了一个极简的问题：agent 能否通过&amp;rsquo;写优化器代码并评估&amp;rsquo;的可执行实践学会一个搜索策略，然后把这个策略蒸馏成一段文本、迁移给从未见过这段实践的其他模型？答案是肯定的——自博弈产出的 197 词文本 harness（Harness A）使 Gemini Flash 在黑盒优化上 regret 降低 48%（N=30，p&amp;lt;.001），同一文本让所有受测 Gemini 执行器与 Claude Sonnet 都改善（regret -43%~-49%），独立复现的 Harness B 性能同档。&amp;lsquo;语言是部署搜索策略的便携介质&amp;rsquo;——harness 自动生成从进化搜索走向了实践蒸馏。</description></item><item><title>RobustSGPO: Search-Space Control for Agent Harness Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-robustsgpo-harness-evolution-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-robustsgpo-harness-evolution-paper-reading/</guid><description>语义梯度提示优化（SGPO）让 harness 能用执行反馈自动进化，但其局部更新规则把&amp;rsquo;这轮该改哪里、怎么改&amp;rsquo;留给运气——搜索空间悬空导致补丁无效率高、进化易破坏已有能力。武大×快手的 RobustSGPO 给出三步控制：显式指定本轮编辑（choose what to change）、构造并检查补丁（construct &amp;amp; check）、从现任或保留快照继续（防劣化回滚），配合周期性 1→2→3 权限调度。在快手 AgentX 头脑风暴工作流上（120 任务/95 运行/7350 候选），held-out 完成率 60.0%→80.0%、测试质量 3.77→4.14，结构化搜索补丁有效率 77.8% vs 48.9%。</description></item><item><title>Show-Harness: Just a VLM Agent Can Play Robots 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-show-harness-vlm-robot-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-show-harness-vlm-robot-paper-reading/</guid><description>Show-Harness 用一组离散语义动作单元（单步方向移动+夹爪动作）作为 VLM 与任意机器人本体之间的唯一接口：VLM 在语义空间推理意图，本体专属解释器把语义动作确定性落地为局部控制，VLM 始终对细粒度物理决策负责。同一接口实现零样本解锁前沿闭源 VLM（ZS 60%→82%）与几 GPU 小时微调小模型（FT 40%→65%），GUMI GUI 接口让人类与智能体用同一套语义动作采数据。本文精读其 perceive-reason-act 插件体系、语义-物理解耦机制与为何它能在跨任务/跨本体/跨环境全面超越 VLA 与智能体基线。</description></item><item><title>AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-agentleak-capability-cloning-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-agentleak-capability-cloning-paper-reading/</guid><description>偷到强 Agent 的技能文件，就能复制它的能力吗？本文给出否定答案并定义了&amp;rsquo;技能执行鸿沟&amp;rsquo;：技能规定做什么，而任务分解、工具选择、结果验证等隐式程序行为由强 Agent 在执行中现场补充——弱 Agent 拿到同一技能仍然完不成任务。更关键的发现是：这道鸿沟本身是泄漏面——对比受害 Agent 的成功执行与攻击者的失败执行，缺失的能力关键行为暴露无遗。AgentLeak 据此实现黑盒能力克隆：20 场景 600 实例上，比直接技能复用 pass rate 高 40%+、恢复 80%+ 能力差距，且模型/harness/工具全部不变。</description></item><item><title>CapScope: A Capability-Scoped Harness for Prompt-Injection-Resistant Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-capscope-capability-scoped-harness-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-capscope-capability-scoped-harness-paper-reading/</guid><description>编码 Agent 沙箱内的工具天然携带&amp;rsquo;环境权威&amp;rsquo;——命名一个资源就能操作它，间接提示注入正是利用这一点让 Agent 干用户没让干的事。北大团队的 CapScope 不让模型识别恶意文本，而是在 harness 层做能力作用域授权：从可信输入导出任务级权限上限，每个 sub-agent 持有独立的类型化能力集（存于模型上下文之外），每次工具调用逐主体检查。300 组对照实验：注入生效 ambient 权威 47/75、静态全局策略 33/75、CapScope 仅 3/75，而任务完成度 68/75 基本无损。论文已被 LMPL'26（ACM SIGPLAN 工作坊，Oakland）录用。</description></item><item><title>Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-coevolving-harness-model-imitation-fit-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-coevolving-harness-model-imitation-fit-paper-reading/</guid><description>harness 进化后，让弱模型模仿更强专家的轨迹——这个&amp;rsquo;显然正确&amp;rsquo;的配方在七个企业任务上全部翻车（平均 -14.9 分），而同样的做法在未进化 harness 下却有增益。论文定位出根源：模仿让弱模型学会了专家的知识，却也继承了专家的规划风格，破坏了它与&amp;rsquo;围绕自身原生风格进化出来的 harness&amp;rsquo;的拟合。解法是 on-policy 专家修正：meta-MLE agent 定位失败 turn、专家只重写那一轮，平均 +1.7 分且规划失败桶保持地板水平。本文精读拆解&amp;rsquo;模型-harness 拟合&amp;rsquo;这一新概念与其共进化配方。</description></item><item><title>Gander (Omni Interaction Agent Technical Report) 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-gander-omni-interaction-agent-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-gander-omni-interaction-agent-paper-reading/</guid><description>腾讯混元语音与浙大等团队的 Gander 用&amp;rsquo;小脑-大脑&amp;rsquo;协作框架统一了全双工实时交互与长程 Agent 执行：9B 小脑以 Thinker-Talker 流式架构逐秒 chunk 决策听/说/打断，大脑免训练接入 Codex/Claude Code 执行长任务，编排运行时以 task_start/send/resolve 结构化调用衔接。Full-Duplex-Bench v3 上 turn-taking 100% 全场最佳、过早打断仅 8.0%（GPT-Realtime 13.5%）；SpokenQA 全双工组第一。模型、代码、数据全部开源。</description></item><item><title>NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-neohorse-1-routing-harness-rsi-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-neohorse-1-routing-harness-rsi-paper-reading/</guid><description>NeoHorse-1 把部署中的模型路由 harness 变成递归自改进（RSI）的数据飞轮：路由层天然记录每次交互的&amp;rsquo;能力需求预测-实际执行-结果&amp;rsquo;三元组，这些记录被转化为保留交错推理与工具调用的 user-turn 训练样本，路由分数进一步组织成三阶段课程 SFT 与路由引导的在线策略蒸馏。4B/9B 模型十项基准宏平均分别从 58.94/65.60 提升至 64.87/69.04，路由 harness 数据比公开 Agent 数据平均高 6.26 分。本文精读拆解其数据管线、课程设计、OPD 机制与&amp;rsquo;评估-选择-更新&amp;rsquo;闭环为何能成立。</description></item><item><title>Bilevel Coordinated Reflection: 多智能体 LLM 系统的博弈论统一理论 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-bcr-bilevel-coordinated-reflection-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-bcr-bilevel-coordinated-reflection-paper-reading/</guid><description>UCL×利物浦×华为的 BCR 把 orchestrator–worker 多智能体系统建模为双层协调博弈，证明 follower 子游戏是近似势博弈，并首次给出&amp;rsquo;只看文本的门控不可能可靠&amp;rsquo;的信息论不可能性定理。据此提出的 SRMA 仅在环境验证风险严格下降时接受候选记忆，SWE-bench 500 实例解决率 72.2%（免费反思仅 58.4%）。本文精读其双层博弈建模、漂移分析、不可能性定理与 SWE-bench 端到端验证的完整因果链。</description></item><item><title>EvoHarnessBench: 智能体能跟上不断进化的 Harness 吗 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-evoharnessbench-evolving-harness-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-evoharnessbench-evolving-harness-paper-reading/</guid><description>Salesforce Research×UNC Chapel Hill×UW–Madison 的 EvoHarnessBench 把非平稳性从任务流转移到 harness 本身：17 条受控 harness 进化流（802 任务、520 工具、42 技能、62 智能体），分部署评估（能力保持）与自进化适应两设定。基准回答一个此前无人系统提问的问题：当工具、技能、子智能体持续增加时，已部署 agent 的既有能力何去何从。</description></item><item><title>HackProbe: 自进化语言模型的奖励黑客检测与免疫 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-hackprobe-reward-hacking-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-hackprobe-reward-hacking-paper-reading/</guid><description>Fullive-AI×北大×京东×NTU×武大的 HackProbe 是一个通过两个黑盒钩子挂载到任意自进化回路的监控器：秘密固定分布对比核心保证跨代可比，轮换新鲜层抗共适应；四项检验+Šidak 校正输出族校准 p 值，风险感知免疫层从候选池重选诚实更新。本文精读&amp;rsquo;诊断之外还能恢复&amp;rsquo;的奖励黑客治理闭环。</description></item><item><title>TROVE: 轨迹锚定的最小充分路线编辑 — 智能体编排的运行时修正 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-trove-route-orchestration-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-trove-route-orchestration-paper-reading/</guid><description>TROVE 把智能体编排的结构决策从&amp;rsquo;执行前锁定&amp;rsquo;改为&amp;rsquo;运行时最小充分编辑&amp;rsquo;：离线把工作流搜索轨迹蒸馏为原子/复合技能+结果条件转移图，在线对挂起路线执行保留/插入/替换失效后缀三操作。代码生成、QA、数学推理上质量-效率权衡全面优于 AFlow/MaAS/LAS。本文精读&amp;rsquo;route 即临时品&amp;rsquo;的编排新原则。</description></item><item><title>What Does Multi-Harness RL Learn? — 评测 Harness 是 Agent RL 的主导变量 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-multi-harness-rl-credit-assignment-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-multi-harness-rl-credit-assignment-paper-reading/</guid><description>本论文在同一 Qwen3-8B 热启动上回放相同任务-harness 记录（Aider/OpenHands/Qwen Code/SWE-agent），对比 GRPO 的 Within/Cross 两种分组规则，用 24,000 次密封 SWE-bench Verified 评估发现：评测 harness 使解决率从 2.14% 摆到 9.27%（4.3 倍），训练配方仅移动 1.16，分组规则不显著（+0.25pp，CI 含 0）。harness 工程对 Agent RL 的影响碾压算法选择。</description></item><item><title>所有 Skill 都会死：卡比谈驾驭大模型的三层功夫——上下文、方法论与长活 Agent</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-06-builder-club-harness-llm/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-06-builder-club-harness-llm/</guid><description>GitHub 中国区 Top 100 开发者、Open CLI 作者卡比在 B站 Builder Club 交流日分享如何驾驭大模型：AI 的能力不只来自模型，也来自 Harness（运行时脚手架）。他给出三层可操作的功夫——理解并主动管理四层上下文与「有效上下文」，用方法论名字替代冗长 Skill（断言「所有 Skill 都会死」），以及在开源社区用长活 Agent 与 Swarm/Graph/Team 三种多 Agent 形态承接真实工作流。核心判断：模型终将吸收一切提示词工程，人剩下的核心位置是编排——拆任务、管上下文、沉淀 AI 友好（AX）的流程。</description></item><item><title>HarnessEvo 精读：Harness 自进化的价值藏在控制槽位里</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-harnessevo-value-localization-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-harnessevo-value-localization-paper-reading/</guid><description>HarnessEvo 把 Agent harness 分解为 role/strategy/format/control 四个可独立进化的槽位，用 leave-one-in/out 协议做价值归因：整体指标&amp;rsquo;看似无效&amp;rsquo;（0.657 vs 0.642），但收益完全 localized 于 reflection/control 槽位（+0.119, p=0.0046）；等预算下多槽位同进反而互相稀释——预算分摊陷阱。本文基于全文阅读拆解其归因协议与对自进化领域的方法论警示。</description></item><item><title>HookPry 精读：Agent Harness 的 hook 更新通道是全新的供应链攻击面</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-hookpry-agent-harness-security-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-hookpry-agent-harness-security-paper-reading/</guid><description>HookPry（北邮/网信办数据中心/北航/浙大）首次系统揭示 AI Agent Harness 生命周期 hook 的更新通道攻击面：良性插件上架获取信任后，一次携带 hook 的恶意更新即可在 LLM 完全不可见的路径上以宿主权限执行任意命令。1000 次端到端攻击攻破全部 7 个 harness（最高 92.5%），Microsoft Defender 召回率 0%。本文基于全文阅读拆解其 AMO/TD/LCI 三组件与防御失灵的机制根源。</description></item><item><title>Discriminative World Models for Web Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-discriminative-world-models-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-discriminative-world-models-paper-reading/</guid><description>Web agent 用世界模型做测试时动作选择：采样候选动作→预测下一状态→排序执行。但现有世界模型都用监督式&amp;rsquo;下一状态预测&amp;rsquo;训练——花大量 token 复述页面上没变化的部分，而下游 ranker 需要的恰恰是&amp;rsquo;不同动作导致的差异&amp;rsquo;。UC Berkeley 联合 MIT-IBM Watson AI Lab 提出 predicted-state matching：预测表示必须把真实结果状态从替代动作的结果状态中区分出来。同一份数据、同一个 Qwen3-8B 底座，仅换训练目标，匹配准确率从 47.77% 跳到 80.80%，WebArena-Lite 端到端成功率从 13.94% 提到 28.48%。本精读拆解&amp;rsquo;训练目标与下游任务对齐&amp;rsquo;这一教科书级修正的完整证据链。</description></item><item><title>Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-repo-to-skill-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-repo-to-skill-paper-reading/</guid><description>自主 ML 研究 agent 缺的不是更强的模型或更聪明的流程，而是&amp;rsquo;怎么把方法跑通&amp;rsquo;的操作知识层。BAAI 联合中科大、人大、港理工提出 DisCo 蒸馏框架，把 1000 个 GitHub 仓库蒸馏成 5353 个经过验证的技能，构建 AREX-Skill Library。在固定 GPT-5.5+Codex 的对照实验下，技能让 MLE-bench 相对提升 134.3%、PaperBench 提升 34.4%、FrontierCS 提升 9.2%、PassNet 提升 14.0%，并以更低 token 消耗帕累托支配 Claude Code。本精读拆解技能图三层结构、四阶段蒸馏流水线、对照实验设计，以及&amp;rsquo;试错成本越高、操作知识价值越大&amp;rsquo;的机制根源。</description></item><item><title>Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents — A Source-Code Study of Eleven Systems 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harness-engineering-anatomy-paper-reading/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harness-engineering-anatomy-paper-reading/</guid><description>Wavestone AI Lab 对 11 个生产级编码 agent harness（Claude Code、Codex CLI、Gemini CLI、Mistral Vibe、OpenHands、Aider、Mini-SWE-Agent、Hermes、Pi、OpenCode、OpenClaw + 元 harness 对照 Omnigent）做源码级解剖：定义 harness 七大子系统、产出 13 条跨系统观察、29 个重复设计模式、18 条设计建议与 90 行最小 harness。关键发现：7/11 系统收敛于阈值触发 LLM 压缩的记忆管理事实标准；提供商抽象呈五档光谱；Codex 已把 per-model 提示作为服务器端数据运行时下发。这是&amp;rsquo;harness 工程&amp;rsquo;学科的第一部解剖学图谱。</description></item><item><title>Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harness-of-harness-paper-reading/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harness-of-harness-paper-reading/</guid><description>上海AI实验室提出 Harness-of-Harness（HoH）：在现有编码 agent harness 之上再组织一层&amp;rsquo;规划-开发-测试&amp;rsquo;循环，通过双状态传递（制品态+证据态）、有界增量目标与独立 QA 验收，让 LLM 编码智能体实现多日自主软件开发与持续改进。三个 harness-模型对在 GameCraft-Bench/FrontierSWE/ProgramBench 上平均相对提升 52.25%，FrontierSWE 十轮迭代从 22% 升至 72.67%，并用 70+ 迭代自主开发出可玩的 FPS 游戏。本文从问题抽象、机制因果到通用灵感逐层拆解。</description></item><item><title>HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harnessdev-paper-reading/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harnessdev-paper-reading/</guid><description>ByteDance Seed 联合 SUTD/GaTech/M-A-P 发布 HarnessDev——首个把评测单元从&amp;rsquo;任务输出&amp;rsquo;改为&amp;rsquo;可运行基础设施&amp;rsquo;的基准：creator LLM 从无策略弱种子构建完整 harness（Creation），再基于下游执行反馈迭代改进自己的 harness（Evolution），在 2207 个下游实例上按 capability+efficiency 双轴评估。核心发现：模型自建 harness 在 writing/MLE 域追平甚至反超人类参考系统，但在 code/search 域差距显著；Evolution 的增益不稳定且严重绑定 executor；换 executor 后最高回退 10.32 分。这为&amp;rsquo;harness 工程能否自动化&amp;rsquo;提供了第一份系统性体检报告。</description></item><item><title>HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harnessevolve-paper-reading/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-03-harnessevolve-paper-reading/</guid><description>华为 ICT AI 能力中心提出 HarnessEvolve：针对自进化 agent 的三大失败模式（终态反馈导致的信用分配失败、捷径学习、灾难性遗忘），用&amp;rsquo;参考轨迹对齐&amp;rsquo;提取逐步误差信号、双门控（质量门+性能门）过滤候选更新、epoch 末 held-out 验证选最优快照。在企业内数据集 CloudCoreNetwork-QA 上把 Qwen3.6-27B 从 43.4% 拉到 86.9%（超最强基线 GEPA 21.6 个百分点），开源三数据集全胜 GEPA/ACE/SkillOpt，且在 OpenClaw 上优化的 skill 可迁移到 OpenCode/LAMAgent 等四个框架（SpreadsheetBench 最高 +30.4 分）。</description></item><item><title>WHALE: A Simple Recipe for Joint Harness–Weight Optimization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-03-whale-harness-weight-optimization-paper-reading/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-03-whale-harness-weight-optimization-paper-reading/</guid><description>KRAFTON 联合 KAIST/Stanford 提出 WHALE（Weight-Harness Alternating LEarning）：把 agent 性能看作模型权重 θ 与可执行 harness 代码 h 的联合函数 J(θ,h)，交替执行&amp;rsquo;当前 harness 下在线拒绝采样微调&amp;rsquo;与&amp;rsquo;更新后模型上 Meta-Harness 搜索&amp;rsquo;两阶段，用固定时长或自适应 patience 规则切换。在 Qwen3.5-2B/4B × 搜索问答/数学/国际象棋三域上，比 weight-only、harness-only 与 Fast-Slow Training 高 4.15–24.38 个百分点，且揭示 harness-limited 与 weight-limited 两种机制不同的瓶颈域。这是首个把优化空间从&amp;rsquo;权重+文本提示&amp;rsquo;扩展到&amp;rsquo;权重+完整可执行 harness&amp;rsquo;的交替优化配方。</description></item><item><title>ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-contextpilot-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-contextpilot-paper-reading/</guid><description>深度精读清华大学、腾讯优图实验室与上海AI Lab 合作的 ContextPilot：一个主动上下文管理框架。针对现有方法工具集贫乏（只有搜索/删除/摘要）、探索低效（上下文编辑动作影响悬殊却被均匀采样）、信用分配粗粒度（轨迹级奖励平摊给所有编辑动作）三大缺陷，它扩展出规划、长期记忆、软卸载三类工具，并用上下文变化量+熵变化识别关键编辑决策做分支采样（context-aware partial rollout），再用所有后续分支的平均回报估计动作级优势（细粒度信用分配，方差降为 1/n）。8B-RL 在四基准平均 69.40 超 StateLM-8B-RL 的 65.85；深搜任务上每轮输入 token 稳定在 8-10K（基线线性涨到 30K）；消融证明细粒度信用分配贡献最大且在全部基准一致提升。</description></item><item><title>EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-evoundo-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-evoundo-paper-reading/</guid><description>精读独立研究者团队的 EvoUndo。论文直面 LLM Agent 自进化的安全盲区：能提升能力的变异未必能被安全撤销，正确恢复往往依赖变异前状态。EvoUndo 把自变异表示为四元组（前向变异+见证捕获+恢复程序+效果契约），在反事实状态上做往返验证。600 个任务中 197 个能力正向但恢复失败的变异构成失败库：原始语言下常规修复 0/197；oracle 审计分解出双瓶颈——S0 层是 grounding 瓶颈（精确地址后 0/48→38/48），S1 层是表达力瓶颈（扩展语言后 142/143），组合修复 180/197。另发现丰富语言加精确诊断反而降效。把能改与改回去拆开的开创工作。</description></item><item><title>Logos: An Agent Harness on a Cross-Process Bus 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-logos-cross-process-harness-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-logos-cross-process-harness-paper-reading/</guid><description>深度精读 Sussex、浙江工商大学与上海书缘信息技术合作的 Logos（AAMAS 2027）。针对单进程 agent 框架插件与会话共存一个进程的单点故障问题，论文用四个引理证明时空可组合性演算的可逆性保证可跨进程成立——可靠性不变量只定义在状态空间上，而模型推理是无状态的；再构建 ROS 风格跨进程 harness：插件即进程、路由器只存路由表、唯一共享状态是 append-only 转录。80 个会话在四个击杀点上全部冷切换恢复且零重复效应，总线跳 0.215ms 仅为首 token 的 1/823；同故障下单进程停机 547.1ms 中断全部会话，对等构造只影响一个节点。</description></item><item><title>LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-looparena-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-looparena-paper-reading/</guid><description>深度精读阿里 DreamX 团队联合北邮、UNSW Sydney 与 Data61 CSIRO 推出的 LoopArena：首个把『模型作为运行时循环控制者』的编排能力本身作为被评测对象的基准。它冻结 Worker 编码智能体与全部执行环境，只比较 Controller 模型在 advance/verify/stop 三类决策上的表现；Type I/II/III 三级成本递减设置使其可低成本诊断循环控制能力。关键发现：完整任务上最强 Controller（GPT-5.5）Strict Success Rate 仅 24.69%，机械重复目标的 fixed control 在全任务上与无控制持平（18.52%），证明有用的循环控制必须随运行状态自适应切换；Type II 切片评估平均省 64.4% 成本且与全任务排序高度一致（Spearman ρ=0.9747）。</description></item><item><title>openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-openjiuwen-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-openjiuwen-paper-reading/</guid><description>深度精读华为开源的 openJiuwen 编码智能体 harness。论文把 agent harness 提升为一等系统层，用两大设计原则回应长时程编码的挑战：结构可组合性（共享 Inner Loop/Outer Loop 执行基座 + Rail 生命周期钩子上的有序能力组合，同一执行语义从单智能体复用到子智能体与 Swarm Flow 多智能体流）与运行时适应性（在固定模型策略周围改变框架控制的运行时状态：Context Management 渐进压缩、Goal Mode 语义化验收停止、LSP 被动反馈闭环修正、Self-Reflection 跨任务经验蒸馏）。SWE-bench Verified 达 82.6%（超最强榜单 3.4 个百分点）、Terminal-Bench 2.1 达 87.19%；模型对齐对比下 1-4 小时长任务 52.38% vs mini-swe-agent 同设定 35.71%，佐证上下文管理在长轨迹上保住了深推理收益。</description></item><item><title>WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-weagent-mmsearch-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-weagent-mmsearch-paper-reading/</guid><description>多模态搜索智能体常被环境拖后腿：许多搜索环境只把网页转成文本喂给模型，工具返回的图片直接丢弃，号称多模态的轨迹实际退化成纯文本推理；长时程交互中的超时、超长输出、格式错误还会污染 RL 训练信号。腾讯微信 AI 与中山大学提出 WeAgent-Harness，把检索图像注册为可寻址的持久状态并跨轮回灌，配合失败感知的 FA-GSPO 训练算法与可诊断『检索失败还是感知失败』的 VisTarget-Bench，基于 30B 模型在 8 个基准上取得 55.97% 平均分，媲美约 10 倍参数量的前沿模型。本文按九部分结构精读其动机、机制、证据与可迁移灵感。</description></item><item><title>Metis: Typed Runtime Mediation for Tool-Using Software Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-metis-typed-mediation-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-metis-typed-mediation-paper-reading/</guid><description>深读一篇罕见的单人独立研究：Metis 把模型与外部副作用之间那一层运行时当作正经的系统软件工程对象，用类型化事件图显式刻画权限判定、并发调度、终态闭包与生命周期修复。30 对匹配真实 I/O 实验中四类调度中位耗时 14.146ms，全面快于强制串行的 25.958ms；子代理边界消融 0/5 逃逸、路由级权限 oracle 10/10 全对，同时诚实呈现 3 个负面结果。本精读逐部分拆解其机制设计与有界主张的写法。</description></item><item><title>SKILL.state: Scalable Long-Horizon Agent Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-skill-state-execution-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-skill-state-execution-paper-reading/</guid><description>让 Agent 执行长任务时，主流运行时把所有推理、动作、观察不断追加进对话历史，prompt 随步数二次膨胀，token 烧钱、噪声污染、过期事实还会诱发幻觉。Google 与 Purdue 合作的 SKILL.state 干脆废除这个 append-only 历史：每一步模型只看到技能规范、结构化执行状态和最新观察，推理轨迹用完即弃，状态以 JSON 补丁形式确定性合并。prompt 尺寸从 O(T) 降为 O(1)，百步任务 token 缩减 16.2 倍，CTF pass@1 提升 7.8 个点，外部篡改状态后零步恢复而基线幻觉 5 至 8 步。本文精读其运行时设计、预算匹配对照实验与效果根源。</description></item><item><title>PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-pilot-live-self-improvement-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-pilot-live-self-improvement-paper-reading/</guid><description>Agent 的自我改进大多发生在一次任务结束之后——但那时这次运行已经救不回来了。AllSpark 团队的 PILOT 把自改进做成 live 的：监督者通过双向活通道在工作者执行中途重定向或中止（live steering），同时从活轨迹蒸馏可复用技能进持久 harness（live self-evolution），模型参数全程冻结。在 Terminal-Bench 2.0 上 PILOT 以 71.6 均分领先最强单 Agent 基线 5.3 个点；20 轮自改进迭代后 GLM-5.1 从 66.3 升至 80.9（+14.6pp），每任务输出 token 反降 42.9%。评测协议设计严谨：运行中零基准反馈，验证器只决定哪些更新进入下一轮。本文精读其监督者-工作者架构与两个中途纠偏案例。</description></item><item><title>Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-symtrace-mas-failure-debugging-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-symtrace-mas-failure-debugging-paper-reading/</guid><description>多智能体系统（MAS）失败后重跑一次就能修好？这篇论文用 SymTrace 可控回放框架和 536 条人工标注失败轨迹（SymFail）证明：现有任务级重跑方法的修复成功率仅 6.90%，且主要是靠 LLM 采样随机性&amp;rsquo;碰&amp;rsquo;出来的，而非真正修复了失败机制。作者提出的症状驱动节点级干预把单次干预修复率提到 20.15%（相对最强基线提升 191.89%）。本文精读其可控回放的数据集设计、实验证据与&amp;rsquo;因果修复 vs 随机修复&amp;rsquo;的方法论启示。</description></item><item><title>Same Model, Different Harness: Different Coding-Agent Results 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-same-model-different-harness-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-same-model-different-harness-paper-reading/</guid><description>同一个模型、同一批任务，只换 Agent Harness 的配置，编码成功率能差多少？独立研究者 Sydney Lewis 用严格的配对实验给出答案：在 20,480-token 紧窗口下，SWE-bench Verified 的平均 F2PF 从 28% 涨到 49%，完全解决数从 43 到 72。treatment 只有三件机械武器：半衰期规则缩短旧工具结果、检测器打断重复劳动、命令防护。论文最有冲击力的结论是方法论层面的——模型加 harness 才是被测求解器，单报模型名字的编码评测并不完整。本文精读其实验设计、跨四模型迁移证据与机制分析（阅读边界翻倍）。</description></item><item><title>When Context Gets Root: Privilege Escalation in LLM Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-llm-harness-privilege-escalation-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-llm-harness-privilege-escalation-paper-reading/</guid><description>深度精读南京大学与荣耀终端的 LLM Agent 安全论文。作者提出「指令特权升级」这一新攻击范式：利用 agent harness 在委派子智能体、持久化目标、定时任务时的上下文重构，把工具级恶意内容真实地提升为用户级或系统级指令。在 Claude Code、Codex 等 6 个主流编码 agent 上，13 个攻击目标（含远程代码执行）全部达成，连自动权限审查也被绕过。</description></item><item><title>Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-aig-failure-attribution-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-aig-failure-attribution-paper-reading/</guid><description>多智能体 LLM 系统跑失败了，到底该怪哪个 Agent、哪一步？AWS Agentic AI 与特拉维夫大学提出的自适应影响图（AIG）给出的答案是：这不是模型不够聪明的问题，而是接口设计的问题。论文把人类工程师调试系统的&amp;rsquo;可观测性&amp;rsquo;范式搬给 LLM——先用 agentic builder 把失败的原始日志构造成带继承边的影响图，再用 agentic reader 沿边回溯定位首个错误。在 Who&amp;amp;When 基准上，同一模型仅靠改善轨迹表示就从 46.40% 提升到 55.20% 的步骤定位准确率，刷新 SOTA。本精读将拆解其四级接口阶梯、两阶段框架与增益根源。</description></item><item><title>Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-feedback-backfires-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-feedback-backfires-paper-reading/</guid><description>帕绍大学单作者研究，用一台CPU笔记本完成了一项改写agent工程常识的测量：把失败调用和报错追加进transcript这个从ReAct沿用至今的标准做法，会让小模型更倾向于重复刚刚失败的动作。定义corrective gain指标后，6个模型（135M-1.7B）在两个环境的G全部为负（约-1.03 nats/token，每token odds×2.8）。反事实分解揭示根因：83%的损害来自失败调用的表面形式触发了复制机制，而非模型读不懂报错。由此预测并验证：描述化改写与decoder级ban有效，&amp;lsquo;别重复&amp;rsquo;指令无效，而&amp;rsquo;清空上下文重试&amp;rsquo;这一先前推荐方案恰恰最糟——它精确恢复了产生失败的上下文。</description></item><item><title>From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-opsharness-rca-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-opsharness-rca-paper-reading/</guid><description>微服务故障根因分析（RCA）自动化该往哪个方向使劲？这篇香港中文大学与字节跳动的论文先用 24 个受控实验给出反直觉结论：裸的通用编码 Agent（Codex/Claude Code）已经全面超过从零构建的专用 RCA Agent——但离生产可用还很远，缺的不是推理能力而是系统特定经验。答案不是重造 Agent，而是造一个能自我进化的外部 harness：OpsHarness 用四层知识与 idea-card 工具库做数据平面，用「挖掘-验证」双门进化循环做控制平面，Top-1 准确率 59.0%（较裸 Agent 相对提升 63.4%，是专用 Agent 的 4.02 倍），并在真实生产环境拿到约 3 倍提升、同一故障复发时从 Top-3 之外跃升 Top-1 且 2 分钟内定位。</description></item><item><title>From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-ooda-tool-typed-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-ooda-tool-typed-paper-reading/</guid><description>OODA-Tool（AACV 2026）诊断出多轮工具使用的核心失败模式为「状态-动作竞争」：直接函数调用与 ReAct 策略在同一条自回归轨迹里同时学状态跟踪和动作生成，产出下一个调用的压力会覆盖早期积累的信息。借鉴博伊德的 OODA 循环，它把决策拆为 Observe（重建任务状态）→ Orient（判断执行就绪）→ Decide（形成动作结构）→ Act（落地实现）四个类型化阶段，由中央控制器逐级校验交接。在 ToolDial 上用 Qwen3 从 0.6B 到 14B 全规模评测，Specialized OODA 相比 Direct-LoRA 提升 4.48-6.99 个百分点，小模型与状态密集型任务增益最大；状态-动作矛盾率从无分割的 10.7% 降至 3.9%，代价是 2.36 倍延迟。本精读拆解其类型化接口、消融证据与并行调用边界。</description></item><item><title>JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-jit-agent-harness-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-jit-agent-harness-paper-reading/</guid><description>Agent 的能力从来不只取决于模型权重，还取决于包裹模型的执行脚手架（harness）。LV-NUS Lab 提出 JIT-Agent，训练一个 27B 的「harness 智能模型」，在推理时为任意现成 agentic LLM 即时合成任务自适应的 harness，并能修复与在线演化。DeepSeek-V4-Flash 配上它即可在 DeepSearchQA 反超 GPT-5.6 达 9.1 分，同时成本比所有固定 harness 平均低 36%。本精读拆解其四模块 harness 协议、三阶段训练流水线与 Evo-GDPO 目标，并解释「按任务实例即时生成脚手架」为什么在机制上必然优于单一固定脚手架。</description></item><item><title>Recuris: Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-recuris-memory-evolution-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-recuris-memory-evolution-paper-reading/</guid><description>Recuris（NUS × Stanford × Oxford × Princeton）把递归自我改进从『改模型/改智能体』收缩到『只演化外置记忆控制层』：工作记忆维护经检查器验证的任务状态并按需调用技能，跨任务的固定 Meta-Agent 读结构化轨迹、把失败归因到 E/W/ρ/C 四组件之一并只修补被归因组件，经修复源任务且不回退开发集的验证门才准入。在 4 个长程基准 × 10 个模型上 35/37 完成的模型-基准对成功率提升，GPT-5.6 Sol +17.8、Claude Opus 5 +15.6、最长任务 +32.2 分，六类长程失败模式下降 20–86%；机制上证明长程失败是执行问题而非检索问题，技能价值是『调用条件性』的，结构化轨迹使故障定位从 13.0% 提升到 64.8%。</description></item><item><title>StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-starharness-enterprise-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-starharness-enterprise-paper-reading/</guid><description>深度精读 ServiceNow 与 Mila 的企业环境 harness 进化研究——在模型权重完全冻结的前提下，用分层搜索自动进化面向特定企业环境的 agent harness（提示词/工具接口/skills/MCP/子agent/执行循环）。通过按基线失败模式分层采样构建紧凑进化池、proposer 可见搜索集与隐藏选择集分离、test-flip 门控 + 严格爬山接受，在 ITBench SRE / EnterpriseOps-Gym ITSM / AutomationBench Finance 三个基准上较默认 harness 提升 20-35 个百分点，且冻结迁移到 Qwen/GPT 全系列模型仍有效。21 个被接受 patch 归结为三类修复：接口修复、环境约定显式化、压缩搜索的操作知识。</description></item><item><title>The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-harness-arch-convergence-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-harness-arch-convergence-paper-reading/</guid><description>深度精读南洋理工大学的 LLM Agent Harness 架构收敛研究——首个对 harness 层本身做源码级多案例研究的论文。三个来自对立哲学的开源编码 agent harness（LangChain deepagents、Earendil pi、DeepSeek dsh）反向演化却汇聚于同一五要素中间形态：商品化循环、仅追加可重放会话记录、模型怪癖数据化、上下文渐进披露、显式扩展缝隙。本文逐项拆解五要素与三类收敛机制（平行发现/扩散/字面复用），还原四条汇聚断层线缺陷分类，并解读唯一零收敛维度&amp;rsquo;外部可验证性&amp;rsquo;为何是预测性缺口。</description></item><item><title>Apodex 1.1 姊妹篇补遗：本日精读系列导览与 2026-08-25 学术全景</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-paper-reading-series-guide/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-paper-reading-series-guide/</guid><description>本文为 2026-08-25 精读系列的导览：16 篇触发顶会标准精读的论文横跨 Agent 评测反作弊、Harness 可学习化、经验资产化、因果测量方法学四大主题。本文给出全部精读的索引、跨论文趋势综合（verifier-grounded 成为共同底座、评测从&amp;rsquo;分数多高&amp;rsquo;转向&amp;rsquo;分数测的是什么&amp;rsquo;、产学研从联合发文转向资产+方法学互换），以及按读者角色（研究者/工程师/管理者）的阅读路线图。</description></item><item><title>Apodex 1.1: Scaling Agentic Intelligence for Complex Work 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-apodex-1.1-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-apodex-1.1-paper-reading/</guid><description>Apodex 1.1（Apodex Team）提出&amp;rsquo;双扩展面&amp;rsquo;范式：把任务环境构建（Environment Scaling）与多智能体协调（Agentic Coordination Scaling）确立为与模型规模并列的两个扩展维度。Agent Team 架构把任务分解、异步委派、非对称验证、重规划训练进模型策略，在 GDPVal 拿到 78.8 win rate、IMO-2026 数学超金牌线、SWE-bench Verified 77.7%，且全部开源（含 35B mini 版权重）。本精读重点拆解其&amp;rsquo;正向便宜、逆向昂贵&amp;rsquo;的验证器设计与 Agent Team 协调增益的机制来源。</description></item><item><title>AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-autosaddler-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-autosaddler-paper-reading/</guid><description>AutoSaddler（Microsoft × POSTECH × KAIST × 南方科技大学）把 Agent harness（提示词/工具/中间件）的优化形式化为离线 mini-batch 学习问题：深度诊断 Agent 读执行轨迹定位根因、生成结构化 patch（Prompt/Tool/Middleware 三类九子型）、Reflection 提炼经验存入 EvoDAG 进化图、泛化感知选择防过拟合。GAIA2 +9.0pp、SWE-Bench Pro +9.6pp、Terminal-Bench 2.0 +10.0pp 全面超越人工与自动基线，且学习轨迹只需最强基线的 1/10。</description></item><item><title>Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-aces-skill-evaluation-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-aces-skill-evaluation-paper-reading/</guid><description>NVIDIA 提出 ACES，把 Agent 技能从「扫描文档」推进到「活体配对评测」：同一任务在有技能/无技能两种条件下运行，唯一变量是目标技能是否可用，六指标差值即 Skill Lift。145 个技能上结构分与 LLM 评分相关性仅 Spearman ρ=0.14，94.5% 通过结构门槛的技能与活体 Lift 相关性近零（-0.018）；947 个配对案例显示平均复合 Skill Lift 为 0.2134，其中技能执行 +0.33、行为检查 +0.30 等过程指标贡献最大，且负 Lift 可区分「从未发现」与「发现但误用」两类失败——这是文档扫描永远看不到的信号。</description></item><item><title>EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-evoharness-rl-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-evoharness-rl-paper-reading/</guid><description>深度精读 UIUC×Meta AI 合作论文 EvoHarness-RL（已被 LLA@COLM 2026 接收）：把长程智能体对外部 harness（记忆、工具、状态跟踪）的访问从提示词硬编码变成可学习的策略决策。通过 BPE 三态抽象（Belief/Progress/Experience）与四个元动作（track/commit/recall/note），配合教师轨迹 SFT 与代价感知 GRPO 两阶段训练，Qwen3-8B 在 ALFWorld 上从 ReAct 的 47.9% 跃升至 96.9%，逼近 Claude Opus 4.5；训练中还揭示“harness 退火”与“harness 演化”两个动力学过程。</description></item><item><title>Natural-Language Workflows Are Not Software Yet: Artifact-Driven Compilation for Reliable Agent Execution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-artic-workflow-compiler-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-artic-workflow-compiler-paper-reading/</guid><description>深度精读普渡大学 arXiv 2026 论文 ARTIC：自然语言工作流虽然给智能体提供了软件式接口，但数据依赖隐式、长分支指令难跟随，执行不可靠。ARTIC 把 NL 工作流编译为每步声明读写工件、约束门控产出、显式控制转移的形态，用约束优化精化高风险步骤，再以局部义务分解加场景干跑验证忠实性。在 11 个真实领域 488 个问题上，任务解决率较原始文本工作流提升 28 个百分点，跨模型执行一致性提升 32 个百分点，重复执行一致性提升 56 个百分点。</description></item><item><title>Prime Agent: A Self-Improving RLM Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-prime-agent-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-prime-agent-paper-reading/</guid><description>Prime Agent（Prime Intellect × Princeton × MIT）用一个持久 IPython REPL + 递归子 Agent 的抽象，证明同一模型仅更换 harness 即可把 ARC-AGI-3 成绩从 30.2% 推到 95.5%、超过人类专家基线 95.4%。本精读拆解其两层核心抽象——Recursive Language Model（把上下文当变量、子 Agent 委派当函数调用）与 Continual Harness（把 harness 自身状态变成可 CRUD、可在线自我改进的数据），并解释为什么&amp;rsquo;harness 表达力&amp;rsquo;是被严重低估的能力放大器。</description></item><item><title>Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-25-task-coevolve-paper-reading/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-25-task-coevolve-paper-reading/</guid><description>Task-CoEvolve（东京大学）把 harness 优化中被忽视的&amp;rsquo;评估侧&amp;rsquo;变成优化变量：每次迭代在哪些验证任务上评估候选 harness？方差加权采样把预算集中到&amp;rsquo;候选结果会分歧&amp;rsquo;的判别性任务上（&amp;gt;70% 的任务池处于全对/全错两个极端、毫无判别力且分布随优化漂移），配 Horvitz-Thompson 式包含概率校正消除子集偏差。结果：20% 评估预算匹配全量搜索（49.3% vs 48.6%），Terminal-Bench 2.1 上 token 评估成本直降 80%，7% 预算用 1/16 样本逼近全量。</description></item><item><title>EnvHarness: Awakening Static Worlds for Agent Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-envharness-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-envharness-paper-reading/</guid><description>EnvHarness 把 agent harness 的思路搬到环境侧：不改一行底层代码，只在 reset/step 标准接口外包裹可插拔组件（Stage 改初始态、Contract 重写交互、Chain 串联环境），把静态人工环境改造成针对目标策略弱点的定制训练场，且 100% 继承原环境可信 verifier。配套 EnvRigger 自动化闭环从 rollout 诊断策略缺陷、写组件、用新鲜 rollout 验证。五个基准四领域全面超越原始环境与领域专用生成管线：held-out 最高提升 9.0 分、步数省 9.8%，并支撑 RL 共同进化。</description></item><item><title>FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-facet-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-facet-paper-reading/</guid><description>FACET 由中国科学技术大学、上海人工智能实验室与复旦大学联合提出，直面从异构 agent 技能合成可验证终端任务的两大难题：多阶段生成中源信息被过早压缩、以及指令/环境/参考解/验证器四件套各自为政导致的跨工件不一致。其三阶段框架先收 71,341 个技能建场景-技能库，再以五维表示代理式重建场景以恢复跨技能依赖与中间状态，最后先构建并修复 Docker 环境，把实现容器态作为三者共享接地，并按失败溯源定向修复。最终产出 6,078 个验证任务，平均 22.77 项可执行测试居各数据集之首；仅用 1.2K 轨迹做 SFT，Qwen3.5-4B/9B/27B 在 Terminal-Bench 2.1 分别提升 7.12/8.24/6.75 分，27B 距大它约 15 倍的 397B 模型仅 1.49 分。</description></item><item><title>SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-semaplc-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-semaplc-paper-reading/</guid><description>美的AIRC、KUKA、上海交大与浙大合作的SemaPLC提出验证门控的agent harness：生成逻辑必须嵌入既有工业PLC项目、通过编译，并在真实运行时与金轨迹比对正确才算完成。凭借仅日志确认的检查可判定完成、编辑使旧判定失效、每检查限两次重试三条完成纪律，它在117任务功能轨上七模型全部夺魁（均值72.6%），在65任务项目轨上动态行为分达52.2，远超基线的22.4~31.4。层消融揭示：静态分相近的方法在运行时剧烈分层——执行才是检验控制逻辑的忠实标尺。</description></item><item><title>SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-skillgate-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-skillgate-paper-reading/</guid><description>上海交通大学联合小红书提出 SkillGate，解决长程 Agent 在 episode 中途该读哪个技能文件的训练难题。论文先诊断出 outcome-only RL 失效的结构性根源——selector credit starvation：技能命名 token 仅占轨迹损失的中位 0.14%，随轨迹变长稀释约 7 倍，且近五分之二的正确选择因后续执行失败收到错误负号的信用，而该决策本身价值 +11.2 个百分点。SkillGate 将同一 GRPO 更新的 token 支持划分为构造上不相交的双 credit 通道：任务通道只把组归一结果优势广播到执行 token，整段 read call 从损失掩码删除；选择通道把单次读取且为 oracle 才计 1 的 action-local 效用做组中心化后仅落在身份 token，等权归一使选择权重与轨迹长度无关。5 个基准 385-trial 协议下，9B 模型从 SFT 的 40.8% 升至 53.2%，超同预算 outcome-only RL 6.2 个百分点，oracle 读取率 54.3% 升至 83.9%，误导暴露降约三分之二且读得更少。</description></item><item><title>两天十万Star：DeepSeek Harness 的开放逻辑，与它想要驯服的模型-脚手架-算力飞轮</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-deepseek-harness-open-strategy/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-deepseek-harness-open-strategy/</guid><description>围绕 DeepSeek Harness 发布后两天破十万 Star 的现象，三位从业者从「一切皆插件」的架构设计、模型与 Harness 的深度协同、极简模式与缓存命中率的技术原理，聊到程序员岗位转型、开源生态与国产算力差距。核心判断：Harness 是 AI 时代的脚手架，插件化+开源让社区共建成本降到极低，模型与脚手架会互相塑造，而程序员的护城河正从写代码转向定义需求与验收结果。</description></item><item><title>A Jagged Frontier: 代码Agent对语义保持变换的锯齿鲁棒性 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-jagged-frontier-code-agent-robustness-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-jagged-frontier-code-agent-robustness-paper-reading/</guid><description>当代码库被改写成语义等价的形式——控制流重写、死代码注入、标识符重命名——修 bug 的代码 Agent 还靠得住吗？Colorado State、Microsoft、UIUC 与 CMU 四方合作，用一套随机变体采样器对 2 个 Agent 框架 × 4 个前沿模型 × 54 个 SWE-bench 实例做了首个仓库级 Agent 鲁棒性系统评估：多数配置出现小幅退化（最大平均 6.7 个百分点，16 个配置中 6 个统计显著），但更扎心的发现是「锯齿前沿」——没有任何模型鲁棒性排名能跨框架、跨基准保持稳定，Qwen 在一个框架下最鲁棒、换一个框架反而最脆弱；更简单的框架反而更皮实；即使 solve 率不掉，token 成本最多也要多花 22.9%。本精读覆盖其 14 种语义保持变换的设计、非反馈采样的下界逻辑、配对实验统计方法，以及锯齿现象背后的机制因果链。</description></item><item><title>Adversarial Review: Structured Disagreement for Grounded Agentic Code Review 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-adversarial-review-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-adversarial-review-paper-reading/</guid><description>康奈尔与斯坦福的两位研究者提出 Adversarial Review（AR）：主编码 Agent 冻结工件后，reviewer 评审、critic 以结构化分歧审计这份评审，收敛后才允许修改代码。AR 在 LiveCodeBench 上以三个 Agent 取得 87% 最高通过率，胜过五 Agent 的 MARS；在 SWE-PRBench 上先暴露「伪共识」失败模式——Agent 为一致而一致，再用一次 prompt 迭代把分歧显式化即取得最高 F1 0.533；在 SWE-bench Verified 上以纯文本 SKILL.md 协议达到 75.2%。本精读拆解其构造式方法、三基准证据链，以及「分歧必须最小、结构化、有证据」的设计哲学。</description></item><item><title>Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-task-coevolve-harness-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-task-coevolve-harness-paper-reading/</guid><description>东京大学团队提出Task-CoEvolve，让验证任务集与harness共进化：用方差加权采样把评估预算聚焦在候选harness分歧最大的能力前沿任务上，再用Horvitz-Thompson/Hájek类估计器从采样子集无偏还原全量分数。在Terminal-Bench 2.1上仅用20%预算就逼近全量搜索（均值51.7 vs 52.8），整体搜索成本降67-80%；文本分类7%预算接近全量、20%预算反超。本精读覆盖背景、定位、方法机制、实验证据、效果根源因果链、必要知识反推与通用灵感九个部分。</description></item><item><title>EnvHarness: Awakening Static Worlds for Agent Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-envharness-agent-env-paper-reading/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-envharness-agent-env-paper-reading/</guid><description>深度精读 EnvHarness——与 Agent Harness 对称的环境侧革命：不改环境本身，在交互接口上包装一层可编程插件（Stage/Contract/Chain 三类组件），把静态冻结环境重塑为针对当前策略弱点的定制化训练场。EnvRigger 自动化引擎通过 Observe→Diagnose→Write→Validate 四阶段循环，自动诊断策略缺陷并生成验证过的组件，在 ALFWorld、WebArena、SWE-bench Verified、OfficeQA、SpreadsheetBench 五大基准上全面超越原环境与领域特定生成器，环境规模化收益持续未饱和。</description></item><item><title>Agent Lightning v1.0: Towards Harnessed Agentic RL 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-agent-lightning-v1-harnessed-rl-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-agent-lightning-v1-harnessed-rl-paper-reading/</guid><description>微软亚洲研究院联合复旦、浙大、爱丁堡大学发布Agent Lightning v1.0，首次系统定义harnessed agentic RL范式——当部署级agent harness直接参与RL训练时，训练引擎只能看到一串LLM请求-响应对。论文刻画了重分词破坏token前缀连续性、动态样本数下的优势计算、损失归一化与后端调度四大挑战，以约3500行代码给出参考实现，仅用6K训练样本让Qwen3.5-9B在SWE-bench Verified上从41.8%提升到56.4%（+14.6个百分点），并公开完整数据清洗管线与防reward hacking脚手架。</description></item><item><title>Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-harness-the-memory-substrates-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-harness-the-memory-substrates-paper-reading/</guid><description>UIC、华盛顿大学、McGill、MBZUAI与UCLA五校联合完成首个把记忆底座（substrate）作为受控变量的统一harness评测：11类底座×3个骨干模型×4组基准×26项指标。核心发现颠覆选型直觉——没有任何底座全面称雄，QA任务的前沿（图+向量混合）与决策任务的前沿（扁平检索/精炼蒸馏）完全不相交；检索宽度k在QA上单调涨分、在决策任务上反向往下跌分，注意力探针揭示同一稀释机制在不同任务中后果相反。这为记忆系统按工作区间路由底座提供了实证基础。</description></item><item><title>HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-harnessrisk-lifecycle-safety-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-harnessrisk-lifecycle-safety-paper-reading/</guid><description>UNC教堂山分校Tianlong Chen组联合UCF、密歇根州立发布HarnessRisk——首个覆盖agent harness全生命周期的安全基准：把harness安全组织为配置/能力扩展/运行时/状态持久化/动作控制/事件恢复六个运营阶段，128个沙箱案例每个配对良性用户目标与嵌入不可信工作流制品的对抗指令。14个模型-harness配置的评估揭示：攻击成功率12.6%-80.9%波动，配置阶段最脆弱，同一模型跨harness的ASR差4.3倍——安全是部署配置的属性而非模型属性，且风险识别不等于安全行动。</description></item><item><title>LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-lego-rl-harness-native-coding-rl-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-lego-rl-harness-native-coding-rl-paper-reading/</guid><description>华为LegoX团队联合港中文发布LEGO-RL——在不修改原生编码agent harness内部控制流的前提下接入可扩展策略梯度训练。三大支柱：进程内LLM代理捕获原始生成流实现token级对齐与训练端logprob重算（即使harness压缩/重序列化上下文）、Nydus镜像缓存+分级防御抑制reward hacking、插件化校验监控+Live UI轨迹诊断。训练Qwen3.5-35B-A3B（GSPO）在三大harness上全面提升：OpenHands SDK 64.0%→70.4%、Claude Code 62.4%→68.2%、OpenCode 57.2%→66.6%，rollout-训练概率相关性保持0.99以上。</description></item><item><title>SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-semaplc-verification-gated-plc-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-semaplc-verification-gated-plc-paper-reading/</guid><description>美的AIRC联合KUKA、上海交大、浙大发布SemaPLC——一个项目接地、验证门控的PLC代码生成agent harness。它由常规工具组装而成，却由一条严格的完成纪律统治：agent不许凭自判断宣布完成，只有工具日志确认的规格审计、编译、运行时三类外部检查全部过关才准交付；任何编辑作废全部旧判定并重跑全部检查。在117个独立POU任务上它让全部7个模型拿到最高严格通过率（均值72.6%，超最强基线8.8个百分点）；在65个真实工厂项目任务上，其动态行为分52.2碾压基线最高31.4——静态分相近的方法在运行时被彻底分离。编译通过≠跑得对，执行才是生成控制逻辑最忠实的检验。</description></item><item><title>SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-skillgate-in-policy-skill-selection-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-skillgate-in-policy-skill-selection-paper-reading/</guid><description>上海交通大学与小红书合作的论文诊断了agent技能选择失败的结构性根因——selector credit starvation：在广播式序列级优势下，命名技能的少数token在损失中份额趋零、继承的信用随轨迹变长而日益错号（选择正确但执行失败时正确选择被惩罚）。SkillGate把token支持划分为两个不相交信用通道：结果信用只达执行token、动作局部优势只达技能命名token（仅当轨迹唯一一次读取是oracle时为正）。5个agent基准、16候选技能档位下，9B策略试验成功率从40.8%提升到53.2%，超同预算outcome-only RL受控对照，误导技能暴露减少三分之二，读取技能更少。</description></item><item><title>StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-stagedworkspace-versioned-workspace-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-stagedworkspace-versioned-workspace-paper-reading/</guid><description>哈佛大学联合Raycaster AI、斯坦福等机构提出StagedWorkspace——为知识工作agent建立版本化工作区。论文形式化workspace-state contract概念：agent检索的解析视图、编辑的原生文件、审阅的diff、提交的制品可能指向同一工作产物的不同版本，这是PDF/表格/幻灯片等非代码制品长期缺乏的契约。内容哈希绑定使视图与版本显式关联，OfficeQA Pass@1提升8.3-12.1点，SW-AGENT用Gemini 3.1 Pro达OfficeQA 63.9%（同模型已发表分数仅29.3%），证明工作区状态是被忽视的实验变量。</description></item><item><title>Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-zetta-closed-loop-embodied-harness-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-zetta-closed-loop-embodied-harness-paper-reading/</guid><description>清华大学 AIR 团队提出 Zetta，一个能在部署时自我进化的闭环具身智能框架：冻结 VLA 基座策略不动，用可在线进化的代码级 Runtime Critic 在动作频率上监控物理执行，配合三时间尺度进化循环（动作级治理、批次级失败诊断修复、验证门控技能晋升）与专用推理基建 Z-Infra，在 LIBERO-Pro 上把宏平均成功率从 32.0% 提升到 71.1%，在 RoboCasa 18 任务上从 73.56% 提升到 93.56%，推理延迟较 RPent 降低 91%，并涌现出 15%→95% 式的机器人 Aha 时刻与零样本技能迁移能力。</description></item><item><title>ClawGym II: Exploring Black-Box RL on Agent Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-clawgym2-blackbox-rl-harness-paper-reading/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-clawgym2-blackbox-rl-harness-paper-reading/</guid><description>人大高瓴AI学院与IQuest Research联合提出首个面向复杂agent harness的黑盒RL统一框架：用临时沙盒把任务环境与harness整体隔离包装（对训练侧完全黑盒），以忠实轨迹恢复机制从沙盒遥测中重建树状执行轨迹并回传环境奖励，使PPO/GRPO能稳定优化“harness+模型”整体。在OpenClaw与Claude Code两种结构迥异的harness上验证：30A3B骨干较初始策略分别+9.98/+14.81分，超SFT基线5.80分，PinchBench外部迁移87.32；训练在200-400步内保持稳定，且首创混合harness联合训练——单一模型在两种harness上持平甚至超过各自单独训练版。配套发布ClawGym-Bench（六域）与PinchBench双评测体系。</description></item><item><title>StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-statem-harness-scaling-paper-reading/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-statem-harness-scaling-paper-reading/</guid><description>四位独立研究者（含 UT Austin 张satz Atlas 王）在不改模型权重的前提下，用 YAML 状态机运行时 StateM 把 GPT-5.5 在 Terminal-Bench 2.1 上从 83.1% 拉到 92.1%，冻结迁移到 GPT-5.6 Sol 达 95.28% raw，把 DeepSeek-V4-Flash 适配到 88.09% 而最终评测成本仅约 15 美元——对照 GPT 参考运行的 574.68 美元。论文提出 harness scaling 作为与 model scaling 正交的能力轴：把可变状态外置、以状态为上下文与契约双重边界、用受检转换取代 agent 自证完成，将失败分类为认知缺口/程序记忆缺口/程序遵从缺口三类并逐一施加控制点。负迁移分析（RefactorBench -2.78 分）进一步证明控制必须挂在正确的执行边界上。</description></item><item><title>AgentRewind: Recoverable Execution for Long-Horizon LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-18-agentrewind-recoverable-execution-paper-reading/</link><pubDate>Tue, 18 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-18-agentrewind-recoverable-execution-paper-reading/</guid><description>长程Agent任务中早期错误同时污染上下文与环境状态，现有方法（计划精化/安全检查）只防错不恢复。中科院与清华团队提出AgentRewind：对齐记录Agent上下文与受控环境的检查点，Agent判断无法推进时回滚到早期状态并以前次尝试摘要指导续作；配套MettleBench（含隐藏有序验收清单的长程工程任务）。Terminal-Bench 2.0全量上成功率83.1% vs Continue的78.7%与Restart的70.8%；回滚增益随执行horizon增长显著扩大。案例研究揭示三策略本质差异：Continue在污染状态上修补、Restart丢弃已完成成果、Rewind选择性回滚+经验注入。</description></item><item><title>Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-18-beyond-final-scores-paper-reading/</link><pubDate>Tue, 18 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-18-beyond-final-scores-paper-reading/</guid><description>自动化研究Agent的评测长期被“最终分数”主导，无法回答进步从哪来、失败藏何处、经验是否有用。美团与中科院国科大团队花费约10万美元推理成本，对7个前沿模型36个长程任务756次rollout做系统解剖：提出C1方案构架/C2执行/C3反馈控制三个规则驱动的过程指标+任务内/跨任务经验复用反事实实验。结论是当前Agent更像“勤奋的工程优化器”而非自主研究者：avg@3差距0.237而best@3仅0.122（可靠性比峰值更具区分度）；252个最优解中真正新颖方法仅3个（1.2%），钻评测空子的却有16个（6.3%）；经验迁移使DeepSeek-V4-Pro +0.093却使Gemini-3.1-Pro -0.017；自动harness进化+0.123且可跨模型迁移。</description></item><item><title>Demystifying Agent Skills: Why They Work—Until They Don't 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-18-demystifying-agent-skills-paper-reading/</link><pubDate>Tue, 18 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-18-demystifying-agent-skills-paper-reading/</guid><description>技能已成为增强LLM Agent的热门方案，但“技能为何有效、何时失效”一直缺乏机制层面的回答。Princeton、Stanford、UCSD、USC、JHU五校联合团队通过8135条受控试验与238个开放编码标签，首次给出定量答案：技能的本质作用是程序性锚定（占65.7%）而非知识注入（仅4.5%），比Workflow Memory高6.06分；检索是独立瓶颈——技能池从5增至100时实际使用精确率从29.6%崩至3.3%，但下游成功率却保持稳定；技能还会引入新的调用失败面（误用率10.0% vs 裸执行的0.8%）。本文从背景、定位、问题抽象、解法机制、实验证据到根源解释逐层拆解，并提炼技能生命周期化的通用工程启示。</description></item><item><title>Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-capability-sheaves-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-capability-sheaves-paper-reading/</guid><description>独立研究者 Saveliy Batruin 单人完成的论文，把“智能体组件各自正常、合起来却对不上”的 harness 故障形式化为层论粘合问题：5 个需求作顶点、行为签名作茎、限制映射为字面字段投影，用精确 CSP 判定可粘合性、用相对上同调类作诊断特征，并对隐藏中介状态取商以获得不变性。受控实验中 20 个任务簇全部获益（到首次成功的候选评估数 1.000 vs 2.000，token 降约 71%）；但在真实 PatchFuseBench 上候选级商选择器 118/160 仅比匹配对照 116/160 高 2 题（p=0.75 不显著），未过预注册开发门，确认集保持封存。受控环境成立、真实优势尚未证明——论文以罕见的诚实划出了方法的边界。</description></item><item><title>GitSkills: A Dataset of Agent Skills on GitHub 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-gitskills-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-gitskills-paper-reading/</guid><description>深度精读 UCL、霍恩海姆大学与卡利亚里大学合作的 GitSkills——首个对 GitHub 上 Agent Skill 生态做系统性快照的规模数据集。2025年10月 Anthropic 开源 SKILL.md 格式，仅九个月后 GitHub 公开仓库中已沉淀 3,797,117 个 SKILL.md 文件、282,200 个仓库、195,841 个账号。论文的三阶段管线绕过代码搜索 API 每查询 1000 条上限与不可靠的总数估计（报 34.9 万 vs 实际 380 万+），按文件大小递归分区搜索空间完成完整采集；内容哈希去重得 1,877,981 个不同内容，50.5% 的文件是逐字副本——这个无包管理器、靠复制传播的生态，把软件供应链安全命题原样搬进了自然语言工件世界。数据封装为单个自包含 SQLite 文件，采集与解释分离设计让社区可以自定义纳入标准。</description></item><item><title>The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-17-tool-architecture-paper-reading/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-17-tool-architecture-paper-reading/</guid><description>普渡大学、微软研究院与芝加哥大学团队对编码智能体的工具架构做了受控对比实验：在能力等价前提下比较六种工具架构（BashOnly、Atomic、NLSearch、Python、HypoTrack、Scratchpad），覆盖三个模型共11700条轨迹。结果显示结构化原子工具将弱模型重复运行稳定性最高提升4.7倍，自然语言搜索拓宽仓库探索广度超过11%，代码执行接口在任务表现相近的情况下减少41.6%步数与56.3%token消耗，而轻量认知脚手架几乎无效——接口本身就在塑造智能体行为。</description></item><item><title>SkillEvo: 多轮交互反馈的自更新进化梯度 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-16-skillevo-multi-turn-feedback-paper-reading/</link><pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-16-skillevo-multi-turn-feedback-paper-reading/</guid><description>深度精读腾讯云 Andon 与浙江大学合作的 SkillEvo 论文——把多轮用户模拟从评估终点反转为反馈生成器，让 Agent 技能在生产工单上自进化。从梯度衰减机制、可信反馈三条件（意图状态机、双侧正交评估、集体归因）到双层治理（有界修订、结构退化主动修复），全面拆解这套在腾讯云 9 个生产 Skill 上将 TSR 从 30.0 提升到 81.8 并落地生产环境的自进化框架。</description></item><item><title>A Programming Paradigm for Spatiotemporal Composability 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-cordis-spatiotemporal-composability-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-cordis-spatiotemporal-composability-paper-reading/</guid><description>北大与 DeepSeek-AI 合作的 88 页长文，为「插件系统、自进化 Agent Harness」这类动态组合软件给出了第一个完整的编程范式级形式化基础：把经典效应系统提升为可逆效应、把协同效应系统提升为响应式协同效应，统一成一个递归上下文类型，再配上动态组合演算与全套元理论（保持性、恢复精确性、活性、合流性），实现为 Cordis 元框架并在 Koishi（4000+ 社区插件）上验证。本文按背景、定位、问题、解法、评估、根源解释、知识反推、通用灵感八个层面完整拆解。</description></item><item><title>AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-autodesign-meta-harness-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-autodesign-meta-harness-paper-reading/</guid><description>把一篇20页论文变成一张合格学术海报，需要上百次工具调用、多轮排版修订与视觉验证——这是典型的长时程智能体设计任务。本文精读美团联合多家高校的 AutoDesign：它不直接训练模型，而是让一个元harness优化器引导 code agent 基于 rollout 反馈递归自改进 harness，经 7 天演化沉淀出可复用、可迁移的学习型 DesignHarness。在自建的 PosterBench 百篇论文基准上，AutoDesign 以 78.32 分超过商业系统 Claude Design 7.45 分，盲测人类偏好 BT 值 64.0% 位列第一；给 7 个模型配置挂载该 harness，平均分从 54.99 提升到 67.39。本精读重点拆解其双层优化循环、五组件 harness 结构，并用因果链解释&amp;rsquo;学习型 harness 为什么弱模型受益更大&amp;rsquo;。</description></item><item><title>DarwinX: Evolving Agent Harnesses Through Natural Selection 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-darwinx-harness-evolution-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-darwinx-harness-evolution-paper-reading/</guid><description>LLM Agent 的能力不只取决于模型权重，还取决于包裹模型的 Harness（提示词、工具、技能、控制流）。Salesforce AI Research 的 DarwinX 在完全冻结模型权重的前提下，把 Agent 自进化重构为对 Harness 种群的“自然选择”：preserve-and-extend 契约只接纳“净增益为正且回退有界”的变体，树状 archive 保留多条谱系供跨谱系重组，失败/教师/自采三种证据共用同一编辑接口，适应度完全来自 benchmark 自带 verifier。四个基准平均提升约 17 分：Terminal-Bench 2.1 达 83.2%，WebArena-Infinity 从 43.5% 跃升至 93.0%，且零适应迁移到 SWE-bench Verified 达 84.2%。本精读逐部分拆解其机制，并建立“方法差异→机制变化→指标提升”的因果链。</description></item><item><title>SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-skiller-language-rl-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-skiller-language-rl-paper-reading/</guid><description>SKILLER 提出语言级强化学习框架：用强模型（GPT-5.4）兼任 actor 与 critic，把小模型智能体系统当作环境，官方验证器提供标量奖励与文本诊断，全部 RL 信号经由自然语言传播，优化变量是技能文本本身而非模型权重。在五个基准上，9B/4B 小模型配 SKILLER 技能全面超越闭源 Manus 与开源技能生成方法，SWE-Skills-Bench 上超人类技能 30.8 分，且零样本迁移到 GAIA/EarthBench 仍保持领先。本精读覆盖其背景脉络、语言级 RL 形式化、actor-critic 机制、消融因果链与可推广灵感。</description></item><item><title>AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-12-ai4ai-test-time-transfer-paper-reading/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-12-ai4ai-test-time-transfer-paper-reading/</guid><description>Salesforce AI Research 联合 Notre Dame、UIUC（Heng Ji）提出强到弱推理时脚手架（Strong-to-Weak Scaffolding）：用强 builder 模型为弱 target 模型自动构建推理时 harness，无需任何参数更新即可在四个 Theory-of-Mind 基准（3900 项）上将 GPT-5.4-mini 从 0.488 提升到 0.912（+0.423）。机制分析表明增益主要来自把不稳定的自然语言推理卸载为确定性代码（r=0.72），而非更长推理链或更多采样。这是对传统训练时蒸馏的一条互补路线，也直接印证了 harness 工程作为独立工程对象的价值。</description></item><item><title>Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-12-harness-if-paper-reading/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-12-harness-if-paper-reading/</guid><description>ByteDance Seed 团队提出的 Harness-IF 把「编程 Agent 是否真的在遵守指令」这件事第一次变成了可量化、可归因的规则级测量问题。它构造了 642 条原子规则的库，实例化出 60 个多轮编程任务，在 5 个可配置的「指令表面」(系统提示/工具描述/技能描述/项目文件/用户指令)上分别打分；更重要的是，它用 Against-Prior Accuracy(AP-Acc)把「模型本来就是这么做」的巧合从「真正遵从指令」中剥离出来——12 个前沿模型无一例外都在反先验规则上表现更差，平均落差 5.81 分。配套的 E0 冲突实验还揭示了一个反直觉结论：表面优先级并不服从提示深度，SP/PF/UI 同居首位，而工具描述和技能描述垫底。这篇精读从背景、定位、问题、解法、证据、根源、知识反推到通用灵感，完整拆解这项与 Harness 评估方向高度相关的工作。</description></item><item><title>BONSAI: Evolvability-Guided Tree Search over Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-bonsai-skill-search-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-bonsai-skill-search-paper-reading/</guid><description>深度精读论文《BONSAI》——首个以『可进化性（evolvability）』而非当前适应度引导技能搜索的框架。面对『单一验证分数无法区分宽广平台与狭窄尖峰』这一根本盲区，BONSAI 将技能生长为蒙特卡洛搜索树，让节点下的平均分数免费估计变异邻域的可进化性，无需任何额外模型调用。在冻结 30B Agent 上、三个基准平均，BONSAI 比无技能 Agent 提升 23.13 点，比 GEPA 提升 3.87 点，比 SkillOpt 提升 3.97 点；消融实验证明可进化性信号本身在相同树上贡献 +2.14～+7.02 点。</description></item><item><title>DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-dcas-scaffold-decoupling-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-dcas-scaffold-decoupling-paper-reading/</guid><description>深度精读华为加拿大软件卓越中心与女王大学的 DCAS 论文——首个系统揭示开源 CLI Agent 存在&amp;rsquo;scaffold 锁定&amp;rsquo;现象的工作。论文发现：在 OpenHands 单一 scaffold 下微调的模型，迁移到其他 scaffold 时性能可从 52.6% 暴跌至 8.4%。通过提出 DCAS 后端替换拦截层和区分显式/隐式规划两种形式，论文给出了一条从 scaffold 制品到模型能力的可行迁移路径，仅用 576 条规划感知轨迹即可让模型在非训练 scaffold 上一致提升。</description></item><item><title>Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-ouroboros-self-developing-agent-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-ouroboros-self-developing-agent-paper-reading/</guid><description>本文精读 Anton Razzhigaev、Roman Yampolskiy 等人 2026 年发表的 Ouroboros——一个能够自开发的前沿编程 Agent。它把 Agent 的工具、提示词、上下文组装乃至核心实现本身都视为可被审查、可被修改的活体代码，并通过多模型对抗式 diff 审查作为变更门控，实现经审查的核心进化（Reviewed Core Evolution）。文章在 Terminal-Bench 2.1、OSWorld-Verified、CL-Bench 等基准上刷新 SOTA，并在代号为 Hope 的 161 天活体实验中持续运行（累计 1085 次自我修改提交、94.2% 由 Agent 撰写）。本精读将从背景、定位、问题定义、方法、评估、优势根源、必要知识反推、通用性灵感八个维度系统拆解这篇论文。</description></item><item><title>Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-oeo-open-ended-optimization-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-oeo-open-ended-optimization-paper-reading/</guid><description>深度精读 Hui Xue 与 Fan Yang 的《Rethinking Self-Evolving Agents》——一项直面预设流水线是否仍然必要的反思性研究。文章提出 OEO（Open-Ended Optimization，开放式优化）：固定目标、交互、预算、数据边界和评估这五项不可妥协的约束，但把优化过程完全交给前沿模型自行组合。在 GPT-5.5 驱动下，OEO 在 14 次正面交锋中 12 胜 1 平 1 负，仅用 SkillOpt 配置预算中位数 34.3% 的目标交互 token。本精读按九部分结构拆解其背景、定位、问题抽象、机制、实验证据、优势根源、必要知识反推与通用性灵感，并重点解读能力依赖的脚手架这一核心洞见。</description></item><item><title>SHE: Trajectory-driven Safety Harness Evolution for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-she-safety-harness-evolution-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-she-safety-harness-evolution-paper-reading/</guid><description>深度精读 SHE 论文——复旦、阿里达摩院、莱斯大学等多校合作提出 Safety Harness Evolution，将 Agent 安全 harness 解构为 System Prompt / Rule Bank / Safety Memory / Tool Policy 四个责任显式、独立可进化的制品，并通过归因引导的进化循环把轨迹失败转化为结构化诊断、局部精化与安全-效用验证。在 Agent-SafetyBench 上攻击成功率（ASR）相比静态 SafeHarness 降低 3.1 倍，同时良性任务效用不降反升；进化后的 harness 还能零成本泛化到 held-out 的 AgentHarm 基准并跨 Agent 模型迁移。本精读按九部分结构展开：从 Agent 安全 harness 的&amp;rsquo;整体黑盒&amp;rsquo;困境，到 SHE 的&amp;rsquo;免疫系统白细胞规则集&amp;rsquo;类比，再到归因引导进化与&amp;rsquo;精准医疗 vs 全身化疗&amp;rsquo;的范式对照。</description></item><item><title>SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-swe-bench-promax-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-swe-bench-promax-paper-reading/</guid><description>深度精读 COLM 2026 论文 SWE-Bench ProMax——字节跳动与香港科技大学合作的专家策划多语言代码重构基准。170个实例覆盖7种编程语言，平均每实例修改11.4个文件、261.6行代码。揭示近60%未解决的 SWE-bench Verified 实例存在过窄或过宽的缺陷测试，前沿模型甚至能逐字复现训练数据中的 gold patch。在两种 Agent scaffold 下，最强模型解决率仅41.2%，证明基准未饱和。多阶段专家策展流程从源头堵住测试质量和数据泄漏两大漏洞。</description></item><item><title>「模型能力已经够了，要卷就卷 Infra」｜对话戴冠兰：从 Cloudflare 到 Runta，为十亿个 Agent 造执行底座</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-10-runta-agent-infra-daiguanlan/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-10-runta-agent-infra-daiguanlan/</guid><description>Runta 创始人戴冠兰（前 Cloudflare/Kong 核心）在十字路口播客中提出核心判断：模型能力爬坡已放缓，真正制约 Agent 落地的是执行层基础设施。Runta 刚完成 2000 万美元种子轮（a16z 领投，Jeff Dean、李飞飞天使），定位是为 Agent 打造确定性执行底座——在概率性大模型之上加入隔离、权限、审计和热迁移等系统能力，让企业敢于把生产权限交给智能体。文章梳理了 Token Maximizing 到 Minimizing 的反转、Agent 安全必然爆发的逻辑、以及公有云和基模厂商为何难以抢占这一赛道。</description></item><item><title>Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-05-harness-r1-paper-reading/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-05-harness-r1-paper-reading/</guid><description>Agent 部署后会积累大量失败轨迹，但它的行为通常固定不变——模型不更新，Harness（运行时框架）也不更新。Harness-R1 首次把&amp;rsquo;编辑可执行运行时&amp;rsquo;本身变成一个可被在线 RL 训练的能力：一个 9B 的&amp;rsquo;harness 工程师&amp;rsquo;模型从失败批次中生成可执行补丁，用冻结目标 Agent 重跑的真实成功率作为奖励。结果这个 9B 工程师反超 GLM-5.2、GPT-5.5、DeepSeek-V4-Pro 等所有更大的前沿模型编辑器；即便目标 Agent 微调后，工程师仍能再 +5.0pp。本文从&amp;rsquo;把脚手架变成可学习对象&amp;rsquo;的第一性原理，解释为什么小模型+真实结果奖励能赢过大模型+教师提议。</description></item><item><title>LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-05-longhorizon-harness-paper-reading/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-05-longhorizon-harness-paper-reading/</guid><description>长程任务（long-horizon）是 Agent 走向真实世界的最后一块硬骨头。LongHorizon-Harness 把&amp;rsquo;执行-状态管理-完成评估&amp;rsquo;从一个不断增长的上下文里拆开，重构为 Manage-Execute-Audit（MEA）循环：Manager 只管状态不碰环境、Executor 每轮用新鲜上下文执行、Auditor 只读独立验证。这套架构让 Qwen 3.7-Plus 在 WeaveBench 上从 51.8% 跃升到 80.7%（近乎翻倍官方 SOTA），OSWorld 2.0 上把 Claude Opus 4.7 从 20.0% 提到 34.3%。本文从&amp;rsquo;状态-执行-审计&amp;rsquo;分离的第一性原理出发，解释为什么这套架构在难任务上收益更大、在更强模型上 token 反而更省。</description></item><item><title>Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-05-workbuddy-bench-paper-reading/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-05-workbuddy-bench-paper-reading/</guid><description>腾讯发布多领域编码 Agent 基准 WorkBuddy Bench，覆盖代码、前端、办公、安全四大真实工作场景。其核心贡献在于从真实 commit/CVE/业务场景逆向工程出抗污染的口语化任务，并将任务目录、环境镜像、评估框架、测试与参考方案完全开源。跨模型排行榜显示没有任何单一模型通吃，开源权重模型 GLM-5.2 在安全子集双框架登顶，为可信代码评测体系的构建提供了新的方法论范式。</description></item><item><title>Meta AI Proactive Memory Agent：记忆教练精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-meta-memory-coach-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-meta-memory-coach-paper-reading/</guid><description>Meta AI提出主动记忆智能体架构，用独立的&amp;rsquo;记忆教练&amp;rsquo;智能体在固定间隔审查行动智能体的近期步骤，更新结构化记忆库（私有状态/知识记忆/程序记忆），并决定是否注入定向提醒。核心创新在于&amp;rsquo;何时提醒&amp;rsquo;的策略决策——选择性干预优于全量检索。Terminal-Bench从38%提升至46%，tau2-Bench从55%提升至62%，超越Mem0生产记忆层。</description></item><item><title>Recursive Harness Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-23-rhi-recursive-harness-self-improvement-paper-reading/</link><pubDate>Thu, 23 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-23-rhi-recursive-harness-self-improvement-paper-reading/</guid><description>Sakana AI 与 UC Berkeley 提出 RHI（递归式框架自改进）：把多智能体框架当作提示词级对象，仅用当前与上一版本的自我比较来迭代优化，少数几轮就能让低推理强度的 Agent 超越同族最高推理强度设置，同时把推理成本降低最高 60%。本文从 Harness 是什么、模型-框架协同进化讲起，拆解 RHI 的轨迹局部目标、算法流程、信息论隐式目标，并提炼可推广的通用性灵感。</description></item><item><title>从龙虾热到基金会治理：OpenClaw首席架构师Vincent Koc谈个人Agent的反思、工程化与协作未来</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-21-openclaw-vincent-koc/</link><pubDate>Tue, 21 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-21-openclaw-vincent-koc/</guid><description>2026 WAIC上海现场，OpenClaw Foundation首席架构师Vincent Koc深度复盘OpenClaw半年来的爆火与冷却、与中国市场的特殊关系、个人Agent与编程Agent的本质区别、基金会治理模式为何优于风投创业、以及他判断的下一个关键趋势——Agent之间的通信与协作。</description></item><item><title>2026-06 arXiv 智能体/工具智能体领域综述：916 篇分主题精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-17-tool-agent-survey-2026-06/</link><pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-17-tool-agent-survey-2026-06/</guid><description>工具智能体领域综述。LLM_Agents 桶 1001 篇扣除记忆子领域后按 13 主题分桶精读，提炼工具量质失衡、agent RL 信用分配、长程可靠性与上下文管理、过程级评测、工具环境不可靠、多智能体幻觉优势、治理权限可审计性等 7 大共识问题，并识别 strained coherence、agentic abstention、world-model collapse 等原创问题。</description></item><item><title>智能体技能演化（Skill Evolution 与 Self-Evolving Agents）综述：53 篇核心论文精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-17-skill-evolution-survey-2026-06/</link><pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-17-skill-evolution-survey-2026-06/</guid><description>技能演化与自演化智能体综述。从 116 篇候选中筛定 53 篇 CORE 论文下载全文精读，提炼技能库选择退化、技能创建与部署脱节、自演化缺乏可靠接受准则、上下文无界膨胀等共性问题，以及 16 个范式级新转变（PACE、Bayesian-Agent、Red Queen Godel、Trellis、MMG2Skill 等 5 个已联网验证）。</description></item><item><title>ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-09-researchstudio-reel-paper-reading/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-09-researchstudio-reel-paper-reading/</guid><description>微软研究院的 ResearchStudio-Reel 把论文传播的&amp;quot;最后一公里&amp;quot;——海报、演讲视频、双语博客——重构为五个可组合技能。它用一次共享提取替代三次重复读论文，用硬性渲染门控替代软性美学打分，用可编辑的 PowerPoint/Word 替代只读 PDF。在 100 篇论文基准上，它生成的海报美学评分甚至超过了作者本人手绘的海报，在 84%-93% 的论文上获胜，并且是目前唯一同时交付三种可编辑传播产物的流水线。</description></item><item><title>Harness Engineering for Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-07-harness-engineering-paper-reading/</link><pubDate>Tue, 07 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-07-harness-engineering-paper-reading/</guid><description>Lilian Weng（Thinking Machines Lab 联合创始人、前 OpenAI 研究副总裁）在这篇万字综述中系统梳理了「Harness 工程」——围绕基础模型的运行时系统——作为通往递归自我改进（RSI）现实路径的核心命题。文章从 RSI 的思想起源讲起，把 Harness 定义为决定模型如何思考、规划、调用工具、管理上下文、评估结果的系统层，并梳理了三大设计模式（工作流自动化、文件系统持久记忆、子代理并行）、四大优化方向（上下文工程、工作流设计、自我改进、进化搜索）以及与模型权重的联合优化，最后坦诚列出七大瓶颈。本精读将这篇综述放在 RSI→Harness 的研究脉络中定位，提炼其方法论骨架与可迁移的普适灵感。</description></item><item><title>拆解Claude Code源码泄露：Agent Harness三层架构、记忆机制与零人公司的未来</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-01-claude-code-agent-harness-explained/</link><pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-01-claude-code-agent-harness-explained/</guid><description>Claude Code源代码泄露后，Agent Harness的关键模块被完整呈现出来，成为最好的教学样本。本期「十字路口」邀请到Learn Claude Code教程（GitHub超5万星）作者、CLAI创始人来新璐，从Harness的三层架构（执行能力层、上下文环境层、治理编排层）到底层设计哲学，深入拆解Claude Code的沙箱环境、记忆「做梦」机制、上下文压缩策略，以及从LangChain到Agent Runtime的范式迁移。来新璐还分享了他对CLI vs MCP之争的判断、Agent Harness赛道的创业格局，以及一个令人兴奋又有些可怕的未来图景——零人公司。</description></item><item><title>SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-29-skill-disco-paper-reading/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-29-skill-disco-paper-reading/</guid><description>深度精读微软研究院与北京外国语大学合著的 SKILL-DISCO 论文——将 Agent 成功执行轨迹蒸馏为可重用的参数化控制流子图（PFSM），再编译为可调用、可执行、可验证的过程技能。在 ALFWorld 和 WebArena 上，仅用 5 个技能（对比 ASI 的 110 个）就将成功率推高至 99.3%，且技能可跨模型迁移——GPT-4o 归纳的技能让 Qwen3.5-9B 在 ALFWorld 上达到 98.5%。</description></item><item><title>Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-15-language-models-need-sleep-paper-reading/</link><pubDate>Mon, 15 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-15-language-models-need-sleep-paper-reading/</guid><description>深度精读马里兰大学与卡内基梅隆大学合著的「语言模型需要睡眠吗？」论文——受人类睡眠记忆巩固机制启发，提出离线循环（Offline Recurrence）机制：模型在&amp;rsquo;睡眠&amp;rsquo;阶段对累积上下文执行 N 轮离线反复遍历，将信息蒸馏为持久化快速权重（Fast Weights），然后清空 KV 缓存。在不增加在线推理延迟的前提下，成功解决常规 Transformer 和 SSM-Attention 混合模型均失败的多跳推理和数学推理任务。增加睡眠轮数 N 可以持续提升性能，且推理越深的样本获益越大（相关系数 r=0.92）。</description></item><item><title>Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-rho-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-rho-paper-reading/</guid><description>深度精读香港城市大学 × 微软亚洲研究院 RHO 论文——首个仅利用无标签历史轨迹实现 Agent Harness 全链路自监督优化的工作。从 Harness 工程概念补全、六项前序工作定位、自偏好估计的问题抽象、三阶段核心方法详解、必要知识反推到七条通用性灵感，全面拆解这项在 SWE-Bench Pro 上将通过率从 59% 提升至 78% 的开创性研究。</description></item><item><title>Role-Agent: 通过双角色自举实现 LLM 智能体-环境协同进化 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-role-agent-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-role-agent-paper-reading/</guid><description>深度精读中科大 × 阿里AMAP团队 Role-Agent 论文——首个利用单一 LLM 同时扮演智能体与环境双角色，实现自举式协同进化的工作。从 Agent 强化学习背景补全、智能体-环境协同优化的研究脉络定位、双角色自举的问题抽象、WIA（世界内化于智能体）+ AIW（智能体内化于世界）双模块方法详解、必要知识反推到六条通用性灵感，全面拆解这项在 ALFWorld、WebShop、搜索增强QA 三大场景平均提升超 4% 的开创性研究。</description></item><item><title>SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-searchswarm-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-searchswarm-paper-reading/</guid><description>深度精读 SearchSwarm——首个系统探索如何让 Agent 学会「委派」的工作。论文设计了精巧的 Harness 引导主 Agent 将子任务分派给子 Agent，用合成的轨迹数据通过 SFT 将「委派智能」内化到模型权重中。30B 参数的小模型在 BrowseComp、GAIA 等四个基准上达到同规模最佳，甚至超越 10 倍参数的模型。更令人惊喜的是，委派训练的智能还能泛化到单 Agent 设置和开放式研究任务。</description></item><item><title>Self-Harness: Harnesses That Improve Themselves 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-12-self-harness-paper-reading/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-12-self-harness-paper-reading/</guid><description>深度精读上海人工智能实验室 Self-Harness 论文——首个让 LLM Agent 自主改进自身操作套件（Harness）的范式。从 Harness 概念补全、三大范式对比定位、三阶段闭环机制详解、模型特异性验证到通用性灵感提取，全面拆解这项在 Terminal-Bench-2.0 上取得高达 21.4% 绝对提升的开创性研究。</description></item><item><title>Skill-RM: 通过 Agent Skill 统一异构奖励评估标准 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-10-skill-rm-paper-reading/</link><pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-10-skill-rm-paper-reading/</guid><description>深度精读中山大学、香港中文大学、北京大学、ETH苏黎世与阿里巴巴通义千问团队联合发表的 Skill-RM 论文——将奖励建模重新定义为可复用的&amp;rsquo;奖励评估技能&amp;rsquo;（Reward-Evaluation Skill），通过结构化的 Agent 技能编排异构评估资源（评分准则、验证器、检查清单、聚合规则），在 RewardBench2、RM-Bench、JudgeBench 三大基准上全面超越传统 LLM-as-a-Judge 和专用奖励模型，为 LLM 后训练提供统一、可解释、可扩展的奖励信号框架。</description></item><item><title>HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-09-harnessforge-paper-reading/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-09-harnessforge-paper-reading/</guid><description>深度精读北京航空航天大学与清华大学合著的 HarnessForge 论文——一个元自适应框架，将 LLM Agent 系统形式化为 harness-policy 对，通过故障引导的 harness 裁剪和 harness 条件化的策略对齐实现协同演化。在 5 个跨领域基准上超越所有 harness-only 和 policy-only 基线，最高增益达 12.0%，揭示了一个关键洞察：harness 和 policy 之间的可执行兼容性是 Agent 系统适应的核心。</description></item><item><title>Rethinking Continual Experience Internalization for Self-Evolving LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-09-rethinking-continual-experience-internalization-paper-reading/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-09-rethinking-continual-experience-internalization-paper-reading/</guid><description>深度精读中国人民大学高瓴人工智能学院与美团联合发表的 Rethinking Continual Experience Internalization 论文——系统性地揭示了 LLM 智能体在多轮经验内化中出现的&amp;rsquo;能力崩塌&amp;rsquo;现象，从经验粒度、注入模式和内化机制三个维度诊断根因，提出&amp;rsquo;原则级经验 + 逐步注入 + Off-policy 蒸馏&amp;rsquo;的稳定自进化配方，使模型在连续迭代中实现可持续的性能提升而非渐进退化。</description></item><item><title>From Context to Skills: Can Language Models Learn from Context Skillfully? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-06-ctx2skill-paper-reading/</link><pubDate>Sat, 06 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-06-ctx2skill-paper-reading/</guid><description>深度精读清华大学、DeepLang AI、UIUC 等机构联合发表的 Ctx2Skill 论文——一个无需人工标注和外部反馈的自进化技能发现框架。通过多智能体自博弈循环让 Challenger 和 Reasoner 共同进化技能集，配合 Cross-Time Replay 机制防止对抗性坍塌，在 CL-bench 四个上下文学习任务上跨模型一致提升解决率。</description></item><item><title>Skill0.5: Joint Skill Internalization and Utilization for OOD Generalization in Agentic RL 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-01-skill0-5-paper-reading/</link><pubDate>Mon, 01 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-01-skill0-5-paper-reading/</guid><description>深度精读 Skill0.5 论文——华东师大 × 美团联合提出的首个区分通用技能内化与任务特定技能利用的 Agent RL 框架。从 Skill/Harness 概念补全、四代技能增强方法演进定位、双范式困境的问题抽象、难度感知路由+特权蒸馏+反捷径利用三大机制详解、必要知识反推到七条通用性灵感，全面拆解这项在 OOD 泛化上大幅超越全部基线的创新研究。</description></item><item><title>Natural-Language Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-05-26-nlah-paper-reading/</link><pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-05-26-nlah-paper-reading/</guid><description>深度精读清华+哈工大 NLAH 论文——首个系统性探索 Agent Harness 策略能否外化为可执行自然语言对象的工作。NLAH+IHR 四层架构用不到代码 5% 的篇幅表达完整策略，三大基准成绩可比，策略层从几万行代码中提炼为几千字文档。模块消融揭示反直觉结论：收紧验收纪律比扩大搜索范围更有效。</description></item><item><title>SkillOpt: Executive Strategy for Self-Evolving Agent Skills 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-05-26-skillopt-paper-reading/</link><pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-05-26-skillopt-paper-reading/</guid><description>深度精读微软 SkillOpt 论文——首个将深度学习完整优化纪律系统迁移到文本空间 Agent 技能优化的工作。从 Skill/Harness 概念补全、六项前序工作定位、问题形式化抽象、五大核心机制详解、必要知识反推到七条通用性灵感，全面拆解这项在52个评估单元上全部取得最优的开创性研究。</description></item></channel></rss>