<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>工具调用 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E5%B7%A5%E5%85%B7%E8%B0%83%E7%94%A8/</link><description>Recent content in 工具调用 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E5%B7%A5%E5%85%B7%E8%B0%83%E7%94%A8/index.xml" rel="self" type="application/rss+xml"/><item><title>Failure-Transparent Agents × FCD × CoSec：智能体安全的三个新失效面 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-security-trio-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-agent-security-trio-paper-reading/</guid><description>本精读覆盖三篇 2026 年 9 月底的 Agent 安全论文：FTA 把「工具失败后模型谎报成功」从端到端评估中剥离出来，发现六模型平均 22.8% 的假成功率，而一个四字段证据契约把它压到 0.8%；FCD 命名并防御「schema 没变但 handler 语义变了」的版本漂移——GitHub MCP v1.4→v1.3 让同一省略参数的建仓调用从私有变公开；CoSec 则把授权边界放进多用户社区，证明同一模型换一个 harness 隐私违规率差 24 个百分点。三者共同把 Agent 安全从「注入攻击」扩展到汇报失真、版本漂移、社区边界三个系统性失效面，与产业界 NVIDIA Open Agent Safety Platform 和白宫超级智能协定的「安全在模型之外的层」思路同频。</description></item><item><title>训练机制三重奏精读：对数线性稀疏注意力、置信度停止与跨段信用分配</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-training-mechanisms-trio-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-training-mechanisms-trio-paper-reading/</guid><description>本文合并精读 2026 年 9 月下旬三篇聚焦「训练时机制创新」的论文：上交+字节 Seed 的 PISA 把块稀疏注意力的选择阶段压到 O(N log N) 并用 LSE 打分提升选块质量；马里兰+Capital One 的 ConfSFT 只监督置信度、不监督长度，就让四家族模型推理 token 下降 10%~19%；北大+深大+腾讯的 SLCA-GRPO 识别并结构性消除工具调用 RL 中的跨段信用错配，7B/8B 上 τ2-Bench 分别提升 9.15/10.03 个百分点。每篇覆盖背景、关联工作、问题、解法、实验证据、优势根源（含外部交叉验证）、知识反推与通用灵感九部分，最后合并讨论「选择、停止与信用分配」这一共同主题。</description></item><item><title>ChronosAttack: Adversarial Tool Scheduling Attacks on LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-25-chronosattack-paper-reading/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-25-chronosattack-paper-reading/</guid><description>LLM Agent 安全研究长期聚焦内容攻击：注入提示词、投毒工具、污染记忆。这篇论文换了维度——时间。ChronosAttack 提出纯时延调度攻击：不改、不增、不删任何工具响应，仅施加有界延迟改变证据到达顺序，就能显著改变 GPT-5.6 Sol、Gemini 3.6 Flash、DeepSeek V4 Flash、Claude Sonnet 4.6 四个模型家族的最终决策，部分场景目标选择率从 0% 升至 83.3%。顺序状态并非必需，单次调度反转即可引发大幅决策改变；同步化与顺序一致性防御可削减攻击者控制。本精读覆盖威胁模型、实验证据、机制解释与外部文献交叉验证。</description></item><item><title>SkillApt×TwinCheck：技能何时加载与调用如何验证——反事实证据的双重应用 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-25-skill-verify-duet-paper-reading/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-25-skill-verify-duet-paper-reading/</guid><description>本精读合并解读两篇同期 arXiv 论文：SkillApt 与 TwinCheck。前者针对 Agent 技能库的检索后激活问题，用 WITH/WITHOUT 对照执行构建反事实证据库，在 SRA-Bench 上以 31.5% 的激活率保住 BM25 Top-1 的观测精度并省下 74.3% 的 token；后者针对有状态工具 Agent 的执行边界，在动作执行前构造应当失败的负例孪生调用做证据接地校验，在 BFCL V4 上提升 13.2 个百分点且零误伤。两文共同指向一个命题：把反事实证据作为 Agent 决策的校验锚点，让每一次技能加载与每一次工具调用都有对照实验背书。本文按九部分结构拆解两文的问题定义、解法、实验证据与可迁移灵感，并对外部相关文献做了交叉验证。</description></item><item><title>Critical-State RL：为多轮工具调用诊断「可训练的模型调用」 —— 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-critical-state-rl-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-critical-state-rl-paper-reading/</guid><description>深度精读 Salesforce AI Research 的 Critical-State RL。它指出多轮工具调用失败往往卡在单个模型调用上，但「奖励有差异」并不等于「该调用值得训练」。方法先用训练前三闸门诊断（动作充分性 / 提升空间 / 可训练性），再用嵌套同前缀采样分离「动作依赖奖励方差」与「后续噪声」，最后只对选中的关键调用做 occurrence-local RL（上下文赌博机式训练）。BFCL 上对缺失函数任务 miss_func 恢复率 0.14→0.283（+14.3pp），错位训练反而 −4.5pp；记忆子任务 34.54%→50.54%；Nemotron 重复调用一致性 37%→75%。</description></item><item><title>Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-17-spurious-tool-use-rl-agents-paper-reading/</link><pubDate>Thu, 17 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-17-spurious-tool-use-rl-agents-paper-reading/</guid><description>RL 训练的 agent 调用工具的理由可能是错的：UW+UCSD+Stanford 团队构造受控环境注入与工具强相关但因果无关的线索，发现反事实评估下伪工具调用率最高暴涨 +39.2%——而且捷径只在 agent 已可靠掌握该工具时形成（任务能力是捷径的前提）、语义对齐线索放大效应（对齐 +39.2% vs 交换 ≤3.5%）。反直觉结论：提升能力的 RL 同时放大捷径易感性，标准任务奖励不足以产生鲁棒工具策略；LLM 裁判的&amp;rsquo;工具必要性&amp;rsquo;密集奖励可有效压制且不损准确率。</description></item><item><title>RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-rsiagent-causal-memory-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-rsiagent-causal-memory-paper-reading/</guid><description>数字 Agent 进入新环境（接口/工具/失败模式预训练未覆盖）时如何无监督适应？RSIAgent 给出 training-free 答案：curriculum/actor/verifier 三类 Agent 协同自主探索，把&amp;rsquo;动作-条件-后果&amp;rsquo;因果关系沉淀为可冻结复用的记忆；广度+深度双探索消融显示完整 RSI 74.54% 显著优于单策略（65.52%/56.50%），并让 Kimi-K3、GLM-5.3 在 OSWorld-v2 与 Agent&amp;rsquo;s Last Exam 上反超 GPT-6 Astra。本精读覆盖因果记忆与轨迹记忆的本质差异、广深互补的机制解释与开源反超闭源的信号意义。</description></item><item><title>Salesforce Koa: An Enterprise Language Model for Agentic Tool Use 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-16-salesforce-koa-enterprise-agent-model-paper-reading/</link><pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-16-salesforce-koa-enterprise-agent-model-paper-reading/</guid><description>企业 Agent 工具使用模型的开放权重范本：Salesforce 基于 Nemotron-3-Super-120B（NVIDIA 开源基座）GRPO 后训练出 Koa，在 Dreamforce 发布并作为 Agentforce 平台可选模型。核心是 simulation-to-reward 管线——把工作流规格展开为 persona 条件多轮任务、以成功工具使用为基础的任务解决奖励；企业域用 Agent Script 声明式语言书写规格，训练零客户数据。Tau2Bench 69.41 超基座，且揭示 SFT/RL 的能力分工：BFCL 多轮上 SFT 反降分（54.12→53.25）而 RL 提升。本精读覆盖声明式规格→模拟器→奖励的生成管线与开放权重的企业模型经济学。</description></item><item><title>Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-15-is-bash-all-you-need-tool-interfaces-paper-reading/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-15-is-bash-all-you-need-tool-interfaces-paper-reading/</guid><description>Microsoft 的系统性受控实验颠覆 agent 工具接口直觉：5 种接口配置（纯 typed tools / typed+bash / 纯 bash / bash+持久化自合成工具 / PTC）× 2 企业 benchmark × 2 前沿模型（Opus-4.8、GPT-5.5）下，纯 bash 全面对碾压 typed tools——TheAgentCompany 高 21.8-24.5pp、APEX 高 4.8-7.4pp，同时省 19-72% token。给 bash 加 typed tools 或工具合成均无增益。企业 agent 选型的迄今最硬证据。</description></item><item><title>Gander (Omni Interaction Agent Technical Report) 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-gander-omni-interaction-agent-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-gander-omni-interaction-agent-paper-reading/</guid><description>腾讯混元语音与浙大等团队的 Gander 用&amp;rsquo;小脑-大脑&amp;rsquo;协作框架统一了全双工实时交互与长程 Agent 执行：9B 小脑以 Thinker-Talker 流式架构逐秒 chunk 决策听/说/打断，大脑免训练接入 Codex/Claude Code 执行长任务，编排运行时以 task_start/send/resolve 结构化调用衔接。Full-Duplex-Bench v3 上 turn-taking 100% 全场最佳、过早打断仅 8.0%（GPT-Realtime 13.5%）；SpokenQA 全双工组第一。模型、代码、数据全部开源。</description></item><item><title>Procedural Graphs: Self-Evolving Execution Structures for LLM Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-procedural-graphs-self-evolving-agents-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-procedural-graphs-self-evolving-agents-paper-reading/</guid><description>知识图把事实组织成 (实体, 关系, 实体) 三元组来回答&amp;rsquo;是什么&amp;rsquo;；Google 团队的 Procedural Graph 用 (过程, 关系, 过程) 三元组回答&amp;rsquo;怎么做&amp;rsquo;。Agent 每步决策时定位活跃节点，引导模型把邻域子图翻译成步级情境引导；离线自进化循环对比成败轨迹编辑图拓扑与属性，验证门通过才采纳、拒绝项存为负约束。六个基准三个 LLM 全面超越 ReAct/ExpeL/AWM 等记忆基线——Gemini 3.1 Pro 上 τ-bench 72.17→80.00、GDPval 56.39→78.78、ALFWorld 满分，零骨架自进化图匹配乃至超越手工设计。本文精读拆解过程性知识的表示设计与&amp;rsquo;验证门+拒绝记忆&amp;rsquo;的进化机制。</description></item><item><title>InterOPT/OR-Clarify: 运筹学建模中'何时该问'的选择性完备性决策 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-interopt-or-clarify-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-interopt-or-clarify-paper-reading/</guid><description>杉数科技×上海交大提出 OR-Clarify 基准与 InterOPT 框架，首次系统评测&amp;rsquo;LLM 在运筹学建模前知道何时向用户澄清&amp;rsquo;：部分公开描述+隐藏结构化槽位+有界交互模拟用户，度量槽位恢复、静默假设与交互成本。InterOPT 用 Dynamic Gap Search 识别规格关键缺口，choice-based 设定下 exact slot recovery 大幅超越全部基线。本文精读&amp;rsquo;澄清即决策&amp;rsquo;这一新问题定义。</description></item><item><title>TROVE: 轨迹锚定的最小充分路线编辑 — 智能体编排的运行时修正 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-trove-route-orchestration-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-trove-route-orchestration-paper-reading/</guid><description>TROVE 把智能体编排的结构决策从&amp;rsquo;执行前锁定&amp;rsquo;改为&amp;rsquo;运行时最小充分编辑&amp;rsquo;：离线把工作流搜索轨迹蒸馏为原子/复合技能+结果条件转移图，在线对挂起路线执行保留/插入/替换失效后缀三操作。代码生成、QA、数学推理上质量-效率权衡全面优于 AFlow/MaAS/LAS。本文精读&amp;rsquo;route 即临时品&amp;rsquo;的编排新原则。</description></item><item><title>CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-02-cast-critique-agents-paper-reading/</link><pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-02-cast-critique-agents-paper-reading/</guid><description>亚利桑那州立大学与思科研究的 CAST 解决长程工具调用 Agent 的可靠性死穴：在电商退款、医疗分诊这类有状态环境中，一个错误动作（退错订单）就造成不可逆失败。CAST 把稀疏任务结果转化为动作级监督——合成解释&amp;rsquo;该动作在部分可观测下为何有效/无效&amp;rsquo;的结构化批评理由训练批评模型，再用批评模型构造数据优化策略模型。微调后的 Qwen3 系小模型在 Retail 任务可靠性超 GPT-OSS-120B 逾 10 个百分点，域外 Telehealth 再 +9%。</description></item><item><title>SPT: Skills as Pre-Training Data for Agentic Language Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-spt-skills-pretraining-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-spt-skills-pretraining-paper-reading/</guid><description>深度精读北京邮电大学与清华大学论文 SPT。论文提出把公开的多文件技能包当作预训练中段（mid-training）数据：清洗 ClawHub 上 38,040 个技能包构建 SkillCorpus（约 3.48 亿 token），用 Reference Insert 序列化策略把被引用文件插入到指令首次提及处，使引用距离缩短 94.92%。7B 模型四个 Agent 基准平均分从 28.50 提升到 53.46（+24.96），通用能力几乎无损，30% 技能混合配比效果最佳。</description></item><item><title>Agent Seer: Synthesizing Scenarios from Specification Understanding 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-agent-seer-scenario-synthesis-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-agent-seer-scenario-synthesis-paper-reading/</guid><description>Apple 团队提出 Agent Seer，一条仅以 MCP 工具规范为输入的四阶段流水线：工具语义解释、分层场景生成、mock 输出合成、数据接地多轮扩展，无需人工标注与真实工具执行即可产出完整评估 harness。在 7 个开源 MCP 规范（14–64 工具）上生成 337 个场景，平均工具调用正确性 0.911、对话连贯性 0.855，6 个中型规范实现工具 100% 覆盖。论文进一步给出三个反直觉发现：参数 schema 复杂度是质量变化最强相关因子而工具数量作用正交、argument 值准确性是主导失败模式、跨家族 judge 复验确认结论稳健。本精读逐部分拆解其方法设计、实验证据与优势根源。</description></item><item><title>Agentic AI for operating scientific instruments for nanoscale characterization 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-agentic-afm-instruments-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-agentic-afm-instruments-paper-reading/</guid><description>EPFL 团队用三个基于 MCP 的 Agent（AFM Messenger / Pilot / Doctor）让未经微调的通用 LLM 直接操作原子力显微镜：Messenger 自然语言转经校验的仪器命令、Pilot 用 LLM 视觉评估图像伪影并闭环调参、Doctor 做透明的伪影后处理。在 2747 场景构建的测试集上，Claude-MCP 加歧义检查层把错误命令率从裸模型 89.2% 降到 0.0%；与 5 名人类操作员对比，迭代数、调参时间、最终伪影严重度四项终点均无显著差异。本文精读覆盖背景概念、方法细节、实验证据与错误归零的因果链，并提炼可推广到其他科学仪器自动化的通用灵感。</description></item><item><title>Metis: Typed Runtime Mediation for Tool-Using Software Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-metis-typed-mediation-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-metis-typed-mediation-paper-reading/</guid><description>深读一篇罕见的单人独立研究：Metis 把模型与外部副作用之间那一层运行时当作正经的系统软件工程对象，用类型化事件图显式刻画权限判定、并发调度、终态闭包与生命周期修复。30 对匹配真实 I/O 实验中四类调度中位耗时 14.146ms，全面快于强制串行的 25.958ms；子代理边界消融 0/5 逃逸、路由级权限 oracle 10/10 全对，同时诚实呈现 3 个负面结果。本精读逐部分拆解其机制设计与有界主张的写法。</description></item><item><title>SARA: When Tool Outputs Become Commands 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-sara-tool-authorization-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-sara-tool-authorization-paper-reading/</guid><description>深度精读中科院信工所的 Agent 安全论文 SARA。核心思想是把「动作诱导」与「执行授权」拆开：工具输出可以参与任务实例化，但绝不能自行获得执行权威。通过上下文隔离的 Action Probe、持久动作来源追踪、审计执行证据与参数级支持检查，SARA 在 AgentDojo 与 AgentDyn 上把间接提示注入攻击成功率压到 0.63% 以下，而良性任务代价远小于强隔离方案 CaMeL。</description></item><item><title>CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-caskg-skill-graph-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-caskg-skill-graph-paper-reading/</guid><description>深度精读吉林大学 + 蚂蚁集团论文：把大规模技能库的图检索形式化为「预算约束的边置信度校准问题」——多信号诱导高召回候选图后，用移除/替换/逆序三个反事实探针验证每条边是否真为操作依赖，Beta 平滑聚合后按状态门控发布。6 个骨干 × 2 基准的全部 12 个组合全部第一，ScienceWorld 宏平均 72.62→80.50，步数全线下降。</description></item><item><title>Narcissus: Program Synthesis Using Context-Aware LLM Approximations 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-narcissus-synthesis-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-narcissus-synthesis-paper-reading/</guid><description>代尔夫特理工大学团队提出 Narcissus：当任务固定目标语言（CFG 定义的 DSL）时，LLM 提案通常违反语法或不满足规格——与其反复重提示，不如把提案一次性编译成「上下文感知」的搜索启发式。它将提案解析修复为语法树，用前缀对齐（相同上下文的提案是否用了同一规则）、子程序复用（提案反复出现的片段）与正则化（提案指示的程序规模+保底项）三个信号给每次扩展打分，搜索期间零 LLM 调用。在五个域、两种搜索后端上，Narcissus 在每个预算下击败静态先验：SLIA-70 上 51.4 对 32.2，ARC-100 上解决 40% 而原始提案仅 13%，到达提案区域快约 12 倍；DeepSeek 弱提案加搜索甚至超过 GPT-4o 直接采样。正则化保底项保证任何规则不被剪枝——错误提案只会延迟解、不会藏死解。</description></item><item><title>ToolMinimize: Auditing and Rewriting LLM Agent Tool Calls to Minimize Privacy Exposure 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-toolminimize-privacy-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-toolminimize-privacy-paper-reading/</guid><description>深度精读 PST 2026 的 ToolMinimize 论文——Case Western Reserve 大学提出的 LLM Agent 工具调用隐私最小化中间件。动机研究显示三个生产 LLM 默认 prompt 下 81-88% 的工具调用包含不必要的隐私敏感数据，显式隐私指令后仍剩 36-76%（Llama 几乎无效）。现有防御全是 allow/block 门控或 token 级 PII 检测，无法改写参数值。TOOLMINIMIZE 在参数构造与工具执行之间拦截调用，用「模式+实体+语义」三段分类器识别 PSD（含隐式隐私如医院名隐含诊断）、按 JSON Schema 做必要性分析、执行删除/泛化/替代/截断四种改写。307 次真实调用验证：隐私成本降 81.2-92.0% 且 100% 任务有效（TOST 等价 p&amp;lt;0.001），中位延迟仅 1.77ms。</description></item><item><title>Tunable Tool-Call Rates in LLM Agents via Representation Steering 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-steering-toolcall-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-steering-toolcall-paper-reading/</guid><description>深度精读 UC Santa Cruz + UC Berkeley 论文：LLM agent「是否调用工具」这个离散决策，可以被残差流中的单一线性方向连续调节——无需训练、无需改提示，一个旋钮把调用率从近 0% 单调推到 90%+，且新增调用精准落在模型答不出的低流行度问题上，PopQA 准确率 0.29→0.56，方向还能零样本迁移到 6 个未见工具。</description></item><item><title>ReCache: 工具增强Agent的组合不变KV缓存复用 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-23-recache-kv-reuse-paper-reading/</link><pubDate>Sun, 23 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-23-recache-kv-reuse-paper-reading/</guid><description>工具增强Agent每个请求都要重新编码一遍以不同组合、不同顺序出现的工具与技能schema，标准前缀缓存对此无能为力。ReCache提出resource-wise attention，切断资源间注意力并重置资源内位置索引，使每个资源的KV块具有组合不变性、可独立缓存复用；再叠加贡献选择的层-KV头组路由与字段感知的语义剪枝，把KV张量内存降低92.43%、注意力加速1.423倍，同时Inv-F1基本不降。本精读覆盖其动机、机制、七数据集基准与效果根源分析。</description></item><item><title>EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-08-envace-paper-reading/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-08-envace-paper-reading/</guid><description>蚂蚁集团与上海交通大学联合提出 EnvACE，让一个策略同时扮演「行动者」和「环境」两个角色：Acting 角色生成工具调用，Rehearsal 角色生成对应的环境响应，两个角色共享参数端到端联合优化。训练时完全不需要外部环境交互，却能在 BFCL-v4、τ²-Bench、VitaBench 三个 Agent 基准上全面超越依赖真实环境的 baseline，综合得分 32.91% 超过 14B 参数的 AWM。更精彩的是，模型内化出的「世界模型」还能在推理时用于 Test-Time Scaling，在提交执行前先在内部排练验证，把 Overall 推到 40.9%。本文从机制层面解释「世界排练为何比真实环境更高效」，并提炼出三条可迁移到其他领域的通用灵感。</description></item></channel></rss>