<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>校准 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%A0%A1%E5%87%86/</link><description>Recent content in 校准 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 11 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%A0%A1%E5%87%86/index.xml" rel="self" type="application/rss+xml"/><item><title>Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-agent-confidence-internal-representations-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-agent-confidence-internal-representations-paper-reading/</guid><description>UMass Amherst 的这篇论文利用 LLM 内部表征预测智能体任务成败：Latent Trajectory Dynamics（LTD）总结交互轨迹上残差流表征的变化动态，Action Representation Probe（ARP）在动作决策点读取表征预测成功。在 InterCode 的 Bash/SQL/Python 三个交互基准 × Qwen-14B/Qwen-7B/DeepSeek-6.7B 三个模型家族上，两方法全面超越表层 token 概率与序列校准基线（漏损泄漏的交叉验证协议），且零额外开销——不改提示、不需多样本 rollout。智能体安全关键应用第一次有了&amp;rsquo;从模型内部读出置信度&amp;rsquo;的免费监视器。</description></item></channel></rss>