<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>多轮对话 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E5%A4%9A%E8%BD%AE%E5%AF%B9%E8%AF%9D/</link><description>Recent content in 多轮对话 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 01 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E5%A4%9A%E8%BD%AE%E5%AF%B9%E8%AF%9D/index.xml" rel="self" type="application/rss+xml"/><item><title>LLMs Get Lost in Evolving User Intent 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-01-evolving-intent-paper-reading/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-01-evolving-intent-paper-reading/</guid><description>本文精读 Microsoft Research 团队发表于 2026 年 7 月的论文《LLMs Get Lost in Evolving User Intent》。论文提出一个将任意静态单轮基准测试转化为动态多轮对话的框架，通过三种意图转移（论点揭示、论点修正、函数切换）模拟用户意图的真实演化过程，同时保留原始评估协议实现免标注的自动验证。跨数学、Text-to-SQL、搜索、编程四个领域的实验揭示了一个一致现象：在单轮设置下表现优异的模型，一旦用户意图动态演化，性能便大幅下降，最严重时直接归零。这一发现暴露了静态评估的盲区，对协作式 Agent 的未来发展具有关键启示。</description></item></channel></rss>