<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>过程式评测 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E8%BF%87%E7%A8%8B%E5%BC%8F%E8%AF%84%E6%B5%8B/</link><description>Recent content in 过程式评测 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 23 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E8%BF%87%E7%A8%8B%E5%BC%8F%E8%AF%84%E6%B5%8B/index.xml" rel="self" type="application/rss+xml"/><item><title>OSWorld-Pro：用过程式评测给 Computer-Use Agent 做「分步体检」</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-osworld-pro-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-osworld-pro-paper-reading/</guid><description>OSWorld-Pro 是 NVIDIA 提出的首个面向 Computer-Use Agent（CUA）的过程式评测基准，用来补 OSWorld 那类「只看最终结果」评测的盲区。它包含 305 个长程任务、2814 个顺序依赖子目标、67,264 条步级人工标注（&amp;gt;5000 人时），覆盖 Diversity（117）/ Coordination（109，需跨 ≥4 个应用）/ Robustness（79，跨 Linux 发行版与 GUI）三类，任务平均 9.2 个顺序子目标、3.45 个应用（OSWorld 仅 1.34）。论文用与人类对齐的 LLM-Judge（GPT-5.6-Sol Max）做子目标完成度判定，其 1-MAE 达 93.0，逼近人类标注的 96.0。核心发现：即便最强模型也很吃力——Claude Opus 4.8 Max 以 77.7% 总完成率居首（OSWorld 同级最强 Opus 为 83.4%）；开源最佳 Qwen3.8 Flash Next 仅 55.1%，且在 Robustness 上骤降到 32.9%；Minimax M3 从 OSWorld 的 75.2% 暴跌到 28.9%。过程式视角还暴露了结果式评测看不到的失败模式：强模型也会陷在 subgoal-irrelevant 动作里（Claude Opus 5 曾卡 59 步做无关操作），弱模型则在 click 坐标这类基础操作上频繁出错。本文按九部分结构拆解其背景、定位、问题定义、数据构建、LLM-Judge 设计、外部交叉验证、核心结果、失败模式分析与启示。</description></item></channel></rss>