<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>评判器 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E8%AF%84%E5%88%A4%E5%99%A8/</link><description>Recent content in 评判器 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 08 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E8%AF%84%E5%88%A4%E5%99%A8/index.xml" rel="self" type="application/rss+xml"/><item><title>OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-08-osreward-paper-reading/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-08-osreward-paper-reading/</guid><description>深度精读港大与腾讯联合出品的 OSReward——首个系统检验计算机使用 Agent（CUA）轨迹评判器可靠性的基准与开源奖励模型工作。论文构建了覆盖 Web/Windows/macOS/Ubuntu/Mobile 五大平台的 1,019 条人工标注轨迹基准，揭示了所有主流 VLM 评判器存在的系统性「宽容偏差」，并训练出成本降低 30-60 倍的开源奖励模型 OS-Shepherd。本精读从背景补全、研究脉络定位、问题抽象、方法机制、评估证据、效果根源、必要知识反推到可推广灵感，九部分完整拆解这项为 CUA 评判建立标准化评估体系的开创性研究。</description></item></channel></rss>