<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>数据标注 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%95%B0%E6%8D%AE%E6%A0%87%E6%B3%A8/</link><description>Recent content in 数据标注 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Mon, 28 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%95%B0%E6%8D%AE%E6%A0%87%E6%B3%A8/index.xml" rel="self" type="application/rss+xml"/><item><title>出题、卖题、判卷：AI数据行业的权力、红线与瓶颈</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-28-ai-data-industry-benchmark/</link><pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-28-ai-data-industry-benchmark/</guid><description>硅谷101对话Scale AI何允中与伯克利博士后孙一铀，拆解AI数据生意：估值半年涨十倍的AfterQuery、百亿美元级的Mercor与Scale背后，交付物已从人工标注进化为专家评分标准（rubric）与强化学习环境；评测的权威性天然通向卖数据生意，但卖评测数据是不可碰的红线；行业真正的瓶颈正从标注产能转向垂直领域的采购与版权；而激励设计决定了数据质量的上限。</description></item><item><title>onPanda: 通过 Token 级纠错高效标注 LLM 与 Agent 的同策略对齐数据 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-onpanda-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-onpanda-paper-reading/</guid><description>onPanda（阶跃星辰 StepFun + 厦门大学）提出以 token 级纠错为核心的新型标注范式：标注者只需定位第一个不合适的 token，从模型候选集中点选或自由改写，系统随即截断后续内容并由模型从修正后的前缀继续生成（locate-correct-continue 循环）。该范式让绝大多数 token 由 rollout 模型原生生成，从而在低成本标注的同时高度保留同策略（on-policy）保真度，并自动产出可精细到位置的监督信号。本文按九部分结构精读其动机、系统设计、实验证据，并通过外部检索交叉验证标注效率、同策略数据价值与 RLHF 标注工具（Argilla/POTATO/Reptile）等相关工作。</description></item></channel></rss>