<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>世界模型 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E4%B8%96%E7%95%8C%E6%A8%A1%E5%9E%8B/</link><description>Recent content in 世界模型 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 26 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E4%B8%96%E7%95%8C%E6%A8%A1%E5%9E%8B/index.xml" rel="self" type="application/rss+xml"/><item><title>Agent-Editing World Model 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-26-agent-editing-world-model-paper-reading/</link><pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-26-agent-editing-world-model-paper-reading/</guid><description>人大高瓴学院 AEWM 论文精读。论文把语言世界模型的预测目标从「重建环境观测」重构为「预测决策效果并直接编辑 agent 状态」，用 Action Judge 三分类（CRITICAL/EXPLORATORY/NOISY）+ State Revision 推理动作联合编辑组成推理时闭环 EditAct，再用 AEWM-RFT 把编辑能力内化回 agent。Action Judge 基准 macro-F1 70.5% 超最强基线 10.6pp；六基准三骨干平均提升 3.2–6.7 分；9B+EditAct 反超 35B+ReAct。本精读覆盖动机、方法、证据链、外部交叉验证与可迁移灵感。</description></item><item><title>Training Object Permanence in World Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-26-object-permanence-world-models-paper-reading/</link><pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-26-object-permanence-world-models-paper-reading/</guid><description>16 所高校联合团队发布 WROP 基准与训练资源，用 150 个 Blender 参数化生成器、150 万样本的系统化合成数据，检验并训练视频世界模型的「客体永久性」与「客体固体性」两类核心认知先验。微调得到的 16B 模型 PWM-WROP 在 20 人盲测成对比较中以 Elo 1679.5 位列真续写模型第一、全场第三，超最强真续写对手 222.5 Elo，且在匹配分辨率下 LPIPS 0.081、MS-SSIM 0.921 全场最优。本精读覆盖背景、定位、问题定义、解法、实验证据、优势根源与外部交叉验证、知识反推与通用灵感九个部分。</description></item><item><title>在像素里复刻世界之前，先给世界定一把尺——云栖2026世界模型分论坛全记录</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-world-model-digital-to-physical/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-world-model-digital-to-physical/</guid><description>云栖2026「世界模型：从构建数字世界到物理世界」分论坛八场演讲与圆桌的完整复盘：阿里联合高校发布W1-W6能力分级、超1600用例的Benchmark与超10万次对战的Elo Arena，给名实混乱的世界模型定了第一把尺；长视频误差累积、GPT-6吞噬具身上层能力、实时会话基建与「消费即创作」的IGC叙事，构成从数字世界通往物理世界的四道门槛。</description></item><item><title>WorldCrafter：用隐式 3D 感知记忆打造可一致探索的视频世界模型 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-23-worldcrafter-paper-reading/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-23-worldcrafter-paper-reading/</guid><description>深度精读腾讯 IEG ARC Lab 与北京大学联合提出的 WorldCrafter。它用一个「隐式 3D 感知记忆」模块让视频世界模型在长时间、可自由探索的交互中保持场景一致性：以 LagerNVS 初始化记忆编码器、用姿态引导读出、与视频 DiT 联合训练，配合最大覆盖历史检索与少步蒸馏，在 4 卡上达到 16fps 实时。重访一致性 LPIPS 从 Lyra 2.0 的 0.487 降到 0.255，相机控制旋转误差 13.536 为全场最低。</description></item><item><title>机器人 Scaling Law 出现了吗？——徐梦迪的答案：有信号，但真正的分水岭是 in-context learning</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-21-xumengdi-scaling-law-signal/</link><pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-21-xumengdi-scaling-law-signal/</guid><description>清华叉院助理教授徐梦迪在「十字路口」提出：具身领域已出现预训练数据从十万到百万小时、held-out loss 随规模下降的 scaling 信号（如 Dyna-2 百万小时人类视频预训练），但 loss 与真机成功率脱钩，真正有意义的是『未见任务成功率随规模上升』的 scaling law；她判断当前主流 VLA 的『预训练+微调』范式对应 GPT-1 时刻，期待的是通过 prompting 适应个体偏好的 GPT-3 时刻。本文拆解她论证中的证据链、与世界模型路线的关系，以及数据定义由模型能力反推的行业机制。</description></item><item><title>JEPA-Anything: Learning Predictive Models across Different Worlds 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-20-jepa-anything-paper-reading/</link><pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-20-jepa-anything-paper-reading/</guid><description>世界模型至今一域一模型：视觉用 V-JEPA、细胞用 Cell-JEPA、控制用 Dreamer——能否用一个学习原理统治所有&amp;rsquo;世界&amp;rsquo;？PhAI Labs 联合八机构的 JEPA-Anything 提出正交预测因子分解（OPF）：把 JEPA 的单一目标嵌入拆成 K 个正交子空间各配专属预测头，再用伪逆合成完整潜状态。同一核心横跨视觉、单细胞、临床、控制、分子动力学、PDE、天气七域，10 项动力学任务全胜匹配 JEPA 基线，Interventional Pong 单干预误差降 34.8%，四分子系统 100 步 rollout 全部最低误差；因子坐标提名的 IL-18+CD73 联合干预在类器官与小鼠实验中获得验证，轨道潜模式恢复开普勒指数 -1.4991。</description></item><item><title>具身智能的四条路线分歧：数据、Astra 与商业化的真问题</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-14-embodied-intelligence-crossroads/</link><pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-14-embodied-intelligence-crossroads/</guid><description>2026 外滩大会圆桌实录：苏度韩铮、蚂蚁灵波沈宇军、自变量王潜、破壳许华哲四位一线创业者正面回答具身智能三场路线之争——仿真还是真机、GPT-6 Astra 是否构成降维打击、跨过泡沫的指标是什么。共识是具身不会复刻语言模型路径，分歧在数据从哪来、智能住在哪里、钱从哪来。</description></item><item><title>Memory as Plans 精读：把记忆从执行期条件重构为规划期证据，机器人非马尔可夫任务 SOTA</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-13-memory-as-plans-map-wam-paper-reading/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-13-memory-as-plans-map-wam-paper-reading/</guid><description>哈工大×NTU×山大提出 MaP-WAM：将记忆依赖的世界-动作建模拆解为记忆锚定规划与计划条件执行两层——情景区段记忆作为规划期证据，因果世界模型（WAN-2.2-5B 微调）生成视觉计划，World-Action-Progress 模型把任务进度升级为一等模态。RMBench 83.3% SOTA、真机 78.0%，执行器延迟随历史增长恒定；Swap T/Press Button 达 96%。</description></item><item><title>Recursive Code World Models 精读：global-local-global 递归构造可执行 3D 世界</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-13-recursive-code-world-models-paper-reading/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-13-recursive-code-world-models-paper-reading/</guid><description>佐治亚理工提出 RCWM：从单张参考图重建复杂 3D 世界为可执行场景代码。核心是递归场景程序（RSP）表示 + 自递归构造求解器——每次调用遵循建立整体→递归重建未解部分→回访整体精炼组合的 global-local-global 循环，参考对齐视图跨层级传播共享相机投影，父级回访修正局部精炼后浮现的边界错误。三级递归较固定二级 whole-frame PSNR 16.8→19.0、local SSIM 0.52→0.60，全面超越 image-to-scene-program 基线。</description></item><item><title>World in World 精读：免训练控制视频世界模型的统一证据接口</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-13-world-in-world-training-free-control-paper-reading/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-13-world-in-world-training-free-control-paper-reading/</guid><description>西湖大学 AGI Lab 提出 World in World：training-free 推理时接口，把四类异构控制证据（源视频观测/目标视角投影/几何渲染/检索生成态）统一转为带相机与时间标注的干净视觉状态，经冻结视频世界模型的原生 self-attention 读入；对应路由器建立 token 对应关系，证据级 attention CFG（EWA）按通道独立调权。同一冻结骨干完成相机控制重渲染/长时程回访/人体动作迁移，重渲染全指标超 ReCamMaster 等训练方法（CLIP-Sim 92.5 vs 83.8）。</description></item><item><title>UniMPA 精读：给 VLA 模型一个“动作锚定”的统一接口，训练 epoch 砍半还涨点</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-unimpa-memory-prediction-action-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-unimpa-memory-prediction-action-paper-reading/</guid><description>南京大学+九天团队提出 UniMPA：统一记忆-预测-动作模型，用共享的&amp;rsquo;动作锚定转移接口&amp;rsquo;解决 VLA 的转移可实现性缺口（转移歧义/预测失准/多阶段混淆）。LIBERO/LIBERO-Plus/RoboTwin 2.0 Hard/真机四线超 π0.5 达 1.7/11.7/18.5/12.6pp，只需 25-50% 训练 epoch。</description></item><item><title>Programmable World Model 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-programmable-world-model-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-programmable-world-model-paper-reading/</guid><description>Programmable World Model（PWM）把视频世界模型拆成两层：agent 把自然语言编译为可执行程序（实体状态+转移规则），轻量引擎执行程序维护显式持久全局状态（含屏外实体与非视觉属性）；状态经增广 3D OBB 中间表示+相机轨迹确定性编译为像素对齐条件信号，驱动预训练视频模型当生成渲染器。自建 CombatStateBench 上 Count Accuracy 94%（超 LingBot-World-V2 达 53.25 分）、State Accuracy 98%（超 90 分），支持连贯长时程生成。本文精读&amp;rsquo;符号引擎管规则、生成模型管外观&amp;rsquo;的解耦架构为何系统性消灭状态漂移。</description></item><item><title>Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentstream-progressive-latent-memory-paper-reading/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-latentstream-progressive-latent-memory-paper-reading/</guid><description>流式视频理解的主流范式是&amp;rsquo;存历史、按需检索&amp;rsquo;，但外部证据永远只是临时上下文。南京理工×蚂蚁×NUS×港中文的 LatentStream 把范式翻转为&amp;rsquo;检索并内化&amp;rsquo;：分层流记忆 + 渐进扩张感受野的 latent token 把历史证据固化进固定长度潜记忆，用熵构造的渐进置信奖励在测试时联合优化。OVO-Bench 64.2%、StreamingBench 76.9%、MLVU +6.1，全部 SOTA。</description></item><item><title>Discriminative World Models for Web Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-04-discriminative-world-models-paper-reading/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-04-discriminative-world-models-paper-reading/</guid><description>Web agent 用世界模型做测试时动作选择：采样候选动作→预测下一状态→排序执行。但现有世界模型都用监督式&amp;rsquo;下一状态预测&amp;rsquo;训练——花大量 token 复述页面上没变化的部分，而下游 ranker 需要的恰恰是&amp;rsquo;不同动作导致的差异&amp;rsquo;。UC Berkeley 联合 MIT-IBM Watson AI Lab 提出 predicted-state matching：预测表示必须把真实结果状态从替代动作的结果状态中区分出来。同一份数据、同一个 Qwen3-8B 底座，仅换训练目标，匹配准确率从 47.77% 跳到 80.80%，WebArena-Lite 端到端成功率从 13.94% 提到 28.48%。本精读拆解&amp;rsquo;训练目标与下游任务对齐&amp;rsquo;这一教科书级修正的完整证据链。</description></item><item><title>Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-code-as-worlds-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-code-as-worlds-paper-reading/</guid><description>深度精读 MirroS 联合清华、北大、南洋理工的技术报告 Code as Worlds。论文提出用可执行代码表示物理世界的组成、演化与外观（EWR 三元组），把&amp;rsquo;从观测恢复世界表示&amp;rsquo;建模为溯因式的 agent 发现环：提出-实例化-执行-渲染-验证迭代修正。再用验证过的世界免费生成带精确物理量标签的 VQA 数据训练 VLM，9B 模型在 QuantiPhy 上 55.4 分超过 Gemini-3.1 Flash 的 54.8，27B 推理变体 58.6 分超过全部基线。</description></item><item><title>GameWAM: A World Action Model for Video Games 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-gamewam-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-gamewam-paper-reading/</guid><description>复旦、腾讯光子与清华深圳研究生院团队提出 GameWAM，首个面向电子游戏原生键鼠闭环控制（游戏玩法+GUI）的世界-动作模型。它用并行 Video-DiT 与 Action-DiT 做块因果联合流匹配，同时生成未来视觉观测与可执行动作；用每步动作路由器区分 gameplay/GUI 两种控制分布，用预测长执行短的块周期控制解耦规划与承诺，用分层历史压缩维持长时程记忆。在 Minecraft MCU 上平均成功率 50.7（次优 36.8）且执行步数全面最少，零样本迁移 VoxeLibre 达 59.2%。论文还发现并命名了 LASI 失效模式：采样动作源的低频分量会系统性牵引生成相机运动，复用可致原地旋转。整篇论文训练仅 2.79B token，是 Game-TARS 配方的 1/200。</description></item><item><title>WM-R1: Training GUI Agents to Reason and Leverage World Models with Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-wm-r1-gui-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-wm-r1-gui-paper-reading/</guid><description>用强化学习训练手机 GUI 智能体，为什么必须忍受昂贵的真实环境交互？华东师范大学提出的 WM-R1 给出了第一个完全相反的答案：把世界模型从推理时的辅助工具升级为训练环境本身。冻结的 Code2World-8B 世界模型生成全部状态转移，Agent 通过 GRPO 在纯模拟环境中学习；更关键的是 &amp;lt;call_wm&amp;gt; 机制把世界模型嵌进思维链，让 Agent 学会『提出候选动作—模拟后果—评估修正—再提交』的推理策略。AndroidWorld 上 WM-R1-7B 达到 39.8 的 SOTA（超 UI-R1 达 9 个百分点），OOD 平均提升 +16.0，训练全程零真实环境交互、单卡 11.2 小时完成。本精读拆解其训练框架、奖励设计与效果根源。</description></item><item><title>PAWBench: How Far Are We from Probabilistically Aligned World Modeling? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-pawbench-world-model-alignment-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-pawbench-world-model-alignment-paper-reading/</guid><description>深度精读上海交大、上海AI实验室、Krea AI、Hugging Face 等多方合作的 PAWBench：把「概率对齐」形式化为视频世界模型的分布级标准——固定初始观测与动作下，模型诱导的未来分布应匹配物理上有效结果的正确概率。50 场景两套件评测 11 个视频生成模型，无一同时做到概率准、覆盖广、场景稳；核心洞见是「一条合理的未来不等于分布对齐」，加大采样预算只提高覆盖率、纠正不了概率分配。</description></item><item><title>Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment（Station v2）精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-station-math-discovery-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-station-math-discovery-paper-reading/</guid><description>Station v2（DualverseAI × 剑桥 × 港大 × UCSD）把 AI 数学发现从『固定管线里的工具』搬进开放世界多智能体环境：6 个跨模型家族的 agent 在无中央协调器的房间制生态里自选方向、跑实验、发论文积累共享文献。在 AlphaEvolve 目录 12 个构造类问题上 5 题产出相对既有文献新颖的结果——kissing 数 d=11 三个精确 604 点构型（两个为新等距类）、Erdős 最小重叠下界 0.37912→0.380552（闭合已发表区间约 82%）、有限域 Kakeya 新无穷族、离散 Kakeya 针 CT(128)≤0.107067、符号不确定性 0.3089；还独立重构 Jacobian 猜想反例。机制归因：高度自主使 agent 能直接追求不可打分的广义数学目标，评估耗时上限倒逼理论引导构造，46.4% 的亮点结果来自跨模型家族协作。</description></item><item><title>Code World Model: Coding Agent as World Brain 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-code-world-model-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-code-world-model-paper-reading/</guid><description>西湖大学 AGI Lab 与南洋理工提出 Code World Model：让编码 Agent 充当「世界大脑」，用可执行代码维护持久世界状态并驱动世界演化，再通过 proxy（粗粒度代理视频）接口把状态翻译成帧级时空约束，交给视频模型渲染高保真画面。该框架把「世界演化」与「视觉实现」解耦，直面视频世界模型只能从画面反推规则、上下文不足一分钟、离屏后果无法延续三大结构性缺陷。在仅 5.6 小时 GTA V 游戏数据上 LoRA 微调 MiniMax-H3 后，模型即可跟随 proxy 指定的角色位置、轨迹、场景布局与相机运动，并泛化到训练之外的角色与风格。本精读覆盖其问题定义、方法组件、数据管线与局限。</description></item><item><title>VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-vbvr-pro-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-vbvr-pro-paper-reading/</guid><description>NTU 牵头、20 个机构 50 余位研究者共建的 VBVR-Pro，为「原生视觉推理」——把视觉生成当推理介质本身——建立了可训练、可验证、可优化、可受控比较的闭环测试床：300 个程序生成任务（347 万图+130 万视频）、100 个任务的确定性奖励评分器（人类对齐超 GPT-5.5 且完全可复现）、30+ 生成器受控比较。训练使 9 个开源模型平均 +0.29 并在 7 个外部基准迁移 +28 分级别；RLVR 优于 RLVLM；三面证据链证明「视觉轨迹比语言 CoT 更关键」是范式级发现。</description></item><item><title>Q-Learning with World Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-qwm-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-qwm-paper-reading/</guid><description>斯坦福与北大合作的 QWM 提出了一条与世界模型协作的新路线：世界模型完全不参与策略与价值函数的训练，只在测试时（在线采样与评估执行）通过树搜索帮助挑选动作。策略提议 N=8 个候选动作，世界模型为每个动作想象 K=8 个未来状态，递归展开至深度 4，节点值由 Q 函数估计器与奖励加未来值估计器等权聚合。由于训练只发生在真实转移上，模型偏差不会复合进策略；在 Robomimic 与 LIBERO 机器人操作基准上，QWM 在样本效率与最终性能上均显著超越强基线，而传统 model-based 方法几乎无法在同等预算内学出非零成功率。</description></item><item><title>世界模型是具身的永动机吗：北京人形谈 VLA 续命、大一统与机器人幼儿园</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-world-model-perpetual-motion-embodied/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-world-model-perpetual-motion-embodied/</guid><description>《晚点聊》WRC 期间对话北京人形创新中心戴勇、张怡与前华为 AI 专家唐都钰。VLA 与世界模型的路线之争被拆到表征层：VLA 泛化差的病根是&amp;quot;特征漏斗+预训练与后训练范式不一致&amp;quot;；世界模型则被戴勇称为&amp;quot;AI 时代的永动机&amp;quot;——指望它生产数据，它本身却缺数据，&amp;ldquo;至少到现在是个童话&amp;rdquo;。北京人形的答案是 Pelican-Unify 大一统强耦合路线，年底 2.0 要拿出具身领域的 scaling law；唐都钰离职创业做&amp;quot;主动式物理因果模型&amp;quot;，并转述图灵奖得主 Sutton 的机器人幼儿园设想。</description></item><item><title>SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-20-semcomp-bench-semantic-video-completion-paper-reading/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-20-semcomp-bench-semantic-video-completion-paper-reading/</guid><description>SemComp-Bench由中国科学技术大学联合FrameX.AI与中山大学提出，定义了结果导向的语义任务完成视频生成任务：给定参考图像与指令，要求生成视频既达成指定结果，又与参考保持任务相关的语义接地（如把钞票折成乌龟，结果必须是那张钞票折成的乌龟）。团队从Koala-36M约2万条视频经四阶段管线构造1273个结构化实例，并设计OA/GR双维度VLM评测协议。七个代表模型中最高OA仅37.8%，I2V全面碾压T2V（37.8% vs 4.4%），brief指令下OA暴跌至1.7%，GR与OA排名显著错位，揭示了当前视频生成模型会生成却不会完成任务的系统性缺口。</description></item><item><title>HarnessEval-W: Agentifying the Evaluation of Visual Worlds 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-harnesseval-w-agentified-benchmark-paper-reading/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-harnesseval-w-agentified-benchmark-paper-reading/</guid><description>北大、清华、上海AI Lab等机构30余位作者联合提出HarnessEval-W，把LLM生态的harness范式首次引入世界模型基准测试：父Agent解释每个评测案例的语境并路由到技能库（记录激活与跳过理由），技能把问题分解为可测子问题、交给配备诊断工具的专职子Agent，证据经校验后聚合为分数——每次评测产出一棵可回溯到具体子问题与工具证据的证据树。在330案例×18个世界模型上，与5000次人类A/B判断拟合的Bradley-Terry排序对比达Spearman 0.93（Intentional）/0.87（Physical）；对照最接近的WBench协议，Physical成对准确率从31.9%提至71.7%、平局率从52.2%降至1.8%。榜单揭示：Seedance 2.0综合75.5居首，视频生成器改造为世界模型会重新分配能力而非均匀提升。</description></item><item><title>物理AI的下一站：让AI发现人类不知道的方程</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-chen-yuntian-physical-ai-paradigm/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-chen-yuntian-physical-ai-paradigm/</guid><description>机器之心对话东方理工陈云天：从&amp;quot;知识嵌入&amp;quot;到&amp;quot;知识发现&amp;quot;的双向耦合范式。用AI从真实实验数据中找出人类未知的控制方程（如海浪破碎方程），用机械臂高通量实验找色谱方程替代耗时实验；他判断AI科学家&amp;quot;一定到了这个节点&amp;quot;，但资本市场节奏与湿实验闭环的天然慢速之间正在撕开一个gap，而通用模型&amp;quot;每个行业都很浅&amp;quot;，真正的机会在把物理一致性嵌进专业模型。</description></item><item><title>AI for Science 爆发前夜：曹原谈验证瓶颈、概念抽象与 AGI 的最后一公里</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-yuancao-unreasonable-labs-ai-for-science/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-yuancao-unreasonable-labs-ai-for-science/</guid><description>Jeff Dean 带三位谷歌元老出走创立 Discovery Loop 当周，前 DeepMind 科学家、Unreasonable Labs 联合创始人曹原在硅谷101拆解 AI for Science 为何在此时爆发：代码与数学能力到位后科研成为下一个智能爆发点，但真正的瓶颈从建模移到了物理世界的验证环节；LLM 无法凭训练数据产生真正新的科学概念，概念抽象可能是 AGI 的&amp;quot;最后一公里&amp;quot;甚至不可计算；他主张 AI and Science 而非 AI for Science——科学难题是驱动 AI 本身进化的催化剂，并给出&amp;quot;AI 拿诺贝尔奖至少还需二三十年&amp;quot;的长期判断。</description></item><item><title>Alaya-EVOKE: From Linear-Scaling Supervision to Endless World 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-alaya-evoke-endless-world-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-alaya-evoke-endless-world-paper-reading/</guid><description>交互式世界模型要同时做到“记得住、答得快、跑得久”，但三者天然冲突。Alaya-EVOKE 给出了一个系统级解法：把持久记忆外部化为按相机位姿索引的世界状态库，让去噪器上下文有界；同时把“教师”本身当作设计变量，用稀疏注意力改造出能看 30 秒的长视野教师，再经 DMD 蒸馏出三步、无 CFG 的学生模型。结果是：WBench 导航拆分三组指标全第一、VBench-2.0 总分 66.77 登顶（超过 Veo 3），单卡 H200 上每 1.5 秒内容只需 2.11 秒生成，并且能连续跑一个多小时不崩。本精读面向初学者，拆解其三大机制、实验证据与效果优势的因果链。</description></item><item><title>PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-15-playworld-world-model-benchmark-paper-reading/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-15-playworld-world-model-benchmark-paper-reading/</guid><description>可交互世界模型（如 Genie 3）正在爆发式涌现，但“每个模型用自家基准自测”使得跨模型公平比较几乎不可能——固定动作序列在不同模型上会走出完全不同的轨迹。PlayWorld 提出 Agent-as-Player 范式：让多模态 Agent 像人类玩家一样，为指定的长时程目标（转一圈看环境是否一致、走进水里看有没有涟漪）主动探索交互，再用四维度 VQA 体系打分。171 个人工标注场景、9 个世界模型的大规模评测显示：最高的 Genie 3 Overall 也只有 2.12/5，且所有模型在“视野外演化”“洞察演化”两类长时程状态维持维度上普遍不超过 2 分——“能演、但不能持续演化”是当前公认瓶颈。本精读覆盖其动机、方法、实验证据与因果链根源解释。</description></item><item><title>在算力最多的地方做世界模型：对话英伟达Cosmos掌舵人刘洺堉</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-13-liu-mingyu-nvidia-cosmos-world-model/</link><pubDate>Thu, 13 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-13-liu-mingyu-nvidia-cosmos-world-model/</guid><description>英伟达研究副总裁、Cosmos Lab负责人刘洺堉的4小时深度访谈。从GAN到Diffusion到世界模型的20年研究者之路，到Cosmos 3为何将语言/视频/音频/动作统一进单一模型，到&amp;quot;模型竞争不是零和游戏&amp;quot;的Low Ego哲学，再到黄仁勋的第一性原理决策与&amp;quot;不裁员&amp;quot;文化。他认为模型能力终将收敛，真正决定胜负的是生态整合；他不想击败任何人，只想帮助Physical AI整个领域成功。</description></item><item><title>ComBodied Agents: a New Paradigm of Human-Centric Agentic AI 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-12-combodied-agents-paper-reading/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-12-combodied-agents-paper-reading/</guid><description>Bang Liu 团队 38 页范式论文，提出继 Digital Agents（数字状态）和 Embodied Agents（物理状态）之后的第三种 Agentic AI 行动基底——Combodied Agents（以人的演化状态为核心）。文章构建了一个以事件级多模态感知、可纠正纵向记忆、Personal World Models、可接受干预策略四模块组成的闭环框架，并把&amp;rsquo;保留并增强人类 agency&amp;rsquo;首次系统化为可评估的指标体系。本文从背景、定位、问题抽象、解法机制、评估体系、优势根源、必要知识反推、通用性灵感八个维度逐层精读。</description></item><item><title>Addressable Memory for Video World Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-worldtrace-video-memory-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-worldtrace-video-memory-paper-reading/</guid><description>交互式视频世界模型在长时程生成中会遇到一个隐蔽的「记忆失效」问题：KV cache 里明明存着过去的画面，模型却读不出来。本文精读 NVIDIA/Princeton/ Toronto 联合提出的 WorldTrace 框架，它精准定位了 RoPE 旋转位置编码超出训练范围导致的「内容不可寻址」根因，并用一套无需训练的虚拟槽位机制，在时间一致性上提升 15.5%、在 LoopBench 情景回忆上提升 19.5%。本精读将从世界模型的记忆机制讲起，逐层揭开位置编码、相位抵消、虚拟槽位、规范 key 平均等关键技术。</description></item><item><title>MASS：用权威共享状态解耦多智能体世界模型——多人游戏架构如何破解视频世界模型的可扩展性瓶颈</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-09-mass-multiplayer-world-models/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-09-mass-multiplayer-world-models/</guid><description>MASS 借鉴在线游戏的权威服务器架构，将多人世界模型分解为推进权威类型化状态的 Logic Engine 与按需生成视角的 Rendering Engine，在匹配 Snake 基准上取得 0.764 的状态恢复（最强视频基线仅 0.128），并支持 1,024 个并发玩家、10,000 个循环 tick 的结构稳定预测。</description></item><item><title>EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-08-envace-paper-reading/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-08-envace-paper-reading/</guid><description>蚂蚁集团与上海交通大学联合提出 EnvACE，让一个策略同时扮演「行动者」和「环境」两个角色：Acting 角色生成工具调用，Rehearsal 角色生成对应的环境响应，两个角色共享参数端到端联合优化。训练时完全不需要外部环境交互，却能在 BFCL-v4、τ²-Bench、VitaBench 三个 Agent 基准上全面超越依赖真实环境的 baseline，综合得分 32.91% 超过 14B 参数的 AWM。更精彩的是，模型内化出的「世界模型」还能在推理时用于 Test-Time Scaling，在提交执行前先在内部排练验证，把 Overall 推到 40.9%。本文从机制层面解释「世界排练为何比真实环境更高效」，并提炼出三条可迁移到其他领域的通用灵感。</description></item><item><title>WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-06-worldcycle-paper-reading/</link><pubDate>Thu, 06 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-06-worldcycle-paper-reading/</guid><description>腾讯提出WorldCycle，利用可逆动作循环的物理对称性作为免标注自验证RL信号。通过空间闭合奖励和时间一致性奖励，将视频世界模型的状态返回漂移降低44%，复合动作准确率提升近4倍。核心突破在于：不需要任何ground-truth未来状态，物理可逆性本身就是密集监督信号。</description></item><item><title>Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-05-video-deep-research-paper-reading/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-05-video-deep-research-paper-reading/</guid><description>上海 AI Lab 等机构联合提出 Video-DeepResearch（Video-DR），把多模态 Deep Research Agent 从静态图像推进到连续视频流。论文诊断出当前 Agent 的两大顽疾——模态偏见（回避视觉工具转向文本搜索）与参数知识泄露（靠内部记忆蒙答案而非真正调用工具），并设计解耦感知-探索流水线 + 阶段式工具解锁 + SFT+GRPO 两阶段训练予以破解。其 35B-A3B 模型以 64.0% 平均准确率刷新 VideoDR-Bench SOTA，超越 Claude-4.5-Sonnet 5.0 分、GPT-5 11.5 分；30B 变体也追平 Claude-4.5-Sonnet。本文从机制因果层面解释：为何一个激活参数仅 3B 的模型能在视频 Deep Research 任务上反超数十倍体量的顶级闭源模型。</description></item><item><title>WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-05-wcm-paper-reading/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-05-wcm-paper-reading/</guid><description>VLA 模型做 RL 后训练时，critic（价值估计器）通常只看单帧画面——这和机器人控制&amp;rsquo;部分可观测&amp;rsquo;的本质根本不匹配。WCM（World Critic Model）基于 LeJEPA 架构，让 critic 同时预测未来潜在状态（世界建模）和估计价值，使 critic 的表示被显式训练去捕捉时序动态。在 149 个任务上，OpenVLA-OFT 的 in-distribution 成功率达 99.0%、out-of-distribution 达 77.9%；从 0.8% 的 zero-shot 提升超 12000%；7 个真实世界任务全面超越基线。本文从&amp;rsquo;纯标量回报回归不足以学时序动态&amp;rsquo;的第一性原理，解释为什么单纯加历史帧没用、必须加世界建模目标。</description></item><item><title>Mental World Modeling 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-03-mental-world-modeling-paper-reading/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-03-mental-world-modeling-paper-reading/</guid><description>世界模型已经能很好地预测物理场景将如何演化，但它常常预测错人类会做什么——因为人不是被物理推动的，而是被自己的信念、目标、情绪和社会规范推动的。本文提出 Mental World Modeling（MWM）框架，把心理变量从&amp;rsquo;事后解释&amp;rsquo;提升为世界状态的&amp;rsquo;一等公民&amp;rsquo;，并用一个无需训练的六阶段基线 MENTIS 验证：在 448 条情境决策数据上，显式建模心智让 8 个大模型的决策预测 F1 平均提升到 87.9%，而删掉心智通道会掉 12.1 分，最大瓶颈是&amp;rsquo;耦合世界状态如何转移&amp;rsquo;。本精读按九部分结构展开，从&amp;rsquo;世界模型是什么&amp;rsquo;讲起，逐层拆解 MWM 的形式化、MENTIS 流水线、评估设计与瓶颈归因。</description></item><item><title>ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-03-odeworld-continuous-world-model-paper-reading/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-03-odeworld-continuous-world-model-paper-reading/</guid><description>深度精读 arXiv:2607.27924——首个用物理时间 ODE 取代离散下一步预测的连续时间潜在世界模型。PT-Flow 把&amp;rsquo;未来预测&amp;rsquo;重新定义为&amp;rsquo;在紧凑潜在空间里对一个连续速度场做积分&amp;rsquo;，靠动力学解耦 + 直接一阶 JVP 监督，一举绕开 JEPA 长期头疼的表示坍缩难题；还能做离散模型做不到的任意时刻查询和后向预测。在 LIBERO 视频生成上 PSNR 比离散 baseline 高 3 分以上，64 帧长程预测延迟仅 0.072 秒，并在 LIBERO-LONG 和真实双臂 AgileX 机器人上把策略成功率推到 SOTA。</description></item><item><title>PhiZero: A World Model Built Around Physical Language 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-phizero-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-phizero-paper-reading/</guid><description>PhiZero由中科院自动化所提出，通过自监督学习从野外视频中提取紧凑离散的&amp;rsquo;物理语言&amp;rsquo;表示世界状态转移，采用&amp;rsquo;先推理后渲染&amp;rsquo;范式：自回归VLM先推理物理语言序列，再由扩散解码器渲染为视频。4秒33帧视频仅需256个离散符号（比Wan2.2 VAE压缩175倍），在Physics-IQ、PhyGround、WorldModelBench、IntPhys2四个基准上物理一致性全面超越Sora 2、Cosmos3等。</description></item><item><title>ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-shadowdancer-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-shadowdancer-paper-reading/</guid><description>ShadowDancer提出影子对（shadow pairs）和跨影子预测（cross-shadow prediction），通过构造方式解决潜在动作模型的外观-动力学耦合问题。同一动力学轨迹在不同外观下重放，预测一个影子所需的表示必然是共享动力学本身。任何演示片段成为可复用动作资产，在新环境中重放无需动作标签、运动估计器或微调，跨五族动力学平均盲测胜率86%。论文揭示了&amp;rsquo;构造性不变量提取&amp;rsquo;的全新自监督范式。</description></item><item><title>VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-videococo-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-videococo-paper-reading/</guid><description>VideoCoCo由港中文Pheng-Ann Heng联合中国科学技术大学等26位作者提出，用可执行的Blender程序作为视频生成的过程级链式思维：编码智能体将文本提示合成为Blender代码，仿真引擎运行产生确定性时空草稿，生成式视频引擎通过草稿条件编辑转化为逼真视频。PhyGenBench从0.475提升至0.558，VBench-2.0从52.18提升至77.88。</description></item><item><title>Momenta IPO 后再访曹旭东：没有尽头的 AI，从智驾到家庭机器人的十年推演</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-24-momenta-cao-xudong-endless-ai/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-24-momenta-cao-xudong-endless-ai/</guid><description>Momenta 创始人曹旭东在 IPO 后接受「晚点聊LateTalk」专访，回顾十年创业历程，系统阐述智驾竞争格局（中国两三家、全球三四家的终局判断）、&amp;ldquo;一个飞轮两条腿&amp;quot;战略、每年十倍的智驾摩尔定律、从自动驾驶向家庭机器人的技术外溢逻辑，以及从 AI 研究员到 CEO 的认知进化——&amp;ldquo;一流的工作不是想出来的，是做出来的&amp;rdquo;。</description></item><item><title>2026 Q2 AI季报：RSI从科幻走向创业赛道，Coding战场大洗牌，强者愈强的未来</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-22-ai-q2-review-rsi-coding/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-22-ai-q2-review-rsi-coding/</guid><description>2026年Q2 AI季报深度解读：Anthropic与OpenAI的模型竞争进入新阶段，GPT 5.6与Claude Maestro/Phable正面交锋；RSI（递归自进化）从科幻概念变成明确的创业方向，Recursive、Miranda等公司涌现；Cursor以600亿美元天价被收购；中国开源模型&amp;quot;四杀&amp;quot;引发全球关注；Anthropic的Cloud Tag与OpenAI的Record and Replay重新定义AI交互。本文基于播客全文转写整理，涵盖竞争格局、RSI、机器人、智能扩散、交互创新和公司动态。</description></item><item><title>具身原生的豪赌：蚂蚁灵波沈宇军，为什么坚持从传感器和视频里重训整个机器人模型？</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-22-embodied-native-model-ant-lingbo/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-22-embodied-native-model-ant-lingbo/</guid><description>蚂蚁灵波首席科学家沈宇军的深度访谈。他从GAN研究起步，经字节、蚂蚁研究院，最终主导蚂蚁灵波做机器人&amp;quot;大脑&amp;quot;。文章梳理了灵波最核心的技术主张——&amp;ldquo;具身原生&amp;rdquo;：不再沿用数字世界的模型做下游适配，而是从传感器、视频时序、单向MoE架构出发，为物理世界从头训练一套完整的机器人基础模型（V-Ren、DEPS、VLA 2.0、Video、Word六件套）。沈宇军也坦率谈到了数据是当前最大瓶颈、灵波为什么不做本体、以及他对&amp;quot;大脑落后于本体&amp;quot;这一行业判断。</description></item><item><title>世界模型这半年：XLR Labs 谈原生路线、4D 数据护城河与物理 AGI 的下半场</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-18-world-model-xlr-labs/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-18-world-model-xlr-labs/</guid><description>《晚点聊》WAIC 期间对话 XLR Labs（拓元智慧）三位核心成员。这家从 2022 年起就押注&amp;quot;原生世界动作模型&amp;quot;的公司，分享了它与 VLA、隐式世界模型的路线分歧，千万小时级 4D 真实交互数据如何构成护城河，以及从智慧零售切入工业物流的&amp;quot;以终为始&amp;quot;商业化逻辑——一个关于&amp;quot;预训练与后训练一致性&amp;quot;的scaling law故事。</description></item><item><title>Qwen-AgentWorld: Language World Models for General Agents 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-06-29-qwen-agentworld-paper-reading/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-06-29-qwen-agentworld-paper-reading/</guid><description>深度精读阿里通义千问团队的 Qwen-AgentWorld——首个覆盖 7 大领域（MCP/Search/Terminal/SWE/Android/Web/OS）的统一语言世界模型。它通过&amp;rsquo;CPT注入→SFT激活→RL锐化&amp;rsquo;三阶段训练管线，以混合 rubric-and-rule 奖励驯服开放环境模拟的强化学习训练，在 AgentWorldBench 上以 58.71 分超越 GPT-5.4（58.25）。更关键的是，世界模型训练可作为智能体基础模型的有效预热——在 7 个下游基准上带来泛化增益，其中 3 个完全分布外基准平均提升 +9~11 分。</description></item></channel></rss>