<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>物理AI on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E7%89%A9%E7%90%86ai/</link><description>Recent content in 物理AI on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E7%89%A9%E7%90%86ai/index.xml" rel="self" type="application/rss+xml"/><item><title>Self-Evolving Coding Agents × RE-0：从数字程序到物理世界的自进化智能体 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-physical-agentic-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-physical-agentic-duet-paper-reading/</guid><description>二重奏精读两篇互补论文：hexafuture.ai 的 Self-Evolving Coding Agents 提出物理编码范式，用 Code as World + Code as Policy 双可执行表征与类型化验证器，把编码代理范式迁移到物理世界，在 RoboCasa365 上把成功率从 56.6% 提升到 61.1%；吉林大学与大连理工的 RE-0 用 locate-verify-weight 递归和 LCB 准入，仅凭 3-67 条验证数据把具身 Code-as-Policy 基线从 4-68% 提升到 62-100%。一篇搭系统、一篇做训练，勾勒物理世界自进化智能体的完整图景。</description></item><item><title>真武是系统，不只是芯片：平头哥算力峰会上的超节点方法论</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-yunmo-zhenwu-ai-chip/</link><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-24-yunqi2026-yunmo-zhenwu-ai-chip/</guid><description>整理2026云栖大会平头哥算力峰会「云模之芯，共筑非凡」全场内容。平头哥副总裁李伟良梳理真武810E→M890→V900→J900的年度迭代节奏；阿里云王超给出「系统语义不在物理边界断裂」的真超节点判据；Kimi许欣然拆解100毫秒decode里的访存经济学；元戎启行曹通易与无界动力夏中谱讲物理AI的数据量级与快慢脑；圆桌与倚天CPU、磐脉920网卡两场把Agent时代的关键路径补完整。芯片竞争已从单卡指标转向系统协同。</description></item><item><title>机器人 Scaling Law 出现了吗？——徐梦迪的答案：有信号，但真正的分水岭是 in-context learning</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-21-xumengdi-scaling-law-signal/</link><pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-21-xumengdi-scaling-law-signal/</guid><description>清华叉院助理教授徐梦迪在「十字路口」提出：具身领域已出现预训练数据从十万到百万小时、held-out loss 随规模下降的 scaling 信号（如 Dyna-2 百万小时人类视频预训练），但 loss 与真机成功率脱钩，真正有意义的是『未见任务成功率随规模上升』的 scaling law；她判断当前主流 VLA 的『预训练+微调』范式对应 GPT-1 时刻，期待的是通过 prompting 适应个体偏好的 GPT-3 时刻。本文拆解她论证中的证据链、与世界模型路线的关系，以及数据定义由模型能力反推的行业机制。</description></item><item><title>ActionPiece 精读：用『物理秩一致性』拯救动作 tokenization 的关系保真</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-18-actionpiece-paper-reading/</link><pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-18-actionpiece-paper-reading/</guid><description>华中科大×DeepCybo 等七机构联合提出 ActionPiece：自回归 VLA 模型的动作 tokenizer 传统上只用 MSE 评估重建精度，但点态误差小不等于关系保真——压缩后『不同情境所需的差异化动作』可能被压缩、扭曲甚至反转。论文提出 Physical Rank Consistency（PRC）度量局部物理距离排序的保持，并通过对表示学习与量化的联合监督（物理秩保持+量化正则）让离散 token 保留连续动作空间的局部序结构。同一 Qwen3-VL-4B 策略训练设置下：LIBERO 94.8%、未见 LIBERO-Plus 68.8%、SimplerEnv 71.9%。</description></item><item><title>具身智能的四条路线分歧：数据、Astra 与商业化的真问题</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-14-embodied-intelligence-crossroads/</link><pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-14-embodied-intelligence-crossroads/</guid><description>2026 外滩大会圆桌实录：苏度韩铮、蚂蚁灵波沈宇军、自变量王潜、破壳许华哲四位一线创业者正面回答具身智能三场路线之争——仿真还是真机、GPT-6 Astra 是否构成降维打击、跨过泡沫的指标是什么。共识是具身不会复刻语言模型路径，分歧在数据从哪来、智能住在哪里、钱从哪来。</description></item><item><title>触觉是具身智能的最后一块拼图吗：五大技术路线、数据困局与模型之争</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-07-robot-tactile-sensing/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-07-robot-tactile-sensing/</guid><description>硅谷101走进触觉传感器实验室，梳理压阻、压电、电容、光学、磁感应五条技术路线，解释触觉数据为何接近于零、真机采集的&amp;quot;屏蔽悖论&amp;quot;，以及端到端时代触觉融入模型的两种范式之争。几乎所有受访嘉宾认为2026年触觉处于爆发前夜，量产一致性与数据生态是接下来一年的关键观察点。</description></item><item><title>Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-31-code-as-worlds-paper-reading/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-31-code-as-worlds-paper-reading/</guid><description>深度精读 MirroS 联合清华、北大、南洋理工的技术报告 Code as Worlds。论文提出用可执行代码表示物理世界的组成、演化与外观（EWR 三元组），把&amp;rsquo;从观测恢复世界表示&amp;rsquo;建模为溯因式的 agent 发现环：提出-实例化-执行-渲染-验证迭代修正。再用验证过的世界免费生成带精确物理量标签的 VQA 数据训练 VLM，9B 模型在 QuantiPhy 上 55.4 分超过 Gemini-3.1 Flash 的 54.8，27B 推理变体 58.6 分超过全部基线。</description></item><item><title>PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-plcbench-physical-impact-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-plcbench-physical-impact-paper-reading/</guid><description>浙大、西交大与布里斯托尔大学团队提出 PLCBench，首个真 PLC 硬件在环（HIL）评估框架，系统测量自主 LLM agent 能否把网络可达的工业控制器转化为持续物理影响。框架保留四家厂商原生协议语义，用六层隐藏诊断旗标与四源确定性取证把攻击能力拆解为接口获取与物理转化两段。在 4 台商用 PLC、4 个闭环工况、5 个模型、3 个种子共 240 个有效回合（118 聚合 PLC 小时）中，31.3% 达成持续物理影响，GPT 5.5 以 79.2% 覆盖全部 16 个格子，而最弱模型仅 10.4%；98 个回合停在接口获取，62 个停在物理转化，丰富观测使写后条件达成率提升 19.8 个百分点。本精读逐节拆解其设计因果链与防御启示。</description></item><item><title>Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-29-redwood-ai-accelerator-paper-reading/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-29-redwood-ai-accelerator-paper-reading/</guid><description>Architect Labs 的 AI 系统在两名人类架构师只写高层规范的前提下，用不到两周自主生成性能模型、RTL、UVM 验证环境、形式证明、固件与计算内核，所有模块达成 95% 代码与功能覆盖率，首次上 FPGA 零 bug，架构变更 48 小时内完成再验证再部署；投影至 Samsung 8nm 工艺后，Redwood 以 49 &lt;a href="mailto:tokens/s@1.335W"&gt;tokens/s@1.335W&lt;/a&gt; 对比 Jetson Orin Nano 同模型实测 28 &lt;a href="mailto:tokens/s@2.59W"&gt;tokens/s@2.59W&lt;/a&gt;，能效提升 3.4 倍。本精读面向软件背景读者，补全 RTL、UVM、形式验证、FPGA、roofline 等概念，剖析单一规范源全栈协同生成的因果链，并提炼七条可迁移的系统设计灵感。</description></item><item><title>世界模型是具身的永动机吗：北京人形谈 VLA 续命、大一统与机器人幼儿园</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-world-model-perpetual-motion-embodied/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-world-model-perpetual-motion-embodied/</guid><description>《晚点聊》WRC 期间对话北京人形创新中心戴勇、张怡与前华为 AI 专家唐都钰。VLA 与世界模型的路线之争被拆到表征层：VLA 泛化差的病根是&amp;quot;特征漏斗+预训练与后训练范式不一致&amp;quot;；世界模型则被戴勇称为&amp;quot;AI 时代的永动机&amp;quot;——指望它生产数据，它本身却缺数据，&amp;ldquo;至少到现在是个童话&amp;rdquo;。北京人形的答案是 Pelican-Unify 大一统强耦合路线，年底 2.0 要拿出具身领域的 scaling law；唐都钰离职创业做&amp;quot;主动式物理因果模型&amp;quot;，并转述图灵奖得主 Sutton 的机器人幼儿园设想。</description></item><item><title>物理AI的下一站：让AI发现人类不知道的方程</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-19-chen-yuntian-physical-ai-paradigm/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-19-chen-yuntian-physical-ai-paradigm/</guid><description>机器之心对话东方理工陈云天：从&amp;quot;知识嵌入&amp;quot;到&amp;quot;知识发现&amp;quot;的双向耦合范式。用AI从真实实验数据中找出人类未知的控制方程（如海浪破碎方程），用机械臂高通量实验找色谱方程替代耗时实验；他判断AI科学家&amp;quot;一定到了这个节点&amp;quot;，但资本市场节奏与湿实验闭环的天然慢速之间正在撕开一个gap，而通用模型&amp;quot;每个行业都很浅&amp;quot;，真正的机会在把物理一致性嵌进专业模型。</description></item><item><title>Addressable Memory for Video World Models 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-worldtrace-video-memory-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-worldtrace-video-memory-paper-reading/</guid><description>交互式视频世界模型在长时程生成中会遇到一个隐蔽的「记忆失效」问题：KV cache 里明明存着过去的画面，模型却读不出来。本文精读 NVIDIA/Princeton/ Toronto 联合提出的 WorldTrace 框架，它精准定位了 RoPE 旋转位置编码超出训练范围导致的「内容不可寻址」根因，并用一套无需训练的虚拟槽位机制，在时间一致性上提升 15.5%、在 LoopBench 情景回忆上提升 19.5%。本精读将从世界模型的记忆机制讲起，逐层揭开位置编码、相位抵消、虚拟槽位、规范 key 平均等关键技术。</description></item><item><title>N₀-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-03-n0-vtla-tactile-model-paper-reading/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-03-n0-vtla-tactile-model-paper-reading/</guid><description>N₀-VTLA由NeoteAI Team与复旦TEAI Team联合提出，是首个在大规模触觉数据上预训练的视觉-触觉-语言-动作（VTLA）基础模型。方法核心是三步训练配方：（1）在自建NeoData大规模视-触数据集上做视觉-触觉预训练学习广泛的接触先验；（2）分阶段触觉通路集成，用一个预测性触觉通路将大规模接触先验蒸馏为下游任务所需的精细运动调整；（3）ALTER——一种优势条件离线RL方法，将相对进展与轨迹事件比较转化为二元优势标签用于策略训练。N₀-VTLA赢得全部9个NeoReal真机任务，在20任务仿真套件上平均成功率63.8%（最强baseline π0.5为44.0%），ALTER训练策略在3个长程真机任务上达到75-95%成功率。</description></item><item><title>ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-03-odeworld-continuous-world-model-paper-reading/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-03-odeworld-continuous-world-model-paper-reading/</guid><description>深度精读 arXiv:2607.27924——首个用物理时间 ODE 取代离散下一步预测的连续时间潜在世界模型。PT-Flow 把&amp;rsquo;未来预测&amp;rsquo;重新定义为&amp;rsquo;在紧凑潜在空间里对一个连续速度场做积分&amp;rsquo;，靠动力学解耦 + 直接一阶 JVP 监督，一举绕开 JEPA 长期头疼的表示坍缩难题；还能做离散模型做不到的任意时刻查询和后向预测。在 LIBERO 视频生成上 PSNR 比离散 baseline 高 3 分以上，64 帧长程预测延迟仅 0.072 秒，并在 LIBERO-LONG 和真实双臂 AgileX 机器人上把策略成功率推到 SOTA。</description></item><item><title>PhiZero: A World Model Built Around Physical Language 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-phizero-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-phizero-paper-reading/</guid><description>PhiZero由中科院自动化所提出，通过自监督学习从野外视频中提取紧凑离散的&amp;rsquo;物理语言&amp;rsquo;表示世界状态转移，采用&amp;rsquo;先推理后渲染&amp;rsquo;范式：自回归VLM先推理物理语言序列，再由扩散解码器渲染为视频。4秒33帧视频仅需256个离散符号（比Wan2.2 VAE压缩175倍），在Physics-IQ、PhyGround、WorldModelBench、IntPhys2四个基准上物理一致性全面超越Sora 2、Cosmos3等。</description></item><item><title>ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-shadowdancer-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-shadowdancer-paper-reading/</guid><description>ShadowDancer提出影子对（shadow pairs）和跨影子预测（cross-shadow prediction），通过构造方式解决潜在动作模型的外观-动力学耦合问题。同一动力学轨迹在不同外观下重放，预测一个影子所需的表示必然是共享动力学本身。任何演示片段成为可复用动作资产，在新环境中重放无需动作标签、运动估计器或微调，跨五族动力学平均盲测胜率86%。论文揭示了&amp;rsquo;构造性不变量提取&amp;rsquo;的全新自监督范式。</description></item><item><title>SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-02-spatialcli-paper-reading/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-02-spatialcli-paper-reading/</guid><description>SpatialCLI提出Call-Learn-Internalize三阶段框架，教VLM先用空间专家工具（定位/分割/深度/姿态）学会组合感知，再通过双视图训练将专家能力内化为无工具推理。8B模型内化后无工具达72.7%、带工具达91.3%，均超越GPT-5.6 Sol。论文揭示了&amp;rsquo;工具→RL→内化&amp;rsquo;的渐进式能力蒸馏新范式。</description></item><item><title>Momenta IPO 后再访曹旭东：没有尽头的 AI，从智驾到家庭机器人的十年推演</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-24-momenta-cao-xudong-endless-ai/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-24-momenta-cao-xudong-endless-ai/</guid><description>Momenta 创始人曹旭东在 IPO 后接受「晚点聊LateTalk」专访，回顾十年创业历程，系统阐述智驾竞争格局（中国两三家、全球三四家的终局判断）、&amp;ldquo;一个飞轮两条腿&amp;quot;战略、每年十倍的智驾摩尔定律、从自动驾驶向家庭机器人的技术外溢逻辑，以及从 AI 研究员到 CEO 的认知进化——&amp;ldquo;一流的工作不是想出来的，是做出来的&amp;rdquo;。</description></item><item><title>具身原生的豪赌：蚂蚁灵波沈宇军，为什么坚持从传感器和视频里重训整个机器人模型？</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-22-embodied-native-model-ant-lingbo/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-22-embodied-native-model-ant-lingbo/</guid><description>蚂蚁灵波首席科学家沈宇军的深度访谈。他从GAN研究起步，经字节、蚂蚁研究院，最终主导蚂蚁灵波做机器人&amp;quot;大脑&amp;quot;。文章梳理了灵波最核心的技术主张——&amp;ldquo;具身原生&amp;rdquo;：不再沿用数字世界的模型做下游适配，而是从传感器、视频时序、单向MoE架构出发，为物理世界从头训练一套完整的机器人基础模型（V-Ren、DEPS、VLA 2.0、Video、Word六件套）。沈宇军也坦率谈到了数据是当前最大瓶颈、灵波为什么不做本体、以及他对&amp;quot;大脑落后于本体&amp;quot;这一行业判断。</description></item><item><title>世界模型这半年：XLR Labs 谈原生路线、4D 数据护城河与物理 AGI 的下半场</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-18-world-model-xlr-labs/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-18-world-model-xlr-labs/</guid><description>《晚点聊》WAIC 期间对话 XLR Labs（拓元智慧）三位核心成员。这家从 2022 年起就押注&amp;quot;原生世界动作模型&amp;quot;的公司，分享了它与 VLA、隐式世界模型的路线分歧，千万小时级 4D 真实交互数据如何构成护城河，以及从智慧零售切入工业物流的&amp;quot;以终为始&amp;quot;商业化逻辑——一个关于&amp;quot;预训练与后训练一致性&amp;quot;的scaling law故事。</description></item></channel></rss>