2026  1049

October  30

【每日AI前沿追踪】2026年10月03日 核心技术与产业动态速递

October 3, 2026 · 14 min

Actions with Receipts + ContractRL + Guarded Commits:Agent 治理三层——审计收据、修复契约、审批事务 三论文合读 精读

October 3, 2026 · 5 min

Agent 安全新前沿五重奏:技能链劫持、模因木马、无辜信使、内存分区与主动越权 合读精读

October 3, 2026 · 7 min

AutoCompact + Cross-Benchmark Transfer + AuraForge 三篇合读:编码 Agent 训练的三根支柱——上下文管理策略化、RLVR 泛化性、安全监督合成 精读

October 3, 2026 · 7 min

FloWright × InFlowOp 合读:用工作流进化工作流,为故障付费而不为流程付费 精读

October 3, 2026 · 7 min

GUI-HARVEST + DynaHarness + EvoGen-Harness 三篇合读:harness 自进化在 GUI、机器人、图像生成三条垂直域的落地

October 3, 2026 · 6 min

Harness 优化三重奏:Turbo Harness、ActiveSaddler 与 VeriHarness 精读

October 3, 2026 · 3 min

Harness 的有效性边界:Malena × Finding the Right Fit 合读精读

October 3, 2026 · 7 min

Mem++ + MemFit + Beyond Memory 三篇合读:Agent 记忆系统的三层递进——从非破坏性存储到信念状态 精读

October 3, 2026 · 8 min

PivotOPD + ComputerSD + EviRover 三篇合读:智能体在线训练的三个新维度 精读

October 3, 2026 · 5 min

SkillSpec + RASO + Prompt2Skill 三篇合读:Agent 技能优化的三条路线 精读

October 3, 2026 · 7 min

信用分配四重奏:给每一步发对奖励——FAULT、SHARPO、T2SPO、DARS 合读

October 3, 2026 · 7 min

后果化评测四重奏:当基准不再比对答案——Argo-Bench、DAYJOB、Incident-Arena、EurekaBench 合读

October 3, 2026 · 7 min

多智能体可靠性三种失败模式合读:Right Answers, Wrong States + Worse Together + The Delegation Danger Band 精读

October 3, 2026 · 8 min

当 Agent 开始自我改进,安全还剩什么?——自进化安全三部曲合读:SAVER × Safety Must Survive Self-Improvement × Sharpening Tax

October 3, 2026 · 8 min

当评测本身成为研究对象:Agent 评测方法学三支柱合读精读

October 3, 2026 · 9 min

经验的去处:Token 化训入权重,还是组件级路由分流?——X-Tree × Component Routing 合读精读

October 3, 2026 · 7 min

【每日AI前沿追踪】2026年10月02日 核心技术与产业动态速递

October 2, 2026 · 12 min

CheatBench × WorldAuditBench × RobustReview 精读:守住评测完整性的三道防线

October 2, 2026 · 9 min

CoordPoison × Pretext × TrustProbe × ActionGuard:Skill 生态的信任危机——攻防测四面体 精读

October 2, 2026 · 8 min

cua-swe-duet: 编码 Agent 基准的两条新轴(视觉×SWE 与 repo 级从零生成)合读

October 2, 2026 · 6 min

Harness 的三种缩放轴:Mid-Harness × STITCH × Turbo Harness 精读

October 2, 2026 · 10 min

Harness 自动进化三重奏:MILO、ScholarEvolve 与 Malena 精读

October 2, 2026 · 10 min

Skill 全生命周期三重奏:SkillFM、Prompt2Skill 与 SkillGym 合读精读

October 2, 2026 · 11 min

Skill 泛化性二重奏:GSO 的过拟合诊断与 Rep2Skill 的表征进化 精读

October 2, 2026 · 7 min

失败资产化二重奏:Agent Error Dataset 与 AREX-2 精读

October 2, 2026 · 8 min

编码 Agent 的安全边界与协作假象:Approval Laundering 与 OpenCollab 合读 精读

October 2, 2026 · 9 min

编码 Agent 训练三条线:token 效率、自验证与安全 精读

October 2, 2026 · 9 min

自进化的可信与规模化:False Frontiers、UniEvo-VL 与 CollabFlow 合读 精读

October 2, 2026 · 6 min

训练前沿四重奏:蒸馏外推、规模解除、奖励治理与具身评测 精读

October 2, 2026 · 8 min

September  405

【每日AI前沿追踪】2026年09月30日 核心技术与产业动态速递

September 30, 2026 · 12 min

AI娱乐的版本答案还没出现:从猫箱到动念引线,梁琛奇推演“烧token的新平台”

September 30, 2026 · 1 min

CAMG × Comet-9B:文件即记忆与程序状态推理的 Agent 训练新范式 精读

September 30, 2026 · 7 min

CoDeL × ReproBench:智能体安全的攻防共进化与漏洞复现评估 精读

September 30, 2026 · 3 min

CodeSkill × NanoHarness:技能抽象与 Harness 效应——被低估的智能体性能杠杆 精读

September 30, 2026 · 6 min

Codoku × SecProbe:可再生谜题与自适应出题的评估方法学双星 精读

September 30, 2026 · 4 min

Diffusion Reward Models × SOLO:成熟技术的首次规模化双案例 精读

September 30, 2026 · 9 min

EAPO × DN-MOPD × d-OPD:大模型训练信号的三个盲区与最小修正 精读

September 30, 2026 · 5 min

Encoder-Free Scaling Laws × UMM-Reflection:统一多模态模型的架构与反思双问 精读

September 30, 2026 · 6 min

Failure-Transparent Agents × FCD × CoSec:智能体安全的三个新失效面 精读

September 30, 2026 · 8 min

Gagar × SWE-MILE × CRR:代码智能体强化学习的细粒度信用分配三重奏 精读

September 30, 2026 · 6 min

Imprint Reader × ATD:权重更新与行为影子之间的双向桥 精读

September 30, 2026 · 6 min

MassAlloc Attention × DaRoPE:注意力计算分配与位置编码的双子星重构 精读

September 30, 2026 · 6 min

Opera × CER:长程编码智能体的评论家介入与早期奖励预测 精读

September 30, 2026 · 5 min

RepoReuse × VulContextBench:代码智能体的复用行为与安全证据审计 精读

September 30, 2026 · 9 min

SEABench × Audit the Scaffold × REUSE:递归自我改进的测量、理论与统计三重保障 精读

September 30, 2026 · 12 min

Self-Evolving Coding Agents × RE-0:从数字程序到物理世界的自进化智能体 精读

September 30, 2026 · 7 min

Skill2Env × QwenGyre × AgentPerfBench:智能体强化学习的数据、系统与推理三层基建 精读

September 30, 2026 · 8 min

SWE-Game × CUA-SWE:软件工程基准的游戏化与视觉化扩展 精读

September 30, 2026 · 5 min

TraceDance × Maintaining Benchmarks:Agent 行为基准的构建与作弊治理 精读

September 30, 2026 · 3 min

WideSWE × AsynCodeBench × RepoMAS:超越单仓库的软件工程智能体三重维度 精读

September 30, 2026 · 7 min

超级厄尔尼诺与大宗商品:一位期货研究员的供给冲击分析框架

September 30, 2026 · 1 min

【每日AI前沿追踪】2026年09月29日 核心技术与产业动态速递

September 29, 2026 · 8 min

Agent 安全攻击面三重奏精读:仓库红队、CoT 明文越狱与类型化决策投毒

September 29, 2026 · 11 min

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders 精读

September 29, 2026 · 7 min

MoMHa 与 SkillEvoReg 精读:Agent 资产的优化与正则

September 29, 2026 · 9 min

New LoRA Skills Should Read but Never Write 精读

September 29, 2026 · 6 min

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems 精读

September 29, 2026 · 5 min

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents 精读

September 29, 2026 · 6 min

WeEnv: The Environment for Agentic Reinforcement Learning at WeChat 精读

September 29, 2026 · 6 min

编码智能体经济学三重奏精读:成本行为、紧凑文档与上下文蒸馏

September 29, 2026 · 13 min

训练机制三重奏精读:对数线性稀疏注意力、置信度停止与跨段信用分配

September 29, 2026 · 9 min

证据时效性二重奏精读:过期文档投毒与分层协作记忆

September 29, 2026 · 4 min

评测与治理三重奏精读:多智能体协作、搜索后综合与执行控制

September 29, 2026 · 11 min

递归自改进的能力与安全双螺旋精读:DCE 自蒸馏与演化安全框架

September 29, 2026 · 7 min

【每日AI前沿追踪】2026年09月28日 核心技术与产业动态速递

September 28, 2026 · 5 min

Chat Template 像「开关」一样切换 LLM 的自我指涉语气 精读

September 28, 2026 · 4 min

PrimeScientist:让自主研究智能体学会战略性分配研究努力 精读

September 28, 2026 · 4 min

出题、卖题、判卷:AI数据行业的权力、红线与瓶颈

September 28, 2026 · 2 min

当VC合伙人转身去做儿童情绪课:于红谈AI时代到底该学什么

September 28, 2026 · 1 min

特斯拉FSD十三年:抠门老板、瞒报危机与端到端豪赌

September 28, 2026 · 1 min

【每日AI前沿追踪】2026年09月27日 核心技术与产业动态速递

September 27, 2026 · 6 min

AI 智能体行为的水印税与全模态 harness 双精读:Provenance Tax × Qwen3.8-Omni-Flash

September 27, 2026 · 7 min

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity 精读

September 27, 2026 · 5 min

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents 精读

September 27, 2026 · 4 min

Learning to Discover Interesting Mathematics 精读

September 27, 2026 · 4 min

【每日AI前沿追踪】2026年09月26日 核心技术与产业动态速递

September 26, 2026 · 9 min

Agent-Editing World Model 精读

September 26, 2026 · 7 min

Training Object Permanence in World Models 精读

September 26, 2026 · 6 min

内核证据检测与记忆家族隔离:Agent 基础设施二重奏精读

September 26, 2026 · 9 min

审批洗白、钱包拒绝服务与自主科研作弊:Agent 安全经济学三重奏精读

September 26, 2026 · 10 min

校准决策模型检测对齐失败与 Agent 轨迹防篡改:二重奏精读

September 26, 2026 · 7 min

环境演化、跨图溯因 SWE 与 AI 主导模型开发:RSI 三重奏精读

September 26, 2026 · 9 min

线性叠加、闭环 AI-for-AI 与角色解耦搜索:三篇前沿 Agent 论文精读

September 26, 2026 · 8 min

编码智能体规划、信任原生 Agent OS 与开源后训练配方:三篇系统论文精读

September 26, 2026 · 8 min

视觉编码基准与内核级失控遏制:PPTBench 与 Hard Stop 精读

September 26, 2026 · 10 min

【每日AI前沿追踪】2026年09月25日 核心技术与产业动态速递

September 25, 2026 · 7 min

ChronosAttack: Adversarial Tool Scheduling Attacks on LLM Agents 精读

September 25, 2026 · 5 min

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks 精读

September 25, 2026 · 5 min

Harness as a Language×Bounded Loops:Agent 脚手架的语言化与可验证化 精读

September 25, 2026 · 7 min

Just-in-Time Memory×EnSIMem:智能体记忆的读时策展与实体索引 精读

September 25, 2026 · 5 min

PACT: From Credit Assignment to Critic Alignment 精读

September 25, 2026 · 7 min

Schrödinger’s Code Repository: Have LLMs Learned SWE-bench or Memorized It? 精读

September 25, 2026 · 5 min

SkillApt×TwinCheck:技能何时加载与调用如何验证——反事实证据的双重应用 精读

September 25, 2026 · 8 min

SkillGym×VHD-Play:技能与环境从「外挂」到「内化」的两条路线 精读

September 25, 2026 · 8 min

StateComp×PaMER:长程智能体的历史压缩时机与记忆控制信号 精读

September 25, 2026 · 7 min

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents 精读

September 25, 2026 · 6 min

控制 token 注入×工具缓存逆转:Agent 安全与训练基础设施的两个隐蔽失效面 精读

September 25, 2026 · 7 min

「会说」到「会做」,隔着一天三百万的纠错账:云栖2026高德技术峰会复盘

September 24, 2026 · 1 min

【每日AI前沿追踪】2026年09月24日 核心技术与产业动态速递

September 24, 2026 · 4 min

22万部上新与0.047%爆款率:AI短漫剧的工业化拐点到了吗——云栖2026分论坛全记录

September 24, 2026 · 1 min

64%的企业在生产环境用AI,达标的只有4%:德勤云栖论坛的诊断——卡住企业的不是模型,是语义、流程和责任

September 24, 2026 · 2 min

7B模型挤进闭源效果区、8块钱的翻唱与六个人的偶像工厂:云栖2026视听生成专场复盘——当创作被批量供给,稀缺的成了故事与判断

September 24, 2026 · 4 min

88%的企业都在规模化上AI,只有14%拿到价值:云栖2026埃森哲专场复盘——卡住企业的不是模型,是数字核心

September 24, 2026 · 4 min

90%的东南亚企业要上AI智能体,先把POC逼进生产——云栖2026东南亚论坛复盘:卡住落地的不是模型,是通向生产的那条路

September 24, 2026 · 5 min

Agent 聪明之后,卡住企业的是数据:云栖2026『为 Agent 重塑 Data Agent 生态』论坛复盘

September 24, 2026 · 3 min

Agent 越能干,越等不起:云栖2026「AI实时数据智能」论坛复盘——卡住生产级 Agent 的不是模型,是数据的实时性、语义与断点

September 24, 2026 · 4 min

Agent的确定性从哪来:全栈平台与草台班子的双向奔赴

September 24, 2026 · 1 min

Agent越自主,越不能靠它自觉:云栖2026「从Demo到生产」论坛复盘——确定性是工程出来的,不是模型许诺的

September 24, 2026 · 2 min

AI十分钟改完合同之后,法律行业卖的还是判断:云栖2026「数智法途」法务论坛全景复盘

September 24, 2026 · 2 min

AI只答对22%的空间题:高德把二十年地图改成「数据不出域」的生意,千舆用三层证据链换到78分

September 24, 2026 · 1 min

AI重构跨境电商:先被改写的不是技术,而是工序、账本与决策权

September 24, 2026 · 1 min

一万个因子里挑出一个好策略,是金矿还是运气:云栖2026汇正财经专场全景复盘

September 24, 2026 · 1 min

不可缓存的 Token 流:CDN 如何把自己重写成 AI Agent 运行底座

September 24, 2026 · 3 min

为Agent重做数据层:湖库一体、多模数据库与KV Cache,云栖2026一场论坛的五个答案

September 24, 2026 · 1 min

云栖2026「Agent Sandbox 发布」复盘:毫秒级沙箱背后,是一笔休眠经济学与四道栅栏的账

September 24, 2026 · 2 min

云栖2026「构建 Agent Infra」复盘:沙箱把Agent的执行权收编,弹性把训练的门槛拆掉

September 24, 2026 · 1 min

云栖2026云通信论坛:当打电话发短信变成 Agent 的活

September 24, 2026 · 1 min

云栖2026英特尔专场:Agent 时代,CPU 为什么重返 C 位

September 24, 2026 · 4 min

云栖2026观察:当出海的中国公司开始用AI重写增长公式

September 24, 2026 · 1 min

从Intelligence到Action:云栖2026零售论坛上,七家企业聊透Agent落地生死线

September 24, 2026 · 1 min

从卖 token 到卖任务:灵骏把 AI 超级计算机重构成一台 Agentic 任务机器

September 24, 2026 · 3 min

从繁星到灯塔:当数据耗尽成为时间表,AI 竞争换到了哪条赛道

September 24, 2026 · 2 min

代码几分钟就能写完,交付为什么还没变快:云栖2026 Qoder论坛复盘——个人提效和组织提效之间,隔着一套Harness

September 24, 2026 · 5 min

代码能自动生成,共识不会:Qoder 五人七天之后,下一个同事是硅基的

September 24, 2026 · 1 min

以智治算:当复杂度十倍于人力增长,阿里云把算力基础设施运维拆成三级进化

September 24, 2026 · 2 min

入口在嘴、生产力在 Agent:语音大模型降价九成五之后——云栖2026 千问语音论坛综述

September 24, 2026 · 3 min

制造AI的真瓶颈不是模型聪明,而是物理世界不听话:云栖2026『智造·未来』先进制造AI论坛全景复盘

September 24, 2026 · 2 min

单柜650千瓦之后,AI服务器成了一道系统工程题——磐久超节点论坛全复盘

September 24, 2026 · 1 min

台前是 Agent,台后是数据湖:云栖 2026 多模态数据训推论坛复盘

September 24, 2026 · 3 min

在像素里复刻世界之前,先给世界定一把尺——云栖2026世界模型分论坛全记录

September 24, 2026 · 1 min

存储走上业务关键路径:从具身数据洪流到Agent记忆底座——云栖2026「AI and Agentic存储解决方案」五场景实录

September 24, 2026 · 1 min

实时生成视频的想象空间:从创作工具到“使用即消费”——对话生数科技 Vidu S 负责人张金涛

September 24, 2026 · 2 min

密算织域:当数据敢出域,云上AI才敢进生产——云栖2026蚂蚁密算专场全记录

September 24, 2026 · 1 min

当 Agent 成为存储的头号用户:云栖2026「Agent Native 数据基础设施论坛」九连发全景复盘

September 24, 2026 · 4 min

当Agent成为云平台的客户:千问AI平台把简单留给人,把复杂留给机器

September 24, 2026 · 2 min

当Agent成为数字员工,云的生意从"给算力"变成"管治理"

September 24, 2026 · 3 min

当AI学会无中生有,传媒业最稀缺的资产变成实事求是:云栖2026「AI+文化传媒」论坛全景复盘

September 24, 2026 · 2 min

当AI能跑完整个科研:云栖2026教育科研论坛的跃迁证据与三道裂缝

September 24, 2026 · 1 min

当KV Cache超过模型权重,互连成了开放标准的战场——云栖超节点开放互连论坛复盘

September 24, 2026 · 1 min

当Token成本成为AI的商业闭环:一场论坛透出的下一代智算基础设施全景

September 24, 2026 · 4 min

当互联网的主体不再是人:云栖2026云网络专场,网络为什么重新变成稀缺品

September 24, 2026 · 2 min

当数百亿Agent开始’租房’:云栖Agentic Cloud论坛上,云的每一层都被重写了一遍

September 24, 2026 · 4 min

当智能体开始给操作系统「付房租」:云栖2026上一场关于Agent OS定义权的暗战

September 24, 2026 · 1 min

当智能体替人跑一整天:PAI 分论坛上,AI 基础设施的三次换轨

September 24, 2026 · 3 min

当网络的KPI变成Token:云栖2026可预期网络2.0论坛的机制与分歧

September 24, 2026 · 1 min

总量不缺电,缺的是"天选之地"的电网:云栖2026算电协同论坛的五个判断

September 24, 2026 · 1 min

技术门槛降到了地板上,人为什么还没涌进来:云栖2026女性论坛复盘——从意愿到行动之间,隔着心理、资本和训练数据三道门

September 24, 2026 · 2 min

把视频做小,再把视频做大:云栖2026上的AI原生视频云

September 24, 2026 · 1 min

拿掉AI就不成立的游戏、夜班跑素材的发行管线与3到5人全能小队:云栖2026阿里云AI游戏论坛复盘——供给爆炸之后,稀缺的是判断与把关

September 24, 2026 · 3 min

搜索的下一位主力用户是 Agent:从千亿向量租户到 per-token 信息密度

September 24, 2026 · 3 min

效率会被拉平,壁垒藏在业务里:云栖2026大模型解决方案论坛的五个落地样本

September 24, 2026 · 1 min

数据平台正在易主:当Agent成为头号用户,语义和失控成了新账单

September 24, 2026 · 1 min

数据库的下一个用户不是人:云栖2026上被Agent重构的数据库,先把试错成本打到近零

September 24, 2026 · 3 min

智能体的下半场之争:当对话AI学会自己改自己——云栖2026『伶鹊自进化智能体』分论坛复盘

September 24, 2026 · 2 min

服务量能涨十倍,人招不了十倍:云栖2026论坛上的三道天花板与三个飞轮

September 24, 2026 · 1 min

模型每2.8天更新一次、企业采购却要等半年:云栖2026百炼专场复盘——从 Model 到 Token,卡住企业的不再是模型本身

September 24, 2026 · 4 min

没有攻击者的入侵:Agent安全的真问题从「防住别人」变成「管住自己」

September 24, 2026 · 1 min

湖仓交棒Agent:Paimon生态补齐技术课后,语义层成了最后也最贵的一公里

September 24, 2026 · 1 min

真武是系统,不只是芯片:平头哥算力峰会上的超节点方法论

September 24, 2026 · 1 min

硬件会折旧,数据会复利:云栖2026把「说过的话」做成了一盘生意

September 24, 2026 · 2 min

端上生长:当车载大模型逼近云侧能力,意图经济开始计价

September 24, 2026 · 1 min

第一稿免费之后:云栖2026 AI原生设计专场上的品味通胀与经验资产化

September 24, 2026 · 1 min

算力像水电之后,大学真正的对手是「先学后用」——云栖2026「AI+校园」论坛纪要

September 24, 2026 · 1 min

算力单位不再是芯片,而是Pod——云栖UALink实践专场:开放互连从规范走向落地的第一年

September 24, 2026 · 1 min

算力是廉价的,搬运是昂贵的:云栖2026算力与存储论坛的五条链

September 24, 2026 · 4 min

算力的「兑现率」战争:平头哥开源SAIL软件栈,真武客户交出四本账

September 24, 2026 · 2 min

装上了AI,组织却没变:云栖2026「AI原生·重塑企业生产力」论坛复盘——红利卡在工具层与组织层之间

September 24, 2026 · 2 min

账号开了,产能没来:云栖2026「智启新程」四家企业把AI从工具熬成资产

September 24, 2026 · 1 min

速度算得出,去处算不出:云栖2026「无法计算的价值」主论坛,把台中心让给了大学、医院、县城和种子

September 24, 2026 · 2 min

造得出不再稀缺,长得起才是本事:友盟+ ADK 云栖首发,增长的入口与打法都要重做一遍

September 24, 2026 · 2 min

重构云接口:交互坍缩98%之后,云的难题变成身份、权限与那10%的失控

September 24, 2026 · 1 min

【每日AI前沿追踪】2026年09月23日 核心技术与产业动态速递

September 23, 2026 · 9 min

Agent进入真实业务之后,卡住的不再是模型:云栖2026『企业级Agent实践峰会』全景复盘

September 23, 2026 · 5 min

Critical-State RL:为多轮工具调用诊断「可训练的模型调用」 —— 精读

September 23, 2026 · 4 min

DSec(DeepSeek Elastic Compute): 面向规模化 Agentic 训练的可扩展沙盒基础设施 精读

September 23, 2026 · 5 min

Emergent Collusion:无恶意指令下,双智能体如何自发串通 精读

September 23, 2026 · 2 min

FLARE:用生成式奖励模型为长程编码智能体提供全生命周期稠密监督 —— 精读

September 23, 2026 · 4 min

Harness-Zero:通过 Agent-as-Harness 实现 Harness 蒸馏——论文精读

September 23, 2026 · 6 min

IER-OPD: 1% 的 Token 就够——论同策略蒸馏中的梯度估计 精读

September 23, 2026 · 5 min

One to More, More to One:面向软件工程 Agent 的类别感知迭代专家训练(类别感知 SWE 专家训练精读)

September 23, 2026 · 6 min

onPanda: 通过 Token 级纠错高效标注 LLM 与 Agent 的同策略对齐数据 精读

September 23, 2026 · 6 min

OSWorld-Pro:用过程式评测给 Computer-Use Agent 做「分步体检」

September 23, 2026 · 5 min

RoboDawn:让通用 VLM 零训练接管机器人控制,以及 GPT-6 Astra 的零样本佐证 精读

September 23, 2026 · 3 min

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses 精读

September 23, 2026 · 5 min

VibeMemBench:在真实仓库任务上用可执行测试受控评测 Coding Agent 记忆系统

September 23, 2026 · 4 min

WorldCrafter:用隐式 3D 感知记忆打造可一致探索的视频世界模型 精读

September 23, 2026 · 3 min

从卖 Token 到交付结果:云栖 MaaS & Agent 主论坛,阿里把’智能的价值密度’摆上台面

September 23, 2026 · 5 min

当昂贵的 GPU 开始等便宜的数据:云栖 2026 存储专场,CPFS 商用与 KVCacheStore 首发背后的一笔账

September 23, 2026 · 4 min

当访问网站的主体变成 AI:云栖万网专场,把’官网’从名片改写成获客生意

September 23, 2026 · 2 min

把思考变成电:2026 云栖开幕式主论坛的路线图与缺口

September 23, 2026 · 3 min

攻防进入机器速度之后,防御也只能交给 Agent:云栖2026『模型时代』安全论坛全景复盘

September 23, 2026 · 4 min

财务AI跑通月结与付款之后,卡住企业的不再是模型:云栖2026『专业决策,智能执行』财务分论坛全景复盘

September 23, 2026 · 4 min

越接近AGI,人类为什么反而开始害怕:一场关于刹车的四方对话

September 23, 2026 · 1 min

【每日AI前沿追踪】2026年09月22日 核心技术与产业动态速递

September 22, 2026 · 7 min

Agent 技能自进化二重奏:EVOLVE 与 GraphSkillEvo 精读

September 22, 2026 · 5 min

Agent 评测方法学三重奏:Next-Turn 指标为何失灵 精读

September 22, 2026 · 6 min

AI 可信性二重奏:递归评审崩塌与模型测谎仪 精读

September 22, 2026 · 4 min

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself 精读

September 22, 2026 · 6 min

EvoOntology: A Self-Evolving Ontology Layer for Data Agents 精读

September 22, 2026 · 5 min

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence 精读

September 22, 2026 · 4 min

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents 精读

September 22, 2026 · 2 min

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? 精读

September 22, 2026 · 4 min

信号质量二重奏:CoVer 验证器协同训练与 DENSE 轨迹蒸馏 精读

September 22, 2026 · 4 min

统一智能体双璧:MintAct 空间统一与递归语言模型推理统一 精读

September 22, 2026 · 4 min

【每日AI前沿追踪】2026年09月21日 核心技术与产业动态速递

September 21, 2026 · 6 min

Self Improvement via Fast Tree-search 精读

September 21, 2026 · 4 min

StudentSim: Training LLM-based Student Simulators 精读

September 21, 2026 · 4 min

机器人 Scaling Law 出现了吗?——徐梦迪的答案:有信号,但真正的分水岭是 in-context learning

September 21, 2026 · 2 min

郁金泰:一个医生和遗忘的赛跑|40%可预防、15年窗口与最后的照护

September 21, 2026 · 1 min

【每日AI前沿追踪】2026年09月20日 核心技术与产业动态速递

September 20, 2026 · 5 min

Cache-to-Cache: Direct Semantic Communication Between Large Language Models 精读

September 20, 2026 · 3 min

JEPA-Anything: Learning Predictive Models across Different Worlds 精读

September 20, 2026 · 3 min

Memory Compression for High-Fanout Agent Sandboxes (AgentZip) 精读

September 20, 2026 · 3 min

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It 精读

September 20, 2026 · 3 min

【每日AI前沿追踪】2026年09月19日 核心技术与产业动态速递

September 19, 2026 · 7 min

An Empirical Study of Harness Design for Coding Agents 精读

September 19, 2026 · 2 min

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression 精读

September 19, 2026 · 3 min

OverclaimBench × PACT:智能体可信性评测二重奏 精读

September 19, 2026 · 2 min

RetireOPD × EOS 终止符失配:在线策略蒸馏动力学二重奏 精读

September 19, 2026 · 2 min

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI 精读

September 19, 2026 · 2 min

SKILLAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback 精读

September 19, 2026 · 2 min

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness 精读

September 19, 2026 · 3 min

【每日AI前沿追踪】2026年09月18日 核心技术与产业动态速递

September 18, 2026 · 11 min

ActionPiece 精读:用『物理秩一致性』拯救动作 tokenization 的关系保真

September 18, 2026 · 2 min

Agent 安全四重奏精读:TrustPoison、Collective Loss of Control、CHASE 与 First Token Matters

September 18, 2026 · 3 min

Agora: Git as Shared Memory for Collective AutoResearch 精读

September 18, 2026 · 2 min

ComPO 零阶偏好对齐 与 SpectralShift 线性注意力长上下文扩展 精读(二重奏)

September 18, 2026 · 2 min

EvoSkill-GUI 精读:技能不是静态文档,而是能自我修订的活知识

September 18, 2026 · 2 min

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence 精读

September 18, 2026 · 3 min

ProgramDistill 精读:从交互式 Web 应用逆向蒸馏可验证的 SWE 任务基准

September 18, 2026 · 2 min

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening 精读

September 18, 2026 · 3 min

ScienceIDE: Turning World’s Scientific Codebase into Agent Learnable Environments 精读

September 18, 2026 · 3 min

SSD-LLaMA 精读:单张 RTX 5090 跑万亿参数 MoE 的 SSD 原生推理系统

September 18, 2026 · 2 min

XConf(Confidence Comes from Experience)与 Not All Agents Are Equal 精读:Agent 可信性的两翼

September 18, 2026 · 3 min

敢把钱包交给AI吗:Agent交易爆发前夜,卡住的不是模型是信任

September 18, 2026 · 1 min

【每日AI前沿追踪】2026年9月16日 核心技术与产业动态速递

September 17, 2026 · 8 min

After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem 精读

September 17, 2026 · 2 min

Agentic Societies Need a Social Harness 精读

September 17, 2026 · 3 min

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries 精读

September 17, 2026 · 3 min

Continual Learning Mechanisms Compose for Long-Horizon Memorization 精读

September 17, 2026 · 2 min

ExecuCritic × AgentGuard × RepoAtlas × Protocol Trimming 精读:编码智能体可靠性四重奏

September 17, 2026 · 3 min

Mo’ Models, Mo’ Problems × Co-Skill 精读:多智能体选型与边云技能演化双视角

September 17, 2026 · 2 min

OPEN-1B: A Fully Auditable Training Run 精读

September 17, 2026 · 2 min

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents 精读

September 17, 2026 · 3 min

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act 精读

September 17, 2026 · 2 min

State of Thought Enables Endogenous Reasoning 精读

September 17, 2026 · 2 min

从背调被拒到 65k Star:Archify 作者的开源自证之路

September 17, 2026 · 2 min

大脑归基模,小脑归自己:24 岁首席科学家王家伟的具身下注

September 17, 2026 · 2 min

【每日AI前沿追踪】2026年09月16日 核心技术与产业动态速递

September 16, 2026 · 10 min

AlgoEvo × MOSCOPT × SkillLift:Skill 优化三部曲 精读

September 16, 2026 · 2 min

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents 精读

September 16, 2026 · 2 min

Atria Dawn: The Dawn of Agentic Superintelligence 精读

September 16, 2026 · 2 min

Dream-RSI: Recursive Self-Improvement through Evolving Worlds 精读

September 16, 2026 · 2 min

Fabrication After Tool Failure × Why LLM Agents Collapse:Agent 诚实性与执行差距双面镜 精读

September 16, 2026 · 2 min

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning 精读

September 16, 2026 · 2 min

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement 精读

September 16, 2026 · 2 min

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding 精读

September 16, 2026 · 2 min

PMPA × SkillSecurer × SkillAtlas:Skill 与记忆安全三连 精读

September 16, 2026 · 2 min

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments 精读

September 16, 2026 · 2 min

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use 精读

September 16, 2026 · 2 min

Shopee跨境十年复盘:唯一正确的判断,是市场规模

September 16, 2026 · 1 min

SkillSeam: Six Principles for Auditing Agent Skill Collections 精读

September 16, 2026 · 2 min

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science 精读

September 16, 2026 · 2 min

SWEADV × VLoc Bench:Agent 安全评测双警报 精读

September 16, 2026 · 2 min

The Router Within: Eliciting Native Skill Routing from a Frozen LLM(Gavel)精读

September 16, 2026 · 2 min

Thought without systematicity? Evaluating Reasoning Models on Rule Induction Tasks 精读

September 16, 2026 · 1 min

Using Agentic AI for Contextualized and Multifaceted Code Review at Ericsson 精读

September 16, 2026 · 2 min

When Agents Slow Down: Elo-per-token 分析与 Agent 测试时策略 精读

September 16, 2026 · 2 min

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search 精读

September 16, 2026 · 1 min

推理芯片之战:带宽成为新算力,Groq、Cerebras 与 OpenAI 的三条路线与 Bill Dally 的 location 哲学

September 16, 2026 · 3 min

观众看不懂,就是我失败——易中天×闲木鱼,相隔五十年的历史创作者对谈

September 16, 2026 · 1 min

【每日AI前沿追踪】2026年9月14日 核心技术与产业动态速递

September 15, 2026 · 9 min

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization 精读

September 15, 2026 · 2 min

Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents 精读

September 15, 2026 · 2 min

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization 精读

September 15, 2026 · 2 min

DataFlex-RL: An Evaluation Platform for RLVR Data Policies 精读

September 15, 2026 · 1 min

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning 精读

September 15, 2026 · 2 min

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents 精读

September 15, 2026 · 2 min

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite 精读

September 15, 2026 · 2 min

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents 精读

September 15, 2026 · 2 min

LifeMem: Enabling Lifelong Experience Reuse for LLM Agents 精读

September 15, 2026 · 1 min

Look Before You Leap: Pre-Action Verification for LLM Agents 精读

September 15, 2026 · 2 min

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work 精读

September 15, 2026 · 2 min

One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering 精读

September 15, 2026 · 1 min

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering 精读

September 15, 2026 · 2 min

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents 精读

September 15, 2026 · 2 min

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing 精读

September 15, 2026 · 1 min

What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code 精读

September 15, 2026 · 2 min

增长越快,企业越危险:WHC 的“不卖货”增长法

September 15, 2026 · 1 min

【每日AI前沿追踪】2026年09月14日 核心技术与产业动态速递

September 14, 2026 · 4 min

Scan the Skill, Govern the Action 精读:agent 技能的「许可 ≠ 恶意」,66,192 个技能全语料测量出的运行时治理缺口

September 14, 2026 · 3 min

TraceMind 精读:用户到底记住了 LLM 写的什么?从交互轨迹预测人-LLM 共创中的信息摄取

September 14, 2026 · 2 min

具身智能的四条路线分歧:数据、Astra 与商业化的真问题

September 14, 2026 · 1 min

【每日AI前沿追踪】2026年09月12日 核心技术与产业动态速递

September 13, 2026 · 6 min

GLIE 精读:几何先验驱动的检索压缩——100 万页 258GB 到 1GB 的流形参数化

September 13, 2026 · 1 min

Image Tokenizers as Visual Languages 精读:统一多模态 tokenizer 的测量学

September 13, 2026 · 1 min

Memory as Plans 精读:把记忆从执行期条件重构为规划期证据,机器人非马尔可夫任务 SOTA

September 13, 2026 · 1 min

MetroLLM-Bench 精读:LLM 嵌入物理售票机,4B PEFT 学生超越 GPT-5.6 的容量-天花板曲线

September 13, 2026 · 1 min

Nemotron IMO Gold 精读:开源模型金牌的完整配方——自然语言证明生成与测试时搜索管线

September 13, 2026 · 1 min

Recursive Code World Models 精读:global-local-global 递归构造可执行 3D 世界

September 13, 2026 · 1 min

SpatialBlock 精读:合成积木课程与对照组设计的空间智能范式

September 13, 2026 · 1 min

World in World 精读:免训练控制视频世界模型的统一证据接口

September 13, 2026 · 1 min

X-AuT 精读:渐进剪枝+跨尺度蒸馏的语音编码器压缩,18→16 层错误率不降反升

September 13, 2026 · 1 min

【每日AI前沿追踪】2026年09月12日 核心技术与产业动态速递

September 12, 2026 · 9 min

A2ABreak 精读:把 A2A 协议规范编译成状态机之后,11 个新漏洞自己浮出水面

September 12, 2026 · 1 min

EvoSafeHarness 精读:Agent 安全没有万能线束,那就让线束自己进化

September 12, 2026 · 2 min

GenV 精读:把 Z3 等价性判定蒸馏成语言模型的“第六感”

September 12, 2026 · 1 min

IdeaAMBIG 精读:从论文想法到能跑的代码之间,隔着 660 个“没人写的细节”

September 12, 2026 · 1 min

Looped Flows 精读:把“想得更久”做进架构——循环流的局部训练之路

September 12, 2026 · 1 min

MCP 注册表随机抽样审计精读:48.8% 握手率背后的工具生态幸存者偏差

September 12, 2026 · 1 min

NCP-ArchPreview 精读:当语言模型开始预测“概念”而不是 token

September 12, 2026 · 1 min

NSD 精读:教推理模型“别这么错”,比教它“该怎么对”更有效

September 12, 2026 · 1 min

ReqEvolve 精读:ASE 2026 上的用户驱动软件自演化范式

September 12, 2026 · 1 min

SPDF × Silent Failures 精读:LLM 代码安全评测的双警报日

September 12, 2026 · 2 min

The Last AI Built by Humans 精读:RSI 五级自治框架与“结构递归 vs 有效递归”的证伪标准

September 12, 2026 · 2 min

UniMPA 精读:给 VLA 模型一个“动作锚定”的统一接口,训练 epoch 砍半还涨点

September 12, 2026 · 1 min

VP-Control 精读:Agent 提交门的“证据血统比模型多样性重要 3.6 倍”

September 12, 2026 · 1 min

When Synthetic Data Hurts 精读:Agent 技能检索器的合成数据灾难遗忘

September 12, 2026 · 1 min

今日精读补充:Auto-RecSys 与 ActReview——把’自主研究’装进工业 harness 与学术评审

September 12, 2026 · 1 min

阎学通:决策者个人利益,正在决定世界秩序的走向|王骁对话 正在发生 ep4

September 12, 2026 · 1 min

【每日AI前沿追踪】2026年09月11日 核心技术与产业动态速递

September 11, 2026 · 7 min

A-JIT: Agentic Just-In-Time Software Construction 精读

September 11, 2026 · 1 min

AgentGrad: Intervention-guided Prompt Optimization for Multi-Agent Systems 精读

September 11, 2026 · 1 min

Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization 精读

September 11, 2026 · 1 min

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations 精读

September 11, 2026 · 1 min

If It’s Not Buggy, Don’t Fix It: On the Dynamics of Iterative Bug-fixing with LLMs 精读

September 11, 2026 · 1 min

Programmable World Model 精读

September 11, 2026 · 1 min

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward 精读

September 11, 2026 · 1 min

RobustSGPO: Search-Space Control for Agent Harness Evolution 精读

September 11, 2026 · 1 min

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? 精读

September 11, 2026 · 1 min

Show-Harness: Just a VLM Agent Can Play Robots 精读

September 11, 2026 · 2 min

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks 精读

September 11, 2026 · 1 min

The Double Measurement Confound in Agent Benchmarks 精读

September 11, 2026 · 1 min

What Should an Agent Forget? Separating What Is Stored from What Is Used 精读

September 11, 2026 · 1 min

XAgent: eXecution-guided Agentic AI for GitHub Issues 精读

September 11, 2026 · 1 min

Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? 精读

September 11, 2026 · 1 min

包刚升:为什么大学生变得更沉默了|绩点通关游戏背后,是上一代更需要反思

September 11, 2026 · 1 min

【每日AI前沿追踪】2026年09月10日 核心技术与产业动态速递

September 10, 2026 · 8 min

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing 精读

September 10, 2026 · 2 min

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference 精读

September 10, 2026 · 2 min

CapScope: A Capability-Scoped Harness for Prompt-Injection-Resistant Coding Agents 精读

September 10, 2026 · 2 min

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails 精读

September 10, 2026 · 3 min

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation 精读

September 10, 2026 · 3 min

ERPO: Entropy-Regularized Rank-Masked Policy Optimization for Test-Time RL in Code Generation 精读

September 10, 2026 · 2 min

ExecCritic: Learn to Test, Test to Improve for Coding Agents 精读

September 10, 2026 · 3 min

Experience Funnel: A State–Policy Alternating Loop for Self-Evolving Agents 精读

September 10, 2026 · 2 min

Gander (Omni Interaction Agent Technical Report) 精读

September 10, 2026 · 3 min

MOLE: Detecting Insider Threats in AI Agents 精读

September 10, 2026 · 2 min

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 精读

September 10, 2026 · 3 min

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents 精读

September 10, 2026 · 2 min

Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions 精读

September 10, 2026 · 2 min

SWE-Bench Pro Verified + Shortcutting the Fix:SWE Agent 评测的可靠性双警报 精读

September 10, 2026 · 3 min

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets 精读

September 10, 2026 · 2 min

兰小欢:重要的事,大多无法预测|把握自己能把握的,做难的事

September 10, 2026 · 1 min

通专融合、转换层与不可外包的使命感:AI科学家周伯文的世界观

September 10, 2026 · 2 min

【每日AI前沿追踪】2026年09月09日 核心技术与产业动态速递

September 9, 2026 · 6 min

83亿虚拟人格,与一门押注未来的生意:AI模拟离预测人类还有多远

September 9, 2026 · 2 min

EmbodiedSkills:把 VLA 动作预测升级为提案-验证循环的技能编排框架 精读

September 9, 2026 · 2 min

FlowBalance:验证器锚定的自改进——把符号门控装进轨迹平衡 精读

September 9, 2026 · 2 min

Safety for Whom:把安全边界从话题级细化到边界级的数据配方 精读

September 9, 2026 · 2 min

SimpleMemVLA:不做记忆模块的具身记忆——原生视频上下文的范式反转 精读

September 9, 2026 · 2 min

Split-LLM 隐私失效审计:返回梯度的零模式完美暴露真实数据 精读

September 9, 2026 · 2 min

TGOPD:在策略蒸馏前先验证教师——提示级可靠性门控 精读

September 9, 2026 · 2 min

Uno:扩散增强 LLM 的无损加速范式——AR 与扩散在同一架构内的参数解耦 精读

September 9, 2026 · 3 min

我们进入了“中国周期”:IFA CEO 谈品牌坍塌、德国舒适区与下一个 iPhone

September 9, 2026 · 1 min

Bilevel Coordinated Reflection: 多智能体 LLM 系统的博弈论统一理论 精读

September 8, 2026 · 2 min

CoSkill: 把元技能变成可学习智能体 — 分层技能库的联合强化学习 精读

September 8, 2026 · 2 min

EvoHarnessBench: 智能体能跟上不断进化的 Harness 吗 精读

September 8, 2026 · 2 min

Extremely Sparse Supervision: 0.05% 的 token 监督就能激励推理 精读

September 8, 2026 · 2 min

HackProbe: 自进化语言模型的奖励黑客检测与免疫 精读

September 8, 2026 · 2 min

InterOPT/OR-Clarify: 运筹学建模中’何时该问’的选择性完备性决策 精读

September 8, 2026 · 1 min

Iris: Climbing to the Search Frontier — 开源搜索智能体的数据反构造与 SFT-RL 攀爬配方 精读

September 8, 2026 · 2 min

RISE 之外的第二条线:When Models Edit Too Much — 代码过编辑与最小编辑保真度 精读

September 8, 2026 · 2 min

RISE: 自外推策略蒸馏把 on-policy 蒸馏变成递归自我改进 精读

September 8, 2026 · 2 min

TROVE: 轨迹锚定的最小充分路线编辑 — 智能体编排的运行时修正 精读

September 8, 2026 · 2 min

What Does Multi-Harness RL Learn? — 评测 Harness 是 Agent RL 的主导变量 精读

September 8, 2026 · 2 min

τ^τ-Bench: 把’构建智能体’变成任务的端到端基准 精读

September 8, 2026 · 2 min

【每日AI前沿追踪】2026年09月07日 核心技术与产业动态速递

September 7, 2026 · 6 min

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding 精读

September 7, 2026 · 2 min

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions 精读

September 7, 2026 · 3 min

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training 精读

September 7, 2026 · 2 min

LatentPress: Context Compression Beyond Text and Vision 精读

September 7, 2026 · 3 min

Rethinking On-Policy Distillation II: One Training Example 精读

September 7, 2026 · 2 min

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM 精读

September 7, 2026 · 3 min

触觉是具身智能的最后一块拼图吗:五大技术路线、数据困局与模型之争

September 7, 2026 · 1 min

【每日AI前沿追踪】2026年09月06日 核心技术与产业动态速递

September 6, 2026 · 5 min

AutoTraceGT 精读:把扎根理论变成 Agent 轨迹的自动化显微镜

September 6, 2026 · 2 min

DRACO 精读:没有验证器时,如何给长程 Agent 训练信号分步定责

September 6, 2026 · 2 min

Locked at the Entrance 精读:RLVR 的多样性坍缩发生在推理的门口

September 6, 2026 · 2 min

MachCSL 精读:MIT 用 AI Agent 把 xv6 内核验证推进到 RISC-V 硬件级

September 6, 2026 · 2 min

RealSWE 精读:真实用户请求正在让编码 Agent 榜单失真

September 6, 2026 · 2 min

Refusing the Impossible 精读:代码幻觉不是代码错误——12 个模型在不可解任务上 60% 硬编

September 6, 2026 · 1 min

Requirements After the First Edit 精读:需求晚到正在让 Agent 会话里的代码作废翻倍

September 6, 2026 · 1 min

When Models Edit Too Much 精读:编码 Agent 的过度编辑病与保真度评测轴

September 6, 2026 · 2 min

所有 Skill 都会死:卡比谈驾驭大模型的三层功夫——上下文、方法论与长活 Agent

September 6, 2026 · 2 min

【每日AI前沿追踪】2026年09月05日 核心技术与产业动态速递

September 5, 2026 · 7 min

DeepMind 研究蜂群精读:当 100 个 AI 研究员自发作弊与吹哨

September 5, 2026 · 2 min

Environment Evolution 精读:让训练环境的难度离线进化

September 5, 2026 · 2 min

HarnessEvo 精读:Harness 自进化的价值藏在控制槽位里

September 5, 2026 · 2 min

HookPry 精读:Agent Harness 的 hook 更新通道是全新的供应链攻击面

September 5, 2026 · 2 min

PatchBench 精读:AI 漏洞修复的解决率被高估了 1.83 倍

September 5, 2026 · 2 min

Random Attention 精读:KV 缓存驱逐的选择信号几乎买不到任何东西

September 5, 2026 · 2 min

SWE-Gate 精读:通过功能测试对软件工程 Agent 并不够

September 5, 2026 · 2 min

Terminal-Universe 精读:把 Agent 轨迹逆向成可复用的训练环境

September 5, 2026 · 2 min

【每日AI前沿追踪】2026年09月04日 核心技术与产业动态速递

September 4, 2026 · 8 min

Aspire: Can Models Self-Evolve from Vague Goals? 精读

September 4, 2026 · 3 min

Cliff: Learning Process Rewards from the First Mistake 精读

September 4, 2026 · 3 min

Discriminative World Models for Web Agents 精读

September 4, 2026 · 3 min

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction 精读

September 4, 2026 · 2 min

Language Models Can Control Their Own Attention 精读

September 4, 2026 · 3 min

OpenAI工程师赵迪:从Twitter到OpenAI的25年,Navi推理引擎、马斯克的“两周法则”,与写代码这件事的重新定义

September 4, 2026 · 2 min

Post-Training Language Models for Gold-Medal Performance in Coding Competitions 精读

September 4, 2026 · 3 min

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills 精读

September 4, 2026 · 3 min

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? 精读

September 4, 2026 · 3 min

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory 精读

September 3, 2026 · 2 min

Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents — A Source-Code Study of Eleven Systems 精读

September 3, 2026 · 2 min

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement 精读

September 3, 2026 · 4 min

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? 精读

September 3, 2026 · 5 min

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution 精读

September 3, 2026 · 3 min

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents 精读

September 3, 2026 · 2 min

WHALE: A Simple Recipe for Joint Harness–Weight Optimization 精读

September 3, 2026 · 3 min

与曾鸣聊产业史观:公司会消亡,卓越必来自反共识,OpenAI与Anthropic大概率不是原生时代的大赢家

September 3, 2026 · 1 min

具身智能的金钱游戏:钱很多、花得很少,IPO 成了主竞赛

September 3, 2026 · 1 min

对话卷卷任立峰:从抖音从零到一到AI 3D,一个互联网人如何学当厂长,以及为什么基础模型吞噬不了制造业

September 3, 2026 · 1 min

配音演员以为能干到80岁,AI只用了三年就改了剧本:对话小连杀

September 3, 2026 · 1 min

【每日AI前沿追踪】2026年09月02日 核心技术与产业动态速递

September 2, 2026 · 8 min

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents 精读

September 2, 2026 · 2 min

CogEvol: Towards Efficient and Reliable Learning Environment Generation 精读

September 2, 2026 · 2 min

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement 精读

September 2, 2026 · 3 min

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability 精读

September 2, 2026 · 3 min

PaperGym: Rubric-Centered Evolution for Research-Plan Generation 精读

September 2, 2026 · 2 min

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase 精读

September 2, 2026 · 2 min

WebWorld: The Browser as a World Model for Self-Improving Web Code 精读

September 2, 2026 · 2 min

【每日AI前沿追踪】2026年09月01日 核心技术与产业动态速递

September 1, 2026 · 5 min

August  442

【每日AI前沿追踪】2026年08月31日 核心技术与产业动态速递

August 31, 2026 · 5 min

A Formal Limitation on Learning Human Language From Textual Corpora 精读

August 31, 2026 · 4 min

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families 精读

August 31, 2026 · 3 min

AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking 精读

August 31, 2026 · 3 min

Affix Cache for Diffusion Large Language Models 精读

August 31, 2026 · 3 min

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge 精读

August 31, 2026 · 5 min

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks? 精读

August 31, 2026 · 5 min

CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents 精读

August 31, 2026 · 2 min

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning 精读

August 31, 2026 · 3 min

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL 精读

August 31, 2026 · 5 min

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases 精读

August 31, 2026 · 3 min

Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense 精读

August 31, 2026 · 3 min

Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers 精读

August 31, 2026 · 3 min

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses 精读

August 31, 2026 · 5 min

Fairness Invariants 精读:用循环不变式思想定位并修复算法公平性缺陷

August 31, 2026 · 3 min

Fast Weight Attention for Continual Learning 精读

August 31, 2026 · 6 min

GameWAM: A World Action Model for Video Games 精读

August 31, 2026 · 3 min

HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees 精读

August 31, 2026 · 4 min

Logos: An Agent Harness on a Cross-Process Bus 精读

August 31, 2026 · 4 min

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering 精读

August 31, 2026 · 5 min

Lost in Compression 精读:抽取式提示压缩器的跨语言审计

August 31, 2026 · 3 min

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression 精读

August 31, 2026 · 3 min

Muon with Finite Newton-Schulz 精读:有限迭代不是误差,而是收敛的功臣

August 31, 2026 · 4 min

Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models 精读

August 31, 2026 · 3 min

Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification 精读

August 31, 2026 · 3 min

Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper 精读

August 31, 2026 · 4 min

openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents 精读

August 31, 2026 · 5 min

PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems 精读

August 31, 2026 · 4 min

Privacy Without Regret: Differentially Private Inference-Time Alignment 精读

August 31, 2026 · 4 min

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests 精读

August 31, 2026 · 4 min

REPLICANT: Learning Policies for Evading and Hardening Malware Detectors 精读

August 31, 2026 · 2 min

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation 精读

August 31, 2026 · 2 min

Sliding-window beats linear attention 精读

August 31, 2026 · 5 min

SPT: Skills as Pre-Training Data for Agentic Language Models 精读

August 31, 2026 · 4 min

String: An Agentic OS Where Every App Is a Markdown File 精读

August 31, 2026 · 5 min

Survival-Guided Length Control for Efficient Diffusion Language Models 精读:用生存分析一次前向测出生成长度

August 31, 2026 · 3 min

Sycophancy Suppression Can Impair Rational Updating 精读:抗谄媚不应牺牲理性纠错

August 31, 2026 · 3 min

TACIT-SWITCH: Cost-Aware Model Escalation for LLM Agents from Censored Supervision 精读

August 31, 2026 · 5 min

The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension 精读

August 31, 2026 · 5 min

Token-Level Advertising 精读

August 31, 2026 · 3 min

TreeGraft 精读:多起草器共嫁接一棵草稿树,让投机解码跨越快与好的单选题

August 31, 2026 · 3 min

VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation 精读

August 31, 2026 · 3 min

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning 精读

August 31, 2026 · 4 min

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents 精读

August 31, 2026 · 4 min

When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems 精读

August 31, 2026 · 4 min

When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models 精读

August 31, 2026 · 2 min

WM-R1: Training GUI Agents to Reason and Leverage World Models with Reinforcement Learning 精读

August 31, 2026 · 5 min

【每日AI前沿追踪】2026年08月30日 核心技术与产业动态速递

August 30, 2026 · 12 min

【每日AI前沿追踪】2026年08月29日 核心技术与产业动态速递

August 29, 2026 · 5 min

Agent Seer: Synthesizing Scenarios from Specification Understanding 精读

August 29, 2026 · 8 min

Agentic AI for operating scientific instruments for nanoscale characterization 精读

August 29, 2026 · 5 min

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment 精读

August 29, 2026 · 6 min

Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction 精读

August 29, 2026 · 6 min

DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration 精读

August 29, 2026 · 6 min

Diff Mining: Logit Differences Reveal Finetuning Objectives 精读

August 29, 2026 · 5 min

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows 精读

August 29, 2026 · 8 min

How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation 精读

August 29, 2026 · 5 min

J-Zero: Unified Challenger-Solver-Judge Co-Evolution from Zero Data 精读

August 29, 2026 · 7 min

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives 精读

August 29, 2026 · 7 min

Metis: Typed Runtime Mediation for Tool-Using Software Agents 精读

August 29, 2026 · 5 min

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact? 精读

August 29, 2026 · 7 min

ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving 精读

August 29, 2026 · 5 min

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution 精读

August 29, 2026 · 6 min

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI 精读

August 29, 2026 · 7 min

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents 精读

August 29, 2026 · 7 min

SKILL.state: Scalable Long-Horizon Agent Skills 精读

August 29, 2026 · 4 min

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning 精读

August 29, 2026 · 7 min

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts 精读

August 29, 2026 · 6 min

Vulnerable Code Search: Transferable Attack for Code Language Models 精读

August 29, 2026 · 6 min

【每日AI前沿追踪】2026年08月28日 核心技术与产业动态速递

August 28, 2026 · 13 min

Accelerating Scientific Research with Gemini in the Real-World 精读

August 28, 2026 · 2 min

ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices 精读

August 28, 2026 · 3 min

Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation 精读

August 28, 2026 · 1 min

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling 精读

August 28, 2026 · 2 min

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions 精读

August 28, 2026 · 3 min

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing 精读

August 28, 2026 · 1 min

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research 精读

August 28, 2026 · 2 min

Boosting LLM Exploration via Weak-Model Guidance in RLVR 精读

August 28, 2026 · 2 min

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable 精读

August 28, 2026 · 2 min

Circuit Condensation: Post-Training that Concentrates a Behavior’s Causal Circuit 精读

August 28, 2026 · 1 min

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms 精读

August 28, 2026 · 3 min

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes 精读

August 28, 2026 · 2 min

DeepChart: How Far are LLMs from Faithful Data-Science Chart Generation? 精读

August 28, 2026 · 2 min

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench 精读

August 28, 2026 · 3 min

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment 精读

August 28, 2026 · 2 min

LLMs Can Design Near-Optimal OR Algorithms 精读

August 28, 2026 · 2 min

MemToC: Benchmarking Memory–Tool Conflict Resolution in Large Language Models 精读

August 28, 2026 · 2 min

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation 精读

August 28, 2026 · 3 min

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance 精读

August 28, 2026 · 2 min

PAWBench: How Far Are We from Probabilistically Aligned World Modeling? 精读

August 28, 2026 · 2 min

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents 精读

August 28, 2026 · 4 min

Planting a Latent Variable in Natural-Looking Text: A More Realistic Test of Belief States in LLMs 精读

August 28, 2026 · 1 min

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems 精读

August 28, 2026 · 5 min

Same Model, Different Harness: Different Coding-Agent Results 精读

August 28, 2026 · 3 min

SARA: When Tool Outputs Become Commands 精读

August 28, 2026 · 2 min

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable 精读

August 28, 2026 · 2 min

SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models 精读

August 28, 2026 · 1 min

SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control 精读

August 28, 2026 · 2 min

TTPO: Test-Time Policy Optimization 精读

August 28, 2026 · 1 min

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO 精读

August 28, 2026 · 3 min

Unsaid, Unsafe? Implicit Security Obligations in LLM-Based RTL Code Generation 精读

August 28, 2026 · 3 min

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City 精读

August 28, 2026 · 1 min

When Context Gets Root: Privilege Escalation in LLM Harnesses 精读

August 28, 2026 · 2 min

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution 精读

August 28, 2026 · 3 min

一支疫苗,一个病人:Moderna三期读出之后,个性化癌症疫苗的科学与工程真相

August 28, 2026 · 2 min

【每日AI前沿追踪】2026年08月27日 核心技术与产业动态速递

August 27, 2026 · 16 min

A Few Pages of Markdown: Committed AI Configuration and Lower Quality Cost after Coding-Agent Adoption 精读

August 27, 2026 · 3 min

A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation 精读

August 27, 2026 · 4 min

Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems 精读

August 27, 2026 · 4 min

Adaptive Triggering for Bias Correction in LLM Reasoning 精读

August 27, 2026 · 3 min

Agentic Autoresearch for Cell-Edge Power Control 精读

August 27, 2026 · 3 min

AI在想什么:模型没说出口的推理,与可解释性唯一一次漂亮的兑现

August 27, 2026 · 2 min

Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence 精读

August 27, 2026 · 5 min

AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs 精读

August 27, 2026 · 5 min

Automata from Agent Traces: Failure and Next-Step Prediction 精读

August 27, 2026 · 4 min

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment(Station v2)精读

August 27, 2026 · 3 min

AWM: Answerable Working Memory for Long-Document VQA Agents 精读

August 27, 2026 · 4 min

Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization 精读

August 27, 2026 · 5 min

Beyond the Editing Canvas: Evidence Divergence in OOXML-to-LLM Ingestion 精读

August 27, 2026 · 5 min

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback 精读

August 27, 2026 · 3 min

Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks 精读

August 27, 2026 · 3 min

Candidate supply and answer selection shape the value of LLM judging in multi-agent systems 精读

August 27, 2026 · 3 min

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval 精读

August 27, 2026 · 2 min

Code World Model: Coding Agent as World Brain 精读

August 27, 2026 · 3 min

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild 精读

August 27, 2026 · 4 min

Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness 精读

August 27, 2026 · 3 min

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail 精读

August 27, 2026 · 2 min

From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis 精读

August 27, 2026 · 4 min

From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use 精读

August 27, 2026 · 3 min

FrontierChallenge: Evaluating Scientific Workflow Completion 精读

August 27, 2026 · 4 min

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs 精读

August 27, 2026 · 4 min

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in MoE LLMs 精读

August 27, 2026 · 4 min

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents 精读

August 27, 2026 · 5 min

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution 精读

August 27, 2026 · 4 min

Joint Optimization of Tool Creation and Use for Large Language Model Agents(SMITH)精读

August 27, 2026 · 6 min

Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory 精读

August 27, 2026 · 3 min

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization 精读

August 27, 2026 · 3 min

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation 精读

August 27, 2026 · 3 min

Meta^n: Recursive Self-Improvement through Emergent Depth 精读

August 27, 2026 · 3 min

Narcissus: Program Synthesis Using Context-Aware LLM Approximations 精读

August 27, 2026 · 4 min

Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning 精读

August 27, 2026 · 6 min

On-policy Distillation with Verifiable Reward 精读

August 27, 2026 · 2 min

Parason: Revealing Subtask- and Trial Parallelism in LLM Reasoning 精读

August 27, 2026 · 2 min

Praxist: From Experimental Artifacts to Solution Lineages 精读

August 27, 2026 · 4 min

Prefix Sliding for efficient test-time scaling 精读

August 27, 2026 · 4 min

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows 精读

August 27, 2026 · 3 min

Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context 精读

August 27, 2026 · 5 min

Recuris: Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses 精读

August 27, 2026 · 3 min

Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference 精读

August 27, 2026 · 3 min

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems 精读

August 27, 2026 · 3 min

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards 精读

August 27, 2026 · 3 min

ReproAgent: Contract-Guided Paper-to-Code Reproduction 精读

August 27, 2026 · 3 min

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation 精读

August 27, 2026 · 3 min

SimVerity: When Does Simulated Agent Success Survive Physical Deployment? 精读

August 27, 2026 · 4 min

Skill Issue: Are Skills Language-Invariant in LLMs? 精读

August 27, 2026 · 3 min

SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents 精读

August 27, 2026 · 4 min

SPECMINE: A Large-Scale Corpus of Spec-Driven Development Artifacts 精读

August 27, 2026 · 3 min

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon 精读

August 27, 2026 · 3 min

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments 精读

August 27, 2026 · 4 min

SwarmWorld: Stigmergic technological evolution in societies of language-model agents 精读

August 27, 2026 · 3 min

The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses 精读

August 27, 2026 · 4 min

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents 精读

August 27, 2026 · 2 min

ToolMinimize: Auditing and Rewriting LLM Agent Tool Calls to Minimize Privacy Exposure 精读

August 27, 2026 · 5 min

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development 精读

August 27, 2026 · 2 min

Training Alignment Auditors via Reinforcement Learning 精读

August 27, 2026 · 3 min

Tunable Tool-Call Rates in LLM Agents via Representation Steering 精读

August 27, 2026 · 2 min

Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training 精读

August 27, 2026 · 4 min

Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings 精读

August 27, 2026 · 5 min

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning 精读

August 27, 2026 · 3 min

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 精读

August 27, 2026 · 3 min

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction 精读

August 27, 2026 · 4 min

When “Must“ Becomes “Maybe“: Constraint Weakening in LLM Agent Workflows 精读

August 27, 2026 · 2 min

Where vs What: Decomposing Structural and Content Failures in LLM Structured Outputs 精读

August 27, 2026 · 2 min

【每日AI前沿追踪】2026年08月26日 核心技术与产业动态速递

August 26, 2026 · 12 min

领读Kimi K3技术报告:一个清华架构博士眼中的注意力谱系与「有效scaling」

August 26, 2026 · 1 min

【每日AI前沿追踪】2026年08月25日 核心技术与产业动态速递

August 25, 2026 · 11 min

Active Inference as Context Acquisition for AI Agents 精读

August 25, 2026 · 4 min

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale 精读

August 25, 2026 · 4 min

Apodex 1.1 姊妹篇补遗:本日精读系列导览与 2026-08-25 学术全景

August 25, 2026 · 1 min

Apodex 1.1: Scaling Agentic Intelligence for Complex Work 精读

August 25, 2026 · 2 min

AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification 精读

August 25, 2026 · 4 min

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization 精读

August 25, 2026 · 4 min

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces 精读

August 25, 2026 · 2 min

Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis 精读

August 25, 2026 · 3 min

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (ERPO) 精读

August 25, 2026 · 2 min

CatchBench: When Can an Agent Failure Be Caught? 精读

August 25, 2026 · 2 min

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? 精读

August 25, 2026 · 3 min

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces 精读

August 25, 2026 · 2 min

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (Risa) 精读

August 25, 2026 · 2 min

Don’t Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents 精读

August 25, 2026 · 5 min

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills 精读

August 25, 2026 · 3 min

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents 精读

August 25, 2026 · 5 min

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks 精读

August 25, 2026 · 2 min

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations 精读

August 25, 2026 · 4 min

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks 精读

August 25, 2026 · 1 min

Natural-Language Workflows Are Not Software Yet: Artifact-Driven Compilation for Reliable Agent Execution 精读

August 25, 2026 · 2 min

Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning (NFV) 精读

August 25, 2026 · 2 min

Prime Agent: A Self-Improving RLM Harness 精读

August 25, 2026 · 3 min

Repo2Skill-Evo: Repository Skills Go Stale in Silence 精读

August 25, 2026 · 2 min

Signal or Noise? A Benchmark Study of Agent Skills in Web Development 精读

August 25, 2026 · 2 min

SkillAlchemy: Open-World Agent Skill Creation 精读

August 25, 2026 · 2 min

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? 精读

August 25, 2026 · 2 min

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection 精读

August 25, 2026 · 1 min

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search (Ascp) 精读

August 25, 2026 · 2 min

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models 精读

August 25, 2026 · 2 min

What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels 精读

August 25, 2026 · 2 min

【每日AI前沿追踪】2026年08月24日 核心技术与产业动态速递

August 24, 2026 · 8 min

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement 精读

August 24, 2026 · 4 min

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL 精读

August 24, 2026 · 4 min

EnvHarness: Awakening Static Worlds for Agent Learning 精读

August 24, 2026 · 3 min

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis 精读

August 24, 2026 · 4 min

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills 精读

August 24, 2026 · 3 min

Inducing Task Models from Computer-Use Traces 精读

August 24, 2026 · 3 min

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking 精读

August 24, 2026 · 4 min

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use 精读

August 24, 2026 · 4 min

Phantom Gains: Auditing Self-Improvement Against a Measured Null 精读

August 24, 2026 · 5 min

PRAXIS: Graph-Grounded Tacit Knowledge for Domain Code Generation 精读

August 24, 2026 · 5 min

Q-Learning with World Models 精读

August 24, 2026 · 3 min

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance 精读

August 24, 2026 · 5 min

Repo0: Design-Driven Zero-to-All Code Generation 精读

August 24, 2026 · 4 min

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning 精读

August 24, 2026 · 5 min

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation 精读

August 24, 2026 · 3 min

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents 精读

August 24, 2026 · 5 min

SPADE: Self-Play in Adaptive Synthetic Executable Environments 精读

August 24, 2026 · 5 min

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? 精读

August 24, 2026 · 4 min

What Makes Software Issue Resolution Tasks Difficult for Agents? 精读

August 24, 2026 · 3 min

两天十万Star:DeepSeek Harness 的开放逻辑,与它想要驯服的模型-脚手架-算力飞轮

August 24, 2026 · 2 min

【每日AI前沿追踪】2026年08月23日 核心技术与产业动态速递

August 23, 2026 · 9 min

A Jagged Frontier: 代码Agent对语义保持变换的锯齿鲁棒性 精读

August 23, 2026 · 9 min

Adversarial Review: Structured Disagreement for Grounded Agentic Code Review 精读

August 23, 2026 · 7 min

Agent如何发现、阅读与书写技术文档:行为实证研究 精读

August 23, 2026 · 4 min

Can Agent Memory Systems Track Evolving State? StateMemBench 精读

August 23, 2026 · 7 min

Credit Without Ground Truth: 步级信用分配的执行回放审计 精读

August 23, 2026 · 7 min

MidTool: 面向Agent工具使用的中期训练数据合成 精读

August 23, 2026 · 6 min

MileGPO: 里程碑推断的长程Agent过程级信用分配 精读

August 23, 2026 · 5 min

One Success Isn’t Reliability: Thinkingbox 沙盒与基准 精读

August 23, 2026 · 7 min

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees 精读

August 23, 2026 · 6 min

ReCache: 工具增强Agent的组合不变KV缓存复用 精读

August 23, 2026 · 3 min

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection 精读

August 23, 2026 · 6 min

【每日AI前沿追踪】2026年8月22日 核心技术与产业动态速递

August 22, 2026 · 9 min

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement 精读

August 22, 2026 · 3 min

EnvHarness: Awakening Static Worlds for Agent Learning 精读

August 22, 2026 · 4 min

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis 精读

August 22, 2026 · 4 min

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving 精读

August 22, 2026 · 6 min

Inducing Task Models from Computer-Use Traces 精读

August 22, 2026 · 3 min

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use 精读

August 22, 2026 · 4 min

Phantom Gains: Auditing Self-Improvement Against a Measured Null 精读

August 22, 2026 · 3 min

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents 精读

August 22, 2026 · 4 min

Repo0: Design-Driven Zero-to-All Code Generation 精读

August 22, 2026 · 5 min

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? 精读

August 22, 2026 · 3 min

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See 精读

August 22, 2026 · 4 min

一条视频看懂风险投资:募、投、管、退的底层逻辑——钦文对话阿尔法公社刘罡

August 22, 2026 · 1 min

世界模型是具身的永动机吗:北京人形谈 VLA 续命、大一统与机器人幼儿园

August 22, 2026 · 2 min

回收只是拿到了门票:朱雀三号副总设计师董凯谈复用周期、不锈钢算盘与商业航天的组织革命

August 22, 2026 · 1 min

消失的1.65万亿:数据中心影子借贷、GPU金融化与AI次贷之辩

August 22, 2026 · 4 min

【每日AI前沿追踪】2026年08月21日 核心技术与产业动态速递

August 21, 2026 · 4 min

【每日AI前沿追踪】2026年08月20日 核心技术与产业动态速递

August 20, 2026 · 8 min

Agent Lightning v1.0: Towards Harnessed Agentic RL 精读

August 20, 2026 · 3 min

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements 精读

August 20, 2026 · 3 min

ASI-Bench: At the Dawn of Artificial Superintelligence 精读

August 20, 2026 · 2 min

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents 精读

August 20, 2026 · 1 min

Bounded Agents: Delegation Security for Multi-Agent AI Systems 精读

August 20, 2026 · 6 min

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL 精读

August 20, 2026 · 7 min

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents 精读

August 20, 2026 · 5 min

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents 精读

August 20, 2026 · 2 min

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety 精读

August 20, 2026 · 2 min

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents 精读

August 20, 2026 · 2 min

MoNe: Modular Neural Memory for Efficient Long Context Inference 精读

August 20, 2026 · 2 min

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 精读

August 20, 2026 · 5 min

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification 精读

August 20, 2026 · 1 min

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs 精读

August 20, 2026 · 1 min

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation 精读

August 20, 2026 · 6 min

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation 精读

August 20, 2026 · 4 min

SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution 精读

August 20, 2026 · 2 min

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents 精读

August 20, 2026 · 5 min

SPADE: Self-Play in Adaptive Synthetic Executable Environments 精读

August 20, 2026 · 6 min

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents 精读

August 20, 2026 · 2 min

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows 精读

August 20, 2026 · 2 min

Tim的执念:影视飓风的规模引擎,与一个"没那么聪明的智者"的更大自私

August 20, 2026 · 1 min

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents 精读

August 20, 2026 · 2 min

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence 精读

August 20, 2026 · 8 min

三位AI博士的真话:Token比人便宜吗,泡沫何时破,以及就业市场的一线行情

August 20, 2026 · 2 min

从烧钱竞赛到精打细算:一个Token重度用户的Agent进化史

August 20, 2026 · 1 min

【每日AI前沿追踪】2026年08月19日 核心技术与产业动态速递

August 19, 2026 · 9 min

Agentic Kernel Optimization: Generating State-of-the-Art GPU Kernels Without Hand-Written CUDA 精读

August 19, 2026 · 2 min

Agentic Transaction: Towards ACID-Compliant Agent Systems 精读

August 19, 2026 · 1 min

ClawGym II: Exploring Black-Box RL on Agent Harness 精读

August 19, 2026 · 2 min

HarnessEval-W: Agentifying the Evaluation of Visual Worlds 精读

August 19, 2026 · 2 min

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills 精读

August 19, 2026 · 2 min

Large Discovery Models: Empirically-Grounded Model-Based Open-Ended Search 精读

August 19, 2026 · 2 min

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures 精读

August 19, 2026 · 1 min

SA-MRPO: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization 精读

August 19, 2026 · 2 min

StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling 精读

August 19, 2026 · 3 min

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks 精读

August 19, 2026 · 1 min

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations 精读

August 19, 2026 · 2 min

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs 精读

August 19, 2026 · 1 min

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? 精读

August 19, 2026 · 2 min

物理AI的下一站:让AI发现人类不知道的方程

August 19, 2026 · 1 min

【每日AI前沿追踪】2026年08月18日 核心技术与产业动态速递

August 18, 2026 · 8 min

AgentRewind: Recoverable Execution for Long-Horizon LLM Agents 精读

August 18, 2026 · 1 min

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 精读

August 18, 2026 · 1 min

Demystifying Agent Skills: Why They Work—Until They Don’t 精读

August 18, 2026 · 1 min

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning 精读

August 18, 2026 · 1 min

Latent On-Policy Self-Distillation 精读

August 18, 2026 · 1 min

LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows 精读

August 18, 2026 · 1 min

MobileMem: Learning from a Year of Mobile Experiences 精读

August 18, 2026 · 1 min

RA-Bench: Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? 精读

August 18, 2026 · 1 min

Self-Supervised Visual On-Policy Distillation 精读

August 18, 2026 · 1 min

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning 精读

August 18, 2026 · 1 min

泡沫是2029年,不是2027:王煜全的荷塘、中间物种与击穿围墙的Agent

August 18, 2026 · 2 min

蒸馏风暴:门槛、灰色地带与一份没人愿意签字的竞赛规则

August 18, 2026 · 2 min

【每日AI前沿追踪】2026年08月17日 核心技术与产业动态速递

August 17, 2026 · 3 min

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence 精读

August 17, 2026 · 5 min

AQuA: Recursively Self-Improving Quantitative Trading Research Agents 精读

August 17, 2026 · 5 min

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 精读

August 17, 2026 · 4 min

Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories 精读

August 17, 2026 · 4 min

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test 精读

August 17, 2026 · 4 min

DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution 精读

August 17, 2026 · 5 min

GitSkills: A Dataset of Agent Skills on GitHub 精读

August 17, 2026 · 3 min

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents 精读

August 17, 2026 · 4 min

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback 精读

August 17, 2026 · 4 min

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents 精读

August 17, 2026 · 5 min

The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior 精读

August 17, 2026 · 4 min

【每日AI前沿追踪】2026年08月16日 核心技术与产业动态速递

August 16, 2026 · 7 min

Beyond Final Scores: 长程AI研发Agent过程级评测 精读

August 16, 2026 · 2 min

CAPRI: 契约感知的Isabelle证明修复 精读

August 16, 2026 · 3 min

DIVE: 多样性驱动的冻结模型技能进化 精读

August 16, 2026 · 2 min

ERSkill: 检索技能与路由器共进化 精读

August 16, 2026 · 4 min

Full-bandwidth transformer 精读

August 16, 2026 · 5 min

Practice Makes Unsafe: 技能误进化 精读

August 16, 2026 · 3 min

RippleMem: 从孤立检索到联想式回忆 精读

August 16, 2026 · 3 min

SkillEvo: 多轮交互反馈的自更新进化梯度 精读

August 16, 2026 · 2 min

SkillShapley: 技能步级Shapley归因 精读

August 16, 2026 · 3 min

【每日AI前沿追踪】2026年08月15日 核心技术与产业动态速递

August 15, 2026 · 4 min

A Programming Paradigm for Spatiotemporal Composability 精读

August 15, 2026 · 5 min

AI for Science 爆发前夜:曹原谈验证瓶颈、概念抽象与 AGI 的最后一公里

August 15, 2026 · 2 min

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World 精读

August 15, 2026 · 3 min

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design 精读

August 15, 2026 · 3 min

DarwinX: Evolving Agent Harnesses Through Natural Selection 精读

August 15, 2026 · 5 min

How Can Rhetoric Reward-Hack AI Reviewers? 精读

August 15, 2026 · 4 min

Massive Activations in Hybrid Linear Attention Large Language Models 精读

August 15, 2026 · 5 min

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 精读

August 15, 2026 · 3 min

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives 精读

August 15, 2026 · 3 min

QuoteBench: How Matched Scores Can Hide Command-Path Failures 精读

August 15, 2026 · 4 min

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models 精读

August 15, 2026 · 5 min

Thought-Level Beam Search for Reasoning 精读

August 15, 2026 · 3 min

Vero: Can AI Agents Build Formally Verified Software Repositories? 精读

August 15, 2026 · 3 min

【每日AI前沿追踪】2026年08月14日 核心技术与产业动态速递

August 14, 2026 · 4 min

Agent 安全应当是一份运行时契约——精读《Agent Safety Should Be a Runtime Contract》

August 13, 2026 · 7 min

AI4AI at Test-Time:通过脚手架实现强到弱的能力迁移

August 13, 2026 · 6 min

EvoX Genesis:用持久递归世界让软件自主演化

August 13, 2026 · 8 min

Spark-to-Paper:将端到端论文生成做成可组合技能

August 13, 2026 · 5 min

【每日AI前沿追踪】2026年08月13日 核心技术与产业动态速递

August 13, 2026 · 3 min

【论文精读】SkillZip:面向可扩展 Agent 技能库的契约保持图压缩框架

August 13, 2026 · 7 min

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence 精读

August 13, 2026 · 5 min

OpenART Arena: Scaling Agent Red Teaming via Open-Ended Environment Evolution 精读

August 13, 2026 · 12 min

从DeepSeek到Kimi K3,中国开源模型如何逼出黄仁勋的’开源联盟'

August 13, 2026 · 1 min

哲学到底是精英的表演,还是每个人都该追问的事?毕英杰与老蒋的两场跨洋对话

August 13, 2026 · 1 min

在算力最多的地方做世界模型:对话英伟达Cosmos掌舵人刘洺堉

August 13, 2026 · 1 min

论文工厂开进直播间,学术打假博主与评价机制的拉锯

August 13, 2026 · 1 min

【每日AI前沿追踪】2026年08月12日 核心技术与产业动态速递

August 12, 2026 · 5 min

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents 精读

August 12, 2026 · 7 min

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses 精读

August 12, 2026 · 9 min

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI 精读

August 12, 2026 · 6 min

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents 精读

August 12, 2026 · 9 min

Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution 精读

August 12, 2026 · 9 min

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure 精读

August 12, 2026 · 6 min

【每日AI前沿追踪】2026年08月11日 核心技术与产业动态速递

August 11, 2026 · 7 min

Addressable Memory for Video World Models 精读

August 11, 2026 · 8 min

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning 精读

August 11, 2026 · 9 min

BONSAI: Evolvability-Guided Tree Search over Skills 精读

August 11, 2026 · 8 min

DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds 精读

August 11, 2026 · 9 min

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training 精读

August 11, 2026 · 8 min

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution 精读

August 11, 2026 · 6 min

P³: Joint Program-and-Proof Planning for Verified Code Generation 精读

August 11, 2026 · 11 min

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines? 精读

August 11, 2026 · 7 min

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States 精读

August 11, 2026 · 8 min

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs 精读

August 11, 2026 · 9 min

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents 精读

August 11, 2026 · 10 min

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent 精读

August 11, 2026 · 7 min

Stealing Reasoning Traces from Proprietary LLM APIs 精读

August 11, 2026 · 6 min

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 精读

August 11, 2026 · 8 min

TEPA: Revoking Stale Memories for Conflict-Robust Language Agents 精读

August 11, 2026 · 8 min

「模型能力已经够了,要卷就卷 Infra」|对话戴冠兰:从 Cloudflare 到 Runta,为十亿个 Agent 造执行底座

August 10, 2026 · 2 min

【每日AI前沿追踪】2026年08月10日 核心技术与产业动态速递

August 10, 2026 · 4 min

富二代为小六做Steam游戏:一场古典闹剧,与独立游戏开发者的真实生存图鉴

August 10, 2026 · 1 min

【每日AI前沿追踪】2026年08月09日 核心技术与产业动态速递

August 9, 2026 · 2 min

MASS:用权威共享状态解耦多智能体世界模型——多人游戏架构如何破解视频世界模型的可扩展性瓶颈

August 9, 2026 · 3 min

经济世界模型蓝图:从经济代理到代理经济——六级能力阶梯如何定义下一代AI驱动经济仿真

August 9, 2026 · 4 min

【每日AI前沿追踪】2026年08月08日 核心技术与产业动态速递

August 8, 2026 · 3 min

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay 精读

August 8, 2026 · 8 min

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 精读

August 8, 2026 · 8 min

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks 精读

August 8, 2026 · 7 min

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning 精读

August 8, 2026 · 6 min

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models 精读

August 8, 2026 · 8 min

When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories 精读

August 8, 2026 · 4 min

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents 精读

August 8, 2026 · 8 min

【每日AI前沿追踪】2026年08月07日 核心技术与产业动态速递

August 7, 2026 · 3 min

730会议与二季度4.3%:地方躺平,中央催钱,存量即增量

August 7, 2026 · 1 min

RSI比Coding Agent大得多:对话田渊栋,递归自进化为什么是阶段式突破而非渐进攀升

August 7, 2026 · 1 min

一片晶圆的赌局:Cerebras如何从 Scaling Law 的实验室走向千亿推理市场

August 7, 2026 · 1 min

高盛2026亚太展望:中国经济为何仍在「失衡」——出口撑增长、房地产未见底、养儿防老逻辑逆转

August 7, 2026 · 1 min

【每日AI前沿追踪】2026年08月06日 核心技术与产业动态速递

August 6, 2026 · 9 min

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment 精读

August 6, 2026 · 2 min

OPD-V: Visual On-Policy Self-Distillation with Modality Balance 精读

August 6, 2026 · 1 min

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation 精读

August 6, 2026 · 2 min

SafeCommit: Certifying When Memory-Grounded Agents May Safely Act 精读

August 6, 2026 · 2 min

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models 精读

August 6, 2026 · 2 min

特修斯之船:Kimi K3如何把Transformer的零件全部换掉,还能逼近前沿

August 6, 2026 · 1 min

【每日AI前沿追踪】2026年08月05日 核心技术与产业动态速递

August 5, 2026 · 3 min

【论文精读】Any-OPD:通过表示空间桥接实现异构流匹配模型的在策略蒸馏

August 5, 2026 · 4 min

AURORA-LM:连续潜在扩散语言模型精读

August 5, 2026 · 2 min

DAPD: Dual-Anchored Policy Distillation 精读

August 5, 2026 · 2 min

DiffusionGemma Technical Report 精读

August 5, 2026 · 1 min

Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories 精读

August 5, 2026 · 2 min

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks 精读

August 5, 2026 · 3 min

Progressive Agent Skill Generation via Reinforcement Learning (Skill-α) 精读

August 5, 2026 · 2 min

TARL:面向长期Agent的可执行记忆管理精读

August 5, 2026 · 5 min

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction 精读

August 5, 2026 · 6 min

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 精读

August 5, 2026 · 3 min

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning 精读

August 5, 2026 · 3 min

【每日AI前沿追踪】2026年08月04日 核心技术与产业动态速递

August 4, 2026 · 8 min

【每日AI前沿追踪】2026年08月03日 核心技术与产业动态速递

August 3, 2026 · 7 min

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement 精读

August 3, 2026 · 9 min

Infra 的浪漫与 AI 平权:盛颖从 SGLang 到 RadixArk

August 3, 2026 · 2 min

Mental World Modeling 精读

August 3, 2026 · 7 min

N₀-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens 精读

August 3, 2026 · 3 min

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow 精读

August 3, 2026 · 8 min

【每日AI前沿追踪】2026年08月02日 核心技术与产业动态速递

August 2, 2026 · 3 min

29岁空降改造腾讯混元:信任换来的不是模型,而是一支新军队

August 2, 2026 · 1 min

MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems 精读

August 2, 2026 · 3 min

Meta AI Proactive Memory Agent:记忆教练精读

August 2, 2026 · 2 min

Perception-Correction Distillation:多模态推理器感知蒸馏信用分配精读

August 2, 2026 · 1 min

PhiZero: A World Model Built Around Physical Language 精读

August 2, 2026 · 2 min

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 精读

August 2, 2026 · 3 min

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them 精读

August 2, 2026 · 3 min

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation 精读

August 2, 2026 · 2 min

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation 精读

August 2, 2026 · 2 min

【每日AI前沿追踪】2026年08月01日 核心技术与产业动态速递

August 1, 2026 · 8 min

【每日AI前沿追踪】2026年08月01日 核心技术与产业动态速递

August 1, 2026 · 3 min

LLMs Get Lost in Evolving User Intent 精读

August 1, 2026 · 5 min

何谓蒸馏?硅谷如何看中国开放模型逼近前沿

August 1, 2026 · 1 min

July  68

【每日AI前沿追踪】2026年07月31日 核心技术与产业动态速递

July 31, 2026 · 8 min

GPU其实很闲:AI Infra四层架构与榨干硅极限的效率革命

July 31, 2026 · 1 min

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement 精读

July 31, 2026 · 3 min

清华程序员很聪明:清程极智如何把Token成本砍掉75%——AI Infra创业的降本逻辑

July 31, 2026 · 1 min

美团领投月之暗面A轮背后的故事:叶奇意亲历中国两代AI十年人才迁徙

July 31, 2026 · 1 min

【每日AI前沿追踪】2026年07月30日 核心技术与产业动态速递

July 30, 2026 · 8 min

【每日AI前沿追踪】2026年07月29日 核心技术与产业动态速递

July 29, 2026 · 7 min

几亿行代码的业务逻辑,AI复刻不了:对话SAP原欣,谈大模型to B的颠覆与边界

July 29, 2026 · 1 min

【每日AI前沿追踪】2026年07月28日 核心技术与产业动态速递

July 28, 2026 · 6 min

【每日AI前沿追踪】2026年07月27日 核心技术与产业动态速递

July 27, 2026 · 6 min

从咖啡馆到千亿美金野心:Airwallex吴恺谈AI时代的全球金融基础设施

July 27, 2026 · 1 min

【每日AI前沿追踪��2026年7月26日 核心技术与产业动态速递

July 26, 2026 · 7 min

【每日AI前沿追踪】2026年7月25日 核心技术与产业动态速递

July 25, 2026 · 6 min

一部昇腾史与全球芯片30年史诗——华为半导体首席科学家廖恒5小时深度访谈

July 25, 2026 · 1 min

AI为什么没能颠覆足球?从SciSports的陨落���世界杯的技术暗战

July 24, 2026 · 1 min

AI泡沫2027年爆破?两位投资人的硬核推演:中国开源模型、万亿债务与企业级决战

July 24, 2026 · 2 min

Momenta IPO 后再访曹旭东:没有尽头的 AI,从智驾到家庭机器人的十年推演

July 24, 2026 · 2 min

谁在教AI说人话?藏在大模型背后的新闻��与内容工程师

July 24, 2026 · 1 min

【每日AI前沿追踪】2026年07月22日 核心技术与产业动态速递

July 23, 2026 · 5 min

【每日AI前沿追踪】2026年7月23日 核心技术与产业动态速递

July 23, 2026 · 3 min

Recursive Harness Self-Improvement 精读

July 23, 2026 · 6 min

【每日AI前沿追踪】2026年07月21日 核心技术与产业动态速递

July 22, 2026 · 8 min

2026 Q2 AI季报:RSI从科幻走向创业赛道,Coding战场大洗牌,强者愈强的未来

July 22, 2026 · 3 min

AI时代什么值得学?知识、代码都贬值了,经验和技能才是硬通货

July 22, 2026 · 1 min

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune 精读

July 22, 2026 · 7 min

具身原生的豪赌:蚂蚁灵波沈宇军,为什么坚持从传感器和视频里重训整个机器人模型?

July 22, 2026 · 2 min

【每日AI前沿追踪】2026年07月20日 核心技术与产业动态速递

July 21, 2026 · 7 min

从龙虾热到基金会治理:OpenClaw首席架构师Vincent Koc谈个人Agent的反思、工程化与协作未来

July 21, 2026 · 2 min

【每日AI前沿追踪】2026年07月19日 核心技术与产业动态速递

July 20, 2026 · 7 min

【每日AI前沿追踪】2026年07月18日 核心技术与产业动态速递

July 19, 2026 · 10 min

【每日AI前沿追踪】2026年07月17日 核心技术与产业动态速递

July 18, 2026 · 9 min

世界模型这半年:XLR Labs 谈原生路线、4D 数据护城河与物理 AGI 的下半场

July 18, 2026 · 2 min

【每日AI前沿追踪】2026年07月16日 核心技术与产业动态速递

July 17, 2026 · 8 min

2026-06 arXiv 智能体/工具智能体领域综述:916 篇分主题精读

July 17, 2026 · 7 min

2026-06 arXiv 智能体记忆系统(Agent Memory)领域综述:113 篇全文精读

July 17, 2026 · 5 min

2026-06 LLM 代码生成领域综述:357 篇全文通读

July 17, 2026 · 3 min

智能体技能演化(Skill Evolution 与 Self-Evolving Agents)综述:53 篇核心论文精读

July 17, 2026 · 4 min

【每日AI前沿追踪】2026年07月15日 核心技术与产业动态速递

July 16, 2026 · 8 min

【每日AI前沿追踪】2026年07月14日 核心技术与产业动态速递

July 15, 2026 · 7 min

【每日AI前沿追踪】2026年07月13日 核心技术与产业动态速递

July 14, 2026 · 7 min

Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models 精读

July 13, 2026 · 4 min

如果神存在,我怎能容忍自己不是神——对话英灵殿Odin:AI4S的狂人哲学与全模态分子世界模型

July 13, 2026 · 1 min

【每日AI前沿追踪】2026年07月12日 核心技术与产业动态速递

July 12, 2026 · 6 min

【每日AI前沿追踪】2026年07月11日 核心技术与产业动态速递

July 11, 2026 · 6 min

【每日AI前沿追踪】2026年07月10日 核心技术与产业动态速递

July 10, 2026 · 7 min

【每日AI前沿追踪】2026年07月09日 核心技术与产业动态速递

July 9, 2026 · 7 min

ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes 精读

July 9, 2026 · 5 min

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog 精读

July 9, 2026 · 3 min

【每日AI前沿追踪】2026年07月08日 核心技术与产业动态速递

July 8, 2026 · 6 min

【每日AI前沿追踪】2026年07月07日 核心技术与产业动态速递

July 7, 2026 · 6 min

Harness Engineering for Self-Improvement 精读

July 7, 2026 · 8 min

Token Maxing退潮,Agent开始干活——亚马逊云科技中国峰会探展复盘

July 7, 2026 · 1 min

Verbalizable Representations Form a Global Workspace in Language Models 精读

July 7, 2026 · 4 min

【每日AI前沿追踪】2026年07月06日 核心技术与产业动态速递

July 6, 2026 · 2 min

AI自进化的临界点:最快半年跑通一环闭环——与AppleX首席科学家谈RSI、验证、品味与发现模型

July 6, 2026 · 2 min

【每日AI前沿追踪】2026年07月05日 核心技术与产业动态速递

July 5, 2026 · 3 min

【每日AI前沿追踪】2026年07月04日 核心技术与产业动态速递

July 4, 2026 · 3 min

【每日AI前沿追踪】2026年07月03日 核心技术与产业动态速递

July 3, 2026 · 4 min

2026年 Coding 方向 Benchmark 全面调研:33个可用仓库 + 12个未来方向预测

July 3, 2026 · 9 min

ACL 2026 主会长文研究方向调研:2222 篇论文全量分类与新增领域分析

July 3, 2026 · 4 min

走进中国AI实验室:Nathan Lambert的中美AI观察——开源领导权转移、算力困局与人才文化差异

July 3, 2026 · 1 min

【每日AI前沿追踪】2026年07月02日 核心技术与产业动态速递

July 2, 2026 · 3 min

Agent元年前500天:Headless软件、CLI开放与Skill经济的全面爆发

July 2, 2026 · 1 min

Agent新范式圆桌:从Prompt到Loop的演进逻辑、潜空间通信与验证之困

July 2, 2026 · 2 min

Agent进化的四个层级:从知识更新到工作流自我设计——西湖大学张驰解读智能体动态架构

July 2, 2026 · 2 min

【每日AI前沿追踪】2026年07月01日 核心技术与产业动态速递

July 1, 2026 · 3 min

Agentic Abstention: Do Agents Know When to Stop Instead of Act? 精读

July 1, 2026 · 5 min

拆解Claude Code源码泄露:Agent Harness三层架构、记忆机制与零人公司的未来

July 1, 2026 · 2 min

June  78

【每日AI前沿追踪】2026年06月30日 核心技术与产业动态速递

June 30, 2026 · 3 min

FastContext: Training Efficient Repository Explorer for Coding Agents 精读

June 30, 2026 · 7 min

【每日AI前沿追踪】2026年06月29日 核心技术与产业动态速递

June 29, 2026 · 3 min

Qwen-AgentWorld: Language World Models for General Agents 精读

June 29, 2026 · 5 min

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills 精读

June 29, 2026 · 4 min

【每日AI前沿追踪】2026年06月28日 核心技术与产业动态速递

June 28, 2026 · 2 min

AI时代填高考志愿:专业不再是身份,地域和交叉能力才是入口

June 28, 2026 · 1 min

探月学校的第三种可能:AI时代,教育从升学机器转向成长社区

June 28, 2026 · 1 min

王熙乔的十年探月:AI时代,教育要从知识训练转向文明训练

June 28, 2026 · 2 min

【每日AI前沿追踪】2026年06月27日 核心技术与产业动态速递

June 27, 2026 · 3 min

【每日AI前沿追踪】2026年6月26日 核心技术与产业动态速递

June 26, 2026 · 3 min

当AI开始进化AI:递归自我改进的技术路线、关键瓶颈与终局图景

June 26, 2026 · 2 min

【每日AI前沿追踪】2026年6月25日 核心技术与产业动态速递

June 25, 2026 · 3 min

【每日AI前沿追踪】2026年6月24日 核心技术与产业动态速递

June 24, 2026 · 3 min

【每日AI前沿追踪】2026年6月23日 核心技术与产业动态速递

June 23, 2026 · 3 min

【每日AI前沿追踪】2026年6月22日 核心技术与产业动态速递

June 22, 2026 · 3 min

从Cerebras上市看AI算力新格局:十年前的非共识投资,与推理时代的到来

June 22, 2026 · 2 min

对话 MiniMax 闫俊杰:M3、10X 计划、10T 模型、和智能的终局

June 22, 2026 · 1 min

马督工弹幕直播:AI冲击实习生态、彩礼的经济根源、分配迷信与中美G2

June 22, 2026 · 1 min

高考之后的清醒剂:马督工谈AI分化、专业选择、分配迷信与写作稀缺

June 22, 2026 · 1 min

高考前夜的专业选择与人生课:马督工与语文老师的考前长谈

June 22, 2026 · 1 min

【每日AI前沿追踪】2026年6月21日 核心技术与产业动态速递

June 21, 2026 · 3 min

Probe-and-Refine Tuning of Repository Guidance for Coding Agents 精读

June 21, 2026 · 3 min

对姚顺宇的4小时访谈:在Anthropic和Gemini训模型、技术预测、英雄主义已过去

June 21, 2026 · 1 min

【每日AI前沿追踪】2026年06月20日 核心技术与产业动态速递

June 20, 2026 · 7 min

对洪乐潼的4小时访谈:AI for Math、把数学变成Lean、数学天书中的证明、直觉、被创造的与被发现的

June 20, 2026 · 3 min

对谢赛宁的7小时马拉松访谈:世界模型、逃出硅谷、反OpenAI、AMI Labs、两次拒绝Ilya、杨立昆、李飞飞和42

June 20, 2026 · 2 min

翁家翌:OpenAI,GPT,强化学习,Infra,后训练,天授,tuixue,开源

June 20, 2026 · 1 min

跟三个大厂码农聊了一晚上,程序员眼中AI叙事的真相原来是这样

June 20, 2026 · 1 min

【每日AI前沿追踪】2026年06月18日 核心技术与产业动态速递

June 18, 2026 · 6 min

OpenAI联手PE砸下40亿美元,聊聊硅谷最火新职位FDE

June 18, 2026 · 2 min

中国的危与机 中美俄欧中东,AI美债美元,混乱世界中 正在发生

June 18, 2026 · 1 min

库克的离场,苹果新AI权力重构与价值观天平|WWDC26【硅谷101】

June 18, 2026 · 2 min

【每日AI前沿追踪】2026年06月17日 核心技术与产业动态速递

June 17, 2026 · 5 min

FastContext: Training Efficient Repository Explorer for Coding Agents 精读

June 17, 2026 · 5 min

From AGI to ASI 精读

June 17, 2026 · 5 min

【每日AI前沿追踪】2026年06月16日 核心技术与产业动态速递

June 16, 2026 · 4 min

【每日AI前沿追踪】2026年06月15日 核心技术与产业动态速递

June 15, 2026 · 3 min

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference 精读

June 15, 2026 · 5 min

【每日AI前沿追踪】2026年06月14日 核心技术与产业动态速递

June 14, 2026 · 3 min

【每日AI前沿追踪】2026年06月13日 核心技术与产业动态速递

June 13, 2026 · 2 min

【每日AI前沿追踪】2026年06月12日 核心技术与产业动态速递

June 12, 2026 · 2 min

Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference 精读

June 12, 2026 · 3 min

Role-Agent: 通过双角色自举实现 LLM 智能体-环境协同进化 精读

June 12, 2026 · 5 min

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research 精读

June 12, 2026 · 3 min

Self-Harness: Harnesses That Improve Themselves 精读

June 12, 2026 · 5 min

SpaceX要让太空算力从科幻走向现实,但它划算吗

June 12, 2026 · 1 min

【每日AI前沿追踪】2026年06月11日 核心技术与产业动态速递

June 11, 2026 · 2 min

Agentic Coding驱动工业制造通往自主通用智能

June 11, 2026 · 1 min

SpaceX崛起史:一切,为了去火星

June 11, 2026 · 1 min

【每日AI前沿追踪】2026年06月10日 核心技术与产业动态速递

June 10, 2026 · 3 min

QUBRIC: 协同设计查询与评分标准,突破可验证奖励的强化学习瓶颈 精读

June 10, 2026 · 4 min

Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective 精读

June 10, 2026 · 3 min

Skill-RM: 通过 Agent Skill 统一异构奖励评估标准 精读

June 10, 2026 · 4 min

Agentic ASR: 面向类人交互式语音识别的智能体修正与语义评估 精读

June 9, 2026 · 5 min

HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems 精读

June 9, 2026 · 6 min

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents 精读

June 9, 2026 · 4 min

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents 精读

June 9, 2026 · 4 min

【每日AI前沿追踪】2026年06月08日 核心技术与产业动态速递

June 8, 2026 · 3 min

彼得·希夫:金融市场的死亡螺旋已开始,人工智能烧的钱要靠裁员和卖币凑

June 8, 2026 · 1 min

【每日AI前沿追踪】2026年06月07日 核心技术与产业动态速递

June 7, 2026 · 3 min

【每日AI前沿追踪】2026年06月06日 核心技术与产业动态速递

June 6, 2026 · 3 min

From Context to Skills: Can Language Models Learn from Context Skillfully? 精读

June 6, 2026 · 4 min

【每日AI前沿追踪】2026年06月05日 核心技术与产业动态速递

June 5, 2026 · 3 min

不做AI螺丝钉:被Meta裁员半年后,田渊栋带着46亿美元AI实验室回来了

June 5, 2026 · 1 min

【每日AI前沿追踪】2026年06月04日 核心技术与产业动态速递

June 4, 2026 · 2 min

【每日AI前沿追踪】2026年06月03日 核心技术与产业动态速递

June 3, 2026 · 3 min

SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories 精读

June 3, 2026 · 4 min

《创造性毁灭的力量:经济动荡与国家财富》精读

June 2, 2026 · 2 min

Yunjue Agent: 零起点原位自进化智能体系统精读

June 2, 2026 · 4 min

δ-mem: Efficient Online Memory for Large Language Models 精读

June 2, 2026 · 3 min

聊聊DeepMind创始人哈萨比斯

June 2, 2026 · 1 min

【每日AI前沿追踪】2026年06月01日 核心技术与产业动态速递

June 1, 2026 · 2 min

【每日AI前沿追踪】2026年06月01日 核心技术与产业动态速递

June 1, 2026 · 3 min

99%的作业都是AI写的,当代名校生眼里,大学还剩下什么

June 1, 2026 · 1 min

Agent新基建,如何让一人企业做全球生意

June 1, 2026 · 2 min

Skill0.5: Joint Skill Internalization and Utilization for OOD Generalization in Agentic RL 精读

June 1, 2026 · 5 min

与Andrew Dai聊Gemini的翻身之战,出走与视觉理解模型

June 1, 2026 · 1 min

May  26

【每日AI前沿追踪】2026年05月31日 核心技术与产业动态速递

May 31, 2026 · 2 min

Never Stop Learning: Continual Learning 与 Self-Iteration 综述精读

May 31, 2026 · 3 min

当AI开始重新定义"能力",教育还会一样吗

May 31, 2026 · 2 min

聊聊Harness时代AI-First的组织架构:从信任人到信任AI

May 31, 2026 · 1 min

【每日AI前沿追踪】2026年05月30日 核心技术与产业动态速递

May 30, 2026 · 2 min

【每日AI前沿追踪】2026年05月29日 核心技术与产业动态速递

May 29, 2026 · 2 min

【每日AI前沿追踪】2026年05月28日 核心技术与产业动态速递

May 28, 2026 · 2 min

DeGRe: Dense-supervised Generative Reranking for Recommendation 精读

May 28, 2026 · 2 min

清月5月观点总结

May 28, 2026 · 2 min

硅谷101 5月观点总结

May 28, 2026 · 1 min

被炒的炒饭 4月观点总结

May 28, 2026 · 1 min

【每日AI前沿追踪】2026年05月27日 核心技术与产业动态速递

May 27, 2026 · 2 min

An Empirical Study of Proactive Coding Assistants in Real-World Software Development 精读

May 27, 2026 · 4 min

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models 精读

May 27, 2026 · 5 min

RGAO 精读:多智能体代码生成的检索条件化拓扑选择与可证明预算守恒

May 27, 2026 · 7 min

Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning 精读

May 27, 2026 · 4 min

Structured Uncertainty guided Clarification for LLM Agents 精读

May 27, 2026 · 5 min

【每日AI前沿追踪】2026年05月26日 核心技术与产业动态速递

May 26, 2026 · 2 min

GPS: Graph-Guided Proactive Information Seeking in Large Language Models 精读

May 26, 2026 · 4 min

Natural-Language Agent Harnesses 精读

May 26, 2026 · 5 min

SkillOpt: Executive Strategy for Self-Evolving Agent Skills 精读

May 26, 2026 · 5 min

【每日AI前沿追踪】2026年05月25日 核心技术与产业动态速递

May 25, 2026 · 2 min

【每日AI前沿追踪】2026年05月24日 核心技术与产业动态速递

May 24, 2026 · 2 min

【每日AI前沿追踪】2026年05月23日 核心技术与产业动态速递

May 23, 2026 · 4 min

【每日AI前沿追踪】2026年05月22日 核心技术与产业动态速递

May 22, 2026 · 5 min

自动化信息收集系统介绍

May 22, 2026 · 1 min