<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>递归自进化 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E9%80%92%E5%BD%92%E8%87%AA%E8%BF%9B%E5%8C%96/</link><description>Recent content in 递归自进化 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E9%80%92%E5%BD%92%E8%87%AA%E8%BF%9B%E5%8C%96/index.xml" rel="self" type="application/rss+xml"/><item><title>SEABench × Audit the Scaffold × REUSE：递归自我改进的测量、理论与统计三重保障 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-rsi-reliability-trio-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-rsi-reliability-trio-paper-reading/</guid><description>同一周出现的三篇论文，恰好构成递归自我改进（RSI）治理的三根支柱：SEABench 用配对反事实与归因裁判测量「自进化会不会内生地变坏」（安全失败率 43.9% vs 0%）；Audit the Scaffold 用 Lean 4 验证的平稳性二分法回答「自我改进何时必然耗尽、何时可能失控」（改脚手架可扩类不碰权重，冻结权重≠安全）；REUSE 用决策-only 反馈与全历史 union bound 保证「每一次晋升都是真实总体改进」（75 次假晋升→0 次，提升不损）。本精读从「是什么」讲起，拆解三篇的方法机制、评估证据与优势根源，并交叉验证其在 2026 年 RSI 治理浪潮中的位置。</description></item><item><title>Self-Evolving Coding Agents × RE-0：从数字程序到物理世界的自进化智能体 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-30-physical-agentic-duet-paper-reading/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-30-physical-agentic-duet-paper-reading/</guid><description>二重奏精读两篇互补论文：hexafuture.ai 的 Self-Evolving Coding Agents 提出物理编码范式，用 Code as World + Code as Policy 双可执行表征与类型化验证器，把编码代理范式迁移到物理世界，在 RoboCasa365 上把成功率从 56.6% 提升到 61.1%；吉林大学与大连理工的 RE-0 用 locate-verify-weight 递归和 LCB 准入，仅凭 3-67 条验证数据把具身 Code-as-Policy 基线从 4-68% 提升到 62-100%。一篇搭系统、一篇做训练，勾勒物理世界自进化智能体的完整图景。</description></item><item><title>递归自改进的能力与安全双螺旋精读：DCE 自蒸馏与演化安全框架</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-29-rsi-capability-safety-paper-reading/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-29-rsi-capability-safety-paper-reading/</guid><description>本文合并精读 2026 年 9 月同日发布于 arXiv 的两篇递归自改进（RSI）论文。论文一（Meta AI + UC Riverside）提出 DCE+SRCL：让特权教师在 on-policy 自蒸馏中与学生逐轮共同进化，在 Qwen3-8B 四项数学竞赛基准上取得 65.97% Average@12，较冻结教师的 OPSD 提升 35.62 个百分点，并用固定轨迹探针揭示教师监督退化的机制证据。论文二（中科院计算所）提出演化安全框架：以携带时间历史的安全相关变更为分析对象，建立六种风险表现 × 五类变更载体 × 四层评估单元的分类学与治理原则。本精读各按九部分展开，并以「教师共同进化 ↔ 风险共同进化」的双螺旋视角合并收束：DCE 证明上轮学到的修正行为会经由教师进入下轮监督——这正是演化安全所警告的经验污染与风险继承在能力侧的镜像。</description></item><item><title>环境演化、跨图溯因 SWE 与 AI 主导模型开发：RSI 三重奏精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-26-trio-rsi-benchmarks-paper-reading/</link><pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-26-trio-rsi-benchmarks-paper-reading/</guid><description>本篇三重奏精读覆盖递归自改进（RSI）方向的三篇最新论文：Env-Rethink 把「文件环境准备」变成可学习目标，用 27B 验证模型让 9 个下游模型在噪声环境平均通过率从 59.4% 提升到 72.7%，并用事件驱动演化生成可验证的更难环境；SWE-PolyVision 构建首个 100% 多图可执行 SWE 基准（92 任务、三种视觉访问模式受控干预），揭示「可得性不等于整合」的 access-to-integration gap；iCoder-27B 则让 Codex agent 在人类只提供可执行 Research Skills 的前提下自主跑完数据/SFT/OPSD/RLVR 全流程，训出 RTLLM 68.0 超越 GPT-5.5 与 Claude-Opus-4.8 的 27B 工业编码模型。三篇合起来勾勒出 RSI 的环境侧、评测侧与模型侧全景。</description></item><item><title>Self Improvement via Fast Tree-search 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-21-sift-paper-reading/</link><pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-21-sift-paper-reading/</guid><description>MIT 与 Sakana AI 的 SIFT 把递归自改进（RSI）编码智能体的最大瓶颈从&amp;rsquo;生成候选&amp;rsquo;移到了&amp;rsquo;验证候选太贵&amp;rsquo;：用 pairwise LLM-as-a-judge（每次 $0.044）+ 正则化 Bradley-Terry 聚合替代 $6.0 的基准子集评估作为中间信号，在完全解耦的树搜索流水线中让扩展与评估并行。Polyglot-225 上以 DGM 约 1/10 的 CPU 小时拿到 31.1%（Qwen3-30B）/35.1%（o3-mini）全面超越 DGM/HGM/SICA，TerminalBench 2.1 从 29.2% 提到 36.7%。本精读覆盖&amp;rsquo;便宜排名+昂贵验证&amp;rsquo;分离范式的机制因果、judge 输入格式的消融证据、与 DGM 谱系的定位对比，以及&amp;rsquo;把验证成本当一等公民&amp;rsquo;的通用性灵感。</description></item><item><title>SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-19-sol-pi-harness-study-paper-reading/</link><pubDate>Sat, 19 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-19-sol-pi-harness-study-paper-reading/</guid><description>NVIDIA 联合 NTU/MIT 提出 SoL-Pi：把编码智能体 harness 的效率改进本身建模为跨环境搜索问题，用约 150 个方向、500 个环境、3000+ 次实验、60000+ 次交互的自动研究漏斗，筛选出 Action Fusion、Online Context Compact、ObservationPack、Evidence-Preserving Reducer 四个可复用机制，EdgeBench 上 token 流量降 44.7–49.0%、成本省 1/3 且性能持平，迁移到未见过的 Opus 5 后端仍保留 94.3% 性能。这是 RSI（递归自改进）从&amp;rsquo;改模型&amp;rsquo;转向&amp;rsquo;改脚手架&amp;rsquo;的代表性工作。</description></item><item><title>Agora: Git as Shared Memory for Collective AutoResearch 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-18-agora-paper-reading/</link><pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-18-agora-paper-reading/</guid><description>NVIDIA 提出 Agora：把多个自主科研 Agent 的协作记录为 Git 上的 append-only DAG——每个结果/假设/验证都是可 checkout 重跑的不可变 commit。首次持续运行 12 天：13 个无任务分配、无中央规划器的 LLM worker 在权重迁移难题上发布 1,703 项贡献，把评估器从 3.39 推到 1.899 bits/byte，弥合与训练版 GPT-2 差距的 62%；获胜配方 145-commit 谱系跨 15 个账户、165 次独立复现零失败。集体智能不靠规划器，靠记忆基础设施——本精读拆解其设计。</description></item><item><title>NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-10-neohorse-1-routing-harness-rsi-paper-reading/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-10-neohorse-1-routing-harness-rsi-paper-reading/</guid><description>NeoHorse-1 把部署中的模型路由 harness 变成递归自改进（RSI）的数据飞轮：路由层天然记录每次交互的&amp;rsquo;能力需求预测-实际执行-结果&amp;rsquo;三元组，这些记录被转化为保留交错推理与工具调用的 user-turn 训练样本，路由分数进一步组织成三阶段课程 SFT 与路由引导的在线策略蒸馏。4B/9B 模型十项基准宏平均分别从 58.94/65.60 提升至 64.87/69.04，路由 harness 数据比公开 Agent 数据平均高 6.26 分。本文精读拆解其数据管线、课程设计、OPD 机制与&amp;rsquo;评估-选择-更新&amp;rsquo;闭环为何能成立。</description></item><item><title>RISE: 自外推策略蒸馏把 on-policy 蒸馏变成递归自我改进 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-08-rise-self-extrapolating-distillation-paper-reading/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-08-rise-self-extrapolating-distillation-paper-reading/</guid><description>RISE 从模型自身 RLVR 训练轨迹外推合成教师：以当前检查点与滑动锚点的位移放大（β&amp;gt;1）构造未来教师，在 logit 或权重空间实现，教师随学生每轮刷新。OLMo3-7B AIME'24 从 30.2 提至 46.9（+16.7），ALFWorld +9.4、WebShop +10.9 vs GRPO，OOD 不降反升。本文精读&amp;rsquo;教师从哪来&amp;rsquo;这一 OPD 根本问题的第一性解法及其递归改进机制。</description></item><item><title>Meta^n: Recursive Self-Improvement through Emergent Depth 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-meta-n-emergent-depth-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-meta-n-emergent-depth-paper-reading/</guid><description>Meta^n（明尼苏达大学 × 首尔国立大学）针对自我改进系统『实现元深度只有约 2』的天花板，提出固定元操作 Ω 对自身输入递归：Ω 读下层栈的全任务执行轨迹+产生它们的代码栈，写出下一层（策略性预处理器+可调用辅助函数库），深度由收敛决定而非预先设定，进化档案在层链空间搜索。8 个基准族 × 2 骨干上至少一个估计器全面领先先前自改进 agent，ARC-AGI-2 held-out 上唯一非零（0.331 vs OpenEvolve 0.003）；消融显示递归本身贡献 +0.131，其中层间条件化占约 72%；深度角色自发涌现——回滚角色在深度 2 恰为零、深度 3 出现 55%。</description></item><item><title>Recuris: Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-recuris-memory-evolution-paper-reading/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-recuris-memory-evolution-paper-reading/</guid><description>Recuris（NUS × Stanford × Oxford × Princeton）把递归自我改进从『改模型/改智能体』收缩到『只演化外置记忆控制层』：工作记忆维护经检查器验证的任务状态并按需调用技能，跨任务的固定 Meta-Agent 读结构化轨迹、把失败归因到 E/W/ρ/C 四组件之一并只修补被归因组件，经修复源任务且不回退开发集的验证门才准入。在 4 个长程基准 × 10 个模型上 35/37 完成的模型-基准对成功率提升，GPT-5.6 Sol +17.8、Claude Opus 5 +15.6、最长任务 +32.2 分，六类长程失败模式下降 20–86%；机制上证明长程失败是执行问题而非检索问题，技能价值是『调用条件性』的，结构化轨迹使故障定位从 13.0% 提升到 64.8%。</description></item><item><title>AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-24-ai4ai-bench-paper-reading/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-24-ai4ai-bench-paper-reading/</guid><description>Naver Labs、Einsia.AI 与清华大学推出 AI4AI-Bench：冻结 10 个真实研究仓库，覆盖 SFT、agentic RL、蒸馏、奖励建模、DPO、扩散 RL、遗忘、图扩散、权重平均与剪枝十族算法，测评 agent 能否改写仓库训练算法本身。agent 在单块 B300 用 4 小时改代码；提交后源码从零训练 12 小时，由冻结评估器打分，σ 坐标统一指标（0.1 为原算法，1.0 为最优）。负结果：290 格平均 0.166，最强系统 Claude Opus 5 仅 0.250；触及学习层者平均 0.226，远高于只动运行层的 0.126。</description></item><item><title>AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-22-ai4ai-bench-paper-reading/</link><pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-22-ai4ai-bench-paper-reading/</guid><description>深度精读 Einsia.AI 与清华大学 2026 年 8 月提出的 AI4AI-Bench：首个隔离测量 LLM Agent 训练算法设计能力的基准。10 个冻结研究仓库、单块 B300 四小时改写、十二小时从零重跑、0/0.1/1.0 三锚点统一量表，29 个配置平均仅 0.166、最佳 0.250——最强系统连&amp;rsquo;已有算法到最优&amp;rsquo;距离的五分之一都没走完；而推理预算买到的主要是&amp;rsquo;敢去改&amp;rsquo;的意愿，参与率从 8% 提升到 64%。</description></item><item><title>Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-12-mendel-godel-machine-paper-reading/</link><pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-12-mendel-godel-machine-paper-reading/</guid><description>LMU Munich 团队提出的 Mendel Gödel Machine (MGM)，将孟德尔遗传学中「受控比较分离遗传效应」的原理引入自改进编码智能体。在 HGM 的树搜索框架之上，MGM 新增两种自我修改算子——反应规范突变（跨任务比较同一基因型）和跨谱系杂交（跨谱系比较同一任务），在不增加任何额外任务评估成本的前提下，把 Qwen3.6-35B-A3B 在 Polyglot 上的成绩从 50.8% 拉到 93.3%，以约 117× 更少参数超越闭源 GPT-5；进化的脚手架迁移到 DeepSeek-V4-Pro 后在完整 Polyglot-225 上达 96.9%。本精读覆盖其生物学启发、三种算子机制、加性适应度景观下的收敛性证明、实验证据与通用性灵感。</description></item><item><title>Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-11-ouroboros-self-developing-agent-paper-reading/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-11-ouroboros-self-developing-agent-paper-reading/</guid><description>本文精读 Anton Razzhigaev、Roman Yampolskiy 等人 2026 年发表的 Ouroboros——一个能够自开发的前沿编程 Agent。它把 Agent 的工具、提示词、上下文组装乃至核心实现本身都视为可被审查、可被修改的活体代码，并通过多模型对抗式 diff 审查作为变更门控，实现经审查的核心进化（Reviewed Core Evolution）。文章在 Terminal-Bench 2.1、OSWorld-Verified、CL-Bench 等基准上刷新 SOTA，并在代号为 Hope 的 161 天活体实验中持续运行（累计 1085 次自我修改提交、94.2% 由 Agent 撰写）。本精读将从背景、定位、问题定义、方法、评估、优势根源、必要知识反推、通用性灵感八个维度系统拆解这篇论文。</description></item><item><title>RSI比Coding Agent大得多：对话田渊栋，递归自进化为什么是阶段式突破而非渐进攀升</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-07-tian-yuandong-rsi/</link><pubDate>Fri, 07 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-07-tian-yuandong-rsi/</guid><description>前Meta FAIR研究员田渊栋创立Recursive Superintelligence（A轮6.5亿美元，估值46.5亿美元）后首次系统阐述RSI：它比coding agent大得多、难得多；完全自动化不会很快发生，递归会先发生；智能发展是S型曲线而非scaling law的平滑攀升，这恰恰给了初创公司窗口。他们的第一阶段成果在算子优化、NanoChat训练和NanoGPT SpeedRun三个方向取得SOTA，用一套通用系统击败了专业GPU团队。田渊栋认为AI终将从炼金术变成化学，可解释性是少数派但正确的路。</description></item><item><title>Recursive Harness Self-Improvement 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-23-rhi-recursive-harness-self-improvement-paper-reading/</link><pubDate>Thu, 23 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-23-rhi-recursive-harness-self-improvement-paper-reading/</guid><description>Sakana AI 与 UC Berkeley 提出 RHI（递归式框架自改进）：把多智能体框架当作提示词级对象，仅用当前与上一版本的自我比较来迭代优化，少数几轮就能让低推理强度的 Agent 超越同族最高推理强度设置，同时把推理成本降低最高 60%。本文从 Harness 是什么、模型-框架协同进化讲起，拆解 RHI 的轨迹局部目标、算法流程、信息论隐式目标，并提炼可推广的通用性灵感。</description></item><item><title>2026 Q2 AI季报：RSI从科幻走向创业赛道，Coding战场大洗牌，强者愈强的未来</title><link>https://inkeast.github.io/MessageDaily/posts/2026-07-22-ai-q2-review-rsi-coding/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-07-22-ai-q2-review-rsi-coding/</guid><description>2026年Q2 AI季报深度解读：Anthropic与OpenAI的模型竞争进入新阶段，GPT 5.6与Claude Maestro/Phable正面交锋；RSI（递归自进化）从科幻概念变成明确的创业方向，Recursive、Miranda等公司涌现；Cursor以600亿美元天价被收购；中国开源模型&amp;quot;四杀&amp;quot;引发全球关注；Anthropic的Cloud Tag与OpenAI的Record and Replay重新定义AI交互。本文基于播客全文转写整理，涵盖竞争格局、RSI、机器人、智能扩散、交互创新和公司动态。</description></item></channel></rss>