<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>腾讯 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E8%85%BE%E8%AE%AF/</link><description>Recent content in 腾讯 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 05 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E8%85%BE%E8%AE%AF/index.xml" rel="self" type="application/rss+xml"/><item><title>Environment Evolution 精读：让训练环境的难度离线进化</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-05-environment-evolution-paper-reading/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-05-environment-evolution-paper-reading/</guid><description>腾讯混元×港科大（广州）的 Environment Evolution 把环境难度演化从 on-policy 共进化中解耦：从多轮学习目标推导三个演化方向，用多 Agent harness 离线逐代提升环境难度，再由谱系调度器持续供给学习信号。Qwen3.6-27B/35B-A3B 经简单长程 RL 在 Terminal-Bench 2.1 分别提升 14.4/18.0 个百分点。本文基于全文阅读拆解演化方向推导与调度器设计。</description></item></channel></rss>