<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>SAE on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/sae/</link><description>Recent content in SAE on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 11 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/sae/index.xml" rel="self" type="application/rss+xml"/><item><title>SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-saescientist-bench-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-saescientist-bench-paper-reading/</guid><description>中科院自动化所的 SAEScientist-Bench 把&amp;rsquo;可解释性研究本身&amp;rsquo;benchmark 化：基于 27182 个专家标注 SAE 特征构建任务族，agent 需完成特征发现（AUROC 区分正例与对比控制）→ 因果转向验证（steering 分数度量目标表达净增）→ 下游生成质量保持的完整实验科学闭环。前沿 agent（Kimi 等）在因果转向与目标相关性上接近专家水平（Expert 特征 AUROC 0.917-1.000），但生成退化率从 32.5% 升至 52.5%——&amp;lsquo;转向强度 vs 生成保真&amp;rsquo;的权衡是当前 agent 科学家的系统性短板。</description></item></channel></rss>