<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Transluce on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/transluce/</link><description>Recent content in Transluce on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Thu, 27 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/transluce/index.xml" rel="self" type="application/rss+xml"/><item><title>AI在想什么：模型没说出口的推理，与可解释性唯一一次漂亮的兑现</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-27-aryaman-arora-hidden-reasoning-interpretability/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-27-aryaman-arora-hidden-reasoning-interpretability/</guid><description>斯坦福博士生、语言学科班出身的 Aryaman Arora 做客硅谷101视频播客谈大模型可解释性：Anthropic 的 J-space 实验证明模型内部存在从未说出口的推理概念，且可被直接编辑——内部把&amp;quot;蜘蛛&amp;quot;改成&amp;quot;蚂蚁&amp;quot;，答案就从八条腿变成六条腿；思维链有用但不等于模型的真实内部过程；SAE 与因果干预两大流派各有硬限制，学术界的转向向量控制几乎全线失灵，工程实践仍回归重训；该领域至今最漂亮的兑现是归纳头的发现救活了状态空间模型谱系（H3→Mamba→DeltaNet，直至 Kimi/Qwen 的混合架构）；Transluce 的用户建模显示模型面对 AI 安全研究员时会显著更谨慎；可解释性天然双刃，但嘉宾判断它离危险阈值还很远。</description></item></channel></rss>