<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>LLM on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/llm/</link><description>Recent content in LLM on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>Looped Flows 精读：把“想得更久”做进架构——循环流的局部训练之路</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-looped-flows-reasoning-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-looped-flows-reasoning-paper-reading/</guid><description>AITHYRA 访问研究者的 looped flows：状态化去噪器每步预测解并更新循环状态、ODE/SDE 步更新流状态，用局部训练目标绕开跨步反传的老大难。六个推理基准整体超先前 looped SOTA，ARC-AGI-1 58.8%、ARC-AGI-2 12.2%——循环模型在抽象推理上首次具备与主流推理范式对话的竞争力。</description></item><item><title>NCP-ArchPreview 精读：当语言模型开始预测“概念”而不是 token</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-ncp-archpreview-latent-space-lm-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-ncp-archpreview-latent-space-lm-paper-reading/</guid><description>上海交大 Intern-NCP 团队发布 8.9B/5.73T tokens 的潜空间语言模型 NCP-ArchPreview：在 next-token prediction 之上引入 Next Concept Prediction 目标，仅用 51.3% 训练 tokens 追平 OLMo-3-7B 最终 loss，GSM8K +5.99 分，17M 参数 VQ 模块即可完成领域适配——迄今最大潜空间 LM 实证。</description></item><item><title>NSD 精读：教推理模型“别这么错”，比教它“该怎么对”更有效</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-12-nsd-negative-self-distillation-paper-reading/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-12-nsd-negative-self-distillation-paper-reading/</guid><description>UVA+Stanford 提出负面自蒸馏（NSD）：针对 on-policy 自蒸馏（OPSD）模仿“带标准答案的伪自信轨迹”导致难推理任务退化的失败模式，构造“负条件”（注入已知缺陷的解）让学生显式规避。token 级自适应门控+gated unlikelihood 在七基准上 1.7B/4B/8B 平均 +2.3%/+7.5%/+6.0%，且保留自纠错行为——模型越大增益越高。</description></item><item><title>Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-phi-bench-llm-infrastructure-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-phi-bench-llm-infrastructure-paper-reading/</guid><description>USTC×StepFun×北大×HKUST×Yale×UPenn 六机构联合的 Φ-Bench 提出一个自指性问题：LLM 推理所依赖的基础设施（kernel、服务栈、集群）能否由 LLM 自己来工程化？85 个任务、三种渐进格式——Kernel 函数补全（KFC）→长程实现（LHI）→端到端优化（E2EO），双轴评分（性能+实现）+内置作弊检测。最强 Claude Opus 5 总分仅 36.53%（KFC 37.16%/LHI 21.60%/E2EO 62.94%），硬件与边缘类最好模型也仅 5.4%——&amp;lsquo;造物者维护造物&amp;rsquo;的能力缺口被量化暴露。</description></item></channel></rss>