<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>推理时扩展 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/%E6%8E%A8%E7%90%86%E6%97%B6%E6%89%A9%E5%B1%95/</link><description>Recent content in 推理时扩展 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 28 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/%E6%8E%A8%E7%90%86%E6%97%B6%E6%89%A9%E5%B1%95/index.xml" rel="self" type="application/rss+xml"/><item><title>CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-08-28-criticl-weak-to-strong-paper-reading/</link><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-08-28-criticl-weak-to-strong-paper-reading/</guid><description>推理时扩展（多次采样投票、自我反思、LLM 裁判）能提升推理精度，但代价是数倍的生成次数与 token 开销。这篇 COLM 2026 论文（俄亥俄州立大学 + 普林斯顿）提出 CritICL：把同族小模型的失败模式当作结构化知识，在推理时以批评式上下文示例引导大模型绕开自身陷阱。其成立基础是一个漂亮的实证发现：同族模型的失败模式分布跨尺度高度一致——Qwen 1.5B 与 72B 的 top-20 失败模式排序与幅度基本稳定，且多小模型聚合分布比单个更逼近大模型。效果上，CritICL-static 让 Qwen2.5-72B 达到 59.2%（超 Consistency@5 的 59.0），而 token 总量仅 3768（1 次生成）vs Consistency@7 的 5440（7 次生成）、Self-Reflection 的 7533。『用失败而非成功做弱到强迁移』，失败是比正确示范更可迁移的信号。</description></item></channel></rss>