<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI 基础设施 on MessageDaily</title><link>https://inkeast.github.io/MessageDaily/tags/ai-%E5%9F%BA%E7%A1%80%E8%AE%BE%E6%96%BD/</link><description>Recent content in AI 基础设施 on MessageDaily</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 11 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://inkeast.github.io/MessageDaily/tags/ai-%E5%9F%BA%E7%A1%80%E8%AE%BE%E6%96%BD/index.xml" rel="self" type="application/rss+xml"/><item><title>Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? 精读</title><link>https://inkeast.github.io/MessageDaily/posts/2026-09-11-phi-bench-llm-infrastructure-paper-reading/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://inkeast.github.io/MessageDaily/posts/2026-09-11-phi-bench-llm-infrastructure-paper-reading/</guid><description>USTC×StepFun×北大×HKUST×Yale×UPenn 六机构联合的 Φ-Bench 提出一个自指性问题：LLM 推理所依赖的基础设施（kernel、服务栈、集群）能否由 LLM 自己来工程化？85 个任务、三种渐进格式——Kernel 函数补全（KFC）→长程实现（LHI）→端到端优化（E2EO），双轴评分（性能+实现）+内置作弊检测。最强 Claude Opus 5 总分仅 36.53%（KFC 37.16%/LHI 21.60%/E2EO 62.94%），硬件与边缘类最好模型也仅 5.4%——&amp;lsquo;造物者维护造物&amp;rsquo;的能力缺口被量化暴露。</description></item></channel></rss>