Import AI 470 No rights for machines; automating environment_全翻译
【文章标题】:Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
【文章标题】:Import AI 470:机器没有权利;用 SPADE 自动化环境生成;以及用 Hawkeye 构建更好的 GPU 内核
【文章正文】: Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.
【文章正文】: 欢迎阅读 Import AI,一份关于 AI 研究的通讯。Import AI 依靠 arXiv、卡布奇诺和读者反馈运转。如果你想支持它,请订阅。
Subscribe now
立即订阅
AI is accelerating some types of progress but not others:
AI 正在加速某些类型的进展,但不是所有类型:
…A nice METR study lays out where acceleration is showing up…
……一项不错的 METR 研究展示了加速出现在哪些地方……
Here’s a little analysis from METR which looks at where AI may be accelerating different types of science and technology. The study looks at three different areas: cyber, math, and AI research, and finds that AI has contributed a lot to cyber, a little bit to math, and it’s hard to say for AI.
以下是 METR 的一项简短分析,考察 AI 可能在哪些方面加速不同类型的科学和技术。该研究考察了三个不同领域:网络、数学和 AI 研究,并发现 AI 对网络贡献很大,对数学贡献一点,对 AI 研究则很难说。
Where have LLMs actually made a difference to scientific discovery?
LLM 究竟在哪些地方对科学发现产生了影响?
Cyber vulnerabilities: Major acceleration.
网络漏洞:重大加速。
“The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, and Microsoft) and for aggregate vulnerability databases (the US NVD, and OSV)”.
“与 2025 年相比,2026 年许多项目报告的漏洞速率急剧加快,无论是具体项目(cURL、OpenSSL、Firefox 和 Microsoft),还是汇总漏洞数据库(美国 NVD 和 OSV)都是如此。”
Mathematics research: Minor acceleration, but harder to measure.
数学研究:轻微加速,但更难衡量。
“AI is clearly contributing to more work being done (arXiv submissions have doubled in some areas in less than 12 months) but quantifying the value of those contributions is difficult.”
“AI 显然正在促成更多工作完成(某些领域的 arXiv 提交量在不到 12 个月内翻了一番),但量化这些贡献的价值很困难。”
Some math problems from prestigious lists have been solved, e.g., “the Jacobian conjecture from Smale’s list, Problem 44 from Green’s list (the halving sieve), and the sofic half of Green’s Problem 100”. However, it may be too early to determine how sustained a trend this is.
一些来自著名清单的数学问题已被解决,例如“斯梅尔清单中的雅可比猜想、格林清单中的第 44 题(减半筛法),以及格林第 100 题的 sofic 一半”。不过,现在判断这一趋势能持续多久可能还为时过早。
Optimization of AI research: No measurable acceleration.
AI 研究的优化:没有可衡量的加速。
When you look at algorithmic progress across seven significant problem areas (CIFAR-10, Hutter compression, Gurobi mixed-integer programming, MIPLIB, nanoGPT, Stockfish, and the matrix-multiplication exponent) there are a couple of these where LLM-attributable contributions have happened (nanoGPT, CIFAR-10), though the rate of increase of usage of AI here is a lot less than with cybersecurity and mathematics.
当你考察七个重要问题领域的算法进展时(CIFAR-10、Hutter 压缩、Gurobi 混合整数规划、MIPLIB、nanoGPT、Stockfish 和矩阵乘法指数),其中有一两个领域出现了可归因于 LLM 的贡献(nanoGPT、CIFAR-10),不过这里 AI 使用率的增长速度远低于网络安全和数学领域。
Why this matters - differential acceleration:
为什么这很重要——差异化加速:
This paper highlights how AI is causing advances in some parts of science and technology, but the effect isn’t unified across fields, rather there are pockets of lumpy acceleration (e.g., cyber) and areas where progress is more gradual (math, AI). My suspicion is that acceleration happens when models go through some kind of ineffable phase change for a given skill, as has evidently happened with day-to-day coding (2025), and cyber (2026). The key question is whether we are going to see phase changes in other parts of science and technology or if we won’t.
这篇论文强调,AI 正在推动科学和技术的某些部分取得进展,但这种影响在各领域并不统一,而是存在一些不均衡加速的区域(如网络)和一些进展更渐进的领域(数学、AI)。我怀疑,当模型在某种技能上经历某种难以言喻的相变时,加速就会发生,日常编程(2025 年)和网络(2026 年)显然已经发生了这种情况。关键问题是,我们是否会在科学和技术的其他部分看到相变,还是不会。
Read more:
阅读更多:
Research note: Have We Seen an Acceleration in Discoveries? (METR)
研究简报:我们看到发现加速了吗?(METR)
.
。
Automating environment generation with SPADE:
用 SPADE 自动化环境生成:
…A crude form of RSI bootstrapping via increasing data breadth…
……一种通过扩大数据广度来实现 RSI 自举的粗略形式……
A multi-university group of researchers have built SPADE, Self-Play in Adaptive Synthetic Executable Environments. SPADE is a “general framework for co-evolving environments synthesis and agentic capability through self-play”, and works as a way to generate synthetic data in the form of game-like environments which LLMs can subsequently be trained in, allowing developers to use a powerful model to bootstrap the creation of data that can then be used to further refine that same model.
一个多大学研究团队构建了 SPADE,即自适应合成可执行环境中的自博弈。SPADE 是一个“通过自博弈共同进化环境合成与智能体能力的通用框架”,其工作方式是生成类似游戏环境的合成数据,随后可在其中训练 LLM,使开发者能够使用强大模型来引导数据创建,而这些数据随后可用于进一步优化同一模型。
Who did it
- 谁做的
-
SPADE was developed by researchers with the University of Washington, Stanford University, Northeastern University, Carnegie Mellon University, Massachusetts Institute of Technology, National University of Singapore, Seoul National University, Stevens Institute of Technology, and the University of Chicago.
:SPADE 由华盛顿大学、斯坦福大学、东北大学、卡内基梅隆大学、麻省理工学院、新加坡国立大学、首尔国立大学、史蒂文斯理工学院和芝加哥大学的研究人员开发。
How it works
- 如何运作
-
SPADE has an LLM alternate between generating executable training environments (e.g., puzzles where a system needs to solve a simulated genetic problem in a biology lab) and having an LLM try to solve them. SPADE has two key roles for the model being used:
:SPADE 让一个 LLM 在生成可执行训练环境(例如,系统需要在生物实验室中解决模拟遗传问题的谜题)和让一个 LLM 尝试解决这些环境之间交替进行。SPADE 为所使用的模型设定了两个关键角色:
Environment Designer;
环境设计者;
writes complete, long-horizon training environments as executable code.
编写完整、长时程的训练环境作为可执行代码。
Reasoning Agent
推理智能体
; learns to act in the environments. The reward for the reasoning agent is estimated using the gap between its reward with and without privileged hints. A privileged hint (h) is “task-relevant information that the Environment Designer attaches to an environment (for example, a partial solution sketch or a key structural observation); revealing
;学习在环境中行动。推理智能体的奖励通过比较其有无特权提示时的奖励差距来估计。特权提示(h)是“环境设计者附加到环境中的任务相关信息(例如,部分解题思路或关键结构观察);将
h
h
to the Reasoning Agent makes the environment easier to solve, and the gap in Reasoning Agent return with versus without
透露给推理智能体会使环境更容易解决,而推理智能体在有
h
h
defines the
与无
Environment Designer’s
时的回报差距定义了
hint-based regret reward”.
环境设计者的基于提示的遗憾奖励”。
It works at the 30B scale:
它在 30B 规模上有效:
The authors train three Qwen3 backbones to test out SPADE: Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507. Unsurprisingly, Qwen3-30B works the best. Each model is tuned via GRPO for 400 rollouts of 25 environments each, then assessed against a variety of benchmarks including AIME, GPQA, LC
作者训练了三个 Qwen3 骨干模型来测试 SPADE:Qwen3-4B-Instruct-2507、Qwen3-8B 和 Qwen3-30B-A3B-Instruct-2507。不出所料,Qwen3-30B 效果最好。每个模型都通过 GRPO 进行调优,进行 400 次 rollout,每次包含 25 个环境,然后在包括 AIME、GPQA、LC 在内的多种基准上进行评估