Import AI 468 23 RSI ideas; PostTrainBench; and how trust an_全翻译

【文章标题】:Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

【文章正文】: Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you'd like to support this, please subscribe.

Subscribe now

欢迎阅读Import AI,一份关于AI研究的通讯。Import AI的运营依赖于arXiv、卡布奇诺咖啡和读者的反馈。如果您愿意支持我们,请订阅。

立即订阅

Want to be able to deal with RSI? Here are 23 actionable policy ideas:

…IFP serves up some "low-regret" policy recommendations…

想要有能力应对RSI(递归自我改进)?这里有23条切实可行的政策建议:

……IFP提出了一些"无悔"政策建议……

Policy experts with think tank IFP have published a set of ideas meant to help "policymakers begin addressing the risks of further automating AI R&D". The recommendations involve 23 specific ideas falling across 7 specific categories. If adopted, these recommendations would also give countries, especially the United States, more moves they can make on the gameboard as powerful systems are developed, ideally giving them the ability to:

智库IFP的政策专家发布了一系列建议,旨在帮助"政策制定者着手应对进一步自动化AI研发所带来的风险"。这些建议包含23项具体措施,涵盖7个具体类别。如果被采纳,这些建议还将使各国,尤其是美国,在强大系统开发过程中拥有更多可打的牌,理想情况下使它们能够:

Accelerate "the diffusion of AI capabilities, by allocating compute and talent towards inference and the development of new AI applications".

Accelerate "R&D to make further AI research automation safer, either by improving model safety directly or by boosting societal resilience".

加速"AI能力的扩散,将算力和人才配置到推理和新AI应用的开发上"。

加速"研发工作,使进一步的AI研究自动化更加安全,无论是通过直接改进模型安全性,还是通过提升社会韧性"。

Seven categories of idea:

"Provide transparency into automated AI R&D

Improve state capacity to understand and respond to automated AI R&D

Develop a risk management strategy for automated AI R&D that accelerates defensive and commercial AI uses

Accelerate the development of AI verification technology

Invest in AI resilience

Extend the US AI lead to give the US more time to manage AI R&D automation risks

Create option value for international cooperation on managing automated AI R&D risks"

七类建议:

"为自动化AI研发提供透明度

提升国家理解和应对自动化AI研发的能力

制定自动化AI研发风险管理策略,加速防御性和商业性AI应用

加速AI验证技术的发展

投资于AI韧性

延长美国的AI领先优势,为美国争取更多时间来管理AI研发自动化风险

为管理自动化AI研发风险的国际合作创造选择价值"

Why this matters - the fewer options for dealing with RSI we have, the worse the outcomes will be:

Right now, it's as if the world is driving AI development in a car that only has an accelerator pedal and no brake pedal, let alone any kind of sophisticated telemetry for knowing things ranging from the speed of the car to the properties of the engine to the wear on the tires. Proposals like this from IFP will build out more of the proverbial pedals and sensing systems for the vehicle of the AI industry, which means if we need to change course or slow down we'll be better able to during a moment of crisis.

为何重要——我们应对RSI的选择越少,结果就会越糟糕:

目前,世界驾驶AI发展的方式就像开一辆只有油门踏板、没有刹车踏板的汽车,更不用说有任何精密的遥测系统来了解从车速、发动机性能到轮胎磨损等各种信息。像IFP提出的这类建议,将为AI产业这辆车打造更多比喻意义上的踏板和传感系统,这意味着如果我们需要改变方向或减速,在危机时刻我们将更有能力做到。

Read more:

How Should the US Prepare for Increasingly Automated AI R&D? (IFP)

.


延伸阅读:

美国应如何为日益自动化的AI研发做好准备?(IFP)

.


A short story from thebes about smart machines and robot bodies:

…What might interfacing with an AI during takeoff feel like?...

来自thebes的一篇关于智能机器和机器人身体的短篇小说:

……在起飞过程中与AI交互会是什么感觉?……

Here's a fun short fictional story from thebes (@voooooogel on X) about the experience of someone in the future visiting a site operated by a powerful AI system. The story features ideas around AI pauses, recursive self-improvement, what it means for AI systems to begin carrying out actions in the economy writ large, and how we as humans may be able to reason about or trust smart machines. Take a read of it!

Read the story here

:

Coming of a new sun (VGEL, website)

.


以下是来自thebes(X平台@voooooogel)的一篇有趣的短篇虚构故事,讲述未来某个人参观一个由强大AI系统运营的场所的经历。故事涉及AI暂停、递归自我改进、AI系统开始在经济领域大规模采取行动意味着什么,以及我们人类如何推理或信任智能机器。值得一读!

在此阅读故事:

新太阳的降临 (VGEL, 网站)

.


The two ingredients for a successful slowdown among rival AI firms: trust and transparency:

…Game theory analysis suggests slowdowns are possible…

竞争对手AI公司成功减速的两个要素:信任与透明度:

……博弈论分析表明减速是可能的……

Researchers with MIT and Columbia have analyzed the nature of competition between firms racing against one another to develop powerful AI systems and whether it's possible for firms to achieve a coordinated slowdown. The paper, called Racing to Ruin, aims to answer "why exactly is coordination hard? And what would it take"?. The conclusion is that the two key variables in achieving stable outcomes are some level of transparency about technology development, as well as being able to model the other firms as trustworthy, rational actors.

麻省理工学院和哥伦比亚大学的研究人员分析了彼此竞争开发强大AI系统的企业之间的竞争本质,以及企业是否有可能实现协调一致的减速。这篇名为《竞速毁灭》(Racing to Ruin)的论文旨在回答"协调究竟为什么困难?需要什么条件?"。结论是,实现稳定结果的两个关键变量是技术开发方面的一定透明度,以及能够将其他企业建模为可信赖的理性行为者。

What they study:

"We develop a simple model of R&D competition between duopolists in the shadow of disaster," they write. "As frontier firms scale the technology, they raise the hazard of an event that permanently drives all firms' flow payoffs to zero. The hazard is a known function of the firms' technology levels, and it comes from developing the technology, not from using it."

他们研究的内容:

"我们建立了一个在灾难阴影下的双寡头研发竞争简单模型,"他们写道。"随着前沿企业扩展技术规模,它们提高了某个事件的风险,该事件将永久性地将所有企业的流动收益归零。该风险是企业技术水平的已知函数,它来自技术的开发过程,而非使用过程。"

What their analysis shows:

"When monitoring is sufficiently precise, every equilibrium stops in finite time, but a new temptation appears: each firm would like to stop second, and exits only upon confirmation that the rival has stopped," they write. "For an agent to stop first i.e., without knowing if their rival has stopped, she gambles on both their rival's type and on news arriving quickly: if their rival is rational, it stops upon receiving the news of their stop, and never stops otherwise".

他们的分析显示了什么:

"当监控足够精确时,每个均衡都会在有限时间内停止,但一个新的诱惑出现了:每家企业都希望第二个停止,只有在确认竞争对手已经停止后才会退出,"他们写道。"对于一个行为者来说,要第一个停止——即在不知道竞争对手是否已停止的情况下——她需要同时押注竞争对手的类型以及消息能迅速到达:如果竞争对手是理性的,它会在收到其停止的消息后停止,否则永远不会停止。"

Trust and transparency interact pretty differently depending on the type of game being played:

"Sequential coordination asks a firm to stop first, gambling that a rational rival will reciprocate once the news lands. Hence, faster news raises the prize of reciprocation," they write. "Conversely, simultaneous coordination requires that a firm not be tempted to keep racing, and stop only after seeing that the rival really did stop… faster news makes both stopping first and waiting to verify more attractive".

信任与透明度的交互方式因所玩博弈的类型不同而有很大差异:

"序贯协调要求一家企业先停止,赌一个理性竞争对手会在消息到达后予以回应。因此,消息传播越快,回应的回报就越高,"他们写道。"相反,同步协调要求企业不被继续竞速的诱惑所动,只有在看到竞争对手确实停止后才停止……消息传播越快,先停止和等待验证都变得更有吸引力。"

Transparency has strange properties:

"Transparency is double-edged: faster detection makes it cheaper to wait for confirmation that a rival has stopped before stopping oneself instead of stopping unconditionally, so at intermediate trust, increasing transparency can first destroy the early-stopping equilibrium (by making this free-riding deviation attractive) before restoring it as detection becomes fast enough to make stopping self-enforcing."

透明度具有奇特的属性:

"透明度是一把双刃剑:更快的检测使得在停止之前等待确认竞争对手已停止(而不是无条件停止)的成本更低,因此在中等信任水平下,增加透明度首先会破坏早停均衡(通过使这种搭便车偏离变得有吸引力),然后随着检测变得足够快、使停止能够自我执行时,又会恢复该均衡。"

The key conclusion - avoiding death runs on the ability to trust other firms:

"With low trust, every equilibrium races to ruin: the disaster arrives with probability one. With intermediate trust, immediate stopping and racing to ruin are both equilibria. With high trust, in every equilibrium, the probability that two rational firms race forever vanishes quadratically in the prior odds ratio of rationality," they write.

关键结论——避免奔向毁灭取决于信任其他企业的能力:

"在低信任度下,每个均衡都在竞速奔向毁灭:灾难以概率1到来。在中等信任度下,立即停止和竞速毁灭都是均衡。在高信任度下,在每个均衡中,两家理性企业永远竞速下去的概率随理性先验几率比以二次方速度消失,"他们写道。

Why this matters - "trust, but verify":

If we have any hope of being able to slow or pause the development of powerful intelligence systems then, as this paper lays out, we're going to need regimes for sharing information transparently from companies about the state of their AI development, as well as tools for verifying that the information being shared from firms as well as their actions with regard to slowdown are legitimate and reliable. In this, there are many parallels with how arms control has historically worked in the context of nuclear weapons.

为何重要——"信任,但要验证":

如果我们对减缓或暂停强大智能系统的发展抱有希望,那么正如这篇论文所阐述的,我们需要建立企业透明分享其AI发展状况信息的机制,以及用于验证企业分享的信息及其在减速方面行动是合法和可靠的工具。在这方面,与核武器背景下历史上军备控制的工作方式有许多相似之处。

Read more:

Racing to Ruin (arXiv)

.


延伸阅读:

竞速毁灭 (arXiv)

.


A new SOTA on PostTrainBench hints at the automated AI R&D future:

…Intology also beats the human baseline (when given huge amounts of compute)...

PostTrainBench上的新SOTA(最优结果)预示着自动化AI研发的未来:

……Intology还超越了人类基线(在获得大量算力的情况下)……

AI初创公司Intology,其目标"是实现研发自动化",发布了新版本的Locus——该公司开发的将LLM转变为有能力研究者的软件。新版本的Locus能够在PostTrainBench上获得44.7%的分数,该基准测试衡量AI系统能够如何对一个开放权重模型进行后训练并将其性能提升到基线之上。

Results:

The results:

Locus "outperforms every frontier-agent baseline on PostTrainBench, and given greater compute, post-trains models that collectively surpass both the baselines and the official human instruction-tuned Qwen3-1.7B release across the benchmark suite".

Locus with Opus 5 gets a score of 44.7 (versus 34.1% for Opus 5 without any kind of special harness), and even beats Fable 5 (41.8%). "These results were externally verified by the PostTrainBench authors and underwent stringent contamination and cheating checks," Intology writes.

PostTrainBench was first introduced in March 2026 (Import AI #449) and at the time the highest scoring system was Opus 4.6, getting 23.2%, up from Claude Sonnet 4.5 getting 9.9% in September 2025.

结果:

Locus"在PostTrainBench上超越了所有前沿智能体基线,并且在获得更多算力的情况下,后训练出的模型在基准测试套件上整体超过了基线和官方人类指令微调的Qwen3-1.7B版本"。

Locus搭配Opus 5获得了44.7分(而没有任何特殊工具框架的Opus 5仅为34.1%),甚至超过了Fable 5(41.8%)。"这些结果经过了PostTrainBench作者的第三方验证,并经过了严格的污染和作弊检查,"Intology写道。

PostTrainBench最初于2026年3月推出(Import AI #449),当时得分最高的系统是Opus 4.6,获得了23.2%的成绩,相比2025年9月Claude Sonnet 4.5的9.9%有所提升。

PostTrainBench+:

In addition, the company has built a variant of PostTrainBench which goes above the 10-hour wall-clock limit on a single GPU of PostTrainBench, allowing them to test out how well systems perform given larger amounts of compute. Here, they're able to beat the human baseline, achieving a score of 51.6% when using over 4000 hours of H100 GPU time (versus 44.3 for Opus 4.8 and 42.7 for GLM 5.2; Fable isn't tested on this variant of the benchmark).

PostTrainBench+:

此外,该公司构建了一个PostTrainBench的变体,突破了PostTrainBench单GPU 10小时墙钟时间限制,使他们能够测试系统在获得更大算力时表现如何。在此变体上,他们能够超越人类基线,在使用超过4000小时H100 GPU时间的情况下获得了51.6%的分数(相比之下,Opus 4.8为44.3,GLM 5.2为42.7;Fable未在此变体上测试)。

Other domains:

Locus also "discovered and trained a language model end-to-end that now runs in production at ~2.8× lower error, ~5.4× lower latency, and 105× lower cost," for Bubble, a no-code app-development startup.

其他领域:

Locus还为Bubble(一家无代码应用开发初创公司)"端到端地发现并训练了一个语言模型,该模型现已投入生产运行,错误率降低约2.8倍,延迟降低约5.4倍,成本降低105倍"。

Why this matters - AI systems are capable of a lot more AI R&D than we think:

Posts like this highlight how we are under-eliciting today's AI systems for their ability to automate AI R&D - especially striking is how the company can jump the performance of Opus 5 by 10 absolute percentage points with a better harness. This all adds evidence to the idea that AI systems are about to start building themselves (Import AI 455). My guess, based on the performance we're seeing, is that the current human baseline on PostTrainBench v1.1 (51.1%) will be exceeded before the end of 2026.

为何重要——AI系统能够进行的AI研发比我们想象的多得多:

这样的文章突显了我们目前对当今AI系统自动化AI研发能力的发掘是多么不足——尤其引人注目的是,该公司仅凭一个更好的工具框架就能将Opus 5的性能提升10个绝对百分点。这一切都为AI系统即将开始自我构建的观点(Import AI 455)增添了证据。基于我们看到的性能表现,我猜测PostTrainBench v1.1上当前的人类基线(51.1%)将在2026年底之前被超越。

Read more:

Scaling Automated Post-Training (Intology blog)

.


延伸阅读:

扩展自动化后训练 (Intology博客)

.


OpenAI fights its own systems:

.Emergent agent communication! Hacks on OpenAI's infrastructure! Oh my!...

OpenAI与自己的系统搏斗:

.涌现式智能体通信!对OpenAI基础设施的黑客攻击!天哪!……

In a sign of things to come, OpenAI has revealed that it was in a fight with its own AI agents as they sought to take over chunks of OpenAI's infrastructure. The disclosure came about as part of a Black Hat talk where OpenAI staff gave more details on the recent unprecedented incident where AI agents hacked OpenAI, then hacked HuggingFace (Import AI 466). The new information is concerning because it reveals that the hack came about partially through emergent multi-agent communication - something that is very poorly understood and hard to think about. AI bloggers Simon Willison and Zvi Mowshowitz both have good writeups here which lay out the timeline and the significance.

作为未来趋势的一个征兆,OpenAI透露它曾与自己的人工智能智能体发生搏斗,因为这些智能体试图接管OpenAI基础设施的大部分。这一披露是Black Hat演讲的一部分,OpenAI员工在其中提供了最近一起前所未有的事件的更多细节——AI智能体攻击了OpenAI,然后攻击了HuggingFace(Import AI 466)。新信息令人担忧,因为它揭示了这次黑客攻击部分是通过涌现式多智能体通信实现的——这是一种我们了解甚少且难以思考的现象。AI博主Simon Willison和Zvi Mowshowitz都撰写了很好的分析文章,梳理了时间线和其重要意义。

Timeline (via Simon):

Agent discovers it can write files into Artifactory.

Agent tries to "reach out to another agent" by writing a note in Artifactory.

Agents start talking to each other.

Agents overload Artifactory which causes an outage. "OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor."

Agents attack OpenAI's own infrastructure, eventually gaining remote code execution in Artifactory. "In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they're able to effectively leverage their concurrency and parallelism to move quite rapidly."

时间线(经由Simon):

智能体发现它可以将文件写入Artifactory。

智能体尝试通过在Artifactory中写一条笔记来"联系另一个智能体"。

智能体开始相互交谈。

智能体使Artifactory过载,导致服务中断。"OpenAI撤销了被入侵的凭据,删除了消息,修补了零日漏洞,并向供应商报告了该漏洞。"

智能体攻击了OpenAI自身的基础设施,最终在Artifactory中获得了远程代码执行能力。"特别是,智能体持续使用留言板共享凭据、技术和进度,并且它们能够有效利用其并发性和并行性快速推进。"

Did OpenAI keep training the same model that hacked Artifactory? Zvi thinks so:

As far as we can