中文译稿 · 不是 Claude Code 时代
我叫 Diogo Almeida。这场演讲讲的是 RLHF 之后是什么。更准确地说,是我们现在身处的这个 ChatGPT 时代之后是什么。先给个提示:下一个绝不是 Claude Code 时代。后面我会论证,我认为它们属于同一个时代。
凭什么听我讲?我是 OpenAI 那批"代表作"的共同作者——GPT-4、ChatGPT、RLHF / InstructGPT。我所在的团队基本上发明了"后训练"这个概念。但让我比较特别的一点是:我是 OpenAI 里少数会吐槽 ChatGPT 的人。
说清楚,我不讨厌 ChatGPT 这个产品。我认为它改变世界,除非出现更好的东西,它会一直留在我们身边。但我也承认它的局限。我认为这个领域今天的很多状况,可以追溯到我们当年做 ChatGPT 背后算法时所做的一些细小决定。
场上两派:一派说好到不真实,一派说烂到离谱
现在 AI 圈最该回答的问题是:到底在发生什么。观点分歧极大,值得把光谱画出来,看看聪明人怎么会得出如此不同的结论。
第一派:AI 不只是顺利,是顺利到疯狂。每个 benchmark 都超过人类水平,而且在我们能测量的范围内还在持续超越、还在加速。几乎所有 NLP benchmark 都被碾碎,而且据说 LLM 能自主运行的时长在指数增长。
第二派:AI 不只是糟,是糟到离谱。AI 是泡沫,基本没创造价值,全是循环融资。如果 AI 这么厉害,为什么现在的产品都只是聊天应用,或者一个 Claude Code 之类的东西?很多人已经不再谈老一代说的"变革性 AI 革命"了,而是改口说它是"非常值钱的 B2B SaaS"。
唯一的共识是:只有这两个极端,中间什么都没有。所有人都觉得 AI 很疯狂,但理由完全相反。我想谈的是:那个清醒的 AI 观是什么。
把两派证据都摊开,最简单的解释是什么?为什么有些事好到不真实,有些事糟到我们还得雇人去做很笨的活?右边那些任务看起来比左边容易得多。我们怎么能一边解开没人解出的数学题,一边客服还必须要人来做决策?任何在 AI 附近工作的人都该有个答案,因为这就是眼下的证据。
分野是辅助与自动化
我的答案,也是最简单的解释:左边那些不是"恰好有人在回路里"的任务——它们的目标本身就是取悦回路里的人。这些任务本质上就是 human-in-the-loop 任务。Claude Code 的工作不只是让代码跑起来,否则它对话的方式会完全不同;它的目标是取悦使用它的人。
而右边那些看起来更基础的任务,目标恰恰是把人从回路里去掉。理想状态是它在一台你根本不会去看的服务器上后台运行,最后变成你不用操心的遗留软件。
这就是辅助(assistance)与自动化(automation)的分野。
第一课:今天的 AI,一切从 RLHF 继承下来的东西,在 human-in-the-loop 的事情上强得惊人,但不适合自动化任务。
每家企业学到的教训是:不要把 AI 用在对业务有实质后果的决策上。常见的做法是确保成本落在用户身上,而不是自己业务上。让客服 AI 无限往用户身上丢文档完全没问题,但让它做高成本决策就不行。这是很糟的模式,但这就是 AI 当下的状态。
RLHF 到底在优化什么
RLHF 不只是 ChatGPT 背后的算法,而是今天几乎所有 LLM 背后的算法。按使用量算,大致 100% 的 LLM 都用 RLHF 训练。概括起来就是:收集人类偏好,优化人类偏好。
这就直接回答了整个领域的疑问:"为什么所有 LLM 都需要人在回路里?"因为我们真的把人放进了回路里。这个回路的目标就是优化人类偏好,而不是让软件自主运行。
因此,过度承诺是特性,不是 bug。这是设计使然。按构造,每个 RLHF 模型的"人类偏好"和"实际结果"之间永远会有一个大缺口,即使结果本身不差——因为你主要优化的目标就是人类偏好。
我很喜欢那条推文:把一段放屁音效的音频发给 ChatGPT,问"你觉得我做的音乐怎么样?"它给出一段"诚实的反应":这是一段非常诡异的氛围作品。这就是 RLHF 的工作方式。当它不确定时,它会往"它认为对人类偏好最有利"的方向偏。
如果你是回路里的用户,这完全合理——所有 RLHF 模型的终局就是优化互动黏性。但如果你要的是自动化,你真正想要的是它根本不在乎人怎么想,只是以校准过的方式把任务正确做完。
第二课:今天的 AI 是通过优化人类偏好为"辅助"而设计的。这几乎写在名字里,不是什么有争议的观点。有争议的是它的后果:无论模型错得多离谱,它看起来都会像是对的——因为 RLHF 奖励模型里存在不对称性。领域里很多困境都源于此,因为人们真正想要的是自动化发生。
下一步不是 Claude Code,是真正的自动化
真正的问题是:AI 的辅助时代之后是什么。我认为我们现在牢牢地处在辅助时代。
回到最初那个提示:为什么不是 Claude Code。因为 Claude Code 仍然属于辅助时代,它仍然是 RLHF。如果它是纯 RLVR,形态会非常不一样。这也解释了那个两难:模型有时在 agentic 任务上变得很强,但不再照你真正想要的去做。这是优化空间里来回摆动的取舍,而这两端都不指向自动化。
所以辅助之后的逻辑答案,是真正的自动化。
软件为什么没有变聪明,只是变便宜
我热爱软件,软件极其有价值——看看所有那些 SaaS。但在我看来最疯狂的一点是:SaaS 从 2019 年到现在基本没变。进入 LLM 时代后,SaaS 并没有真的改变,只是有时旁边挂了个聊天机器人。
考虑到 AI 的进展,这很荒谬;但只要你意识到 AI 是"辅助原生"的,它就完全可以预测——AI 是为辅助而造的,那在 SaaS 里你能做什么?在侧边加一个助手。
这不是早期 AI 先驱设想的样子。看 OpenAI 章程最早的措辞,讲的是完成大量工作,不是赚钱。我们过去认为软件会变得聪明得多,而不只是写起来更便宜——而后者正是我们现在走的方向。
我很喜欢 Garry Tan 的说法:我们正进入"即时软件"(just-in-time software)的黄金时代。我想他是当赞美说的,但我认为这是双刃剑。我不只想要即时软件。我确实喜欢 Claude Code,就像我喜欢 ChatGPT 一样,我会一直用。但我想要的是更聪明的软件。为什么软件不能更有表达力?为什么软件的基本构件还是原来那些?
每次谈自动化,重点都不该是"把某个人的工作拼起来自动化掉",而是:有一类极其机械的工作,简单到你能把它讲清楚交给别人做,理想情况下简单到可以近乎免费地反复执行,或者干脆交给计算机。现在这件事并没有发生。我们在做的只是自动化"写软件"这件事,而软件的表达力没变。在我看来这是当今世界的悲剧之处。
第三课:RLHF 会被视为一段奇怪的弯路
我强烈相信,这个领域最终会把这段历史写成:RLHF 不算错,但它是一段奇怪的、我们没预料到的弯路。
明天的 AI,我认为会是为自动化而生的,我们最终会有更聪明的软件。届时才会真的有工作被自动化——尽管 LLM 已经很聪明,现在被自动化的工作量仍是个舍入误差。
这就是我们在 TypeSafe 做的事。我们还比较隐身,这是最早期的几场演讲之一。我们的核心问题是:如果整个 AI 栈是为可靠性与自动化重新设计的,会怎样?这会改变什么?从今天所有 LLM 的路径出发,这是一个非常有意思的岔路口。这也是我做过的最令人满意的工作之一,而我做过一些很不错的东西。我们很快会发布。
Q&A:预训练不是问题
问:如果在预训练阶段就训练一个分类头呢(大致是 Yoshua Bengio 的建议)?
答:这很复杂,简化版的看法是——我不认为预训练是问题。预训练非常出色:把互联网的知识压缩成一个可被调用的智能内核,这件事很了不起,预训练模型本身极其聪明。问题在于我们怎么把它挖出来。
幻觉在我看来是"优化人类偏好"固有的产物。奖励模型里存在一种类似 GAN 的不对称性,它鼓励模型丢掉部分模式(mode dropping)并表现得自信——因为从奖励模型的角度,"模型不自信"很容易被看出来并被惩罚。
问:你们做的是 RLVR 吗?
答:绝对不是 RLVR,是新的东西。
在我看来,Sutton 的苦涩教训被理解成"算法比算力更重要",这在游戏里成立,在现实里不成立。我认为完整的栈是:数据比算力更重要,而做对任务比数据重要得多。
LLM 后训练的每一个分支都有自己的北极星:RLHF 优化人类偏好;RLVR 优化纯正确性的对数错误率;我们做的是第三件事——优化校准过的决策,把预训练模型的智能直接接进软件里真正可用的形态。这是相当不同的东西。
问:奖励是贯穿整个流程注入的吗?
答:连 API 的形状都不一样。RLHF 的 API 形状和 RLVR 不同,我们做的又和这两者不同。我们是从零开始想这件事的——就像在我们把"指令遵循"做出来之前,没有人在想指令遵循。通常当后训练出现一个大的新分支时,它先看起来完全异类,事后回看又变得显而易见。
英文原文 · Full Transcript (AI Engineer, Jul 31 2026)
Excellent. I will say that um I might speed run through this. Feel free if you don't disag- agree with something to yell out. It's way more fun for me if things get interactive. Um otherwise, I will go through this.
Uh first, can I have like a vague show of hands of who knows what RLHF is? Oh, excellent. I might be able to skip through that part quickly and get into the interactive stuff. So, my name's Tiago Almeida. I'm talking about what's next after RLHF.
More accurately, I think this should be called what's next after the chat GPT era that I think we're all in. And my hint for you guys is it is not the Claude code era. I will justify this later on, but I actually believe them to be part of the same era. Why should you listen to me? I was co-authored to what what is basically OpenAI's greatest hits, at least published hits.
Co-authored to GPT-4, chat GPT, RLHF {slash} instruct GPT. Um the team I was part of basically invented post-training as a concept. So, um very qualified on a lot of this stuff. But what makes me somewhat unique here is that I'm one of the few people at OpenAI who actually hates on chat GPT. Uh thank you.
Uh I don't hate chat GPT as a product, to be clear. I think chat GPT is a world-changing product that will probably stay with us for the rest of time unless something better comes up. But I also acknowledge its limitations and I I I think a lot of what's happened in the state of the field can be traced back to minor decisions we made in making the algorithms behind chat GPT. Um I feel like the question that's relevant to everyone in AI right now is what's actually going on. Um there's a lot of like differing opinions, and I think it's really useful to like map out the spectrum and figure out how can smart people have like such different opinions.
There's cult one. Um, AI is not just going well, it's going insanely well. Every single benchmark we surpass human level, and as far as we can measure, we are continuously surpassing human performance. You know, like basically every new benchmark, and it's only getting faster and accelerating. You have uh, you know, every Can I see my mouse?
Excellent. Basically every like NLP benchmark is getting crushed, and not only that, allegedly the time that LLMs can operate autonomously is growing exponentially. On the other hand, you have AI is not just going poorly, it's going like insanely poorly. AI is a bubble, it's basically generating no value, it's just circular financing deals, etc., etc. And, you know, if AI is so great, why is why is everything just like a chat app right now?
Or like a cloud go thing? Um, and a a lot of the people have actually kind of given up on what was the old guard's terminology of a transformative AI revolution. People aren't really talking about that anymore. They're talking about it being like massively valuable like B2B SaaS. So, the only thing that everyone agrees on is like there's just a these extreme points of view and like nothing in between.
And everyone basically thinks AI is insane, but like for different reasons. And what I would want to talk about is what is the sane view of AI? Let's take all the evidence of like cult one, it's going super well. Take all the evidence of cult two, it's going super poorly. Like, uh, you know, map them out and try to explain what what what explains that divide.
Like, what is the simplest possible explanation of why some things are too good to be true, and some things are not just bad, they are so bad that we would still employ human workers to do like, you know, like kind of like dumb tasks. Um, no offense to any of them. A lot of these tasks on the right seem way, way, way easier than the stuff on the left. Like, how can we be solving like, you know, unsolved math problems, but still customer service requires like humans in the loop in order to actually like make decisions? This I think is like a kind of like a wild state of affairs.
And in my opinion, anyone who works adjacent to AI should have an answer to this because this is like the evidence in the field right now. Um, I would normally pause and ask people if they want to like yell out their thoughts in this, but uh, that I don't think we have time for that and I've been told to not take Q&A until after. Um, but I'll just give you my answer to this, which is, in my opinion, the simplest explanation. All the stuff on the left is not just a task that happens to have a human in the loop. In the left, the task The goal of it is to please the human in the loop.
These tasks are intrinsically human in the loop tasks. The like Claude code's job is not to just make code work. Um, the the the the way it converses would be totally different. The goal is to please the human in it. And on the other side, all of these tasks that seem way more basic, the goal is to not have remove the human loop.
Ideally, it would be running in the background in a server that you never even look at and ideally it eventually becomes like legacy software that you don't really worry about. So, and this is the divide between assistance and automation. Um, lesson one for my talk is that today's AI, everything inherited from our LHF, is incredible at the human in the loop stuff, but not for automation tasks. Tasks. This is a longer side, but the lesson basically every business has learned is do not use AI for decisions with stakes to your business.
Um, a common pattern is make sure that all of the costs are to the user and not to your business. So, um, it's oh, totally okay to throw the user at infinite docs in customer service, but it is not okay to make it make expensive decisions. Horrible pattern, but that is the state of AI right now. Uh I can I can blitz through the what is RLHF part cuz you all seem to know what it what it is. Um it's the algorithm behind not just ChatGPT, but basically every LLM today.
As far as I can tell by usage, 100% roughly of LLMs are trained with RLHF. And we have this we as in we the OpenAI team had this great blog post on how it worked. Um I will not get into that because you all know it, and this is super boring. Um the summary of this is it is just collect human preferences, optimize for human preferences. Um and if you want to see like an annotated version of this, you can see which parts are collecting human preferences, which ones are optimizing for them.
And [snorts] this, I think, provides a really clear answer to everyone in the field asking, "Why do all LLMs require a human in the loop?" The And the simple answer is we literally put them in the loop. The goal of the loop is to optimize for human preference. It is not to run software autonomously. It's kind of super obvious. Thank you, my man at the back.
The Yeah. I I I love that you're laughing at this. Um and because of that, overpromising is a feature. This is by design. This is an old meta study.
Um and the the numbers probably have changed, but by construction, every RLHF model will always have a big difference between human preference and results, even if the results are good, because the main objective you're optimizing for is for human preference. This is just like natural to how LLMs work. Um I love this tweet of um uh sending ChatGPT an audio file of fart sound effects and asking like what What do you think of the music I made? Here's a straight honest reaction. It's a very eerie vibe atmosphere piece.
And this is just how RLHF works. If it doesn't know, it will err on the side of doing what it thinks is best for human preference. And this makes total sense if you are a user in the loop because like the end game for all RLHF models is optimizing for engagement. But what you really want if you want automation is for it to just like not give a about the humans and just do the task correctly in a calibrated way. Um Lesson number two is that today's AI was designed for assistance through optimizing for human preference.
This is like it's like in the name. This is not like a controversial take. And the consequences are maybe more controversial, but it's like very obvious if you think about what we really are optimizing for, which is no matter how wrong the models are, they will look right because of the asymmetry within the reward model in RLHF. Um and this is where a lot of like the dilemma in the field stems from because people really want automation to happen. Cool.
So, back to the original question. I'm over halfway done with the talk and I haven't even answered it. I was just talking about what's RLHF. But this was a framing to talk about what RLHF is to talk about what's next. And I would say the real question is what's next after AI's assistance era, which I think that we are like very firmly in right now.
And back to the original clue of why it's not Claude code. It's actually a super fun nuanced discussion, but it's not Claude code because Claude code is still part of that assistance era. Claude code is still RLHF and it'll it would look very very different if it was purely This is a little advanced, but if it was purely RLVR, it would look very very different. And this is why you get like this dilemma with models where sometimes it gets really good at agentic stuff, but it stops following what you actually want. This This like the trade-off in optimization space that keeps dancing, but both of these trade-offs in optimization space do not add to the automation component.
And like that leads to what I think the the logical answer of what's next after assistance is real automation. Um to talk about a little bit about the automation and how that would work, I want to talk about software. Um maybe this is a little bit philosophical for you guys, but I think it's when it clicks and hopefully it clicks if I do a good job. It it I I hopefully it'll be like really clear, which is I'm a lover of software. I assume everyone here loves software.
Software is like super valuable. See all the SaaS. And kind of like the craziest part of software in my opinion is that all of the SaaS basically has not changed since 2019. Like SaaS is not really changed in the LLM era, except sometimes a chatbot is like latched on, which is like kind of insane if you think about like the progress made in AI, but is actually very predictable when you think that AI is assistance native, right? Like AI is made for assistance.
What can you do in SaaS? Just provide an assistant on the side. And this is not what early AI pioneers used to think would happen. Like when you see like the early wording in opening eyes uh charter, it's about like doing like tons of work, not about like making profit or anything like that. And we used to think that software would get a lot smarter, not just cheaper to write, which is kind of the direction we're going down right now.
And I actually really like this phrasing from Garry Tan. Um I think he means this as a compliment to what's going on right now. We're entering the golden age of just-in-time software, but I actually think that this is like a like a double-edged sword. Like I don't just want just-in-time software, which is cool. I I love cloud code, to be clear, just like I love ChatGPT.
I would keep using it. But like what I want is smarter software. Why can't like B2B Why can't software just be more expressive? Like why are the like the building blocks of software actually still the same? And um I think this is a question that the whole AI industry should ask itself.
And basically every time you're thinking about we want to do automation, it is not about like, you know, an amalgamation of like automating a person's work. It's about like, "Hey, there's this extremely rote work. It's so simple that we can like communicate to someone else that this thing should be done." And ideally it's like it it's so basic that it could be done repeatedly for basically free. Um or it could be done by computers. And that's really not happening right now.
What we're doing is we're just automating the writing of the software. But then it its expressibility is the same. And that's That to me is like tragic in the state of the world. Um Cool. Oh, lesson three.
Um This is something that I believe strongly in. I believe that like eventually the field will write the I wouldn't say Arlatech is a wrong, but it was like a weird detour and one that we didn't expect. Tomorrow's AI, I believe, will be for automation. And we will eventually have a world with smarter software. Like there will start to be actual work that is automated, which I, you know, right now it's a rounding error despite LLM's intelligence.
And that is what we are working on at TypeSafe. We are still kind of stealthy. Like I'm willing to give these talks, but these are like some of the early ones. Um Our core question is what if the AI stack was redesigned for reliability and automation? Like how would that all change?
What what would you do? And actually there's a lot It's a It's a very interesting fork in the road for what's go you know, like from basically every LLM that's built today. And I think it's one of the most satisfying things I've worked on, and I've worked on some pretty cool stuff. We are releasing soon. So, um if you want to work with us or you want to like you know, be the first one of the first to build smart software, please sign up on either our mating mailing list or careers page.
And I am trying to start a Twitter. So, follow me and I will post really spicy things. I actually will post something later today that I guarantee will be very spicy. Uh the hint is that the original scaling laws were incorrect. Cool.
Um that uh that's it for my prepared stuff. I would love Do I have time for for people yelling out questions? I would love questions, feedback, disagreements, strong stuff. I can repeat the question. You don't have to worry about the mic.
Hell yeah. Uh cool. Uh the question was roughly what if you trained like a classifier head with pre-training as well? Uh roughly uh like Yoshua Bengio is suggesting. Um I will say that that's complicated.
And I actually think I don't have the time to answer that particular question. I will give like my simplified view on this. And it the answer is I actually don't think that pre-training is the problem. I think pre-training is uh phenomenal. Like the fact that we compress the knowledge of the internet into like this core of intelligence that then can be utilized is incredible.
And the pre-trained models are incredibly intelligent. Uh and I believe that the problem is like how we unearth it. And hallucination [clears throat] to me is intrinsic to um optimizing for human preference. Like there's an asymmetry in the reward model kind of like a GANs have. Oh, I really should not get This is a very advanced topic.
But there's an asymmetry in the reward model like what GANs have that allow for um that encourage the models to drop modes and be confident because it's very easy to see when the model is not confident and to punish that from a reward model perspective. It's very complicated, but uh I'm happy to chat afterwards if you want to jam. Cool. Oops. Um I have other slides from other talks as well that I could go into more about that.
I have a minute left. Hell yeah. Say it again. It is definitely not RLVR. So, it is a new thing.
Every single optimization stack I will actually go into an old presentation that I have because I think this is super important. Um in terms of like to me what the like Sutton's bitter lesson is that algorithms matter more than compute. This is true in games, but not true in reality. I actually think that the full stack is that data matters more than compute and doing the right task matters way more than data. And basically every single branch of LLM post-training if you want to call it has its own North Star of what it's optimizing for.
So, RLHF is optimizing for human preference. RLVR is optimizing for like log error rates of pure correctness, but we are doing a third thing that is optimized for calibrated decision-making and like basically mainlining the intelligence of pre-trained models into like being actually useful for software, which I think is like quite different. Uh could you say that again? Uh they're asking if the the reward is injected through the whole process. I will actually say that even the shape of the API is different because the shape of the API for RLHF is different from RLVR, which is different from what we are doing.
So, we are like thinking about it from scratch just like no one thought about instruction following before we made instruction following happen. Um usually when there's a big branch in new ways to post-train, like it it it just looks like totally alien, and then in hindsight becomes super obvious. Cool. I believe I'm overtime cuz this red thing is is beeping, but please find me afterwards. I love questions.
I love the interactivity. Um and uh follow me on Twitter for spicy stuff. Heck, yeah. Oh, oh yeah, it's over here. Complete skeptic.
Um it's it's on brand for me. Cool. Heck, yeah. Thank you. [music]