The Latent Space “Cooking with…” series put OpenAI’s Chief Research Officer Mark Chen in a kitchen and got him to talk through the things research-org leaders rarely say on the record: where scaling laws actually live in 2026, why post-training and RL are the real bottleneck now, how OpenAI structures evals against a moving frontier, and what “AGI” means when you’re inside the org that named the goal.
This is one of the higher-signal AI Engineering interviews of the year — partly because Chen is unusually specific, partly because the format (informal, no slides, no PR minder) catches him in mid-thought.
“Scaling laws are still alive”
The headline disagreement in the field right now is whether pre-training scaling has plateaued. Chen’s position is unambiguous:
“I firmly believe in being on the exponential and in scaling laws. So I think any of these bear takes [are wrong].”
What’s shifted is which axis scales. Pre-training compute is still useful but is no longer where the marginal gain is largest. The current S-curve is RL post-training and inference-time compute — the o1/o3 line. Chen’s framing: pre-training gives the model world knowledge; RL teaches it how to reason with that knowledge. The “scaling continues” claim only sounds controversial because most observers are still looking at the pre-training axis.
RL where the reward signal is cold
The most engineer-relevant section. Chen breaks down where RL post-training works and where it doesn’t:
- Works well: domains with cheap, verifiable signal — math, code, structured games, anything where you can grade an answer mechanically.
- Hard: domains where the reward signal is “cold” — open-ended creative work, long-horizon planning with no ground truth, soft-skill domains (negotiation, taste, judgment).
His prediction: the next 12–18 months of capability work is people figuring out how to manufacture reward signal in cold-signal domains — synthetic verifiers, model-graded rubrics, multi-step process rewards. This is exactly the part of the stack that AI Engineers building production agents are already wrestling with.
Evals at the frontier: “every capability is an eval”
“Every capability on the [shopping list] is an eval. You need some [way to measure it].”
OpenAI’s internal process, in Chen’s framing: list the capabilities you want to push, define an eval for each, drive the eval up, ship. Trivial-sounding, but Chen makes the harder point — at the frontier, the evals themselves are the bottleneck. For superhuman capabilities (novel theorem-proving, frontier science discovery), no human grader can score the output. You either invent a verifier, or you don’t measure, and if you don’t measure you don’t ship.
This is the same problem the AI Engineering community is hitting in production: as agents get better at hard tasks, judging their work becomes the gating step.
Running the research org
A handful of org-design notes worth keeping:
- Pre-training and RL/post-training are distinct disciplines with different research cadences, different infrastructure, different talent profiles. Coupling them too tightly slows both.
- “Cold-signal” research needs slack time. RL on hard sciences doesn’t have a fast feedback loop, so you can’t manage it like a product team.
- Bringing soup to researchers is, somehow, a real strategy. Chen confirms the Mark Zuckerberg story and says taking care of researchers concretely matters for retention at the level OpenAI competes at.
On AGI as the goal-post
Chen is unusually casual about the term. AGI for him is less a magic threshold and more “the point at which my hobbies become the interesting thing again.” The technical content underneath: he expects continued capability growth on the current trajectory (pre-train + RL + inference-time compute), with the open question being whether the cold-signal domains can be cracked or whether progress narrows to verifiable domains.
If you’re building on top of these models, the practical takeaway is: assume the verifiable-domain capability curve keeps going up steeply, and plan for slower progress in soft domains for at least 12 months.
Key takeaways
- Scaling laws aren’t dead — they moved axes. Pre-training compute returns are diminishing relative to RL post-training and inference-time compute. Bear takes are looking at the wrong axis.
- RL needs reward signal, period. Domains with cheap verifiers (math, code) are racing ahead; cold-signal domains (taste, judgment, soft skills) lag and will keep lagging until verifiers are invented.
- Evals are the new bottleneck. At the frontier, you can’t grade what no human can verify. Manufacturing eval signal is the meta-research problem of 2026.
- Pre-training vs. RL are different orgs. Different cadences, different talent. Coupling them slows both.
- Reasoning > knowledge. Pre-training installs world knowledge; RL teaches the model how to use it. Both are required; only the second is currently scaling fast.
- For builders: assume rapid gains in verifiable-domain agent capability over the next 12 months; plan around slower gains in soft/creative domains.
- Org notes that matter: treat researchers well, let cold-signal projects breathe, don’t run RL teams like product teams.
Source
- Talk: Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
- Speaker: Mark Chen, Chief Research Officer, OpenAI
- Host: Latent Space (“Cooking with…” series)
- Duration: ~41 minutes
- URL: https://www.youtube.com/watch?v=fpAthTtha8c