The Latent Space “Cooking with…” series put OpenAI’s Chief Research Officer Mark Chen in a kitchen and got him to talk through the things research-org leaders rarely say on the record: where scaling laws actually live in 2026, why post-training and RL are the real bottleneck now, how OpenAI structures evals against a moving frontier, and what “AGI” means when you’re inside the org that named the goal.

This is one of the higher-signal AI Engineering interviews of the year — partly because Chen is unusually specific, partly because the format (informal, no slides, no PR minder) catches him in mid-thought.

“Scaling laws are still alive”

Scaling laws — Chen’s stance

The headline disagreement in the field right now is whether pre-training scaling has plateaued. Chen’s position is unambiguous:

“I firmly believe in being on the exponential and in scaling laws. So I think any of these bear takes [are wrong].”

What’s shifted is which axis scales. Pre-training compute is still useful but is no longer where the marginal gain is largest. The current S-curve is RL post-training and inference-time compute — the o1/o3 line. Chen’s framing: pre-training gives the model world knowledge; RL teaches it how to reason with that knowledge. The “scaling continues” claim only sounds controversial because most observers are still looking at the pre-training axis.

RL where the reward signal is cold

Where RL has the least to grip

The most engineer-relevant section. Chen breaks down where RL post-training works and where it doesn’t:

  • Works well: domains with cheap, verifiable signal — math, code, structured games, anything where you can grade an answer mechanically.
  • Hard: domains where the reward signal is “cold” — open-ended creative work, long-horizon planning with no ground truth, soft-skill domains (negotiation, taste, judgment).

His prediction: the next 12–18 months of capability work is people figuring out how to manufacture reward signal in cold-signal domains — synthetic verifiers, model-graded rubrics, multi-step process rewards. This is exactly the part of the stack that AI Engineers building production agents are already wrestling with.

Evals at the frontier: “every capability is an eval”

Evaluating frontier model capabilities

“Every capability on the [shopping list] is an eval. You need some [way to measure it].”

OpenAI’s internal process, in Chen’s framing: list the capabilities you want to push, define an eval for each, drive the eval up, ship. Trivial-sounding, but Chen makes the harder point — at the frontier, the evals themselves are the bottleneck. For superhuman capabilities (novel theorem-proving, frontier science discovery), no human grader can score the output. You either invent a verifier, or you don’t measure, and if you don’t measure you don’t ship.

This is the same problem the AI Engineering community is hitting in production: as agents get better at hard tasks, judging their work becomes the gating step.

Running the research org

Inside the research org

A handful of org-design notes worth keeping:

  • Pre-training and RL/post-training are distinct disciplines with different research cadences, different infrastructure, different talent profiles. Coupling them too tightly slows both.
  • “Cold-signal” research needs slack time. RL on hard sciences doesn’t have a fast feedback loop, so you can’t manage it like a product team.
  • Bringing soup to researchers is, somehow, a real strategy. Chen confirms the Mark Zuckerberg story and says taking care of researchers concretely matters for retention at the level OpenAI competes at.

On AGI as the goal-post

AGI trajectory

Chen is unusually casual about the term. AGI for him is less a magic threshold and more “the point at which my hobbies become the interesting thing again.” The technical content underneath: he expects continued capability growth on the current trajectory (pre-train + RL + inference-time compute), with the open question being whether the cold-signal domains can be cracked or whether progress narrows to verifiable domains.

If you’re building on top of these models, the practical takeaway is: assume the verifiable-domain capability curve keeps going up steeply, and plan for slower progress in soft domains for at least 12 months.

Key takeaways

  1. Scaling laws aren’t dead — they moved axes. Pre-training compute returns are diminishing relative to RL post-training and inference-time compute. Bear takes are looking at the wrong axis.
  2. RL needs reward signal, period. Domains with cheap verifiers (math, code) are racing ahead; cold-signal domains (taste, judgment, soft skills) lag and will keep lagging until verifiers are invented.
  3. Evals are the new bottleneck. At the frontier, you can’t grade what no human can verify. Manufacturing eval signal is the meta-research problem of 2026.
  4. Pre-training vs. RL are different orgs. Different cadences, different talent. Coupling them slows both.
  5. Reasoning > knowledge. Pre-training installs world knowledge; RL teaches the model how to use it. Both are required; only the second is currently scaling fast.
  6. For builders: assume rapid gains in verifiable-domain agent capability over the next 12 months; plan around slower gains in soft/creative domains.
  7. Org notes that matter: treat researchers well, let cold-signal projects breathe, don’t run RL teams like product teams.

Source

  • Talk: Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
  • Speaker: Mark Chen, Chief Research Officer, OpenAI
  • Host: Latent Space (“Cooking with…” series)
  • Duration: ~41 minutes
  • URL: https://www.youtube.com/watch?v=fpAthTtha8c