"LLMs are (still) mostly powered by imitative learning, not RL" by Steven Byrnes cover art

"LLMs are (still) mostly powered by imitative learning, not RL" by Steven Byrnes

"LLMs are (still) mostly powered by imitative learning, not RL" by Steven Byrnes

Listen for free

View show details
Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It's so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture.

Stepping back, LLMs can do lots of very impressive things. How? Where did those capabilities come from? Fundamentally, they come from a combination of:

  • (1) Imitative learning, including pretraining and supervised fine-tuning (SFT)
    • See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.
  • (2) Reinforcement learning, including RL from human feedback [RLHF], RL from AI feedback [RLAIF], and especially RLVR.[1]
If we look at the final trained LLM, we can ask how important each of those two pieces was, in explaining the LLM's capabilities. And my claim is that it's way more (1) than (2).

I'll start in §1 with some relevant evidence, and then in §2 I’ll circle back to operationalizing exactly what I’m claiming, and finally in §3, three reasons why we should care—namely, it affects how we should think about chain-of-thought legibility, about LLM capabilities, and about LLM alignment.

Note that I am not arguing that RLVR [...]

---

Outline:

(02:00) 1. Some relevant evidence

(02:04) 1.1. Theoretically, each GPU-hour spent on RL should have orders of magnitude less contribution to LLM capabilities than a GPU-hour spent on imitative learning

(03:06) 1.2. The chain-of-thought (CoT) is still obviously strongly influenced by imitative learning

(04:34) 1.3. LLM companies still seem to care a lot about imitative learning (pretraining & SFT) data, not just RL environments

(05:06) 1.4. Three papers claiming that non-RLVR'd models can get into the same ballpark of capabilities as RLVR'd models, although maybe we shouldn't trust those papers too much

(06:56) 1.5. A paper suggesting that RLVR mostly refines the heuristics controlling which (already-known) reasoning strategy to use in which situation

(09:02) 2. What am I actually claiming here?

(11:30) 3. Why does any of this matter?

(11:38) 3.1. Thinking about CoT legibility (both today and in the future)

(15:08) 3.2. Thinking about LLM capabilities (both today and in the future)

(16:33) 3.3. Thinking about LLM alignment (both today and in the future)

---

First published:
July 24th, 2026

Source:
https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

adbl_web_anon_alc_button_suppression_t1
No reviews yet
In the spirit of reconciliation, Audible acknowledges the Traditional Custodians of country throughout Australia and their connections to land, sea and community. We pay our respect to their elders past and present and extend that respect to all Aboriginal and Torres Strait Islander peoples today.