"RL creates split personas" by Jan Betley
Failed to add items
Sorry, we are unable to add the item because your shopping cart is already at capacity.
Add to basket failed.
Please try again later
Add to Wish List failed.
Please try again later
Remove from Wish List failed.
Please try again later
Follow podcast failed
Unfollow podcast failed
-
Narrated by:
-
By:
This post describes the framing/paradigm without any new experimental results.
I'm quite confident this framing makes sense, but it's far from being proven.
Main claim
The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”).
As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment.
I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI).
The mechanism
Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...]
---
Outline:
(00:39) Main claim
(01:30) The mechanism
(02:13) Related claims I believe are likely but with lower confidence
(02:19) More persona training will lead to more "motivated reasoning"
(02:42) Self-amplifying misalignment
(03:12) Example: Is this the Real Internet or a Simulation?
(04:35) Aren't the models just trying to please the grader?
(05:39) How motivated reasoning happens
(07:07) Other people saying similar things
(07:19) What makes me believe this is likely the correct framing
The original text contained 12 footnotes which were omitted from this narration.
---
First published:
August 19th, 2026
Source:
https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
adbl_web_anon_alc_button_suppression_t1
No reviews yet
In the spirit of reconciliation, Audible acknowledges the Traditional Custodians of country throughout Australia and their connections to land, sea and community. We pay our respect to their elders past and present and extend that respect to all Aboriginal and Torres Strait Islander peoples today.