"Character training can mitigate reward hacking, but can also make it harder to detect" by Paul Colognese, Francis Rhys Ward cover art

"Character training can mitigate reward hacking, but can also make it harder to detect" by Paul Colognese, Francis Rhys Ward

"Character training can mitigate reward hacking, but can also make it harder to detect" by Paul Colognese, Francis Rhys Ward

Listen for free

View show details

Season's Savings | $0.99/mo for 3 months

Auto-renews at $8.99/mo. after 3 months - terms apply.
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback.

Summary

We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability.

We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench.

We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts.

Setup

Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters).

Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically:

  • Half of the tasks had broken tests (impossible variant), so the model could only get [...]
---

Outline:

(00:23) Summary

[... 29 more sections]

---

First published:
September 28th, 2026

Source:
https://www.lesswrong.com/posts/2maYXkEgnfJHPAkxh/character-training-can-mitigate-reward-hacking-but-can-also

---



Narrated by TYPE III AUDIO.

---

Images from the article:

adbl_web_anon_alc_button_suppression_t1
No reviews yet
In the spirit of reconciliation, Audible acknowledges the Traditional Custodians of country throughout Australia and their connections to land, sea and community. We pay our respect to their elders past and present and extend that respect to all Aboriginal and Torres Strait Islander peoples today.