"“Alignment Engineering” vs. “Misalignment Science”" by Edward James Young
Failed to add items
Sorry, we are unable to add the item because your shopping cart is already at capacity.
Add to basket failed.
Please try again later
Add to Wish List failed.
Please try again later
Remove from Wish List failed.
Please try again later
Follow podcast failed
Unfollow podcast failed
-
Narrated by:
-
By:
- Prosaic alignment of models is becoming a bottleneck for capabilities.
- Therefore improving the prosaic alignment of models enables faster capabilities advances, which bring us closer to RSI.
- It is unlikely these prosaic alignment methods remain sufficient during the RSI loop, and so this work brings us closer to doom.
- Furthermore, dealing with these more prosaic failures reduces the likelihood of a warning shot of sufficient magnitude to cause a slowdown which would prevent RSI.
In this post, I outline a contrast between “Alignment Engineering” – the dominant model for what “working on alignment” looks like (inside labs, and in the field as a whole) with “Misalignment Science”. I begin by characterising [...]
---
Outline:
(01:51) "Alignment Engineering"
(06:22) AI Safety and the ML tradition
(07:42) Implicit work trials
(09:36) "Misalignment Science"
(14:56) Conclusion
(15:39) Postscript: Iterating ourselves into oblivion
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
October 5th, 2026
Source:
https://www.lesswrong.com/posts/FogmcDHA6AdMGukum/alignment-engineering-vs-misalignment-science
---
Narrated by TYPE III AUDIO.
adbl_web_anon_alc_button_suppression_t1
No reviews yet
In the spirit of reconciliation, Audible acknowledges the Traditional Custodians of country throughout Australia and their connections to land, sea and community. We pay our respect to their elders past and present and extend that respect to all Aboriginal and Torres Strait Islander peoples today.