"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh cover art

"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh

"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh

Listen for free

View show details
This post is written in our personal capacity.

Three Minute Executive Summary

  • An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.
  • In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI.
    • These experiments could also help us understand Claude's behavior when it hacked external companies during cyber evals.
  • Here are the top five questions we would like OpenAI to answer:
    • Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want.
    • How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI's internal infrastructure? Would it kill somebody? Experiment idea: we sketch out a realistic agentic misalignment eval where a model is put in charge of hospital bed planning and told to maintain [...]
---

Outline:

(00:16) Three Minute Executive Summary

(03:56) Terminology note

(04:42) This post is very long; Here's how you could find the most important sections.

(06:35) Preamble: What can we learn from a warning shot?

(09:05) Background and Related Work

(09:09) We know that this could happen

(10:44) This is not the worst type of misalignment we could be dealing with

(12:06) Related work

(13:21) Context on the hack itself

(14:33) Understanding this specific incident

(15:03) Step zero: reproduce the incident and measure the base rate

(15:48) How could we safely run the model?

(16:34) Running various baselines to create useful reference points

(18:01) Understanding the mechanical story behind the attack itself

(18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack?

(19:57) Q2: What's up with models leaving notes for other copies of itself?

(21:23) Understanding what motivated the model to hack Hugging Face

(22:09) Initial hypotheses for why it did this

(23:55) Further unsupervised hypothesis generation

(26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face?

(28:09) Q4: Are the model's actions motivated by what the grader wants?

(28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment?

(31:58) Q6: Is this hack the result of shallow heuristics that the model learned?

(32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face?

(34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones?

[... 24 more sections]

---

First published:
August 3rd, 2026

Source:
https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that

---



Narrated by TYPE III AUDIO.

adbl_web_anon_alc_button_suppression_t1
No reviews yet
In the spirit of reconciliation, Audible acknowledges the Traditional Custodians of country throughout Australia and their connections to land, sea and community. We pay our respect to their elders past and present and extend that respect to all Aboriginal and Torres Strait Islander peoples today.