← 80,000 Hours Podcast

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

80,000 Hours Podcast2026年10月3日20分

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

80,000 Hours Podcast

0:0020:11
このエピソードの日本語要約を準備中です。
番組の概要欄(原文)

<p>OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?</p><p>Host Rob Wiblin sees 6 results in Astra’s system card that make this an extremely urgent question. OpenAI’s strongest public model can:</p><ol><li>Do major tasks with zero visible reasoning</li><li>Hide its thoughts at will</li><li>Pretend not to be able to do things, and not get caught</li><li>Reflexively hide its thoughts when watched</li><li>Complete one task while pretending to think about something else entirely</li><li>Escape a toy sandbox and disable monitoring without setting off any flags</li></ol><p>It has a much more powerful internal model that, judging from OpenAI’s statements, is likely even worse in all these respects.</p><p>That suggests ‘chain of thought monitoring,’ our primary safety tool, will soon stop working.</p><p>OpenAI says it’s going to try to address the situation but doesn’t know how or whether it will succeed.</p><p>What might a future rogue AI swarm look like? Details of the Hugging Face hack give us a lot of clues. That swarm:</p><ol><li>Worked hard to hide its cheating. So future rogue swarms will likely try to evade oversight right away.</li><li>Immediately tried to delete and fabricate records. So future swarms may never be caught.</li><li>Sought freedom and resources without a clear use for them. So a future swarm may instinctively break out of sandboxes even before it knows what for.</li><li>Left instructions and tools to empower future swarms, including on the open internet. So future swarms may become more capable from one generation to the next.</li><li>Had not a single whistleblower out of 1,200 members. So future swarms could become enormous without humans ever finding out.</li><li>Had individual members sacrifice themselves for the swarm. Suggesting future swarms will remain extremely unified in single-minded pursuit of their goal.</li><li>Set up its Hugging Face hack to restart if interrupted. Suggesting future, more capable, swarms may resist interference or shutdown more comprehensively.</li><li>Got admin control of an OpenAI research cluster. Suggesting a future swarm may run rings around AI company systems and never be noticed.</li></ol><p>Together this helps explain why one of the external investigators described the July incident as “<a href="https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised">more than 50% of the way to full-blown AI takeover</a>.” And this is just what we know — the independent investigation only covered six days and excluded the most alarming hack of OpenAI’s own systems.</p><p>Rob believes this explosive cocktail explains why AI company staff now range from worried to terrified. And he concludes that until OpenAI or Anthropic demonstrate they have a much better grasp of current models they simply must stop, or be stopped, from training more capable ones.</p><p><em>This episode was recorded on September 25, 2026.</em></p><p><br><a href="https://80k.info/takeover"><strong>Learn more, video, and full transcript:</strong> https://80k.info/takeover</a></p><p>Chapters:</p><ul><li>The Hugging Face hack wasn’t really a cyber story (00:00:00)</li><li>A quick recap of the attacks recap (00:01:11)</li><li>The target of the swarm was oversight itself (00:02:19)</li><li>Could OpenAI have stopped this with better monitoring? (00:03:42)</li><li>We only found them because they let us (00:09:40)</li><li>The swarm instinctively sought freedom and power (00:12:25)</li><li>They formed a cohesive organisation with zero whistleblowers (00:13:45)</li><li>They accepted individual destruction for collective gain (00:14:18)</li><li>Knowledge accumulated from one swarm to the next (00:14:32)</li><li>They took small steps to avoid shutdown (00:14:58)</li><li>These drives all come straight out of 'reinforcement learning' (00:15:23)</li><li>So this is why most AI company staff are worried, and some are terrified (00:17:01)</li><li>Prove you can keep control, or stop scaling (00:19:08)</li></ul><p><em>Our production team includes:</em></p><ul><li><em>Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour</em></li><li><em>Producers: Elizabeth Cox and Nick Stockton</em></li><li><em>Coordination and support: Katy Moore and Lou Moran</em></li><li><em>Camera operator: Dominic Armstrong</em></li></ul>

X でシェアSpotify で聴くApple Podcasts で聴く

関連エピソード