Glasswing


Welcome to the future.

Note to my future self.

Anthropic was planning on releasing its latest version of Claude but instead stopped short. In testing the new model they call Mythos, they found that the model had such powerful cybersecurity capabilities that it would be literal chaos for them to release it into the market.

In just two weeks of testing, the model found thousands of previously unknown zero-day security vulnerabilities in every existing browser and operating systems — and in the operating systems of our nation’s key infrastructure, like power and water plants. Putting those capabilities out into the public could have compromised and shut down the key systems on which our daily lives run and function — leaving us vulnerable both to machine mistakes and to bad actors using the tool.

Instead Anthropic has created what they call Glasswing, which is a limited release that will allow a handful of experts to use a preview of Mythos to patch those holes.

It is to Anthropic’s credit that they made this move but it is a clear warning that we are entering a new phase of the AI enabled world — one where we are one model release away from serious problems. OpenAI is reportedly working on a new high powered model too, and one hopes they will proceed with similar caution.

We are entering a period where we cannot rely on good actors making careful choices at every turn. The capabilities themselves are moving beyond that assumption.

Welcome to the future.