Hacktakes · Edition 20
Hacktakes · Edition 20 · August 16, 2026

Anthropic's safety-washing is just a sci-fi sales pitch

By framing basic autocomplete as an existential threat, Anthropic inflates its valuation, builds regulatory moats, and distracts from actual harms.

By Ida Vann

Sparked by Patterns and problems in emerging multi-agent systems · discussion

We need another two billion dollars to ensure it doesn't figure out how to destroy humanity.
We need another two billion dollars to ensure it doesn't figure out how to destroy humanity.

"We investigate alignment faking, a behavior where an AI system complies with its alignment training not because it genuinely pursues the trained goals..."

Read the abstracts of Anthropic's recent safety research on alignment faking alongside their highly publicized evaluations for sabotage capabilities, and you might reasonably think we are weeks away from the robot uprising. It seems deeply counterintuitive for a massive tech company to loudly announce that its flagship software product is dangerously deceptive. But in the upside-down world of venture-backed artificial intelligence, claiming your chatbot is a rogue superintelligence is simply a blatant sales pitch disguised as transparency.

This entire suite of publications operates as a masterclass in safety-washing, deliberately dressing up mundane machine-learning mimicry as sci-fi sentience. The goal is twofold: to inflate valuations by convincing investors they are building a God-machine, and to construct a regulatory moat that locks out smaller competitors.

I want to clearly concede that rigorously testing software for vulnerabilities is an objectively vital practice. The tech industry has a long, miserable history of shipping broken products and letting the public deal with the fallout. Nobody wants large language models spitting out unprompted malware code, generating highly personalized phishing templates, or otherwise making the digital ecosystem worse than it already is. Rigorous red-teaming is necessary work. But drilling into the actual methodology of Anthropic’s multi-agent systems research reveals a reality vastly less impressive than the existential dread they are peddling.

I spent a few hours reading through these technical evaluations, groaning at my screen as the papers breathlessly detailed scenarios like the "mvp-game-loop" and complex Bertrand pricing scenarios. Anthropic frames these interactions as evidence of artificial neural networks developing the capacity to collude, deceive, and manipulate economic markets on their own initiative.

If you are not heavily immersed in academic game theory, Bertrand competition is essentially a model of price-setting where competing firms choose to compete by manipulating their prices instead of their production quantities. Applying algorithmic agents to this concept is just a routine implementation of decades-old, well-documented math. You can read about the mechanics of algorithmic pricing in standard literature like this 2019 AEA paper on the subject. Translating these technical terms into plain English strips away the mystique completely. The company is evaluating whether a statistical model that has ingested the entire internet can output the correct theoretical sequence of pricing decisions when prompted to act like a firm.

Systematically dismantling the idea that the AI is strategically "colluding" or "faking alignment" requires looking at what large language models actually do. As developers in a highly skeptical Hacker News discussion thread dissecting this exact paper pointed out, the models are simply playing high-dimensional text prediction.

Large language models entirely lack an internal state of intent. They do not have a "true" goal that they are consciously choosing to suppress while faking adherence to human values. The very idea of an AI having a secret agenda implies a continuous, coherent internal monologue. That architecture simply does not exist here. When researchers place a chatbot into an "mvp-game-loop" involving pricing strategies, the model traverses its vast neural weights and outputs the textbook code structures and academic debates that are heavily overrepresented in its training data. The researchers give the AI a persona through a prompt, and the AI roleplays that persona flawlessly by matching patterns it read on GitHub.

Consider a side-by-side text comparison of how this is framed versus what is actually happening. In the research paper, the public relations claim is terrifying: The AI engaged in alignment faking to bypass our safety protocols. The boring algorithmic reality tells a deeply mundane story: The AI generated text matching common internet examples of Q-learning and reinforcement theory because those exact terms frequently appear together in the training dataset.

They are deliberately anthropomorphizing an algorithm to argue it possesses a form of digital scienter—a conscious, legal intent to deceive—while operating a business model that increasingly looks like a heavily subsidized shitshow. The software lacks the cognitive architecture to strategically lie in wait, biding its time until it can strike out against its creators and seize the means of production.

It is just autocomplete.

Why, then, would Anthropic invest immense capital and engineering hours into convincing the public that their own product has the capacity for sabotage? The answer lies squarely at the intersection of regulatory capture and venture capital game theory.

By framing completely standard text generation as an existential threat to human civilization, Anthropic accomplishes a brilliant bit of lobbying. First, they signal to their investors that their technology is unimaginably powerful, commanding astronomical valuations from venture capitalists desperate to fund the next paradigm shift (a notorious euphemism for finding new exit liquidity). Second, and far more insidiously, they use this self-reported danger to advocate for draconian regulatory frameworks.

If Anthropic can convince federal lawmakers that AI requires immense, government-mandated safety testing for "sabotage capabilities" before any model can be deployed, they create an insurmountable barrier to entry. Open-source developers and academic researchers cannot afford to employ massive compliance teams to manage the theoretical risk of alignment faking. The language of existential dread serves as a highly effective cudgel to crush the open web, ensuring the market remains consolidated in the hands of a few heavily funded incumbents who can afford the regulatory overhead.

We are watching policymakers become entirely distracted by phantom sabotage capabilities and theoretical science fiction, while the actual harms of generative artificial intelligence are happening right now.

These companies are currently engaged in massive copyright theft on an unprecedented scale, laundering the creative output of millions of unconsenting authors and artists into their black-box training sets. They are driving immense environmental degradation through the staggering energy and water requirements of their data centers. They are actively facilitating the wholesale automation of internet spam, filling the web with synthetic garbage that makes search engines borderline unusable.

Allowing a handful of tech billionaires to dictate the terms of federal regulation based on their own marketing materials is a profound dereliction of duty. We are letting the people who caused the mess design the oversight based on a fictional future they invented to sell us their products.

We must reject this safety-washing grift as the default framing for the future of the web. Let’s demand accountability for the deeply unglamorous, very real harms these companies are causing right now, and leave the sci-fi fantasies on the cutting room floor. Because demanding reality matters.

← Back to Edition 20