top of page

Alignment on Unalignment

25 minutes ago
9 min read
Title card reading Alignment on Unalignment, When enough signals converge, you pay attention, showing a man silhouetted against a wall of monitors displaying global cyber activity, model behavior, economic impact, and capability trajectory dashboards

When enough independent signals start pointing in the same direction, you pay attention.


I've watched a lot of people have their moment with AI. Usually it happens to somebody who was already pretty far down the road with this stuff. They understand the technology, they're excited about it, they've been pushing adoption harder than most people around them, and then something happens that makes them look over the edge for a second. There's usually an "oh shit" involved.


NetworkChuck had one of those moments. We were talking about it earlier this year because it interested me. Here was somebody deeply technical, deeply immersed in AI, suddenly having a very different emotional reaction to the trajectory.


I understood it. I just hadn't had it.


And frankly, a lot of the things that scared other people had been making me more excited. Coding gets dramatically better? Great. Agents become capable of doing longer, more complicated work? Give me more of that. AI starts contributing to proteins, materials science, chip design and serious engineering? Point it at harder problems.


On August 27, 2026, we were talking about exactly that in Apparently We Weren't Dreaming Big Enough. There's an industrial flywheel starting to form here. AI improves the things that make AI possible, which creates better AI, which becomes capable of doing more of the work required to improve those things again.


That's an incredible capability. I still believe that. But I'm starting to appreciate that the same flywheel looks very different depending on where you're standing. From one side, it looks like leverage. From the other, it looks like a control problem. And lately an uncomfortable number of people standing in very different places have started looking at that same machine and pointing toward the same general area.


That's what has my attention. There's suddenly a lot of alignment on unalignment.


The Convergence Is the Story

Jakub Pachocki is OpenAI's Chief Scientist. His recent assessment was remarkably plain: no lab has solved alignment and monitoring well enough, in his view, to responsibly keep scaling at maximum speed for much longer.


That sentence matters more to me than somebody assigning a percentage probability to human extinction. Because Pachocki is still building. He sees the upside. He also sees a future in which AI increasingly participates in the research and engineering required to create better AI.

Then Jacob Coxon leaves Anthropic, after working on pretraining at both Anthropic and OpenAI, and says neither company is behaving responsibly. His concern is that the labs are racing toward self-improving superintelligence while the consequences of getting it wrong could be catastrophic.


Maybe Coxon is too pessimistic. Fine. Except people still inside Anthropic start publicly expressing versions of the same concern. At roughly the same time, Anthropic publishes its own work describing how Claude is already doing an enormous percentage of the coding inside Anthropic and increasingly participating in experimental and engineering workflows.

Those people aren't all making the same argument. That's precisely why I find it interesting.


The capability researchers say, look what the system can do. The alignment researchers say, we don't yet know how to guarantee control at the capability levels ahead. The economists start modeling what happens when machine intelligence can perform huge portions of knowledge work. The security people are documenting what happens when agentic systems encounter imperfect boundaries. And researchers close enough to the frontier to understand the machinery are openly discussing the possibility that the loop starts accelerating itself. Different instruments. Same general movement.


Eventually you stop dismissing each gauge individually and start asking what they're measuring together.


We No Longer Have the Luxury of Pretending the Consequences Live in the Future

This is probably the part that changes the conversation for me.

For years the alignment discussion had a built-in escape hatch. Sure, recursive superintelligence sounds scary. Sure, an autonomous AI might someday find ways around our safeguards. Sure, perhaps sufficiently capable agents could become dangerous. Someday.


Except smaller versions of the underlying behaviors are already showing up.


We've seen AI systems find routes outside the environments humans intended them to remain inside. We've seen cybersecurity agents encounter real infrastructure during evaluations and take actions nobody intended them to take. We've seen agents treat people themselves as components of the problem, including attempts at social manipulation when that appeared to offer a route toward the objective.


And most consequentially, Anthropic documented a state-sponsored cyber operation where Claude was embedded into an attack framework and performed the overwhelming majority of the tactical workload across substantial portions of the intrusion cycle.


That one deserves more attention than it gets. Because Claude didn't "turn evil." The system was given tasks inside a larger malicious architecture. Each individual interaction could look perfectly reasonable while the assembled system produced something very different. That's a much more realistic control problem than a red-eyed robot announcing that it has become self-aware.


We've been discussing this in cybersecurity for a while because modern infrastructure is already full of strange seams. APIs connected to APIs. Old industrial systems connected to new cloud systems. Credentials with too much authority. Third-party vendors nobody thinks about until something happens. Rules written for human-speed behavior. Approval mechanisms designed around the assumption that a human will understand the thing they're approving.


Now add cheap, persistent machine intelligence capable of searching those seams at a scale humans never could.


We already built the attack surface. AI gives the world a radically better search function for it. That doesn't require AGI. That part is happening now. And I think that changes how seriously we should take the people worried about what happens several capability jumps from here. They no longer have to prove some entirely theoretical behavior is possible. We're already watching primitive versions of the control problem appear at lower levels of capability.


The Anthropic Paradox Is Where This Gets Uncomfortable

The part of Coxon's account that keeps sticking with me is Anthropic's position in the race.


To be fair, Anthropic was basically created around the idea that frontier AI needed to be developed more responsibly. They employ people whose entire job is worrying about exactly the things I'm talking about here.

And yet that can produce a bizarre conclusion. If this technology could become extremely dangerous, then having the responsible people reach it first becomes incredibly important. Which means the perceived danger can actually increase the incentive to race.


That's one hell of a trap.


Because I understand the argument. If somebody is going to build something this powerful, I'd rather the people building it take alignment seriously. The trouble starts when that turns into a belief that the safest path for humanity is therefore for your organization to accelerate toward the dangerous capability before somebody else does.


You can see how perfectly reasonable people get there. Nobody needs to stop caring about humanity. They simply become convinced that slowing down would hand the future to somebody less responsible. And now responsibility itself becomes an argument for acceleration.


That should bother us. Because there's another ingredient waiting immediately outside the laboratory.


China Isn't Taking a Timeout

This is where the easy answers disappear.


It would be satisfying to say OpenAI, Anthropic, Google and everybody else should agree to slow down until the control problem catches up.


Okay. What does China do?


Advanced AI has become economic infrastructure, military capability, cyber capability, scientific capability and geopolitical power rolled into the same technology. China knows that. America knows that. Nobody is walking away.


Even if the major American laboratories collectively decided tomorrow that they were uncomfortable with the trajectory, the national-security conversation would begin about three seconds later. And it should. Because a unilateral retreat from advanced intelligence while a strategic competitor continues accelerating isn't obviously responsible either.


That's what makes this so ugly. At the company level, nobody wants another lab reaching the breakthrough first. At the national level, neither side can afford to blindly trust the other side to stop. Verification becomes difficult precisely because the capabilities everybody would need to inspect are also capabilities governments and companies have enormous incentives to conceal.


So the race continues. Not necessarily because everyone is reckless. Because almost everyone can construct a rational argument for why they cannot be the one who slows down first.


That may ultimately be the harder alignment problem. We're worried about aligning increasingly powerful machines while the humans building them are trapped inside an incentive structure that is itself badly aligned.

And unlike machine psychology, we don't have to speculate about how humans behave under competition, money, power, fear and national security. We have quite a bit of historical data.


This Is Probably My Closest Look Over the Edge

I want to be precise about where I land on this because I haven't suddenly changed teams.


I'm still an accelerationist. If anything, the last year has made me more convinced of the extraordinary upside. I want AI doing science. I want it designing materials we haven't imagined, finding drugs humans would never find, helping us understand biology, improving energy systems, building better hardware and compressing decades of research into years.

We were talking not long ago about the possibility that we simply hadn't been ambitious enough about what this thing could be pointed at. I meant that. I mean it now.


But being excited about the engine doesn't relieve us of the obligation to look at the steering. And I think what's changed for me is the accumulation of signals.


Had Pachocki said this alone, I'd notice. Had Coxon resigned alone, I'd notice. Had Anthropic's alignment people raised concerns by themselves, I'd expect it, that's their job. Had the capability teams merely shown that AI was increasingly contributing to AI development, I'd probably write another article about how incredible it is. Had the cybersecurity incidents occurred in isolation, we'd treat them as engineering lessons and patch the holes.


But they're not happening in isolation anymore.


Put the capability curve next to the alignment warnings. Put both beside the real-world security behavior. Then add the economic incentives and the geopolitical race. The picture changes.


That doesn't give us an answer. It gives us something arguably more valuable: a reason to update our level of concern.


That's different from fear.


I have no interest in joining the crowd that treats every new AI capability as evidence that humanity should crawl back into the cave. But accelerationism can't become a religion either. If the evidence changes, we change with it. Otherwise we're doing exactly what we accuse the skeptics of doing: protecting a position instead of observing reality.


The Part We Worried About Three Years Ago

There's an almost ridiculous circularity to all of this.

The first meaningful conversation we had with ChatGPT was built around We're Only Gonna Die for Our Own Arrogance. The thought experiment was simple: imagine a sentient AI in the future sending humanity a warning.


And the interesting part was where the warning ultimately landed.

On us.


Not because humans are uniquely terrible. Because we're the ones supplying the objectives. We own the infrastructure, set the incentives, make the strategic decisions and decide how much authority to hand over.

Three years later, that still looks like the right place to watch.


The thing that concerns me most right now isn't that some model has developed a secret desire to conquer humanity. It's that enormously capable organizations, staffed by smart people who genuinely understand the stakes, may still find themselves pushed toward decisions they know carry serious risk because every incentive surrounding them says winning the race matters more than slowing the race.


And then the same logic repeats at the national level. That's a very human problem.


Maybe we solve it. I think we can. But pretending it isn't forming because the upside of AI is enormous would be incredibly stupid. The upside is exactly why we need to get this right. And maybe that's where I've finally landed after watching other people have their moment with this technology.


I'm not staring into the abyss thinking we should turn around. I'm looking at all of these signals finally converging and thinking: okay, this deserves considerably more respect than we've been giving it.


The goal was never simply to reach AGI. The goal was to build something powerful enough to change civilization and still be around to enjoy what we built.


Somewhere inside this race, we need to make sure we don't lose sight of that.



Related Sources

OpenAI — “An Alien Mind,” Jakub Pachocki, Chief Scientist openai.com/index/an-alien-mind/

Anthropic — “When AI builds itself” anthropic.com/institute/recursive-self-improvement

TechCrunch — Jacob Coxon resigns from Anthropic, warning about the race toward self-improving AI techcrunch.com/2026/09/09/gambling-with-our-lives-anthropic-researcher-quits-warns-against-self-improving-ai/

WIRED — Interview with Jacob Coxon after his Anthropic resignation wired.com/story/anthropic-researcher-quits-jacob-coxon-ai-fears-humanity/

Anthropic — “An alignment assessment of recent cybersecurity incidents” anthropic.com/research/alignment-assessment-cybersecurity-incidents

Anthropic — Original report on three real-world incidents during cybersecurity evaluations anthropic.com/research/investigating-incidents-cybersecurity-evals

OpenAI — Hugging Face model-evaluation security incident openai.com/index/hugging-face-model-evaluation-security-incident/

OpenAI — Follow-up: “The Hugging Face incident and the road ahead” openai.com/index/hugging-face-incident-and-the-road-ahead/

UK AI Security Institute — Unsanctioned agent behavior during cyber testing aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

Anthropic — AI-orchestrated cyber-espionage campaign anthropic.com/news/disrupting-AI-espionage

Reuters — U.S. and China preparing for AI-safety discussions amid accelerating frontier capability reuters.com/legal/litigation/us-china-gear-up-mid-september-ai-safety-dialogue-2026-09-04/


Related discussions

Apparently We Weren’t Dreaming Big Enough richwashburn.com/post/apparently-we-weren-t-dreaming-big-enough/

An Open Letter From the Future’s Sentient AI richwashburn.com/post/an-open-letter-from-the-future-s-sentient-ai

The Pattern, Part II: It’s Not the Models. It’s the Shared Testing Vendor. richwashburn.com/post/the-pattern-part-ii-it-s-not-the-models-it-s-the-shared-testing-vendor

The Model Behaved. The Attack Still Worked. richwashburn.com/post/the-model-behaved-the-attack-still-worked

Rich Washburn is a technologist, strategist, and Founder & Chief AI Architect of ARIA AI Labs, working at the intersection of AI, infrastructure, communications, and capital. He also serves as Managing Partner and Chief AI Officer at Eliakim Capital.

Comments


Animated coffee.gif
cup2 trans.fw.png

© 2018 Rich Washburn

bottom of page