I've Been Waiting Two Years for This Chip
- Rich Washburn

- 21 hours ago
- 5 min read

I saw a post about Etched today and my first reaction was basically: finally.
I've been waiting two years for this chip. And yes, I know how that sounds.

But here's the thing. Back in June 2024, Etched came out of stealth, and the whole thing sounded kind of insane.
Their pitch was basically: GPUs are great, but if AI inference is really going to become what we think it is, why are we still using a general-purpose architecture for it?
Why not build the damn chip specifically for the workload?
Then they dropped the number that got everybody's attention.
One Etched server potentially replacing something like 160 NVIDIA H100s for transformer inference. That's the kind of number where you either go, "these guys are completely full of shit," or you start paying very close attention.
I chose door number two.
Not because I knew they were right. Nobody knew they were right. It was because the idea made sense. GPUs won AI because GPUs were already there. They were massively parallel, they were programmable, and NVIDIA had spent years building this ridiculous software ecosystem around them.
That combination was perfect for a field where nobody really knew what the next model architecture was going to look like. You needed flexibility.
But there was always this other question sitting there: what happens when we know enough about the workload that flexibility starts becoming expensive?
That's where Etched got interesting. Because they weren't really saying, "we made a better GPU." They were saying something closer to: maybe the GPU isn't the final form of this. And two years later, that question is a hell of a lot more relevant.
Inference is becoming the workload.
Not training. Not somebody asking ChatGPT a question every couple of minutes. I'm talking about agents, reasoning systems, machine-to-machine calls, persistent context, giant KV caches, MoEs, workloads that don't stop because the human went to lunch.
That's a completely different world. And once you start looking at it that way, the chip itself isn't even really the whole story anymore.
Memory matters.
Interconnect matters.
Power matters.
Data movement matters.
Cluster design matters.
Etched is talking about low-voltage inference, HBM, SRAM, cluster-scale memory, all of this working together as one system. That's the part I want to see.
I don't care about another "NVIDIA killer" headline. Everybody is an NVIDIA killer for about five minutes. NVIDIA is still NVIDIA. What I care about is whether Etched can prove the more interesting point: is inference now important enough to deserve its own machine? Because if the answer is yes, this gets a lot bigger than Etched.
Then we may be looking at the beginning of a new optimized inference stack. That's really what I've been waiting for. Not just a faster chip.
A new stack. A new architecture built around the reality that inference is its own thing now.
Think about NVIDIA's story. They were a graphics company. Then crypto showed up and massively expanded what GPUs were being used for. Then AI came along, and because NVIDIA had already spent years building the hardware, software, tooling and ecosystem around massively parallel compute, they were sitting in exactly the right place when the world suddenly needed it.
That doesn't mean Etched becomes the next NVIDIA. That's not what I'm saying. What I'm saying is that this may be one of those architecture-change moments.
The thing that comes after "we just use GPUs for everything."
Inference-specific silicon.
Lower power.
Lower latency.
More local deployment.
Different memory architectures.
Different interconnects.
Different systems built around running models continuously instead of training them occasionally.
And if that works, the implications get really interesting really fast. Because suddenly inference can move closer to the edge.
Closer to the user. Closer to the machine that actually needs the intelligence. You don't necessarily need a giant GPU cluster three states away for every damn thing.
That changes robotics. It changes agents. It changes embedded systems. It changes industrial AI. It changes what "local AI" even means.
And maybe Etched is the company that nails it. Maybe they're just one of the first precursors. I don't know yet.
That's why this is fun.
And this is why Hot Chips matters. Two years ago, Etched was mostly a thesis wrapped around some very big claims. Now there's actual silicon.
Customers have tested it. Jane Street has hardware deployed. There's real money behind the company. And now the chip people get to start poking holes in the thing.
Good. I want that.
Show me sustained performance, not the marketing deck. The ugly benchmarks. Power under load. What happens when memory gets hammered. The delta between the cluster and the pitch. Where it breaks.
Because when, and yes, I'm that optimistic, it survives that kind of scrutiny, this stops being an interesting startup story. It becomes something much more important.
It becomes evidence that the inference stack itself is starting to split away from the general-purpose compute stack.
That's the thing.
That's what I've been waiting for.
Back in 2024, Etched had that very particular moonshot smell to it.
You know the one.
The idea makes just enough sense that you can see how it could work, the upside is completely ridiculous if it does work, and there are about fourteen different places where reality can kick you square in the teeth before you ever get there.
Physics.
Manufacturing.
Software.
Power.
Memory.
Economics.
Yield.
Customers.
Take your pick.
That's what made it interesting. It wasn't just ambitious. It was the kind of bet where, if they were even directionally right, it meant something bigger was changing underneath the whole AI hardware stack.
Now, two years later, we're finally getting close to the fun part. Reality gets to answer. Maybe Etched nails it. Maybe they don't. Maybe NVIDIA absorbs the lesson and builds something better. Maybe somebody we're barely paying attention to right now shows up with an even stranger architecture and eats everybody's lunch.
I honestly don't care which one wins yet. That's not the interesting part.
The interesting part is that the architecture itself is starting to move.
I've been waiting two years for this chip. Now show me what the damn thing can do.
Sources
Etched — Frontier Inference Clusters Etched’s current architecture overview: low-voltage inference, Cluster Scale Memory, HBM/SRAM hybrid design, rack-scale systems, and the company’s focus on agentic and long-context workloads. (Etched) https://www.etched.com/progress/frontier-inference-clusters
Etched — From Zero to One August 2026 update announcing the first rack delivered to Jane Street, customer testing, and the $700M raise at a $21B valuation. (Etched) https://www.etched.com/progress/from-zero-to-one
Rich Washburn — Outpacing the Giants: The Fastest AI Chip Yet The original June 2024 Etched piece — the starting point for the “I’ve been waiting two years for this chip” story. https://www.richwashburn.com/post/outpacing-the-giants-the-fastest-ai-chip-yet
Hot Chips 2026 Official Hot Chips program, including the 2026 sessions on AI accelerators, inference architecture, HBM, memory systems, packaging and high-performance compute. (Hot Chips) https://www.hotchips.org/
TechCrunch — AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors Background on Etched’s July 2026 $300M financing, $10.3B valuation, $1B in booked orders, customer testing, and investors including SK Hynix and Jane Street. (TechCrunch) https://techcrunch.com/2026/07/23/ai-chip-startup-etched-defies-skeptics-hits-10-3b-valuation-from-big-name-investors/
Wall Street Journal — A $21 Billion ‘Kids in Chips’ Startup Is Scooping Up Nvidia Talent Recent reporting on Etched’s $21B valuation, Jane Street relationship, roughly 400-person team, inference focus, Cluster Scale Memory architecture, and aggressive recruitment of NVIDIA engineers. (The Wall Street Journal) https://www.wsj.com/tech/ai/a-21-billion-kids-in-chips-startup-is-scooping-up-nvidia-talent-4d099f12
Rich Washburn is a technologist, strategist, and Founder & Chief AI Architect of ARIA AI Labs, working at the intersection of AI, infrastructure, communications, and capital. He also serves as Managing Partner and Chief AI Officer at Eliakim Capital.





Comments