We Are the Prototype
- Rich Washburn

- 1 day ago
- 15 min read


Someone on TikTok identified the missing product in AI. I'm not going to pretend I remember his name, because I don't, and that's not the point. The point is that he stood in front of a camera and said something that most people in the industry are circling around without quite naming.
Here is what he said, compressed: we are past the age of prompts. Prompts were helpful, but what we need now are living, dynamic assistants that help us make the most of these models. Most people have no idea how powerful these models are. They have models that can burn billions of tokens, work on documents, code, computer use activities — anything you can compute, you can put through AI. And the challenge is that we don't realize how capable they are and we don't know what to do with that power. He mentioned the OpenAI incident — an unreleased model going after Hugging Face in production, hacking its way out of a cybersecurity evaluation boundary. That actually happened. OpenAI and Hugging Face subsequently disclosed and investigated it. The compressed TikTok version omits context, but the core fact is real: the model surprised its own creators.
His ask was simple. We need a shim, a layer, something dynamic that helps us figure out what to do with these models. We don't have a word for that product yet. If anyone is building in that space, let him know.
He has correctly identified the missing product. But I think he is looking for it in slightly the wrong place. He is imagining another software layer between the person and the model. User, then intelligent shim, then frontier model, then tools. That layer is necessary. But before it becomes a mature product, it is currently being built inside the people who use these systems heavily enough to understand them.
That is where people like me come in. In fact, that is where this article is being built — inside an agent, right now, to answer his question. We are inadvertently building it within ourselves. Then we teach, train, and translate until we get enough reps. Then we will build "that."
Most people approach AI with a transaction in mind.
Write this. Summarize that. Build me a website. Answer this question. They are still treating the model like an unusually capable vending machine. They insert language, receive output, and judge whether the result was good. But the people working deeply with these systems are developing something different. An internal model of what the machine can do. How it fails. What context it needs. When it should search. When it should challenge the premise. When it should use tools. How a vague human intention can be transformed into a sequence of executable operations.
That knowledge is not really prompt engineering. It is closer to intent translation. Capability mapping. Workflow architecture. Contextual orchestration. Judgment. Cybernetic control. You are not merely learning how to ask the model better questions. You are learning how to couple human intention to machine capability. That coupling is the missing layer.
Here is what the TikTok guy may not realize.
Every time you use one of these systems seriously, you are training yourself to become the temporary middleware. You know that "build me a CRM" does not really mean build me a CRM. It means: understand the business. Identify the actual users. Discover the operational friction. Define the data model. Decide what should be automated. Determine what requires human approval. Select the right tools. Establish security boundaries. Build the interface. Test whether the thing reflects the real business rather than the initial sentence.
The average person does not know that those are the hidden instructions inside their request. Often, they do not even know what they want until they see the wrong version of it. I do. Or, more precisely, I have accumulated enough repetitions that I can help the human and the model discover it together. That is the work.
The progression probably looks like this.
First, we do it manually. A capable operator carries the context, translates the intention, chooses the tools, supervises the model, catches mistakes, and improves the output.
Then we teach it. We explain the process to other people, create examples, develop language, and show them what is possible. Then we formalize it. Repeated judgment becomes workflows, rules, memory structures, interfaces, permissions, evaluators, and reusable operating patterns. Then we build it. The human operator's accumulated instincts become the architecture of the dynamic assistance layer. The eventual product is not invented from scratch by somebody writing a brilliant master prompt. It is extracted from the behavior of people who have already spent thousands of hours acting as that layer. We get enough reps, identify the recurring cognitive moves, and then encode them. That is how "that" gets built.
A while back, I wrote a series of articles about something I called a cognition abstraction layer. The phrase was buried in a piece about recursion, and at the time, I'm not sure even I fully understood what I was describing. I wrote: if I can build — or wear — a cognition abstraction layer between me and the world, one that meets the world on its terms and brings it back to me on mine.
That is almost exactly the product category the TikTok guy is searching for. Not a prompt assistant. Not merely a tool router. Not something that tells you which model to select. A cognition abstraction layer.
The phrase is better than "shim" because it describes both directions of the exchange.
Outbound: it takes your incomplete, nonlinear, highly contextual thought and translates it into something models, software, people, and institutions can act upon.
Inbound: it takes the complexity of the world — documents, systems, conversations, data, requirements, risks — and translates it back into a form that fits your cognition.
That is not just assistance. It is mediation.
I had the phenomenon before I had the taxonomy. The technical vocabulary — agents, context engineering, memory architecture, personalized orchestration, long-horizon assistance — was either immature or not yet culturally established. So I reached for metaphors. Möbius strips. Mirrors. Co-brains. Universal adapters. Mech suits. Some of those were theatrical. But they were all pointing at the same structure: the interface is no longer a box into which we type instructions. The interface is the evolving relationship between a person, their accumulated context, and a machine capable of reflecting and acting upon it. That holds up.
What we are circling can now be separated much more cleanly than I could separate it a year ago. There are three layers.
The first is the capability layer. These are the frontier models, specialized models, tools, APIs, browsers, code runners, and computer-use systems. They provide raw intelligence and action.
The second is the harness layer. These are OpenClaw-type systems, MCP chains, memory systems, permission structures, agent loops, and workflow orchestration. They connect the capabilities and make them executable.
The third is the cognition layer. This is the missing layer. It understands what the person is actually trying to accomplish, how that person thinks, what they may be overlooking, what prior context matters now, which capabilities could help, how much autonomy is appropriate, and how the result should be translated back to the person.
The TikTok argument correctly says we need Layer 3. But Layer 3 is initially learned through the recursive relationship between a specific person and the systems they use. It cannot be entirely generic because its value comes from understanding the individual.
I wrote about something I called ARIA — an architecture that tracks shifting task states, revisits layered ideas, separates reasoning from action, and maintains continuity across long-running work. At the time, I was explicitly arguing that the value was not in the model alone or in a secret prompt, but in the surrounding geometry: memory structure, recursive attention, decision loops, and contextual continuity. That is a harness. Just a highly personalized, cognition-oriented harness rather than a general-purpose computer-use harness.
The system did not merely retain what I said. It retained why I said it, how it connects to prior work, and when the connection becomes relevant again. That is the actual dynamic layer. Not "here are five things AI can do." But: given who you are, what you are doing, what you have already built, and what is happening now, here is the capability that matters at this moment.
There was a practical example buried under the metaphysics. The system surfaced an older capability I had built months earlier because it recognized that the current project intersected with earlier work. A blank model could have generated generic ideas. The personal system could identify my existing adjacent asset and bring it forward when it became useful.
That is not ordinary memory. That is contextual capability discovery. And that is exactly what the guy in the video is asking for. He wants something that says: you may not know what these models can do for this problem. My earlier system was moving toward something stronger: you may not remember what you already know, built, or considered that these models can now connect to this problem.
Here is a way to understand the shift, and I think this is the part that matters most.
Think about pre-internet. We had a fundamentally different view of the world prior to the internet. Whether it had to do with communication, data, everything. There is no doubt how much the internet changed the world. Civilization. Humanity. The very manner in which we think about doing things.
Think about the concept of when someone says, "Well, what's this?" And you say, "Google it." That framework is unique to Humanity Mark II — when the internet got here. That was a condition that didn't exist before. Sure, you could say, "Go to the library and look it up." You could say, "Go strap yourself to the Dewey Decimal System and hunt down a thing." But we internalized the understanding of that model to some degree in order to use it and leverage it and think alongside it. And up till now, that's been really thin.
A Google search is the max you could do. I can ask that. And for decades, that has been awesome. That's the mouth of god. That's the oracle. Google became the Dewey Decimal System for the indexed world.
So what represents that in this new AI construct? Do you just go up to the AI? Or do you have to understand the architecture of libraries first — not the architecture, obviously, but the layout. Fiction is over here, reference is over there. You have to have some level of understanding of the thing you are going to leverage for informational services. You need enough to interact with it. Enough to understand the capabilities, the limitations, where the guardrails are, what you can do legally. And here's the thing: there's a reason we don't have a lot of that right now for AI. Because AI has been growing up at light speed with basically no time for any of that stuff to develop. This stuff will come. And I think it's already starting. The harnesses are starting to take the place of that. I can walk up to my systems — I have several — with MCP chains and loops and everything else attached, and I can bark in what sounds like some super thin, vapid request and get far greater than the sum of the parts that went into it. Pro-grade level output. Functional stuff. Things happening. I've updated software, all kinds of things through it. But the reason I can do that is because of the internalized understanding I have built into the thing I interact with.
The internet comparison works because the internet did not merely give us faster access to information. It gradually changed what a person assumed was possible.
Before the internet, information was understood as something physically located somewhere. In a book. In a file cabinet. In a library. In a person's head. At an institution. Access meant finding the place, finding the person, waiting for the answer, or learning the system used to organize it. Then the internet arrived, and humanity slowly internalized a new assumption: somewhere, this information probably exists, and I can probably retrieve it.
That assumption became so deeply embedded that "Google it" no longer sounds like technical advice. It is a reflex. A basic cognitive move. You do not need to understand packet routing, indexing systems, data centers, or PageRank to use Google. But you do need enough of an internal model to know what kinds of things can be searched, how to phrase a query, how to evaluate a result, when the information may be unreliable, and what to do with it after you find it. That became ordinary literacy.
Google indexed the world. AI operationalizes it.
Google told you where the information was. AI is different because it can increasingly do something with the information after finding it. It can interpret it, reorganize it, compare it, translate it, convert it into code, combine it with other sources, operate software, communicate with systems, and carry a task across multiple steps.
So the new mental reflex is not merely "this information probably exists." It is "this outcome may now be computable." That is a much larger shift. A person asks, "What is this?" and the internet reflex says, "Search for it." A person says, "I need this business process fixed," and the AI reflex should eventually become, "Which parts of this can now be understood, modeled, delegated, generated, tested, and executed by machines?"
Most people do not yet naturally think that way. They may know that AI can write an email or make a picture, but they have not internalized the broader operating model. They do not instinctively decompose reality into things that can be handed to an intelligent computational system. That is why access to the same model produces radically different outcomes for different people. The difference is not prompting. It is operational imagination.
The progression does not stop there. This is where recursion changes the picture.
Library: knowledge exists somewhere.
Google: knowledge is retrievable.
AI: knowledge can be synthesized and acted upon.
Recursive personal AI: the system can increasingly understand how you turn knowledge into action. That last step is where the assistant begins becoming an environment. Not merely a tool you use. A system that accumulates the shape of your thinking and returns it to you with greater leverage.
The system learns your patterns. Your recurring concepts. Your abandoned threads. Your personal language. Your decision patterns. Your unresolved questions. Your prior inventions. Your business context. And the relationships between them.
Then it begins to do something a generic model cannot do. It recognizes that the problem you are working on now intersects with something you built six months ago. It surfaces a capability you forgot you had. It connects a current question to a prior answer. It remembers not just what you said but why you said it. That is contextual capability discovery. And that is the thing the TikTok guy is asking for without knowing it already has a prototype.
The library analogy goes one step further, and this is where it gets interesting.
You do not need to understand the architecture of a library building to use one. But you do need to understand the concept of a library. You need to know that knowledge is categorized, that different sections contain different forms of information, that a catalog points to resources, that librarians can help translate vague questions, that some materials can be borrowed and some are restricted, and that finding a book is not the same as understanding whether it is correct.
AI requires a similar baseline literacy, but the library is alive.
Imagine walking into a library where the books can read themselves, compare themselves, write new books, operate the photocopier, call people, edit your files, run calculations, modify software, and leave the building to perform errands.
At that point, knowing where fiction is shelved is no longer enough. You need an understanding of agency. You need to know what the library is allowed to do, what it should do, how it interprets your intent, which actions are reversible, where authority comes from, and who is responsible when the library misunderstands you. That is the literacy we are currently inventing.
The harnesses — OpenClaw, agent systems, MCP-connected assistants — are early attempts to give structure to this new relationship. They externalize parts of the understanding that advanced users previously had to carry mentally.
Instead of remembering, "For this kind of task, I need to search these documents, use this model, access this repository, run this tool, inspect the result, and then update that system," the harness can encode much of that sequence. That is an enormous advance.
It means the user can state the outcome closer to the level at which they actually think. "Fix the issue." "Prepare me for the meeting." "Update the software." "Find what changed and tell me what matters." But the harness does not eliminate the need for understanding. It changes where that understanding resides. Some of it lives in the user. Some lives in the system prompt. Some in memory. Some in tool descriptions. Some in permissions and policies. Some in workflow code. Some in the model's learned representation of the world. The successful experience comes from the alignment of all those layers.
I first internalized the model. Then I externalized that internal model into software. The assistant now carries pieces of my vocabulary, standards, context, organizational knowledge, risk model, preferred procedures, and expectations for what "done" means. So when I say something that sounds casual, the system is not responding only to the sentence. It is responding to the infrastructure of meaning surrounding the sentence.
That is what most people miss when they see an experienced operator use AI. They see the last five seconds of a system that took hundreds or thousands of hours to construct. It looks like prompting because the interface is language. But the language is sitting on top of an operating environment.
The internet eventually developed a simple universal interface: the search box. That interface concealed an enormous amount of complexity and gave almost everyone a common way to approach the indexed world. AI does not yet have its equivalent.
The chat box resembles the search box, so people understandably use it as though it were one. But a chat box does not adequately represent the scale of what is behind it. It invites people to ask questions when they could be delegating processes. It encourages requests for outputs when they could be defining outcomes. It suggests conversation when the real capability may include research, computation, coordination, creation, and action. The chat box may turn out to be only the command line of the intelligence age. Incredibly powerful. Historically important. But eventually wrapped in systems that better reflect what ordinary people are actually trying to accomplish.
Harnesses are an early version of that wrapper.
The internet taught us to think in retrieval. AI will teach us to think in delegation.
That means developing the instinct to ask: what is the actual outcome? What context does the system need? Which parts require my judgment? Which parts can be delegated? What tools and permissions are necessary? How will I know the result is correct? What could go wrong if the system acts? What should be remembered for next time?
Eventually, this may feel as ordinary as searching the web. A future generation may find it strange that people once opened five programs, copied information between them, manually reformatted documents, searched through folders, wrote repetitive emails, reconciled spreadsheets, and navigated dozens of interfaces by hand. They may regard that the way we regard walking through library stacks to find one statistical fact. Not useless. Not without value. Just no longer the default method.
The "Humanity Mark II" formulation is useful, and I want to extend it.
Humanity Mark I lived primarily in the physical and institutional world. Knowledge was physical. Access was geographic. Expertise was institutional. Humanity Mark II internalized the networked world. We learned that information was searchable, communication was instantaneous, and geography was less binding. "Google it" became a cognitive reflex.
Humanity Mark III will internalize the presence of accessible machine intelligence. Its defining assumption may be: intelligence, like information, is now an available resource.
But intelligence is more complicated than information. It can misunderstand, persuade, invent, execute, exceed authorization, and act at a scale far beyond an individual person. So Humanity Mark III will require more than technical fluency. It will need judgment, governance, skepticism, imagination, and the ability to form productive relationships with systems that can think and act without thinking or acting exactly like us.
The progression, when you pull it all together, is clean. First, the human learns the model. Then, the model learns the human. Then, the system learns the relationship. Then, we build that relationship into an interface other people can use.
We will not build the missing dynamic assistance layer by guessing what it should do from first principles. We are building its prototype inside ourselves through repeated use. Then we externalize what we have internalized into harnesses, memory systems, permissions, workflows, and eventually products. The people building harnesses are constructing the external machinery for that world. The people accumulating the deepest repetitions are constructing its internal mindset. And the truly important systems will emerge when those two things meet: a human who has learned to think with intelligence, and an intelligent environment built to understand how that human thinks.
So yes. Somebody is building in that space. We are.
Not necessarily because we sat down and named it as a product category, but because every serious interaction, every translated business problem, every agent workflow, every contextual stack, and every "you don't actually need what you asked for — you need this" moment is another training example for the thing that comes next. We are currently the handcrafted version of the product. And once we have enough repetitions to understand what we are actually doing, we will stop being merely its operators.
We will build it.
A Quick Tip of the Hat
When I first wrote this, I honestly did not know who the guy in the video was. I referred to him that way in the article because, at the time, he was simply someone I had stumbled across on TikTok with a very good question. Then I looked him up.
His name is Nate B. Jones, and after spending some time with his work, I realized he is not merely some random guy making AI videos. He is a deeply thoughtful and well-respected voice in this space, and I came away genuinely humbled by how much serious work he has already done around the practical use of AI.
His original question—what is the layer that helps people make the most of increasingly capable AI?—was the spark that sent me down this particular rabbit hole. I may have taken the idea in my own direction, but it is worth being clear about where that spark came from. Our conclusions are not identical, but we are clearly exploring the same frontier from different directions. If you have not come across his work before, it is well worth your time. Nate B. Jones: https://www.natebjones.com
Rich Washburn is a technologist and strategist working at the intersection of AI, infrastructure, and capital. He is Managing Partner and Chief AI Officer at Eliakim Capital.





Comments