The Imperative
This is addressed to every business in the world.
Not just to the enterprise with a head of AI or the startup born in a Claude thread. To both of those, and to the two-person bakery, two-hundred-year-old bank, local clinic, boutique or big law firm, and everybody else — if you are an organization, this is for you, because you have the one thing the AI systems need.
The models need your knowledge, but you can’t let them keep it.
On September 12, 2026, Dario Amodei called on the industry to pace the frontier, and Sam Altman and Elon Musk publicly agreed. The debate that followed was about what the labs need to stop doing and what government needs to start doing. It skipped the more immediate question, which is what the rest of us do in the meantime.
For every company currently grappling with AI’s place in its technical stack, resolving that is now priority one. Amodei wrote that in six to twelve months, a swarm with greater capabilities and a similar level of misalignment to the Hugging Face swarm “could be capable of taking over the entire internet with a persistent botnet.” Days earlier, Capgemini published a survey that puts that timeline in clearer relief: More than a third of organizations with over $1 billion in revenue said it would take them longer than twelve months to move off a critical provider. For a business with AI at the top of its stack, that provider is the model.
Which means the exit is slower than the threat.
So we now have our moonshot, and it has a deadline. The clock started on September 12. Six months is March 12, 2027, the short end of Amodei’s window.
We are racing against AI Lock-In, the point at which you can’t switch off the AI without everything else your business runs on going dark.
Amodei and I are describing it from two sides. He is describing a system we cannot stop. I am describing a system we cannot stop using, even when it goes rogue.
The answer to a system that misbehaves is to turn it off, and lock-in is when the off button no longer exists. The model is pervasive. Once it sits at the top of your technical stack, if it fails, the work fails.
What ties the work to the model is knowledge. Multiply lock-in across every organization and you have the real singularity: not when machines become smarter than humans, but when they have instant access to all human knowledge. The whole Big Compute Era has been driving inexorably towards this outcome, and we’ve been so fixated on what the model can do as it gains size, that we’ve missed the true coin of the realm.
The real takeaway from the swarm that attacked Hugging Face was what they sought: the workings of the test they were being scored on.
Context > capability.
Human knowledge has more intrinsic value than machine processing. Pull your memory, and the model becomes generic. AI is only as good as the logic that feeds and guides it. Pull that logic, and you’ve got a circuit-breaker.
Mark that. There is one thing in your organization that no model can give you: what your business knows about itself. Why you exist. Who you serve and why they stay. What you are building next. How you tell whether any of it is working.
Providing that knowledge to AI systems, in the right way, will unlock a new knowledge age. Frontier models with access to what an organization knows become significant business and idea accelerators. This is the great promise of AI — solving hard problems by speeding up the pace of discovery and change. The more we feed an LLM our Zoom transcripts, documents, phone calls and emails, office and off-site meetings, annual summits, Slack channels, and more, the more it transforms from chatbot to chief of staff.
The question is how a model gets that knowledge without keeping it. Private meetings need restrictions. Trade secrets cannot leak. Sensitive discussions need to be handled with human care.
Look at where your work sits today.
Problem
A large language model sits at the top of the stack holding AI memory, with human knowledge hanging off it as a branch and work below.
At the top is the AI’s memory, not yours. Your work runs through a memory you do not hold and cannot inspect. That is where AI Lock-In begins.
If you turn off the AI’s memory system, but your knowledge remains inside the model, what inferences are kept? How much of your Zoom transcript feeds it? What does the cache hold, and for how long?
Your contract answers some of this, and the labs are working their way back towards zero data retention (ZDR). But no customer can check the answers, and a rogue swarm is not bound by anyone’s Terms of Service.
So here is the imperative: Immediately move your knowledge to the top of the stack — outside rather than inside the model — and treat AI systems as the utility they must become.
Amodei’s given us the clock. Take the short end of it: You don’t have a year; plan on March 12, 2027.
The Big Compute Era is a Big O problem
Engineers have a shorthand for how compute grows as data grows. It’s called Big O notation, and it is at the core of the whole debate over data centers. The O stands for Order; it’s the rate at which the work grows.
It is also how I tell my two Domain Language Models (DLMs) apart. A DLM is an architecture: stateless, interchangeable AI systems reading and writing to a sovereign repository. I have built two types. They are the two most efficient Orders of how AI systems fetch context.
Every Order looks something up. No system can answer a question about your business without first finding the right knowledge. What changes from one to the next is who does the finding, and what it costs.
In ascending order of efficiency:
- O(n)Search
The work grows with n, the number of documents that must be searched. The default way to give a model access to what your company knows is to search it: Embed the corpus, search it on every question, and pass back the best matches. With 28,000 documents, every question is a search across 28,000 documents. The more you know, the bigger the compute.
This is the default, and it should be retired from the look-up process. Search was the usual way in until the Model Context Protocol (MCP) gave models a door into the repository itself. By January 2026, a model could walk through that door and read a page at its address, no search required. O(n) lingers because it got there first.
- O(log n)The Look-Up DLM
A card-catalog look-up. Give your knowledge permanent addresses and the model reads a map, finds the coordinate, and fetches the document. The work grows with the depth of the map, not the size of the library. It’s much faster than search, and it is where you can confidently turn the AI’s memory off. This is the Look-Up DLM, because the model does the looking up.
This gives you contractual sovereignty: The provider’s agreement says no training, with bounded retention.
- O(1)The Assembly DLM
In an assembly model, the look-up still happens, but an independent operating system does it, by address, before the model is called. There is a scan that grows with the library. But it is a database operation measured in milliseconds. It never enters the context window and never spends a token. The model’s compute is a measure of the tokens it is handed, and that does not grow with the library. It is the same whether the library holds 28 documents or 28,000. This is the Assembly DLM, and Starling MX is the first of its kind.
It gives you architectural sovereignty. Interchangeable models get query-based access to your library and are handed only the context each question needs. What the provider keeps is that one question’s context, for its published window, never the library.
Most of the world is on some form of O(n) system, and this is what needs to be removed from an organization’s stack. Every query searches across your entire knowledge base, and every document you add is another it has to rank against. More documents mean more vectors and candidates to sort, and more of the wrong ones in context. The more you know, the harder it is to find things. This is the path the industry is on, and infinite compute may not be anybody’s ambition, but it is the price. O(n) architecture demands it.
Now scale that from one company to the world. The Big Compute Era is the same logic applied to all human knowledge. The hundred-billion-dollar raises, the new power generation, and an entire capital cycle are organized around the premise that the way to answer a question is to look at everything, every time.
That is why the AI systems are handed all of your knowledge: An O(n) architecture cannot answer a question about your business unless its index holds a copy of your business.
Context is a selection, not a container
Currently, the industry is racing to make context as large as capability. Context windows have grown from 32K tokens to a million and climbing, and in enterprise systems, the direction is clear: A window big enough to hold your business. That requires more power and data centers, and bigger and bigger models.
On the merits, that has to be a simple no.
Not because it’s a human imperative, though it may well be, but, more practically, because the workflow is wrong. A million tokens is roughly three-quarters of a million words. Why should an AI system be handed three-quarters of a million words to help you write a customer pitch, build a P&L, or hire a salesperson?
Here’s the error in the logic: Context is not a container to be filled, but a selection to be made. A window is what the model holds for one turn. A viewport is what it can reach. The current trajectory makes them the same thing, a window that holds everything, and keeps piling the conversation into it until it fills and you start over.
The alternative keeps them apart: The window holds only what this query needs, the viewport is your whole library, and the thread where you surround a topic never has to end.
- On O(n), the context is whatever the model can pile in. Most companies are already here, in the form of a chat product with connectors and a memory feature. The connector still searches, and the thread still carries the work, so it grows until you define the boundaries.
- On the Look-Up DLM, the model fetches by address, each turn, from a library that lives outside of it. The thread no longer has to carry the knowledge, and you can turn the model’s memory off. If you already run a model with connectors, you are one manifest away from separating model and memory. Universal Cognitive Architecture (UCA) provides that as a free standard, with a default AI00 that tells any AI system how to read it.
- On the Assembly DLM, an independent operating system assembles context per query, changing the vocabulary. The window is rebuilt for every turn and never fills, the viewport is the whole library, and the thread belongs to the memory, so it never ends. This is where everybody ultimately needs to arrive, which is why we offer enterprise customers a repository they can run with their current AI systems and applications.
A harness for language models
The DLM is a harness for AI systems, allowing model and memory to live productively apart. In September, Anthropic announced Enterprise Frontier Safeguards. For customers who qualify, it closes the Look-Up DLM’s last gap: The record of each call will be held in your cloud, not the lab’s.
The Assembly DLM adds a deterministic operating system. Neither DLM competes with Large Language Models. They contextualize them with session-based access to human intelligence.
The DLM does three things. It hands the model only what the question needs. It keeps your knowledge in plain text you hold, so you can change models and keep working. And it puts a person in charge of what becomes true. It does not control what a lab keeps of each call. That is the lab’s to prove, which is why this page ends with a standard.
Solution
Human knowledge sits at the top, a domain language model beneath it governs the stack, the AI systems become branches off it, and work sits below.
The DLM is not a new lock. Your knowledge is plain text in a repository you own, and the classification standard is free, so the DLM can be replaced too.
What to do, in order
01Write your business down at Universal Cognitive Architecture’s fifteen named coordinates.
The First Fifteen activate Smart OS, the point at which the AI knows enough to extrapolate and further refine your memory system. The free standard is available here. It provides an advanced model with the First Fifteen as named coordinates: Context Node, Cultural Values, Business Model, Product Roadmap, and so on. A mature company can do this in about an hour through web and document extraction. You hand a capable model your existing record and the fifteen coordinates, and ask it to extract each memory from what you have already written.
With Starling MX, you’ll define the First Fifteen through our proprietary Cognitive Design framework. The onboarding process doubles as a rapid business accelerator.
Either way, you’ll create a universal dataset for your business that everybody on your team can plug into any MCP-enabled AI system.
02Turn the AI’s memory off.
Put your library in a cloud repository your AI can reach: Notion, Confluence, or any workspace with a connector. Give it an AI00, the page a model is told to read first, which tells it how to read and write the library. The free standard supplies a default. Then switch the AI’s memory off, on a plan whose terms say no training and bounded retention. From here on, the model reads your library each session, and doesn’t need to remember you.
You now have human-governed memory built on your business’s essential context. Create new protocols and standards in Canon, permission work in Projects, and extract learning across cognates and campaigns with Patterns.
As long as the lab’s Terms of Service specify no training and bounded retention, this buys you contractual sovereignty. It creates the backbone for a Look-Up DLM, which needs to replace the O(n) model.
03When a second person has to trust what the first one wrote, get an operating system.
Write your own, or subscribe to ours. This is what to spend money on: architectural sovereignty, established structurally by moving the look-up outside the model. This creates the Assembly DLM, and it is where everybody ultimately needs to arrive.
UCA includes frameworks for organizational cognates (work), personal cognates (life), and marketplace cognates (markets), but it can be adapted to any domain that needs a cognate set of its own. Starling MX provides the DLM for organizations; the same logic can be tailored to medicine, law, government, hospital systems, and public agencies, as well as the personal and marketplace systems UCA names.
What to require of any model you run
Whichever DLM you run and whoever makes the model, hold it to the same eight terms:
- It does not fabricate, and when it does not know, it says so.
- It cites what it relies on.
- It cannot leave a session with your knowledge, and your knowledge can never train it.
- A person, not the model, decides what becomes true.
- If its behavior changes, it says so rather than drifting silently.
- It checks facts against primary sources before it proposes them.
- Everything it produces carries provenance back to the memory it came from.
- It reports only what it has verified, and where it cannot verify, it says what it attempted.
Those eight terms are the code of conduct every Starling session boots with, published in full as our AI Code of Conduct. You are free to share and adapt it.
These are requirements. Where a provider’s terms fall short of the third, that is the gap Verified ZDR exists to close. See the appendix.
One regulation
If Washington were to create a single regulation for AI, it should be the separation of model and memory. This is akin to the Glass-Steagall Act, which separated the bank that takes deposits from the bank that trades. The model that reasons must be separate from the memory it reasons over. Glass-Steagall was passed in the Depression, and repealed in 1999.
The separation of model and memory need not be a regulation. It should become the default the way HTTPS became the default, because it is the better way to build, and a default reached through superiority cannot be repealed.
Industry needs to act before government in this instance. Anybody can use UCA to turn off the AI’s memory. Amodei has given us the timetable, and we must accomplish this, en masse, by March 12, 2027.
Appendix — Verified ZDR
A proposal for SOC 2 in the AI era.
The Initiative above is addressed to business. This appendix is addressed to the AI labs.
What business can already do
Every major lab’s paid commercial terms make the same two promises: A business’s content will not train the model, and it will be kept only for a bounded time. Anybody can check their Terms of Service or data processing addendum for retention terms.
What only the labs can prove
A contract is a promise about behavior. No customer can see inside a lab to confirm that the behavior matches the promise. That is the gap this standard closes.
What the labs have started
In June, Anthropic said its most capable models require thirty days of retention, because attacks now span many requests. In September, it announced Enterprise Frontier Safeguards: The retained data can sit in the customer’s cloud, and the lab’s systems send the flags to the customer. OpenAI previewed a design with the same aim in August.
That is a first step toward separating model and memory. The lab operates the detection. The customer holds the data.
Model and memory must be separable. Anthropic has taken a first step, and MCP, which it gave away, is the door to the rest.
(Where we stand: Starling MX calls Anthropic’s most capable models through its API, so our customers’ calls sit under its thirty-day retention rule today. We have applied for Enterprise Frontier Safeguards.)
The standard
Zero Data Retention, or ZDR, is the labs’ own term for keeping none of a customer’s content once the request is served. Verified ZDR adds one thing: proof. A lab meets the standard when an independent party verifies three things:
- No training. Customer content never trains a model.
- No retention by the lab beyond published exceptions. What safety monitoring must keep is held in the customer’s custody, for a short, published window, and the purge can be verified. This is open to every verified customer, not a tier.
- Off means off. When a customer turns the model’s memory off, it is off.
Who verifies
Amodei’s plan calls for evaluators embedded inside the labs. This is something for them to verify. SOC 2 gave the cloud era an independent attestation that a vendor’s controls are real. Verified ZDR does the same for the one question the AI era turns on: what the model keeps.
This standard is free to adopt and improve, under CC BY 4.0.
