The Missing Layer
Every institution has already tried to build a world model
Last month, I spent time at the Society Incubator, Edge Esmeralda, in Healdsburg, California. It felt like a monthlong conference focused on the frontier of tech, science, art, and everything in between. I’ll share my reflection on that experience at a later date. Every Wednesday, I would volunteer at their co-working space to teach local businesses how to use various AI tools that piqued their interest. After one of these sessions, I encountered a man named Neil outside the co-working building. We spent over an hour and thirty minutes sharing stories, our backgrounds, philosophies, and our views on the future of democratic governance. After our conversation, he sent me this piece, Fei-Fei Li published, that lays out a taxonomy of world models in AI. After reading her article, I saw my own work on institutions reflected in it and wrote this up as a thinking exercise.
What is a World Model?
She breaks world models into three functions. A renderer produces what something looks like. It outputs pixels, images, and scenes. It optimizes for visual plausibility. A planner outputs actions. Given what an agent can observe, it decides what to do next. A simulator outputs state: a faithful representation of what is actually happening beneath the surface, accurate enough for both humans and machines to act on it with confidence. The definition I got from NVIDIA is: a machine learning system that builds an internal representation of an environment. Instead of merely predicting the next word or frame probabilistically, it understands physics, causality, and spatial dynamics to predict how the environment will change in response to specific actions.
Her central claim is that simulation is the linchpin. You cannot build a reliable renderer or planner without an honest simulation underneath. A renderer without simulation produces outputs that look right but fall apart the moment you need them to hold up under scrutiny. A planner without simulation is making decisions on a model that has never been checked against reality. The simulator is the bridge. Without it, both rendering and planning are built on overconfident models that have stopped improving.
I look at world models a bit differently than they are traditionally viewed. The traditional framing treats a world model as something you engineer, a machine learning system inside a robot or an agent. I think the structure is more general than that. Any conscious or semi-conscious entity that has a perspective is already running a world model, whether it engineered one or not. A perspective is made of the same three components. The clearest way I can put this is in terms of memory. A memory carries a frame of reference, which is the rendering: how the moment looked from where you stood. It carries the actions that led into it and the actions it set in motion, which is the planning. And it shapes the perspective of how you expect the world to behave and how future memories get formed, which is the simulation. A world model, at its most fundamental form, is memory doing all three jobs at once. That is why the taxonomy applies to more than just robots. A person has a world model, and so does an institution.
By moving beyond simple data archival, the Governance Memory System functions as a State-Fidelity Simulator. Following Fei-Fei Li’s “World Models” thesis, while most institutional reports act as “Renderers,” producing a plausible appearance to satisfy optics, GMS is designed to simulate the faithful underlying state of governance, at times contradicting the popular narrative to surface uncomfortable truths. This may come with institutional resistance because transparency is often viewed as ammunition for “Schadenfreude” rather than a tool for improvement. The Governance Memory System must be framed as a “State Simulator” that protects the institution from its own amnesia.
What does it look like when only two of the three requirements are satisfied?
The architect without the engineer is pure rendering. The vision is compelling and the image coherent, but its integrity has never been tested against physical constraints. It looks right until someone has to build it. Take invisible architecture, these are buildings designed with extremely reflective materials that disappear into their surroundings. The renders are stunning. In practice, the structures become hazards to birds, accumulate dirt that requires constant, unsustainable maintenance, and create glare and heat problems for everything around them. An accurate simulation could have prevented all of this.

Next we have the PM without engineers. This is planning. Decisions get made, roadmaps get written, priorities get set, but none of it has been checked against what is actually feasible. The engineers come back and say this takes six months, not two weeks, or that it breaks three other things, and that the plan was never possible in the first place.
Last is something less obvious and more dangerous. The simulation layer is present and running, but it is simulating the wrong world. GDP is the example most people live inside without conscious acknowledgment. It is a simulation of economic health that excludes unpaid labor, environmental degradation, inequality, and long-term sustainability. Governments treat it as progress or failure; policies are planned in response to it, and the model technically runs correctly. But GDP was never built to represent the thing it got promoted to measure. So the rendering and planning built on top of it are coherent relative to the model and systematically wrong relative to actual human welfare.
Education rankings do the same thing. A school gets compressed into a single score, and then districts plan around that number and reports treat it as proof of improvement or decline, even though the score was never a real picture of what the school is. The score was never an accurate component of a world model to begin with. It measured a sliver of the thing, and the whole system contorts itself toward exceeding this target. The simulation is running but isn't grounded against a healthy, nuanced metric.
Financial models are the most consequential case. Faulty financial models were a primary catalyst in the 2008 crisis. The models were seen as sophisticated; they were widely used, and built on assumptions about housing prices and correlated risk that had never been tested against the scenario that actually materialized. There was nothing in the financial system tasked with challenging the model’s own assumptions, nothing actively trying to falsify it. Confidence compounded until the gap between the model and reality could no longer be absorbed.
This is Goodhart’s law at work across all three examples: when a measure becomes a target, it ceases to be a good measure. A proxy gets promoted into a definition of reality, and once that promotion happens, nobody goes back to check whether the proxy still tracks the thing it was supposed to stand in for. Any simulation you build will lose fidelity somewhere. It is almost impossible to model everything, so you have to simulate with the understanding that your model cannot capture everything and accept a tradeoff. That tradeoff should be survivable. What is not manageable is the loss of the mechanism that checks the model against the full physical reality it claims to represent. At that point, it is a closed system mistaken for an open one, and it appears rigorous because it uses math and produces plausible outputs. As we know, there are whole industries built on top of this.
More recent examples are less subtle. DOGE claimed roughly $214 billion in savings from staff reductions and contract terminations. A Politico investigation found that of the roughly $145 billion claimed through canceled contracts, less than 1% were real, verifiable cash savings. Many staff cuts were later reversed in court or reinstated. Federal spending in fiscal year 2025 came in higher than the year before. The planning was aggressive, and the rendering was confident. The simulation, the honest accounting of what those roles and contracts actually did and what removing them would actually cost, was not present. Cuts were made faster than anyone could model and verify their consequences.
In each of these cases, the failure is the same: the three layers exist, but they are not accountable to one another.
What do institutions have to do with world models?
Every institution is already doing all three. The rendering is the annual report, the board minutes, the press release, the strategic plan. The institution produces an image of itself for internal and external consumption. It optimizes for plausibility and tells a coherent story. The planning is every budget cycle, every vote, and every policy decision. Given what leadership observes, it decides what to do next. The simulation is supposed to be the accounting of what is actually happening: which commitments are being fulfilled, which decisions are producing their intended outcomes, and where the stated model of the institution diverges from the physical record.
Most institutions mistake monitoring for simulation. Monitoring tells you that a bill passed. Simulation tells you the bill passed unanimously, that similar legislation was introduced and abandoned in two previous parliaments, that the sponsoring parliament member sits on three committees despite never publicly advocating for the bill, and that the budget area the bill affects is shrinking. Monitoring gives you the explicit record. Simulation tells you how to analyze the record and what it means for the bigger picture.
Almost every institution I have worked with has robust rendering. Almost all of them have some form of planning. Almost none of them have honest simulation. And for the reason Li describes, both the rendering and the planning meander from reality over time, slowly at first, then all at once.
Things GMS can tell you about your institution
Earlier this year I deployed the Governance Memory System on a school board in Jersey City. What I found was corroborated against a state auditor.
The board had approved 484 commitments totaling roughly $514 million. Of those, approximately 3% had signed contracts covering about 7.6% of the approved dollars. The meeting minutes were documented, and the commitments were voted upon, so the rendering and planning were occurring. The simulation was not, though.
The gap between what the institution said it was doing and what it had actually formalized into binding agreements was a missing layer. There was no mechanism that checks the institution’s stated model against its own physical record. So the rendering kept compounding atop a foundation that had never been tested.
The pattern holds across contexts. I ran the same analysis on the national parliament of a small island nation, using only publicly available data and with no prior knowledge of its politics. GMS found that legislative activity was heavily concentrated in capital constituencies, while outlying islands received disproportionately little attention. It flagged the media bill as anomalous before anyone told me it was the most controversial piece of legislation in the current parliament. It detected internal factions within the ruling supermajority rather than treating the party as a monolith. It identified a single MP who voted against his party on every recorded vote across all sessions. None of these findings came from background knowledge. They came from data that contradicted the institution’s own stated model of itself.
Explainer and Breaker
As I was working through Li’s taxonomy, I was reminded of S.A. Senchal, who introduced me to his Observer Theory Extension. This paper formalizes what any observer must do to construct a world model at all. His work sits one level of abstraction above Li’s essay. Where Li asks which functional components constitute world-modeling capacity, Sam asks what the architecture of any system must look like if it is capable of discovery rather than merely recombining what it already knows. The Extension makes a specific prediction that genuine discovery requires a layered architecture built from two kinds of instructions. In a more recent piece, “Critical Observer Theory Predictions,” Sam argues that this prediction has now shown up in the wild, in a system built for entirely unrelated reasons. The two kinds of instructions are the Explainer and the Breaker.
He answers that you need two things working together.
The Explainer - A layer that compresses what the system knows into governing principles.
The Breaker - A layer that actively breaks those principles when the evidence contradicts them.
Here is where the two frameworks click together. A simulator that checks itself against reality is an Explainer and a Breaker working in tandem. A simulator that has stopped checking itself is an Explainer running on its own. That is the failure in every example above. GDP, education rankings, the 2008 risk models, DOGE’s accounting: each had an Explainer producing confident, coherent outputs, and no Breaker holding those outputs against the physical record. These systems had learned to generate plausible answers by blending what they had already seen. It could not produce the answer that requires a fundamentally different framework. The anomaly that only shows up when you check the model against real life.
Sam pointed to the work of Markus Buehler at MIT, who independently built an architecture with this structure for protein design, arriving at it through engineering optimization rather than theory, two principles that govern billions of parameters. The layer above was sparser, but each rule in it carried more weight across everything below. Buehler’s system produces molecular structures that have never existed in four billion years of biology. It does this because the Breaker keeps running. Now imagine the innovation and opportunity that could bubble up if the Breaker was placed in other institutions.
What does the Governance Memory System have to do with World Models?
Coming back to Li’s taxonomy: GMS is technically all three, but not equally. The rendering is what GMS surfaces to the people in the room. The dashboards, the gap analysis, and the findings. That is how humans engage with what GMS has found.
The planning is implicit and intentional. GMS does not tell institutions what to do. When a board member sees that 3% of their commitments have been formalized and decides to fix the contract execution process, GMS did not make that decision. It made that decision legible, and the planner stays exogenous.
The simulation is the integral layer. GMS sits above the institution’s dense operational base, the proposal lifecycles, the outcome data, the relationships, the tacit knowledge, compresses that into governing principles, and then holds those principles up against the institution’s own physical record so the institution can see where it contradicts itself. That is how the Explainer and the Breaker work together. The institution’s actual record feeds up into GMS. GMS’s findings push back down into how the institution understands and organizes itself.
Without that layer, the rendering compounds false confidence and the planning builds on a model that has never been tested. With it, the institution is not being judged by an outside observer. It is being shown itself.
Conclusion
Li’s essay is about robots and spatial intelligence, but the structural argument she makes is older and more general than any specific technology. Any system that tries to understand itself needs a layer that can be falsified. It is the only mechanism that forces the model to keep updating past the point where it became comfortable.
The problem in most institutions is that all three layers exist without accountability to one another. The renders are not being pressure tested. The plans are not being checked against physical reality. The simulations, where they exist, are not accurate models of the world in which the institution actually operates. They are closed systems running confidently inside their own assumptions. The ones that will survive this new technological paradigm are the ones that connect the layers before the gap becomes too wide to close. The rest will keep rendering a world that no longer exists.
The Governance Memory System is institutional memory infrastructure developed by OCC Research. Learn more at occresearch.org.

