There is an old storybook illustration of a kingdom built in the clouds: white spires floating above, an ordinary town on the ground below. To turn it into a picture of the AI economy in 2026, you only need to add two details. The town pays rent upward, by the token. And from that height, the kingdom can see everything: which shops are busy, which trades are profitable, which merchant is one good season away from mattering.

The story of this year is that the town has started to notice.

Start with the one man who lives on both levels. Elon Musk owns an AI lab, xAI. He also runs Tesla, which buys AI tools like every other company in the town. Last week, Tesla capped employee spending on AI at roughly 800 dollars a month, manager approval required above that. And the memo contained one exemption: the cap excludes xAI's own products.

Hold both halves of that policy in your head. Spending on rented models: capped. Usage of the models he controls: uncapped. There are two ways to read it, and both of them are the thesis of this piece. If the cap is cost discipline, then a man who has bet his entire empire on AI has concluded that token spend does not equal value. If the exemption is Musk steering employees toward his own lab, then one of the best informed operators in the industry believes that controlling the model matters more than renting the best one. Cost discipline or sovereignty play: either way, the expense policy tells you what the keynotes will not.

This piece is about the gap that policy reveals: what AI labs are selling, what enterprises actually need, and why the two are diverging fast. The evidence comes from this year's spending data, two product launches, and an interview where Alex Karp said the quiet part loud.

One boundary before we begin. This is about generative AI: the economy of tokens, large language models, and the harnesses around them. Other AI already has defined economics. The vision stack in a car, industrial robotics, fraud detection, ranking systems: those are sold inside products and priced by the value they deliver. Nobody bills you by the token for a car that drives itself. The broken economics live in the generative layer, and the generative layer is what the frontier labs sell.

I — Nobody pays for intelligence

Karp on CNBC last week: "I am paying for tokens that create no value."

He is one of the most polarizing figures in tech, and Palantir profits directly from the thesis he is selling. Discount him however much you want. Then check his claims against what has actually happened over the last few months, because the data keeps agreeing with him.

Start with the core economic fact. A raw language model is general capability. It can reason, write, and predict. But nobody pays for general capability, the same way nobody pays for a brilliant employee who has no phone, no computer, and no permission to touch anything. Value appears when capability gets pointed at a job: tools it can call, loops it can run, permissions defining what it can touch, stop conditions defining when it is done.

That layer is the harness. Palantir calls theirs "ontology" and built a company on it. But you do not need Palantir's word for it: the same model frequently powers both Claude Code and Cursor, and everything that makes them different products is harness.

Now the asymmetry that changes the strategy for every company. Pretraining costs billions and belongs to the labs. Harness building and post training cost ordinary engineering, and they are where the gains have moved as pretraining hits diminishing returns. The layer that costs billions is commoditizing. The layer that creates the value is one an enterprise can own.

II — Tokenmaxxing: the loops are the business model

This year Silicon Valley coined a word for its own behavior: tokenmaxxing.

The maximalist position has a spokesman, and it is not Musk. In March, on the All-In Podcast, Nvidia CEO Jensen Huang said he would be deeply alarmed if an engineer paid 500,000 dollars a year consumed less than 250,000 dollars in tokens, and confirmed that Nvidia is trying to spend 2 billion dollars a year on tokens for its own engineering team. Notice who is talking. Nvidia sells the compute that produces every token in the industry. When the man selling shovels announces that every miner should spend half a salary on shovels, that is not analysis. That is a sales target. For Nvidia, the math is wonderful. For everyone who is not Nvidia, it has to be checked against output, and that is exactly the check that failed this year.

Companies started treating token consumption as a productivity metric. Meta employees built an internal leaderboard, nicknamed Claudenomics, ranking colleagues by tokens burned, with titles like Token Legend. Meta went through 60 trillion tokens in a month. Executives publicly told engineers to spend with no limit.

Smart engineers did what smart engineers always do with a visible number: they optimized it. Agents looping for hours. Prompts stuffed with context nobody reads. The biggest model selected by default.

Then the research caught up. Agentic tasks can consume up to a thousand times more tokens than a simple chat. The same task can vary in cost by thirty times. More tokens does not reliably mean better output. And models underestimate their own token costs, which means the worker deciding how much work to do is also a bad accountant, and the accountant works for the company selling the meter.

Then the invoices arrived. One company spent 500 million dollars in a single month on uncapped AI licenses. Uber gave Claude Code to five thousand engineers, burned its entire 2026 budget in four months, and imposed a 1,500 dollar monthly cap while its COO admitted the link between token spend and output is not proven. Meta went from "no limit" to a cost warning memo to six thousand employees. And Tesla, as we saw, capped the rented meter while leaving the lab its CEO owns outside the cap.

The pattern across all of them is the argument in miniature. When you rent intelligence by the token, spending is a cost to contain. When the model is yours, usage is an asset to encourage. Same technology, opposite economics. The only difference is who owns the weights.

III — The cloud comparison is backwards

The standard defense of token billing is the cloud analogy: AWS also bills by consumption, and cloud created enormous value. But why does cloud have value? Because a cloud provider is a custodian. Your data sits encrypted in its buckets, and AWS does not earn one cent more by understanding what your data means. It is a landlord.

An AI lab is a participant. Every API call passes your prompts, your documents, and your workflows through the provider's systems at inference time. The major labs promise contractually not to train on your data, and there is no reason to doubt they honor it. It does not matter. Even a perfectly honest lab unavoidably accumulates something no landlord ever gets: a real time map of where economic value is being created with its product. Which use cases are exploding. Which workflows are sticky. Which of its customers' businesses are thin layers it could absorb.

You are not just paying for tokens. You are financing your supplier's market research.

And the cloud's own history shows what visibility does. When AWS could see which open source software its customers ran, it launched competing managed services and forced Elastic and MongoDB to change their licenses in self defense. Amazon's visibility into marketplace seller data drew a congressional investigation. Wherever a platform can see its customers succeed, it eventually competes with them. The cloud analogy does not defend the AI stack. It predicts what happens to it.

IV — The supplier becomes the competitor

This is not a prediction anymore. It has already happened twice, to two of the best product companies in the world, at the hands of the same supplier.

Cursor built one of the fastest growing developer products in history on top of Claude models, spending years proving what the product should be and what developers would pay. Then, in early 2025, Anthropic shipped Claude Code, a direct competitor to its own customer, and grew it to roughly 2.5 billion dollars in annualized revenue by early 2026. Cursor responded the only way a dependent company can: it built its own model and started routing around its supplier.

Figma was collaborating with Anthropic as recently as February of this year. In April, Anthropic's chief product officer resigned from Figma's board, and three days later Anthropic launched Claude Design into Figma's market. Figma's stock fell 7 percent that day.

Be precise about the mechanism, because the precision is what makes it unfixable. Nothing was stolen. No contract was violated. None of that was necessary. Cursor and Figma de-risked entire product categories in public, in full view of a supplier with frontier capability, enormous capital, and a 30 billion dollar revenue run rate that has to keep growing. The supplier watched the categories succeed and vertically integrated into them. The kingdom saw which shops in the town were thriving, and opened its own.

A risk that requires no misconduct cannot be negotiated away. No contract fixes an incentive built into the structure. It can only be architected away.

V — The second accumulation

Zoom out and the shape of the whole thing becomes visible.

The frontier models were built on the first great accumulation: essentially the entire written output of humanity, scraped and trained on, a process still being litigated across dozens of suits, and one that already cost Anthropic a 1.5 billion dollar settlement in 2025 with authors whose books were pirated for training. Whatever you conclude about the ethics, the result stands. A handful of firms hold a compressed copy of public human knowledge.

The token based enterprise model is the solicitation of the second accumulation: the private layer. Proprietary workflows, clinical data, trading logic, manufacturing processes. The accumulated alpha of every company that routes its operations through a frontier API. Not necessarily as training data. As visibility, dependency, and position. The kingdom does not need to raid the town. The town delivers its treasure by subscription.

Follow it to the endpoint. Two or three firms holding all public human knowledge, a real time map of private economic activity across every industry, frontier capability nobody can replicate, and the capital to enter any market they observe. The same entity as your supplier, your competitor, and the observer of your industry. Markets work because competitors face each other on something like equal information. A market where one player sees everyone's hand is not a market. And opting out is not an answer when they hold the frontier, because then the choice becomes dependency or irrelevance.

VI — The sovereignty stack

To be fair about scope: token metered frontier access is genuinely excellent for individuals, for prototyping, and for work where your data carries no alpha. I use frontier models daily and they are remarkable. The argument is aimed at exactly where the money is: companies whose moat is proprietary data, process, or IP. For them, the token model carries a hidden line item, the compounding transfer of strategic position to a supplier structurally incentivized to eventually use it, and no contract zeroes that out.

The durable architecture is already visible, and it is simpler than the industry wants you to believe.

Take an open weight model. The gap to the frontier keeps narrowing, and inside a bounded business context, reliability within a schema beats frontier generality.

Post train it on your domain, because post training is where the gains live now.

Build the harness yourself: the ontology of your business, the tool permissions, the control flow, the stop conditions that keep an agent from burning a month of budget in an afternoon.

Run it on infrastructure you control, or with a cloud custodian holding your encrypted weights. A landlord is fine. A participant is not.

Own the weights. Own the data path. Own the alpha.

The owners already run this playbook. Musk caps what Tesla rents and exempts the lab he controls. Nvidia's CEO tells the world every engineer should burn half a salary in tokens, and Nvidia sells the compute underneath every one of them. The frontier labs pour their best engineering into harnesses while selling you tokens. Watch what they do, not what they sell.

The town does not need to storm the kingdom. It needs to stop mistaking rent for investment: a model you control, wrapped in a harness you own, pointed at problems only you understand, on ground that belongs to you.