Earlier this month Stripe agreed to acquire OpenRouter for $7.5B.

This kicked off a flurry of commentary. Slow’s Sam Lessin claimed

“there’s NO technology in OpenRouter. It’s a key swap with great marketing.”

And @patrickmandia

“The tsunami of routers launched in the past few weeks is puzzling to me. Every modern app-layer business will become a router. It’s too fundamental to their existence to outsource to another company.”

I intuitively felt like Sam and Patrick Didn’t Get It. I’ve recently developed conviction that Outcomes Pricing best sits at the router layer (and possibly model layer, depending on the market structure of compute), not the app layer. If that’s correct, it’s not much of a jump to imagine that the router market power would follow––indeed!!––market power of systems in ad-tech.

Outcomes Pricing belongs to the Infrastructure-Layer Link to heading

Ten months ago I wrote a takedown on app-layer outcomes pricing. I felt like their motivating comparison to ad-tech was dumb–in ad-tech you’re selling scarce attention while intelligence will soon be Too Cheap To Meter. And it felt like the price of outcomes should be determined by the market e.g., from compute availability, not for a single price on an annual contract set after a polite discussion with ex-Google ex-Salesforce Brett Taylor.

I felt that, contrary to the exposition of Sierra, app layer incentives were not aligned with their customer. As the app layer develops better insight into agent performance on customer problems, and models become cheaper and more performant, there’s no apparent mechanism for customers to participate in this surplus. It all accrues to the app layer. Switching costs are poised to be high and customer visibility into underlying agent operations are poised to be low, so customers really have little leverage to negotiate better deals.

Outcomes pricing at a neutral infrastructure layer is poised to fix this. The model layer could offer outcomes pricing––it’s reported OpenAI is experimenting with this––but it may have the same troubles as the app layer, seeking lock in or prioritizing its own models v only optimizing for the most cost-efficient delivery of outcomes.

In this way, a router layer or inference provider/neocloud that moves up into the router layer, is politically most aligned with customers. They just want to sell more compute––they do not care which model is used.

Even if this is true, however, it’s not clear whether the router layer has genuine right to become a centralized platform. The market structure development of ad-tech is instructive.

Market Power in Ad-Tech Link to heading

The relation between ad-tech and outcomes-pricing infrastructure is simple. Ad-tech instruments outcomes on customer sites or applications. For instance, marketers use Meta’s Conversions API to tell Meta when a customer visits their landing page or buys a product. Marketers do this––indeed, share private operating context!––so that Meta can deliver desired outcomes for the best price.

It’s also true that providing this data to Meta enables competitors. For instance, if you’re Neiman Marcus, telling Meta that a given customer, whose type Meta knows, bought something, you’re helping Meta improve its ad engine to more efficiently deliver ads to your competitor, Saks Fifth Avenue. This would apparently violate e.g., Satya’s learning loop mandate

This means the real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital and token capital compound.

If it’s true that centralized router or model layers could deliver outcomes more efficiently than a business could on its own, the candidacy of every business building its own learning loops starts to diminish. This outcome would be a bit unnerving, as it violates a decade+ of e.g., Databricks-led narratives that all private business data is the lifeblood of every business, and giving some of it up destroys alpha, sovereignty, future growth etc., etc.,

And ad-tech’s development isn’t kind to Satya et al.’s narrative. There are a variety of reasons why the Open Internet is dying, but there’s no doubt that the ad-tech infrastructure of the Walled Gardens is more elegant––primarily for its integrated simplicity-–than that of the fragmented Open Internet. Meta remains one of the best places to buy ads on the internet, and their integrated ‘continual learning’ stack is a reason why. To be clear, it’s not a clean comparison, but marketers who attempt to ‘roll it themselves’ tend to do worse on the open internet than letting Meta or Google do it for them.

Taking a step back, there is one important difference between ad-tech market dynamics and outcomes pricing: Entropy. Ad-tech directly sits on top of a high-entropy stream of changing preferences and businesses serving products to meet these changing preferences. If preferences were more fixed, while it’s true Meta has captured attention, which indeed is scarce, it may be true too that the infrastructure or data advantage Meta has developed would be less valuable, because as soon as you create a decent representation of preferences and they don’t change, you’re basically done. Anyone else who assembles attention could replicate the ad engine by sniffing out never changing preferences.

This is related to my long-running curiosity whether the future of the labs hinges on the ‘amount of surprise’ in the world. I later formally modeled it here.

It now seems like ‘amount of world entropy’ is actually the quantity that dictates whether superior economics accrue to a centralized router platform developing better learning over a stream of new models and problem forms, or, as @patrickmandia supposed, routers are commodity infra deployed at the edge and, because the world actually isn’t that complicated, quickly learns what agent runtime (model, harness, tools, etc.,) work best for edge problems and third party routers have no learning advantage.

I’m making my way to a simulation study of this question, but first it’s worth considering why a router is better positioned to ship outcomes pricing than a lab.

Why router v lab in outcomes pricing? Link to heading

As above, this week it was reported that OpenAI is trialing outcomes pricing with a few of its customers; The Information confirmed. This is in addition to Anthropic’s Define Outcomes documentation that has been online for a few months.

Advertising platforms like Meta and Google managed to convince apps and sites to hand them their telemetry––outcomes!––data, contrary to Satya’s advice. Why couldn’t or wouldn’t private labs follow the same playbook?

I think the difference is that advertising platforms captured something scarce––attention––that the labs really haven’t–-unless compute counts?? Advertisers like apps and sites are compelled to share this data with ad platforms because they have captured scarce attention that brands need. While it might be true that doing so creates a lock-in effect, advertisers may not have a choice. But AI users do have a choice––they can use open models on a host of inference providers.

OpenAI and Anthropic want to sell high-margin inference. Routers and inference providers just want to sell compute. They don’t care what model is used. The open-model platform maintains a more credible position from which to sell “outcomes at the best price” than the labs.

So if we accept that, does a centralized platform that aggregates verified outcomes streams accrue advantage over rolling a router at the edge?

Simulating platform v local router value proposition Link to heading

A mechanism to buy agent outcomes Link to heading

Concretely, consider a centralized router like OpenRouter or Router.com selling outcomes. It’d operate like ad-tech, inviting outcomes buyers to install telemetry (think Meta Conversions API) on their applications to communicate to the router if a promised outcome actually happened. The outcome buyer only pays when an outcome occurs as communicated by the installed telemetry. The outcome buyer specifies a definition of done much like in the Anthropic outcomes spec.

It’s true this invites a scenario where the outcomes buyer could be incentivized to abuse the telemetry and purposefully under-report outcomes. But this scenario is mitigated because doing so would lead to less efficient (more expensive!) outcomes delivery as the router actively optimizes against customer reported outcomes. Specifically, it could be that the router begins rejecting more agent labor (because it cannot be sure that spending compute on an apparently finicky customer is profitable) or begins charging much more to keep margins positive for cases it believes it can perform but is worried about false positives.

Indeed, the agent labor case is different in that when you buy conversions from Meta you do not actually buy outcomes but rather scarce attention, it’s just that the ad mechanism optimizes for conversions. I expect to have a blog exploring these dynamics soon.

Economics of centralized routers Link to heading

So can such a centralized router deliver value to the application layer?

Such a centralized router contradicts Satya: instead of internalizing learning loops, the outcomes buyer sends verified outcomes to the router in exchange for outcomes at the best possible price.

The basic idea is that a centralized router can absorb uncertainty in customer problems and new model developments faster and cheaper than can its customers. So while it takes time for a candidate customer to observe a full distribution of problems, a router platform sees problems across its customer set. So long as customers observe problems at different times, pooling provides anticipatory benefits to customers. For instance, your competitor might see a set of problems before you do; the router platform can use the verified outcomes stream from your competitor to help you; always with the promise that the router price is cheaper than building yourself minus your perceived opportunity cost of revealing this information.

I visualize the idea below:

Router Pool
How router platforms can use aggregation to learn before customers encounter new problems.

If the lab business model depends on the amount of surprise in the world, the same holds for centralized routers. For a router to retain its value proposition to customers, it should be the case that the world ‘remains surprising enough’ such that buying managed outcomes is economical relative to building agents yourself. If the world is not that surprising, there’s little reason to pay someone else margin to manage it for you. On the other hand, if the world is changing a lot, it can be economical to let someone else manage it for you.

We might call this economic value residual transferable uncertainty: the amount of uncertanty that remains after local learning and that which can be transfered across customers.

To see this, I simulate router platform that observes outcomes across many customers and check whether it can learn to deliver these outcomes more efficiently than each customer can on its own, for now assuming optimization only at the ‘agent runtime’ (model, harness, tools) layer. I further simulate whether a never-ending stream of transferable surprises can make this advantage persistent, and whether a persistent technical advantage can become market power.

Specifically, the core questions I evaluate

  1. Does an agent get better over observed outcomes (or e.g., can an agent be one-shotted by a strong model)?
  2. Does what an agent learns for one company work for another (similar? different?) company?
  3. Does a customer’s own data hold anything that another (say, competitor) company can’t supply––i.e., is Satya’s “every company is a learning loop” or Applied Compute’s Specific Intelligence defensible?
  4. Is a router platform’s edge a function of how much is new to a company, so that it lasts so long as the company continues to encounter new problems?
  5. Do router economies of scale allow it to more quickly or easily adopt cheaper or specialized models v a customer operating on its own?

To evaluate these questions I use Sierra’s tau2-bench

a simulation framework for evaluating customer service agents across domains. Each domain specifies a policy the agent must follow, a set of tools the agent can use, a set of tasks to evaluate agent performance

across domains airline, retail, and telecom. I set each task to be a support ticket: a simulated customer with a hidden fault calls the agent that has the company’s policy and tools; the ticket either resolves or it doesn’t. I measure agents in points––each point representing one percentage point in tickets resolved. Customers are synthetic companies that share telecom policy and tools but differ in problem distributions–different customers, problem types––so that, for instance, one company can end up never having seen a collection of problems another sees regularly. Tickets arrive in batches and an agent learns from these tickets absolute labels.

Learning happens in the harness around a fixed model in two ways: (i) a playbook learner hands failed transcripts to a stronger model that writes or refines a short operating memo into the agent’s instructions and (ii) a case retrieval learner shows the agent, before every turn, the three most similar tickets, two resolved and one not. Each learner comes in versions that vary in whose history they may read (1) the company’s own, (2) everyone’s pooled, or (3) only other companies’.

Now to the results. Not particularly surprising –– learning does work, but its effect gets muted as models get smarter. A mid-tier model (here gpt-5.6-luna) resolves 17 of 100 hard tickets, but gets up to 53 when trained on its own failures.

Learning on your own data is worse than learning from other company’s data: A harness learned only from other companies’ outcomes beat one learned from a company’s own history by 6 points, even when

  • companies had genuinely different customer sets
  • and even for problems a company has already seen before!

The small 6 point improvement speaks to the importance of residual transferable uncertainty to routers: for problem spaces that don’t change much, once you’ve already seen most of the problems, the opportunity for a platform to deliver value deminishes. The same is true for model improvements (this is a guess about other domains, but likely in highly legible or simulate-able spaces like customer service)––as models get better, responsibility may transfer from––here––the harness or post-trained model to the new model release. But the result also follows from the fact that other companies’ histories tend to contain more varied mistakes that provide clearer learnings than a single company’s does on its own.

This is a direct test of Satya’s “every company is a compounding learning loop” and Applied Compute’s Specific Intelligence thesis. A customer-specific moat does not show up: nothing a company learned from its own ticket stream was something a competitor’s tickets couldn’t have taught it. Where a business’s problems are the same problems its competitors have, the Specific Intelligence meme appears to be a larp designed to soothe an anxious executive class now clinging to “Your Data. Your Moat.” as weighty strategy.

For new classes of problems (in tau2 “issues” e.g., in telecom, roaming v. broken messaging setting), a router platform is 11 percentage points more likely to solve them, but for ones a customer has already seen, the platform only has small advantage. We might refer to this as a novelty advantage. This cleanly specifies router platform’s advantage as one of world entropy: residual transferable uncertainty exists so long as there is a stream of new problem types.

It’s true that a stream of new, more performant and cheaper models could also supply entropy a router platform needs. Indeed, “smarter” models appear to need “less harness”, and managing transitions to new models (or models and chips!) could be an expensive exercise for an individual application but economical for a platform with economies of scale. Experiments here showed new models delivered significant savings on agent labor but the relative value of platform v locally developed agents was unclear.

AI is confusing because it has strong decentralization autarky-maxxing forces (“everything is a dark pool”: perform all the intelligence you can out of sight while aggregating sensors of an important reality only you control) while also strong centralization forces (whole-stack efficiency-maxxing your way to inevitability).

For me, this exercise makes clear that naive applications of Specific Intelligence will not work. Most businesses do not have data only they observe and in the experiments here teach businesses nothing their competitors’ data could not. Building Specific Intelligence ‘just cuz Satya said so’ will likely not play out well unless you’ve a clear view of what value propositions your firm distinctly owns and how differentiated intelligence uniquely powers it.

Legal AI is unlikely to be a strong candidate for Specific Intelligence. Firm activity is observed by counter-parties and partners switch firms all the time. Customer service similarly may not be a strong candidate: strong-NPS-producing service methods likely generalize. If customer service is viewed as a mere cost center, efficiency-maxxing (v autarky-maxxing) may be the better choice.

In sum, a centralized router functions as a learning-loop business whenever the residual transferable uncertainty remains high or economies of scale for operating it dominate. These economies of scale will include model routing over a stream of new models and eventually chips. Specific Intelligence exists where a firm’s problems are genuinely its own, likely in scarce sensing regimes e.g., as operated by OpenEvidence, Waymo, Instagram, or Oura. Most firms do not operate in these regimes, and likely are better served by efficiency-maxxing v autarky-maxxing as Satya or Applied Compute might suggest. If your competitor’s problem stream could teach the same lessons your own might, applications of Specific Intelligence are a folly.

You’re just building a bad router.

[Send me a note on X if you’d like access to the simulation code.]