Email’s New P&L: From Cost-Per-Send to a Revenue, Data, Outcome and Media Surface (Part 3)

Email Earns from Outcomes

Who carries the risk

The first income line changes what an email can do. The second changes who is responsible when it does not work.

That is the bigger change and the more uncomfortable one, because every incumbent commercial model in marketing is built to avoid it. The ESP sells the send and is paid whether it converts. The martech platform sells the seat and is paid whether it is used. The agency sells the retainer and is paid whether the campaign lands. The ad platform sells the impression and is paid whether the customer was going to buy anyway. In every case the brand carries the entire outcome risk of a system it did not build and cannot fully see.

Progency inverts that. It is not a fifth product and not a services wrapper. It is an accountable operating layer that sits after the CRM and before the auction: it takes responsibility for a defined customer state, runs the interventions on the surface, measures against a control, and earns only on the lift it can prove.

Three mandates, three counterfactuals

Progency runs three mandates, one per attention state. Each is a bet against a different counterfactual — a different answer to the question what would this customer have done anyway?

Recover is for customers gone dark. The counterfactual is adtech: the money the brand would otherwise spend buying that person back through a platform which already holds their email address. This is the wedge, because it attacks the most expensive leak in the P&L and it is the one almost nobody else is fixing. Inside the email the sequence is connection first, then recovered attention, and only then conversion. A dark customer asked to buy on first contact simply stays dark.

Protect is for valuable customers whose attention is cooling. The counterfactual is drift: left alone they become lost, and a lost customer is later reacquired at several times the cost of having kept them. The email here is Digest and Relate, not Sell. The job is to arrest the slide while arresting it is still cheap.

Grow is for the attentive. The counterfactual is a slower next purchase and margin left on the table. The email is Sell and Notify made live — the right next thing, completed in the inbox, with the customer suppressed from paid retargeting because they are already reachable at no cost.

These are not marketing segments. They are attention states, and the point of measuring them is that a customer moves between them continuously and the correct intervention changes when they do.

One floor separation has to be held here, because it is exactly the kind of drift that quietly corrupts the vocabulary. The three mandates are not the four NeoMarketing zones — Retain, Finish, Recover, Acquire. Zones answer where the work sits in the operating architecture. Mandates answer what Progency is trying to change in the customer’s state. Finish, for example, is a zone: a KYC, a quote, a renewal or a purchase left incomplete. That unfinished job might sit inside Recover or inside Grow depending on whether the customer is dark or attentive. Recover appears on both floors and means different things on each. The vocabulary only works while the floors stay apart.

Figure 5 — Three mandates, one per attention state, each priced against a different counterfactual.

Beta, Alpha, Carry

The pricing follows from the mandates and is simple enough to survive a CFO meeting.

Beta is what would have happened anyway — the baseline, agreed in advance, measured rather than asserted. Alpha is the verified lift above that baseline. Carry is Progency’s share of the Alpha, and only the Alpha.

No lift, no fee. The old service-provider invoice was tied to usage. The old agency retainer was tied to activity. The adtech bill is tied to rented reach. This payout is tied to profit improvement and to nothing else.

The economic unit is worth stating precisely, because it is where outcome pricing usually cheats: the unit is not an email, an open, a click or an attributed conversion. It is the incremental completed outcome above the agreed baseline.

What Run leaves behind

There is a second thing Progency produces, and over time it is the more valuable one.

Every intervention writes a Decision Trace: the customer context, the eligible pool, the treatment chosen, the channel, the job the message was doing, the holdout status, the expected outcome, the cost, the actual outcome and the resulting state. Individually each is a row. Accumulated, they are something a competitor cannot buy or copy — a growing memory of decisions tied to their measured consequences.

This matters because the obvious moat is the wrong one. The intelligence that decides what to try — the models, the archetype priors, the agent confidence — will commoditise, and quickly. Everyone will have capable models. What does not commoditise is the corpus that says which decisions, on which customers, in which states, moved the outcome against a control. Progency earns revenue today and, in the same motion, builds the memory that makes tomorrow’s outcomes cheaper and more reliable to produce. Run is the point where the email business stops being a set of messages and starts being an underwritten system for moving customer states — one that gets better at it every time it is paid.

The holdout is not a detail. It is the apparatus.

Everything above is a claim about money, which means it stands or falls on the measurement. Three disciplines, none of them negotiable.

The holdout is concurrent, not historical. A comparison against last quarter measures the season, the pricing, the competitor’s campaign and the weather. A comparison against a randomly withheld group running at the same time measures the intervention. Historical baselines are how agencies have been claiming lift for thirty years, and they are why nobody believes the claims.

The holdout is enforced in the system, not in the contract. It is a hard gate in the automation layer: if the control group is not held, the campaign does not run. A discipline that depends on an operator remembering to apply it under quarter-end pressure is not a discipline.

Simulated judgement and measured Alpha never share a currency. The prediction layer decides what to send and to whom; it gets no vote on what is paid. The holdout decides what is paid; it has no opinion on what to send. Keeping the two separate is not conservatism. It is the difference between a measurement system and a marketing claim. A system must never promote itself using its own predictions as evidence.

Why this rung has to come before the next one

There is a structural argument for the ordering that is easy to miss, and it is the most important sentence in this essay.

The fourth income line — media — is the one that can destroy the asset. Ad load is a dial, and every dial in marketing is eventually turned up by somebody with a quarterly number to hit. The usual protections against this are policy documents and good intentions, and they fail.

An operator paid on carry cannot over-monetise the surface, because the carry depends on the attention surviving. If ad load rises to the point where opens fall, Recover misses, Protect misses, Grow misses, and Progency’s own income falls with them. The commercial model is the governor. It is not a promise not to spoil the surface; it is an arrangement in which spoiling the surface is immediately and directly expensive to the party holding the dial.

That is why Run comes before Network. The third income line is what makes the fourth one safe.

The honest constraint

One thing this model does not solve, named rather than buried: the throttle on outcome pricing is working capital, not demand.

An operator paid only on verified lift funds the interventions before it is paid for them, and the measurement window is counted in weeks or months. Appetite for a model where the brand pays only for improvement is not the scarce input. The scarce input is a balance sheet that can carry the gap between doing the work and proving it. Any plan that assumes otherwise will hit the wall at precisely the point where the model starts working.

Key points: (a) Capability revenue is paid for what an email can do. Outcome revenue is paid for what it caused. (b) Every incumbent model in marketing is arranged so the vendor is paid whether or not it worked. Correcting that is the whole of the second income line.

Thinks 2054

FT: “We need humanities more than ever in the age of AI…Subjects like literature, languages and history are vital to a thoughtful and prosperous society.”

WSJ: “Credit cards are now the sharpest competitive weapon airlines have in their never-ending fight for travelers. Plastic is increasingly essential to carriers’ bottom lines and central to nearly every decision they make, from where they build lounges to the routes they fly.”

Business Standard: “The Centre and states plan to replace $189 billion of imports by boosting domestic production of 1,272 products through targeted manufacturing and industrial clusters.”

NYTimes: “A.I. agents are tools, but you can learn a lot by prompting them for feedback. Try asking your A.I. agent how you frustrate it. I asked mine, and requested that it answer with as much human expression as possible. What did I hear? “You are so exhausting. Please, for the love of computing power, just tell me what you actually want.” Our team’s research seems to suggest this is what we should all do for our agents. Managing A.I. well, it turns out, looks a lot like managing people well.”

Andrey Mir: “Never before have so many people been simultaneously absorbed in issues that have nothing to do with their daily lives. This unnatural involvement with distant events draws us into shared concerns, creating a sense of community more typical of a small village, but now on a global scale. That’s how McLuhan arrived at his ideas of the Global Village and “retribalization.”…Social media took it much further by exposing people not just to the same news but directly to one another. That is real village life on a global scale: everyone watching, judging, and reacting to everyone else in real time.” [via Arnold Kling]

WSJ: “A record level of private-equity investments are stuck in funds limping along past their intended lifespans. Often known as zombie funds, these funds are no longer raising money or making new acquisitions, in part because fund managers haven’t been able to sell their remaining assets. The net asset value of U.S. private-equity assets stuck in funds at least a decade old reached an all-time high of $348.5 billion at the end of 2025, according to PitchBook data. That is 3.5 times the amount in 2015 and more than 100 times that of 2005.”

Email’s New P&L: From Cost-Per-Send to a Revenue, Data, Outcome and Media Surface (Part 2)

Email Earns from Actions

The click-through was never a step. It was the leak.

Consider what happens when an email works.

A customer opens it. Something in the message is relevant enough that they decide to act. At that instant — and it is an instant, not a state — intent is at the highest point it will reach in the entire life of that message. Then the email asks them to leave.

They tap. A browser opens. A page loads, or partly loads. There may be a login. There may be a basket that has forgotten what was in it. There may be an address form, a card form, a one-time password arriving in a different application. Somewhere in that sequence, most of them stop.

The industry has a name for the gap between the click and the conversion and treats it as a fact of nature. It is not. It is a design decision made when email could not do anything except link out, and it has survived long past the constraint that produced it.

Pay-in-Email removes the leaving. Its principle is one line: the transaction should end where the attention begins. The customer completes inside the message — paying by UPI or a stored card, confirming a renewal, booking a slot, accepting a quote, finishing a verification step, making a donation, taking an upgrade. For most repeat relationships the brand already knows the customer, the product, the address and the payment method; the action does not need another journey, it needs a confirmation. Intent is captured at its peak rather than at the end of a sequence designed to lose it.

What completes in an inbox

The list is longer than most people expect, because a great deal of commercial activity is not discovery at all. It is completion of something already decided.

  1. payment of a bill, an instalment or an invoice
  2. renewal of a subscription, a policy or a plan
  3. repeat purchase of a known item at a known price
  4. booking or rescheduling against live availability
  5. completion of an application already begun
  6. a document, consent or verification step
  7. an upgrade, an add-on or a top-up
  8. a donation or a pledge

None of these needs a shop. All of them currently require one.

There is a second saving hidden inside the first, and it is larger than it looks. A customer who completes in the inbox has just proved they are reachable for free. They should leave the paid retargeting pool the same minute. The brand saves the transaction leakage and the reacquisition spend in a single action. That is Never Pay Twice, operating at the level of one customer on one Tuesday.

The second output: declared data

The same surface that takes a payment can take an answer.

Tell-in-Email captures declared data at the moment the customer is already engaged: a preference, a size, a renewal date, an interest, a consent, a piece of feedback, a choice between two service options. One tap, inside the message, with no form and no landing page.

Declared data is a different substance from inferred data, and the difference compounds. Inferred data is probabilistic, decays quickly, and grows more expensive to obtain every year as the identity layer of the open web degrades further. Declared data is exact, durable, consented — and, the part that matters commercially, it is the input the entire personalisation side of the architecture runs on. In the TWIN framework it is the I: individual data, given rather than guessed.

Key points: Pay-in-Email earns revenue now. Tell-in-Email earns the data that makes the next decision smarter. They are two outputs of one action surface, which is why they sit on the same rung and belong in the same section of the statement.

This is the literal content of the phrase Email for Revenue & Data. It is not a slogan about email mattering. It is a description of what a single Living Email produces on every open: a transaction and a signal.

The law that keeps the rung honest

There is a trap here and it should be closed before it opens.

The same Living Email can be sold two ways. Sold as a capability: the brand operates it, the vendor supplies the surface, and the vendor is paid for what the email can do. Or operated as an outcome: Progency runs it against a concurrent holdout and is paid on the lift it can prove. Both are legitimate. Both appear in the equation. They are different rungs of the same ladder — Act and Run.

What is not legitimate is charging for both on the same work.

Key points: On the same audience and the same intervention, charge for the capability or participate in the outcome. Never both. If the vendor is paid for the surface and paid again for the result the surface produced, it has invoiced twice for one email. Never Pay Twice is a rule we hold ourselves to, not only a charge we lay against adtech.

The law also disciplines the sale. A commercial team forced to choose has to be clear about what it is selling — tooling or accountability — and the customer is able to tell which one they bought. Ambiguity here is where outcome pricing usually dies quietly: an operator paid a fee regardless has no exposure, and an operator with no exposure is an agency with a dashboard.

Figure 4 — One Living Email, two commercial models, and the wall between them.

The technical condition

None of this works on a static email.

A payment needs a live balance, a current price and an authorisation window. A booking needs live availability. A renewal needs a real date. A declared preference needs somewhere it will be written back and later read. All of that is state at the moment of opening, not state at the moment of sending.

That is the composed-at-open inflection, and it is why it is the highest de-risking priority in the venture. A static email is a prediction made at send. A Living Email is a decision made at open. Everything in this section depends on which of those two things the email is.

One honest constraint, stated plainly because it is a condition of the bet rather than an objection to it. Interactive rendering is not universally supported: it is Gmail, Yahoo Mail and Mail.ru, behind sender registration and a DMARC policy most brands have not yet set. Fallback-first design is therefore not a courtesy. For a large share of any real list the fallback is the primary experience. An email programme that only works in its interactive form is a demonstration, not a product.

Key points: (a) The click-through was never a step in the journey. It was the leak. (b) Capability revenue is not new demand. It is demand the brand already created and then handed to whoever owned the next page.

Thinks 2053

WSJ: “Sociology, the discipline tasked with pursuing the answers, would seem well poised to take these issues on. For decades, its analyses informed policy on issues ranging from criminal justice to immigration to urban development. It introduced groundbreaking concepts like “white-collar crime,” “gentrification,” “unintended consequences” and “cultural capital.” At various points, sociologists proposed and critiqued affirmative action, inspired and lambasted welfare reform, diagnosed and dismissed the breakdown of the nuclear family. Yet today, sociologists are embroiled in an ideological and existential battle of their own—one that is increasingly public and at least in part of their own making. The study of people, in other words, has a people problem.”

Fareed Zakaria: “A revolution is unfolding. Voters everywhere are furious…Free societies often look weakest in the midst of great technological upheaval because they move slowly and argue constantly. Yet over time they have usually adapted better than autocracies precisely because they can learn, correct mistakes and change course without violence or repression. Democratic leaders should not try to outrace technology; they can’t. But they should prove, through relentless focus and competence, that democracies can deliver. Then people will understand that the friction of accountability and the weight of legitimacy are not bugs in the system. They are the features that keep us free.”

NYTimes: “Given its patchwork construction and graphic patterns, a soccer ball can have two types of symmetry: reflection symmetry and rotation symmetry. Reflection symmetry means what it sounds like: that the object “has two identical halves that are mirror images of each other,” in Dr. Schattschneider’s words. Rotation symmetry means an object looks exactly the same as its starting position after a rotation of less than 360 degrees. For instance, a square looks identical four times when rotated every 90 degrees about its center. A perfect cube has a total of 24 such rotational symmetries — rotating around various axes piercing face-to-face, corner-to-corner, and edge-to-edge. The Platonic solids — the tetrahedron, cube, octahedron, dodecahedron and icosahedron, from left, below — are the foundation of mathematical symmetry, especially in spheres.”

SaaStr: “Pull the last 100+ emails your team sent to your best-fit accounts. Skip the ones that closed. Read the ones that never got a reply. If they read like the two I got, the problem isn’t the market, the pricing, or the competition. The problem is a team spending your most expensive asset, qualified attention, on emails any AI would flag as forgettable. Those are winnable deals, lost at the cheapest and most fixable step in the funnel. You don’t need more leads to fix it. You need the team to treat the leads already raising their hands like they’re worth ten minutes.”

Email’s New P&L: From Cost-Per-Send to a Revenue, Data, Outcome and Media Surface (Part 1)

Every other channel in marketing is priced on what it produces. Email alone is still priced on what it consumes. This essay makes one new claim: the three new forms of email income belong on a single statement, that statement can be audited, and the order of its terms is the order in which the business has to be built.

**

The Inversion

The line item nobody argues about

Somewhere in every consumer brand’s monthly close there is a line for email. It is small. It is stable. It sits near the bottom of the marketing block, beneath a media spend thirty or fifty times its size, and it has looked roughly the same for fifteen years: a per-thousand rate multiplied by a volume, plus a platform fee.

Nobody argues about it, and that is the problem. The line is uncontested because it is understood to be a cost, and the only question a cost invites is whether it could be slightly smaller. Every procurement cycle asks that question. Every vendor answers it the same way, by shaving the rate. And so the most valuable relationship asset a brand owns — a list of people who once gave it permission to write to them directly, with no algorithm standing in between — is managed like a utility bill.

The rest of the marketing block moved on years ago. Search is priced on a click. Social is priced on an impression that increasingly resolves to a conversion. Affiliate is priced on a sale. Retail media is priced on attributed revenue. Every one of those channels, whatever else is wrong with it, is priced against something the channel produced. Email is priced against something the channel consumed: a send, a message, a unit of contact.

That asymmetry is not a pricing quirk. It is the reason email has been strategically neglected for a decade.

What the old statement rewards

Follow the money through the old model and the incentive structure is uncomfortable.

The brand pays per send, so the vendor’s revenue rises with volume. Volume, past a threshold that varies by category but always exists, destroys the attention the list is made of. Opens fall, clicks fall faster, complaints rise, and the engaged base shrinks. The asset degrades — and the supplier’s revenue goes up while it degrades.

Meanwhile the value the email created is captured somewhere else. The email produces the intent. The website closes it. The payment provider clips it. And when the customer does not close, the retargeting platform is paid to chase a person the brand had already reached, for free, an hour earlier. The email did the work; four other parties booked the revenue.

Key point: The old email statement has one line and it points the wrong way. The supplier profits from the behaviour that destroys the asset, and the value the channel creates is booked by somebody else.

The escape is not a better email. It is a different question.

It is worth being precise about what the escape is not, because the obvious answer is wrong and expensive.

The escape is not more interactive email. Interactivity is a capability, not a business model. A calculator, a form, a wheel, a poll or a checkout placed inside a message makes the artefact modern while leaving the economics exactly where they were, because all of it can still be sold as one more custom campaign priced by the send. A brand can buy the most advanced email in the market and still be paying for volume. The decisive question was never what the email contains. It is what the provider is paid for.

That question has an answer, and it is a ladder rather than a feature list. EARN is not a product roadmap; it is four rungs of rising accountability, with the same surface underneath all of them.

  • Email — deliver a reliable, identity-linked owned surface.
  • Act — let the customer complete useful actions and declare useful data inside that surface.
  • Run — operate the surface for measurable incremental outcomes.
  • Network — turn trusted attention into carefully governed media and cooperative acquisition.

The reason the ladder matters is not that it sorts the products tidily. It is that each rung changes who values the provider. Procurement prices infrastructure and benchmarks it downward — that is the Email rung’s ceiling and the trap the whole category is stuck in. Marketing values a capability, which is the Act rung. The CFO trusts a measured outcome, which is the Run rung. Advertisers and partners value scalable attention, which is Network. The ascent is not a feature upgrade. It is a change in who signs the cheque and what they think they are buying.

Figure 1 — The EARN ladder. One surface underneath, rising accountability, and a different buyer at every rung.

Four lines instead of one

The ladder resolves into the statement. The correction is not a better rate; it is a better statement.

An email programme that has been rebuilt properly does not have one line. It has four. The first is the one that exists today — the cost of delivering the message and producing what is inside it. The other three are new, and each corresponds to a different thing an email can now be paid for.

Capability revenue is what the email earns when the customer completes something inside it: a payment, a renewal, a booking, a declared preference. This is revenue that already existed and leaked on the way to a landing page.

Outcome revenue is what the email earns when somebody takes accountability for a customer state and is paid on the measured improvement. This is not revenue from the email; it is revenue from what the email caused, proved against a control.

Media revenue is what the email earns when the attention it has rebuilt becomes inventory another advertiser will pay to reach. This is the brand being paid for attention rather than paying a platform for it.

Set the three against the cost and there is an equation.

Net Email Cost  =  Delivery and Content Cost  –  Capability Revenue  –  Outcome Revenue  –  Media Revenue

When the result reaches zero, the programme has achieved ZeroCPM. When it goes below zero, email has stopped being a cost centre and become a profit centre.

Figure 2 — The two statements. The old line points outward; the new one has three inward terms.

Why an equation and not a manifesto

The equation is the most useful thing in this essay, and it is worth being explicit about why.

It is auditable. A doctrine cannot be checked at the month end; a line can. Each of the three revenue terms resolves to a number somebody in finance can trace to a transaction, a contract or an invoice. This matters more than it sounds. The single greatest obstacle to any of this being adopted is that it usually arrives as narrative, and narrative is not something a CFO can approve.

It is sequenced. The terms are not interchangeable. Capability revenue can be earned by a brand with a modest list and a decent product, this quarter. Outcome revenue requires an operator, a baseline and a holdout. Media revenue requires a surface people voluntarily open, repeatedly, which is the hardest item on the list and takes the longest to build. The order of the terms is the order of the build.

It is falsifiable. If the three revenue terms do not, in practice, add up to the first, the equation says so plainly and the thesis is wrong. That is a feature. Most marketing frameworks are constructed so that they cannot fail.

What the equation forbids

Two things, and both are worth stating before anything else is attempted.

The first: you cannot reach zero by shrinking the first term. Every ESP in the market competes on delivery cost. It is the most contested, most commoditised, lowest-margin number in the category, and it is also the smallest number in the equation. A brand that negotiates its send rate down by twenty per cent has moved the least movable term by the least meaningful amount. ZeroCPM cannot be bought from a vendor. It is what happens when the other three terms are built.

The second: you cannot start at the media term. This is the more expensive mistake, because it looks like the fastest route. A brand with a large list and a monetisation target can put sponsored units into its emails next week and book media revenue immediately. It can also, in the same quarter, destroy the only thing that made the inventory worth anything. A list is not an audience. Attention that has not been earned cannot be sold at all, let alone twice.

That constraint has a name in this corpus and it has not changed: earn attention first, monetise it later, network it last. The equation restates it as arithmetic. The media term is last because it depends on the other two having been built first.

Figure 3 — The three revenue terms taking the cost line down to zero, and past it.

Key points: ZeroCPM is not a pricing promise. It is the structural outcome of the reversal — the scoreboard, not the product. A brand that has not built the inversion cannot buy ZeroCPM from a vendor. A brand that has built it cannot avoid ZeroCPM as the consequence of revenue exceeding send cost. One statement, four lines, and only one of them is a cost.

Thinks 2052

Tyler Cowen: “An AI maniac is someone who is obsessed with working with the latest AI models. They try out new models as soon as they can, they spend hours and hours trying to master them, and they use them to regulate both their workflows and their personal lives. I know one person who has his AI agent text him if he is not drinking enough water, for which he’s placed cameras around his house. One online anecdote tells of a man who canceled a date to spend more time playing around with Claude Fable 5 after Anthropic (where I am a member of the economic advisory board) extended the model’s availability for a few days. Many AI maniacs are using AI tools to start companies of smaller size, and thus of smaller expense, than ever before. For those companies, the humans must set in motion and then monitor a large number of AI tools and agents. Those individuals then stand to reap outsize profits as their companies grow and succeed. Stripe, the payments company, recently issued customer data showing that the number of single-person companies earning $10 million or more has doubled in the past two years. There is no firm estimate how much of that improvement is due to AI, but it stands to reason that AI is a main driver of the trend…”

NYTimes: “Workers often focus solely on their retirement savings balance. But figuring out how you want to spend the rest of your life could be job No. 1.”

WSJ: “To record or not to record? That is no longer a question for people in tech. The answer is a resounding yes. A Zoom call isn’t complete without an artificial-intelligence note taker. Phones are out at meetings, capturing every word. During impromptu conversations with co-workers, someone might turn on the Granola transcription app, which can turn the interactions into one-page summaries or a list of action items. Even at bars and on dates, people are using AI-infused listening apps to analyze conversations later on. Gone are the days of any interaction being private by default. Asking for permission is out, even though in some situations, that could mean violating the law. It isn’t creepy if it’s in the name of productivity, say the people doing the recording.”

Jeff Sommer: “I have no doubt that artificial intelligence is an important technology. Great fortunes are already being made. But I’m also certain that there will be many losers, as there were in two other episodes of mammoth infrastructure investments in budding technologies: the railroads in the 19th century and the various early internet companies of the dot-com era. Well-run, diversified and deep-pocketed companies have a better chance of survival in epochs like these than those that take on inordinate risk with their capital investments. Even so, the future champions may not be any of the early giants. A great winnowing is coming, and prudent investors will accept that they cannot know in advance who the winners and losers will be.”

The Foundry Price: The Number Every Software Vendor Will Soon Have to Explain

Published August 14, 2026

This series has moved from the supply side to the demand side to the factory floor. The Software Foundry named the revolution: AI changes the production function of software. The App-Stack Tax named the waste on the buyer’s bill. Inside the Foundry walked the production system that removes it. But one part of the argument has stayed deliberately unfinished: if software is produced differently, what should it cost? A production revolution becomes economically important only when its gains reach the buyer. A faster factory that preserves the old price has improved the producer’s margin; it has not created an affordability revolution. The revolution becomes visible when the invoice changes — and, more importantly, when the changed invoice becomes the number every other producer must explain.

This essay names that number. And it is a number, not a product — which is what makes it the most contagious idea in the series.

The claim of this essay: markets do not change when one producer becomes cheaper. They change when buyers begin treating the cheaper price as normal. Manufacturing learned this as “the China price.” Software is about to learn it as the foundry price — and the most important property of a reference price is that it does its work in negotiations it never attends.

1

The Price Everyone Knows Before the Quote Arrives

The renewal arrives on a Tuesday afternoon, and it is rarely dramatic. The account manager has prepared the explanation before the buyer asks: the customer crossed a contact threshold; three more employees need access; a capability that used to sit inside the package has moved into a premium module; the annual increase is in line with the market. A discount is offered pre-emptively. The final number is close enough to last year’s to feel inevitable and large enough to require a meeting — and the meeting begins too late in the argument. Procurement debates the percentage. The business owner debates which modules to trim. Finance asks about contract length. Everyone negotiates around the quote, because everyone has accepted the invisible number beneath it: what software in this category is *supposed* to cost. The most powerful price in any market is not the amount printed on the invoice. It is the amount the buyer has stopped questioning.

Reference prices form slowly and then govern almost everything. A washing machine made in a high-cost economy is not judged against its own labour and materials; it is judged against the price the global factory system created. A technology project proposed entirely onshore is judged against the offshore rate even when no offshore supplier is in the room. A medicine whose exclusivity has ended is judged against the generic price however eloquently the branded manufacturer explains its history. In the early 2000s, Western manufacturing boardrooms learned the purest version of the pattern: “the China price” was not a quote from any particular factory — it was a number everyone knew before any quote arrived. And notice what it did not require. It did not require the buyer to move production to China; most never did. It did not require the Chinese producer to bid; it usually hadn’t. The number worked at a distance: once a credible alternative existed at a visible price, every incumbent quote had to be *justified* rather than merely *renewed*. Once the reference moves, the incumbent no longer describes its own economics in isolation. It must explain the spread.

The number does its work before any deal is signed.

Business software is the great exception — the largest spending category in the modern company with no reference price at all. Ask a finance director the market rate for a marketing suite, a service desk, or a workflow platform for two hundred people, and there is no number to give: only what the incumbents charge, which is a different thing entirely. The anchor in every software negotiation today is the vendor’s own last invoice; “market rate” means “last year, plus the uplift.” Discounting is theatre performed against a list-price fiction nobody has ever paid. What buyers have instead is a substitute that answers the wrong question — analyst quadrants and review sites that rank vendors by features and vision: comparisons of *what*, never of *what it should cost*. In no other major category does the seller supply both the price and the standard the price is judged against.

One clarification before the mechanism, because software has seen many low prices and very few price revolutions. Open source lowered the licence and moved integration and operation elsewhere. Freemium lowered the entry price and recovered the economics through upgrades. Venture-subsidised challengers entered below cost and raised prices once dependence formed. Those were changes in *commercial strategy* — and buyers learned, correctly, to distrust them. A reference price must be sustainable: it must remain true when the second product ships, the hundredth customer onboards, a connector fails at midnight and the support queue fills on a Monday morning. The foundry price qualifies for one reason only — it is a change in production possibility. The useful product can be made, operated and supported for less because less of it must be invented, repeated and manually carried each time. That is the difference between a cheap bid and a new normal.

2

How Software Escaped Its Cost Curve

To feel why the foundry price will land so hard, trace how software’s price became detached from its production curve — and begin with fairness, because the story is not one of villainy. The cloud bargain was real and overwhelmingly good. Before software-as-a-service, a customer bought licences, ran infrastructure, carried versions, hired specialists and absorbed long deployments; central operation removed most of that burden, and one product could improve continuously for every customer at once. The industry deserved a large share of that dividend. The question is what happened to the rest of it.

What happened was a doctrine — the most successful pricing doctrine in modern business, and it had a name: value pricing. Price the software not on what it costs to make and run, but on what it is worth to the customer. Reasonable on its face; software’s worth is real and often enormous. But notice what the doctrine quietly abolished: any relationship between the number on the invoice and the cost of the thing invoiced. Once price anchors to “value” — unmeasurable, negotiable, asserted by the seller — there is nothing for it to be compared *against*. Value pricing did not merely raise software prices. It removed the standard by which they could ever be questioned. The pricing grammar that followed acquired a peculiar direction: seats, contacts, events, orders, usage bands, editions — almost every sign of the customer’s success became a reason for the invoice to rise. The previous essay watched the merchant’s stack charge her for succeeding; the same mechanism lives inside the individual application, where the bill expands with the customer’s activity while the product performs essentially the same job on the same centrally operated system. The price became less a reflection of the work the software does and more a claim on the growth the customer creates.

The industry’s scoreboards then locked it in. Annual recurring revenue measured the base; net revenue retention measured how reliably the installed base paid more; customer-success organisations were built not only to preserve use but to locate expansion; packaging became a commercial architecture — enough value in the lower edition to enter, enough friction to make the next edition inevitable. None of it was foolish. It was the rational optimisation of a business whose delivery cost had collapsed and whose investors rewarded predictable expansion. The result was the strangest chart in modern economics: the cost of serving a customer falling towards zero while the price of being one rose every year — a gap with a respectable name, gross margin, and a public-market religion, NRR, that made widening it mandatory.

The cost of delivery fell towards zero. The price rose. Nothing forced them back together.

Why was this stable for twenty years? Because the cloud industrialised only half the problem. It industrialised *delivery* — and left *creation* artisanal. Every serious application still required a rare organisation: a team to discover the category, design the product, build identity and permissions, implement workflows, manage connectors, test releases, onboard customers and staff support. The marginal cost of serving another customer fell; the fixed cost of *becoming and remaining a software company* stayed high — and the successful vendor priced to protect the organisation built around that cost. The first essay named the deeper consequence in a different costume: the cost of creating software was the economic patent, and the patent protected more than the product. It protected the price. For a rival number to exist, someone had to build a credible alternative, and nobody could afford to — so no alternative price ever became visible, and the incumbent’s invoice remained the only number in the room.

The foundry attacks the neglected half. Specifications as control surfaces, proven components as tooling, inspection as software, operation AI-native — and the machinery compounding across products, so the second application does not have to buy a second company. When both curves fall — creation and delivery — the old reference price loses its final production defence, and the trillion-dollar repricing of software equities in early 2026 reads correctly at last: not a judgement about chatbots, but the market pricing in the death of a pricing regime — the end of the twenty-year vacuum in which no number ever questioned the invoice. The cloud made software recurring. The foundry makes its old price contestable.

3

What the Foundry Price Is — and How It Spreads

The phrase will be misunderstood unless its boundary is drawn sharply, so draw it from both sides. The foundry price is not the price of code. Code is becoming abundant, but customers do not buy code; they buy a dependable job completed over time — records that stay correct, permissions that hold, changes that do not break yesterday, support when reality finds the edge case. A repository can be generated in an afternoon. A product is a promise that survives contact with the world. So the foundry price *includes* the full cost of that promise: secure identity, tested behaviour, controlled releases, observability, backup and recovery, maintained connectors, documentation that stays aligned, AI-native support with human judgement for consequential exceptions. What it *removes* is everything the third essay showed being manufactured away: duplicated engineering — the same identity layer, workflow engine and connector framework rebuilt for every application; unused complexity — the accumulated features carried by customers who need the common job done well; and avoidable operating labour — the manual onboarding, the support agent asking for information the system already holds. The double Pareto cut, reaching the invoice: one cut removes feature waste, the other removes foundation waste, and the foundry price is what remains when both cuts land.

The definition: the foundry price is the lowest sustainable price for dependable, focused utility produced on reusable machinery. Lowest forces production discipline. Sustainable excludes subsidised theatre. Dependable preserves the trust floor. Focused rejects the bloated suite. Reusable machinery names the economic source. Remove any word and the phrase collapses into ordinary discounting.

Two honesty boundaries keep the definition serious. First, the foundry price is a licence price, not a total-cost guarantee: migration, integration, training and governance remain real and vary by customer; a foundry reduces them — the migration factory exists for exactly this — but it measures and publishes them separately, because a reference price that overclaims dies at its first audit, and the China price never claimed to include shipping. Second, not every category converges to the same ratio. Payments, systems of record, networks with deep liquidity, products built on irreplaceable data, software that underwrites consequential outcomes — that world holds value beyond repeated engineering, can stay expensive, and deserves to. The foundry price attacks the large exposed middle: mature code, understood workflows, accumulated features, and a premium protected mainly by the fear of leaving. In practice: arithmetic lands it at roughly one-tenth to one-fifth of the incumbent licence for the equivalent daily-use utility — with honest margin inside it. Price is the visible output. Production discipline is the warranty.

Now the mechanism — how a number becomes a market force. Three ingredients, and only three. A credible producer: at least one foundry delivering the useful core reliably in a category — not a demo; a running product with customers who stayed. A visible price: published, simple, self-serve — a number a finance director can screenshot into a board pack. A safe exit: migration measured in days and priced in advance, because a reference price backed by an unpriceable switching cost is a bluff, and procurement can smell a bluff. Assemble the three and the propagation begins:

Three ingredients — then the number spreads to negotiations it never attends.

The foundry price does its damage without a single customer switching. It enters renewals as the question the vendor must answer, board packs as the line that makes the current bill look strange, procurement as the benchmark. A market changes not when one producer becomes cheaper — it changes when buyers begin treating the cheaper price as normal.

And it will move faster than its manufacturing ancestor. The China price spread at the speed of trade shows, supplier visits and shipping lanes — a decade to become a boardroom fixture. A software reference price spreads at the speed of a pricing page: the moment one credible foundry publishes its number in one category, every buyer in that category can see it the same afternoon, and every renewal that quarter arrives with the number already in the room. The propagation infrastructure — comparison sites, procurement platforms, analyst notes, one viral screenshot — already exists and is bored. What manufacturing needed a decade to normalise, software can normalise in a renewal cycle or two.

4

The Incumbent’s Six Binds

The obvious question follows: why don’t the incumbents simply match it? They have the engineers, the customers, the data, the brand — and now the same AI models. They will certainly adopt the models: generate more code, automate support, ship simpler interfaces, launch AI features weekly. The mistake is to assume that access to the same technology creates access to the same economics. A new producer asks one question: what price can our production system sustain? The incumbent must ask a harder one: what happens to everything we already are if we admit that price is sustainable? Six binds, and they interlock.

Every response costs the incumbent something it cannot afford.

The valuation bind is the deepest, and a worked example shows why it travels backwards. Suppose the category leader sells at $100 and a foundry offers the common utility at $20. The leader can launch a $20 edition — but every existing customer using only the common utility now has a question, and the revenue risk is not confined to new deals; it propagates through the installed base, which is where the share price lives. Net revenue retention — the promise that existing customers pay more every year — collapses long before any challenger does. The cost bind: the organisation was built to be funded by the old price — enterprise sales, solution consulting, success hierarchies, twenty years of accumulated estate that must be operated precisely because it justifies the tiers; the price cannot fall without the organisation falling with it, which is why incumbents add AI to the product faster than they remove labour from the customer journey. Features are easier to change than organisations. The channel bind: a commissioned sales motion structurally cannot carry a product priced to need no salespeople; the people who would sell the new price are the people it makes redundant, and they know it. The completeness bind: the feature list is the pricing architecture — a focused, honest, cheap edition indicts the suite it sits beside, after years of describing breadth as value. The signal bind: a price cut is a confession that the old price was padding, and it invites the simultaneous renegotiation of the entire book of business — the one event a subscription company cannot survive. And the timing bind closes the cage: respond early and legitimise a challenger nobody had heard of; respond late and the reference price is already normal, at which point matching it merely confirms it.

The binds interlock, which is what makes them a trap rather than a list: escaping one tightens the others. Cut the price and the signal bind detonates; launch the honest edition and the completeness bind indicts the suite; build the separate low-cost brand — the one genuine escape — and the cost and channel binds fight it from inside the building, because the new unit’s success is, by construction, the old organisation’s obituary. None of this makes incumbents helpless: their trust, distribution, data and installed workflows are real, and some will navigate the transition — usually by becoming foundries themselves under separate brands, and usually only after the third bad quarter, because the innovator’s dilemma was never about ignorance; it was about permission. But the asymmetry is now precise. A lower tier is a product decision. A new reference price is a business-model decision. Features invite a roadmap response; price forces an identity response. The incumbent can copy the feature. It cannot easily copy the economics — because its economics are the thing its investors bought.

5

The Renewal Playbook

Everything above becomes practical at one moment: the renewal. So this part changes audience. It is written for the buyer — the founder, the finance director, the operations head with a contract expiring this quarter — and its advice fits in one sentence: negotiate as though the foundry price already exists. In some categories it already does; in the rest, its arrival is a matter of quarters. Seven questions, in order.

Seven questions to ask before signing.

One: which capabilities did we use every week this year? Not which features were enabled or demonstrated — which jobs would cause real pain if removed tomorrow. Pull the usage report; the vendor has it, and reluctance to share it is itself an answer. The answer is usually shorter than the contract — and know the difference between optionality you value and complexity you merely carry: paying for a fire extinguisher is rational; paying for an entire fire brigade inside every room is not. Two: what is the price connected to? If the honest answer is “your headcount and your contacts” rather than “our cost to serve you,” then your growth is being taxed — separate genuine variable cost from a convenient staircase; a contact sitting unused in a database does not cost the vendor what the tier jump charges for it. Three: what does this year’s uplift buy that last year’s did not? An uplift justified by a roadmap you never requested is a habit, not a price. Four: has anyone priced leaving? In most companies the switching cost is a fear, not a figure. Get a migration quote even with no intention of migrating — the moment leaving has a price, staying has a negotiation. Five: what value sits beyond the code? This question protects you from simplistic price aggression: network effects, irreplaceable data, regulatory trust, underwritten outcomes can justify a real premium. The test that separates them from padding: what would remain defensible if migration became safe, common and reversible? Dependence is not the same as value, though it often stands beside it. Six: will the vendor price the used core, alone? The revealing question — a vendor who refuses to quote the fifteen per cent you use is telling you what the other eighty-five per cent is for. Seven, the anchor: if a focused alternative existed at one-tenth the price, what would we pay to stay? Answer it internally, in a number, before the meeting — then negotiate from that number rather than from last year’s invoice. The vendor’s anchor is history. Yours should be the future.

The output of the seven questions is a one-page renewal memo with four numbers on it: the used core, the growth tax, the exit price, and the anchor from question seven. Four numbers, one page — and the meeting is a different meeting, because for the first time both sides arrive with a standard of comparison. Two disciplines keep it honest: this is not advice to buy the cheapest thing — the series’ standard is trusted affordability, and a good incumbent will have answers where a vulnerable one has packaging; and the playbook serves something beyond any one bill. Every buyer who makes the vendor explain the price against the used core is helping construct the reference itself, the way every manufacturer who asked about the China price helped make it a fact. Reference prices are not announced. They are asked into existence, one renewal at a time.

6

When the Price Falls, the Market Expands

Price revolutions are usually described as wars over an existing pool of customers: the entrant undercuts, the incumbent bleeds share, industry revenue shrinks. That is the first-order view, and it is the least interesting thing that happens. The deeper pattern, every time, is expansion. China’s factories did not merely move the same purchases to cheaper suppliers; lower prices put appliances and electronics within reach of hundreds of millions who had never been customers. India’s delivery machine did not only replace onshore projects; it made technology work viable that would otherwise have been deferred forever. Generic medicines did not change the supplier; they changed who could be treated. The market after every reference-price reset was larger than the market before it.

Software is unusually ready for the same effect, because its non-consumption is invisible. The clinic coordinating through spreadsheets, the school running on messaging groups, the manufacturer whose workflow lives in one employee’s memory, the merchant whose customer database is a phone’s contact list — none of them appears in any vendor’s lost-deal report. They did not choose a competitor. They were never candidates at the old price and the old operating burden. The foundry changes all four terms that excluded them at once: creation becomes cheap (the machinery is reused), distribution stays near-zero (the cloud solved it), onboarding and support become AI-native (no consultant required to begin), and narrow categories become viable (no product must fund a complete software company). The addressable market expands downwards in company size and outwards into languages, local practices and specialised jobs.

The market after the reset is larger than the market before it.

The arithmetic that looks alarming is the arithmetic that matters. A product at one-fifth the price needs five times the customers for the same revenue — and conventional analysis stops there, which is why conventional companies will not do this. Foundry analysis continues: can the production system serve ten times the customers at far below one-fifth the total cost, when onboarding, support and operation scale with machinery rather than headcount — and how much of that volume is revenue that did not exist before? Add the portfolio effect from the third essay: a narrow workflow for one profession is too small to justify a standalone company and entirely viable as the seventh product in a foundry, where localisation and category logic sit on already-paid-for tooling. Lower price and larger market build a larger company, not a smaller one — which is the difference between discounting and abundance. Discounting asks how much less a seller will accept for yesterday’s product. Abundance asks what becomes possible when the cost structure itself changes.

One warning closes the argument, because the trap ahead is the industry’s oldest. Successful foundries will feel the pull of the old climb: win with affordability, add features, move upmarket, build the sales and services organisation — and discover, a decade on, that the reference price they disrupted has quietly reassembled itself inside them. The discipline must survive success: products thin, machinery strong, usage transparent, premium reserved for value that truly sits beyond code. The foundry price is not a launch tactic. It is a constitutional constraint. And when the constraint holds, the most important customer is not the enterprise that saves eighty dollars. It is the small business that can finally spend twenty; the specialist whose narrow workflow finally supports a product; the merchant who moves from messages and memory to a dependable system. The affordability dividend is not that today’s buyers pay less. It is that tomorrow’s buyers finally exist.

Closing: When the Revolution Reaches the Invoice

 The series can now be stated in one line: a production revolution (the foundry) removes a category of waste (the app-stack tax) through a machine (specifications, tooling, industrialised quality) — and transmits itself to the entire market through a number. The number is the final piece and the most contagious one: products must be adopted one customer at a time, but a reference price, once credible and visible, changes the behaviour of buyers who never adopt anything. It will not arrive everywhere at once; some categories hold value beyond code and will keep their premium with justification; some challengers will fail by confusing generated code with a dependable product; the reference will form through evidence — products that stay reliable, migrations that become safe, portfolios in which every product makes the next cheaper. Including, as the previous essay promised, evidence of our own, published favourable and unfavourable alike.

But once the evidence accumulates, the market’s question changes permanently. Buyers stop asking whether the new product is suspiciously cheap and start asking why the old one is inexplicably expensive. Vendors will not have to lose a deal to feel it; they will simply, one renewal at a time, have to explain a number they never had to explain before. The quote will no longer begin the negotiation. The reference price will.

**

The old software price was built on scarcity: scarce engineers, repeated foundations, large operating teams. The foundry price is built on abundance — abundant code disciplined by specifications, reusable machinery, industrialised quality, and human judgement concentrated where consequence demands it.

A market changes not when one producer becomes cheaper, but when buyers begin treating the cheaper price as normal.

The software foundry changes how software is made. The foundry price is what happens when the revolution reaches the invoice.

Thinks 2051

NYTimes: “Americans filed 5.7 million applications last year to start new businesses, according to the Census Bureau, the most in the two decades the government has kept track. New business applications through the first half of this year continued to climb. The strong run of business creation is one of the most surprising and welcome economic developments of the post-pandemic era. New businesses help drive innovation and productivity growth. Although many fail or remain small, some could develop into giants that spur job growth for years to come. “The sustained high rate of both main street and growth-oriented entrepreneurship over the past five years is a piece of super good news about the future of the economy,” said Scott Stern, an economist at the Massachusetts Institute of Technology.”

Noam Brown: “2023: LLMs struggle with 4th grade word problems 2024: LLMs can do high school math 2025: LLMs get a gold medal at the IMO Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn’t even news. Where will we be next year?”

Sandeep Goyal: “For years, marketers believed the key to strong branding was simple: Tell better stories. Storytelling helped brands build emotional connections with audiences. But in today’s digital world, attention is limited and competition intense. Customers don’t just want stories anymore. They respond to stories that help them make decisions. In fact, better decisions. This is where storyselling is becoming more and more powerful. Storytelling entertains audiences. Storyselling motivates action. Brands that succeed today are not simply sharing narratives. They are building stories that guide customers towards solutions, clarity, and measurable results.”

WSJ: “What many of the new winners have in common: they’ve embraced a new model of investing. Today, high-growth companies remain private for much longer—often well over a decade—locking the public out of key, wealth-creation phases of their growth. Many of these investors are investing patiently in private companies and writing check after check, helping the companies scale and establish an edge over rivals. Paying up for stakes in the hottest startups is a riskier strategy than venture capital’s traditional approach, one that can lead to disappointment if companies fail to live up to lofty expectations. Today’s venture capitalists often don’t have the influence to guide or shape companies the way they once did because of heightened competition. Instead, many are willing to invest in companies later on, sometimes years after they were started, betting that there’s more growth ahead.”

Inside the Foundry: The Machine That Makes the Machines

Published August 13, 2026

The Software Foundry named the revolution: AI has changed the production function of software, and a third affordability revolution follows. The App-Stack Tax named what the revolution abolishes: the repeated foundations every business pays for again and again, and the architecture — one core, many products — that ends the tax. Both essays lean on a claim I have so far asserted rather than opened: that there is such a thing as a software production system, as distinct from software development done quickly.

This essay opens the doors and walks the floor. It is written for a particular reader: the builder — the person deciding what to do with the most powerful construction tools ever handed to our industry. Because the awkward truth of this moment is that those tools are available to everyone, and most of what gets built with them will not be a foundry. It will be a faster workshop.

The thesis: AI-generated code is not a software foundry. A foundry is a production system that turns precise intent into dependable software, repeatedly — using specifications as control surfaces, reusable components as tooling, automated quality as inspection, and humans as accountable designers of the system. The model is the ore. The system is the mill.

1

The Workshop and the Foundry

Picture two small teams. Both are talented. Both use the same frontier AI models — the same coding agents, the same context windows, the same per-token prices. Both can produce an impressive working prototype in days. Watch them for six months, and they become different species.

The first team runs an AI workshop. Every product starts from a prompt and a conversation. The agents generate remarkable volumes of code, fast. But each product begins again: a new conversation creates a new repository; the agent chooses a slightly different architecture because the prompt was slightly different; identity is rebuilt, permissions are re-interpreted, logging arrives late, and the tests describe the happy path. The decisions that shaped the product live in chat histories nobody rereads. The team celebrates speed — and speed is real — while debt accumulates at machine pace, because the agents that write code quickly write inconsistency just as quickly. The workshop’s tell is what happens on product two: it starts from a prompt again. Nothing carried over except the people.

The second team runs a foundry. Intent enters as a specification precise enough for machines to execute and other machines to verify. The agents build — but they are instructed to search the existing inventory before creating anything new, they build *from proven components*, and what they generate must pass common gates before any customer sees it. Failures are replayable; operating knowledge flows back into the system; and when product two begins, most of it already exists. The foundry’s tell is the inverse of the workshop’s: the second product is where it gets faster.

The model is identical. The system around it is not.

Here is what makes the fork treacherous: at the demo, the two teams are indistinguishable. Both show a working product; both credit the same models; both are telling the truth. The demo measures the model, and the model is the same. Only repetition measures the system — which means the most important property of a software company is now invisible at exactly the moment everyone is looking. Investors, customers and the teams themselves will spend the next two years confusing impressive first products with production systems, because the evidence that separates them does not exist until product two.

This is the distinction the current discourse keeps missing. The public argument about AI and software is stuck on the model — which model codes best, how many tokens, what percentage of commits. But the model is available to everyone, which means the model is precisely the thing that cannot be the advantage. In the language of the first essay: the generated code is the ore. What separates the two teams is the mill — and the mill has a one-line definition: the difference between a workshop and a foundry is not how much code AI writes; it is how little of the production system must be reinvented for the next product.

A century ago, industrialists learned this same lesson about physical production: the machine tool mattered less than the system of jigs, gauges, tolerances and process discipline built around it. The factories that treated the new machines as faster craftsmen stayed workshops with better tools. The ones that redesigned production itself became something new. Software now stands at the same fork.

Standing at it, a foundry must resist two opposite temptations. The first is to treat every new request as a custom project — that builds an AI services company, fast and unrepeatable. The second is to build the universal platform before shipping a real product — that builds an internal infrastructure programme in search of a customer. The correct sequence is narrower than either: build one real thing, instrument how it was built and operated, then extract only the machinery that proves reusable. A foundry is not designed in advance. It is discovered through disciplined repetition — and the rest of this essay is a tour of the four rooms where the discovery happens: the specification, the tooling, the quality system, and the people.

2

The Specification Becomes the Control Surface

Begin with the most consequential change, and the least visible one. When agents write more of the code, the most important human-authored artefact stops being the code. Code remains the executable artefact. The specification becomes the source of intent — the primary control surface through which humans govern what gets built.

Consider how intent has historically travelled into software: meetings, product documents, tickets, whiteboard photographs, a developer’s interpretation, code review, testing, and a long argument about what was meant. Every hop lost information; the code that shipped was the residue of a negotiation. Slow human execution acted, perversely, as a safety mechanism — there was time for misunderstandings to surface. Agents remove that safety mechanism. The machine is not lazy; it does not pause because the requirement is imprecise. It executes — and a weak specification now generates a large volume of syntactically correct but strategically wrong software, faster than anyone can read it. The faster the machine executes, the more expensive ambiguity becomes.

So the foundry’s first discipline is not prompting. It is executable precision: expressing intent in a form machines can implement and other machines can verify. A serious foundry specification is layered, and each layer answers a different question. Purpose — what job is being completed, and for whom. Behaviour — what the product must do in normal use, stated as observable outcomes. State — what it knows, remembers and changes. Boundaries — what it must never do: the security lines, the data it may not touch, the actions requiring a human. Acceptance — how another machine will know the work is correct: the tests written before the code exists. Operation — how failure will be detected, diagnosed and reversed. The layers are not paperwork; they are the settings on the machine. Change the specification and the production run changes with it.

A worked example makes the difference concrete. “Let managers approve expenses” is a request, not a specification. Which managers? What counts as an expense? What happens when the approver is on leave, the currency is unrecognised, the amount crosses a threshold, the receipt is unreadable, or the employee changes teams mid-claim? Who may override a rejection, and what is written to the audit trail when they do? Which failures block payment, and which merely raise a warning? An agent can generate the screens for this in minutes; the screens were never the hard part. The behavioural truth of the workflow is the product — and every one of those questions is a decision a machine will otherwise make silently, at speed, in whichever direction the ambiguity happened to lean.

Not a pipeline that ends — a loop that learns.

Notice what the figure is and is not. It is not a pipeline with a beginning and an end; it is a loop. The specification drives the build; the build faces inspection; release is controlled; operation produces a trace of what the software did in the world; and the trace flows back — refining the specification, converting failures into permanent tests, teaching the next run what this run learned. In the workshop, the specification (when one exists) helps a developer begin. In the foundry, the specification governs the entire production loop — it is where intent enters, and where learning returns.

This discipline reaches product management too. The traditional product document could remain persuasive while being operationally vague, because a skilled team translated it; the foundry specification cannot hide behind prose. It must separate assumption from fact, normal path from exception, outcome from implementation preference — and state which decisions the machine may make alone, which require verification, and which remain human. It is versioned, diffable, and traceable to the behaviour it produced. All of which creates a craft that does not yet have a name or a curriculum: the writing of specifications that machines execute well. It sits somewhere between product management, systems design and law — the drafting of documents whose every ambiguity will be exploited at machine speed, whose every precision compounds. The people who master it will be to this era what great architects of code were to the last one. And it explains a prediction worth making plainly: in the foundry, the intellectual property that matters will migrate from the codebase — increasingly generated, increasingly commodity — to the specification library: the accumulated, tested, executable statements of what dependable products must do.

3

Components Are the Tooling

In a physical factory, the tooling — the jigs, dies, fixtures and gauges — determines what can be produced repeatedly and to what tolerance. Tooling is expensive to create and priceless once proven, because every subsequent product inherits its precision for free. The software foundry has an exact equivalent, and it is not the AI model. It is the library of proven components: identity and permissions, data models, workflow engines, connector frameworks, interface components, testing harnesses, observability, deployment machinery, billing, support capabilities, security patterns. One clarification before going further, to keep two ideas from blurring: this production tooling is not the shared product core described in *The App-Stack Tax*. One Core, Many Products is an architecture a foundry might *build* — a product decision. The tooling described here is what a foundry *builds with* — the machinery behind every product it will ever make, whatever the architecture.

Here is the accounting change that follows, and it is an unfamiliar one: the code repository is not the balance sheet. Repositories fill up in any AI-equipped team; volume of generated code is nearly free and nearly meaningless. The asset that matters is the reuse ledger: which components exist, which products use them, how reliable each has proven in operation, what each carries in dependencies, what each costs to run — and, the column that justifies the whole ledger, how much production time each one saves when the next product draws on it. A component used once is an expense. A component proven across three products is machinery. Every proven component is stored production time — and the foundry’s net worth is the sum of that column.

The repository is not the balance sheet. This is.

The ledger changes how the agents themselves work. Instead of beginning from a blank context, an agent is required to search the ledger, compare the new requirement against proven patterns, and justify every departure — and it selects on operating evidence, not novelty: which component produced fewer incidents, which connector has known limits, which pattern reduced support questions. The foundry acquires the thing workshops structurally cannot have: institutional memory that survives individual engineers and model upgrades.

The ledger enforces honesty in both directions. It exposes components that were built speculatively and never reused — the foundry’s dead inventory. And it prevents the opposite failure, which has killed more platform efforts than any technology ever has: premature abstraction. The seductive mistake is to build the universal machinery first — the grand internal platform, the framework for every future need — before any product exists to need it. Platforms built in advance of products are speculation wearing an architecture diagram; they generalise from zero examples and are wrong in ways only real products reveal. The foundry’s promotion rule runs the other way: build concretely; reuse deliberately; generalise only after the second real need. A component is extracted from a real product, hardened by a second real use, and only then promoted into the common tooling. The ledger records the promotion — and the temptation it guards against has grown teeth, because agents make frameworks nearly free to generate: a team can now produce an impressive universal layer long before it has discovered the real commonality. A mature ledger therefore records negative knowledge with equal care — the component that is secure but too expensive to operate, the workflow that fits one regulatory context and fails in another, the connector with hidden rate limits. Knowing where reuse is unsafe is as valuable as knowing where it compounds. A foundry’s maturity shows not in the size of its catalogue but in the quality of its boundaries.

Watch what this does to the economics of the portfolio. The first product pays full price: it builds its own tooling as it goes, and most of what it builds is candidate machinery, not proven machinery. The second product is the moment of truth — the essay returns to this at the end. But by the fifth product, something structural has happened: most of a new product is drawn from the ledger, and what remains to build is mostly the product’s one distinctive job. The threshold for viable software falls with every promotion. Products for narrow niches, small categories and local practices — products that could never have carried a full engineering organisation — become economical, because they no longer have to buy their own tooling. The tooling was already paid for, by every product before them. And this is where the durable advantage forms — the answer to the objection that every competitor has the same models. They do. A competitor can reproduce a screen or a workflow in an afternoon. What it cannot reproduce in an afternoon is years of accumulated evidence about which components behave dependably together, under which conditions, at what operating cost. The reuse ledger is not glamorous. Neither is factory tooling. Both are where repetition turns into economics.

4

Quality Must Be Manufactured

Now the objection that every serious reader has been holding since Part 1, and that deserves to be stated at full strength: AI can increase output faster than any organisation increases judgement. Left ungoverned, agents produce brittle code, duplicated logic, insecure dependencies, inconsistent interfaces, and failure modes nobody has mapped — technical debt manufactured at machine speed. The critics who say most AI-built software will be unreliable are not wrong about the workshop. They are describing it accurately. The question is whether the foundry has a structurally different answer — and it does, though it is not the answer people expect.

The expected answer is human review: have experienced engineers read what the machines wrote. At foundry volumes this is arithmetic nonsense — humans reading machine-speed output either become the bottleneck that erases the production advantage, or become a rubber stamp that erases the safety. The foundry’s answer is the one manufacturing found a century ago, when production outran craftsman inspection: industrialise the inspection itself. Machine tools did not make quality control obsolete; they made *statistical* quality control necessary — specifications with tolerances, gauges at every station, sampling, traceability, the discipline Cusumano documented the Japanese software factories borrowing from their own assembly lines. The foundry completes what those factories started, with the ingredient they lacked: generated code makes inspection more important, not less — but inspection must itself become software.

Six gates between generated code and the customer — and a loop that makes each failure permanent knowledge.

The gates come in layers, each catching what the previous cannot. Static controls first: schemas, types, dependency policies and security rules that reject bad structure before anything runs — the cheapest gate, so it runs on everything. Behavioural tests next: unit, integration, contract and end-to-end — and note that in a foundry these are largely written *before* the code, because they are the specification’s acceptance layer made executable. Adversarial tests third: unexpected inputs, permission probing, corrupted data, misuse and attack paths — machines are tireless red-teamers of other machines’ work. Then the gates that operate in the world: controlled release (flags, staged rollout, shadow operation, instant rollback — the assumption that something will eventually be wrong, built into the delivery mechanism); operational inspection (observability, anomaly and drift detection, cost monitoring — the software watched as closely as it was tested); and finally failure replay, the gate that makes the whole system compound: every significant incident captured with its state, inputs and component versions, reproduced, understood, and converted into a permanent test, so the same failure can never ship twice. In a workshop, an incident is repaired. In a foundry, it becomes production machinery.

Consequence, not convenience, sets the level of autonomy at each gate. A cosmetic change moves through the automated path end to end. A permission change, a payment decision, a deletion of customer data — anything whose failure cannot be cheaply reversed — carries heavier gates and a named human approver. The foundry classifies every change by what it could cost, and grants the machines exactly as much independence as the evidence supports: autonomy is earned through evidence, not declared through ambition. And the same production system extends past release into operation, because at radically lower prices the old support model — large success teams, manual diagnostics, repeated explanations — cannot survive. The product must explain itself: structured logs, replayable failures, documentation that updates with the specification, diagnostics that identify likely causes. Human experts remain, concentrated on the novel and the consequential. AI-native operation is part of the production system, not an economy measure bolted on to protect margins.

Two properties make this a production system rather than a checklist. First, the gates are common: every product passes the same inspection, which is what makes quality a property of the foundry rather than a property of whichever team was careful. Second, the loop learns: each replayed failure hardens the gates for every product at once — one product’s incident becomes the whole portfolio’s immunity. This is the honest answer to the slop objection, and it is also the honest standard from the first essay, restated for engineers: trusted affordability means the price fell because waste was removed — duplicated tooling, unused features, manual inspection — not because responsibility was removed. The question a foundry must answer is never whether every line was manually perfect. It is whether the production system detects variation before the customer pays for it.

5

The Craftsman Is Elevated

End where the anxiety is: the people. The shallow version of this era’s story says agents write the code, so fewer engineers are needed — a subtraction story. The foundry’s story is a relocation: the human moves to the level where judgement has more leverage. Nothing in the four rooms above is unmanned. Someone chose which problem was worth a product. Someone wrote the specification’s boundaries — decided what the software must never do. Someone set the architecture, decided what gets promoted to common tooling, judged the exception the gates flagged at 2 a.m., and accepted responsibility for the consequential release. Someone maintained taste — the difficult ability to distinguish a feature that can be generated from a product that should exist. The machines run the loop. Humans design and govern the system — and governing a production system is a larger act of engineering than operating inside one.

This produces a role that existing job titles describe badly. The foundry builder combines product judgement (which job, which user, what sufficiency means) with systems design, specification craft, agent orchestration, quality engineering, and — unusually for an engineering role — economic thinking: what a component costs to run, what reuse is worth, when a product’s margin makes it real. The temperament matters as much as the skills, and it is specific: low ego about who wrote the code, high ownership of the outcome. Comfort moving between product and engineering without treating the border as a wall. The instinct for reuse without the vice of premature abstraction. An obsession with measurement — because in a foundry, the scoreboard is real and public. A willingness to let machines do everything routine, joined to an absolute refusal to delegate judgement. Some of the best people for this work will come from conventional engineering; some will not have written production code for years; a few will be product people who discovered they can now build. What they share is the ability to be accountable for a system rather than proud of a component. The organisation reshapes itself around them: the traditional decomposition — product, design, frontend, backend, QA, infrastructure, support, and the coordination overhead between them — does not vanish, but most of it moves *into the production system*, leaving a small number of people concentrated on end-to-end outcomes. The culture that makes this safe is unusually explicit about responsibility: every product has named human owners for its purpose, its safety boundaries and its economics; high-consequence changes have clear approval; exceptions are reviewed as opportunities to strengthen the system, not as embarrassments. The team is small enough that nobody can hide behind a function, and disciplined enough that nobody must remember everything personally.

The craftsman is not eliminated. The craftsman is promoted — to designer of the system.

And here the essay must be honest about its own boldest claim. Everything above implies that a small team governing a strong production system can outperform a much larger organisation coordinating specialised workshops — that five people with a foundry beat fifty with better tools. I believe it. I cannot yet prove it, and neither can anyone else, because the claim is not a philosophy; it is a wager with numerical terms. The wager: that the small team wins on elapsed build time, on quality reaching the customer, on support load per product, and on the economics of each successive product — measured, not assumed from headcount. A small team can fail faster too; if five people produce fragile products that require an invisible army of rescuers, nothing has been transformed. Headcount reduction is not proof. Reliable utility per unit of human judgement is proof. And it is worth stating what the wager is not: it is not a claim that engineers have become unnecessary, nor that judgement has become cheap. The foundry team is small because its leverage is high — every person governs machinery that multiplies them — and the moment the numbers say otherwise, the honest response is to change the system, not the scoreboard. Every era’s production revolution attracted true believers before it produced evidence; the believers who mattered were the ones who instrumented the factory and published the numbers. That is the temperament this work rewards, stated one final way: the foundry does not ask anyone to believe. It asks them to measure.

The foundry also changes what mastery means, and this may be the deepest cultural shift of all. In the workshop, mastery is visible in the artefact — elegant code, a clever algorithm, a difficult integration pulled off. In the foundry, mastery is increasingly visible in what no longer requires heroics: the dependency standardised, the class of defect caught automatically, the release that rolled back before a customer noticed, the support problem that diagnosed itself, the new product that inherited weeks of work without inheriting yesterday’s mistakes. The highest craft is embodied in the system — which is why the question a foundry asks of every finished product is not “was it impressive?” but “what did it teach the system never to build from scratch again?”

Closing: Product Two

Which is why this essay must end by disqualifying its own most likely misreading. The first product to come out of any foundry — mine or anyone’s — will prove almost nothing about this essay. A talented team with powerful agents can build one impressive application; the workshop can do that too, and 2026 will be full of impressive first applications. The first product proves demand. The test begins with repetition. Does the second product draw on the ledger instead of starting from a prompt? Does it inherit the gates, connect to the existing machinery, and reach dependable operation in a fraction of the time? Does the third become easier still?

Those questions have numerical answers: elapsed build time, percentage of components reused, defects escaping to customers, support load per product, connector effort, operating cost. When those numbers exist for our own foundry, I will publish them — the favourable ones and the unfavourable ones, because public pre-commitment is what stops a metaphor from grading its own homework. If the curves do not move — if the second product turns out to be another act of concentrated craftsmanship — the honest conclusion will be that a new software company was built, not a new production system. A production system deserves to be measured by production. Until then, everything in this essay is exactly what it claims to be: a description of the machine, offered before the machine has run long enough to be judged. The first essay in this series named the revolution. The second named the waste it removes. This one has shown the machine that removes it. The next will show the dials.

Illustrative — the shape is the claim; the numbers are the promised sequel.

**

AI-generated code is not a software foundry. A foundry is specifications as control surfaces, components as tooling, inspection as software, and humans as accountable designers of the system — judged by one measurement: whether every product makes the next one cheaper, faster and safer to produce.

The workshop celebrates what it built. The foundry measures what it will never need to build again.

Not code without craftsmen. Craft embodied in a system — and multiplied.

Thinks 2050

NYTimes: “…Kids these days — Gen Z and Alpha — aren’t talking about Googling things. They may still be using Google, but they’re not “Googling it.” Instead they’re saying “search it up,” as in: Who is that Norwegian soccer player with the long blond hair? Search it up. The “it” is dropped if you have a direct object, as when you “search up Erling Haaland.” The shift toward using this phrasal verb has been profound in the youngest generations.”

WSJ: “The newest generation of companies, infused with AI from the start, offer a vision of how work could soon be structured elsewhere in American corporations: fewer co-workers; more on-staff engineers; and a flatter structure in which nearly everyone is a player-coach instead of strictly overseeing teams.”

Naomi Kanakia: “The Great Books concept is about placing faith in those who came before us. They all rely on trusting other people’s judgment.”

WSJ: “So-called frontier AI models, or the most capable systems made by companies like OpenAI and Anthropic, can be expensive to use partly because they require a lot of compute and process large numbers of tokens, AI’s basic unit of measurement. But these state-of-the-art models are considered the best because they can “reason” through complex, multistep problems and are capable of supporting a variety of tasks, including powering autonomous AI agents. The calculus often is as much a business decision as an engineering decision. If paying a premium for a frontier model means a better product or an upper hand over rivals, many companies say it’s worth it.”