The Vertical Test (Software Foundry Series #9)

Why depth beats breadth, and how to choose a first market

If software can now be produced cheaply, the obvious move is to produce a lot of it — many products, many categories, wherever an incumbent looks expensive. This essay argues the opposite, and offers a test.

The claim of this essay: the smaller the business, the less of its own operating truth it has ever written down — which means affordable software cannot be general. It has to arrive already knowing how the work is done, and ask only what is different here.

1

A Vertical Is Not a Function

Start by clearing away a confusion that makes this question harder than it is. Customer engagement, finance, project management, service and human resources are useful product domains. None of them is a vertical. They are functions — jobs that recur across every industry, which is exactly why they look like large markets.

A vertical is the operating world in which a function acquires specific entities, rules, exceptions and consequences. Commerce has customers, products, orders, fulfilment, returns, consent and replenishment. A professional firm has clients, engagements, people, capacity, deliverables, time and invoices. A field-service business has jobs, technicians, routes, assets, estimates and parts. Project management describes something all three do. What it means in each is a different subject.

Figure 1. The same job, two different products.

The distinction matters because the architecture is One Core, Many Products. The first choice is therefore not which application to rebuild more cheaply. It is which operating core should become the shared memory for several adjacent jobs — and only a vertical has one. A function has customers in every industry and a core in none.

A function tells the buyer what the product does. A vertical tells the product what it already knows. Which turns out to be the difference that decides whether the price is reachable, for a reason the next section sets out.

2

The Specification Asymmetry

The Expensive Last 10% argued that the scarce input in software is not code and not generic domain knowledge, but a current, sourced, organisation-specific account of how one business operates. That argument has a consequence which points directly at market choice, and it took me some time to notice it.

The businesses least able to produce that account are precisely the ones affordable software exists to serve.

A large company has a compliance function, documented policies, a process owner and an internal system of record. Its operating truth is partial and often stale, but it exists in written form and somebody is responsible for it. A forty-person business has none of that. Its rules live in the heads of five or six experienced people, in a spreadsheet, and in the configuration of whatever system it bought four years ago. Ask it to specify its own returns policy across categories and jurisdictions and you will get a thoughtful answer that is roughly sixty per cent complete, with the missing forty per cent being exactly the exceptions that matter.

This is not a criticism of small businesses. It is a description of what having no spare capacity means. But it disposes of a comfortable assumption — that cheap software plus a capable AI equals a solved problem for the small buyer. Generation was never the obstacle for that buyer. Specification was.

3

Horizontal Asks. Vertical Arrives.

Horizontal software resolves this by handing the problem back. A general work-management tool, a general database, a general automation canvas: each is powerful, and each asks the customer to define its own workflow, its own fields, its own approvals, its own exceptions and its own reports. The product is a capable blank surface, and the customer supplies the operating truth.

That trade works for a company with the capacity to specify. It fails predictably for one without — which is why the graveyard of small-business software is full of tools that were bought, configured halfway, and abandoned at the point where the configuration required a decision nobody had authority to make.

The alternative is a product that arrives with a position:

Here is how a business like yours normally operates. Tell us only what is different about you.

That sentence is the whole argument for vertical depth, and it is worth being precise about what it requires. Not a template. Not an industry-flavoured demo. A default operating specification — the rules, the exceptions, the escalation paths, the things that must never happen, expressed the way Software After Code described, so that the customer’s deviations become edits to something rather than authorship from nothing.

Note where this puts the effort. The expensive work is no longer building the application; it is knowing the domain well enough to write its default. That knowledge is what a vertical producer has and a horizontal one structurally cannot — not because horizontal companies are less capable, but because a default that fits every industry fits none of them.

And it inverts a familiar assumption about scale. Conventional wisdom says horizontal software is the larger opportunity because the market is bigger. In a world where production is cheap and specification is scarce, breadth is the constraint rather than the advantage. The general product must ask. The specific product can already know.

4

Six Tests

Vertical depth narrows the question but does not answer it. Some verticals suit this model and some are traps, and the difference is not the size of the incumbent’s bill. Six tests, and a market has to pass all six.

Figure 2. Six tests for a market. The last one decides.

One: is the price beyond the job? Not whether the incumbent is expensive, but whether the expense is explained by what the software does. The previous essay in this series located the answer: look at what proportion of revenue goes to being found and believed rather than to delivering the product. Where that block is large, the price is funding an organisation rather than a capability.

Two: is there a visible stack? Several products fragmenting one connected workflow, each carrying its own copy of the same foundations. Where the buyer has one tool and it works, there is no App-Stack Tax to remove and the case rests on price alone — which is a weaker case than it appears.

Three: can the data get out? This is the test most likely to be waved through and the one that kills quietly. If the incumbent traps the data, the migration factory does not function, and without migration the price is theoretical: the customer agrees it is better and cannot reach it. A market where switching is impossible is not a cheap market waiting to happen. It is a closed one.

Four: is the execution repeatable? The two-plane architecture requires that most work runs deterministically after intelligence has configured it. A domain where every case is novel — where genuine judgement is required on each transaction rather than on the policy behind it — cannot reach the price, because the cost scales with usage and never comes down.

Five: is the accountability tail bounded? Every domain carries consequence when software acts wrongly. The question is whether that consequence can be carried by a small team with good machinery, or whether it requires the apparatus — the compliance function, the legal exposure, the certifications, the human in the room. Payroll, clinical care and regulated finance all fail this test for a company doing it for the first time. That is not a permanent verdict. It is a sequencing one.

Six: is there distribution, and domain truth? Can this producer reach thousands of relevant buyers at near-zero cost, and does it understand the work well enough to write the default specification? This test decides, because it is the only one AI did not make easier to pass. Production cost collapsed. The cost of being found, believed and understood did not move.

One methodological point, because it is where this framework would most easily be misused. The six are an intersection, not a score. Averaging them produces a number in which a fatal absence is offset by strength elsewhere, which is how a market with spectacular margins and no route to the buyer comes to look attractive. A market like that is a research project. A market with familiar buyers and no shared operating core is a services trap. A market has to pass all six.

5

The Worked Example

Applying the tests without flattery produces an uncomfortable result, and publishing it is more useful than publishing a ranked list of attractive markets.

On price and fragmentation, several verticals score better than commerce engagement. Field service is a clear case: per-seat pricing that compounds as a business grows, and a standard small-operator stack of scheduling software plus accounting plus payroll plus messaging plus a review tool — five products, one connected workflow from enquiry to payment. Professional services scores similarly. Both have fatter margins to attack than commerce.

Both fail test six. No route to the buyer, and no domain truth in the building.

So the first market chosen here is commerce engagement for smaller merchants, and the honest reason is the sixth test rather than the first. There is an existing customer base, a marketplace where those buyers already shop, and real operating knowledge of how merchants work. It is not the fattest market available. It is the one that can be reached without paying for the privilege — and at one-tenth pricing, a market that must be bought into is not a market at all.

State that plainly rather than dressing it up, because the alternative invites a fair question. A producer who claims its first vertical is objectively the best opportunity in software is either lucky or not being straight, and the second is more likely.

6

Attractive Markets That Are Not First Markets

Applying the tests across other categories is useful for a reason that has nothing to do with picking the next one. Each failure is a different kind, and seeing four of them makes the framework legible in a way the abstract version is not.

Field service and trades. An obvious stack across lead capture, scheduling, dispatch, estimates, payments, communication and reviews, with per-seat pricing that compounds as the business grows. It passes tests one and two more clearly than commerce does. It fails on distribution and domain truth, and the work is mobile and physical in ways that raise the support burden considerably.

Small-business finance operations. Large bills and highly repeatable workflows in receivables, payables, reconciliation and cash visibility. The failure is test five: correctness requirements are severe, incumbent trust is deep, and the consequences of a wrong automated action are immediate and legal. The general ledger is a poor place to learn accountability.

Project and work management. Easy to build, easy to distribute, apparently enormous. That is the problem. These are blank canvases — the customer becomes the specifier, which is the failure mode of section three — the category is crowded, and general-purpose agents will absorb much of the surface. A function without a vertical underneath it.

Regulated professional practice — healthcare, legal, clinical. Rich margins and strong vertical data models, and they fail on almost everything the previous essay described: relationship selling, formal validation, specialised liability and human support expectations. The company the price builds cannot serve them, which is not a criticism of either party.

A fat incumbent margin does not create a right to win, and a right to build does not create a right to enter. These are sequencing verdicts rather than permanent ones — several become reasonable for a producer that has already proved its machinery somewhere else.

7

Why the Second Vertical Is a Trap

The tests above are a method for choosing markets, which makes it tempting to run them across a dozen categories and build a sequence. That would be a mistake, and the reason is the one thing this series has committed to measuring.

Inside the Foundry set the standard: a production system is judged by whether each product makes the next one cheaper, faster and safer to produce. A second vertical resets that measurement to zero while looking like progress. New domain truth, new connectors, new regulatory context, new buyers, new default specification. Revenue grows, headcount grows, and the number that decides whether a foundry exists never gets tested.

So products two and three are adjacent jobs on the same core — different jobs, same customers, same identity, catalogue, orders, consent and workflow. That is the arrangement The App-Stack Tax called One Core, Many Products, and it is the only arrangement in which reuse can be observed rather than asserted.

Expansion into a second market is not the proof that the machine works. It is the reward for having proved it. A company that expands first has chosen the version of the story that cannot be falsified — which is comfortable, and worth nothing.

Which is why this essay publishes a method and only one market. The method is durable and belongs to any reader who wants it. The ranking of markets, wedges, prices and timing is a plan, and plans belong where they can change when the first ten customers teach you something.

The discipline of a vertical strategy is not how many markets it can identify. It is how many attractive ones it can decline while the first is still teaching the machine.

Thinks 2080

Jill Lepore: “By the artificial state, I mean a kind of state that is replacing the liberal democratic nation-state in the United States and around the world. It’s both a real thing, a construct, but it’s also an idea. And so, in this book “The Rise and Fall of the Artificial State,” I trace the rise of the idea that we should live under an artificial state or government by machines. I also trace the notion that this is an inevitable failure, that the artificial state cannot survive, and I trace that idea through science fiction.”

Mint: “Household debt ratios across emerging markets have largely plateaued post-covid. In India, however, they have kept climbing, reaching a record 48% of gross domestic product by December 2025, up from 38% before the pandemic. The Reserve Bank of India’s latest financial stability report underscores the nature of this expansion: Non-housing credit accounts for nearly three-fifths of total household borrowing, with half driven purely by consumption.”

Tim O’Reilly: “It may be a mistake to assume that the AI race is about who builds the best intelligence. It may turn out to be about who builds the electrical grid.”

Tarek Mansour: “The beauty of prediction markets is they take a debate that is subjective, emotional, partisan and put it in a place where it’s mathematical, objective and the incentive structure is very clear. If you do research, you analyze things, you’re smart and you put in the effort to truth-seek, you will probably get rewarded by making money. If you have an opinion that’s too biased, non-calibrated, too partisan or too polarized, you probably will lose money. There’s a certain elegance in markets where you know for sure why someone is having the opinion that they have. They’re truth-seeking because they are trying to make money.”

The Migration Factory (Software Foundry Series #8)

Why the next software battle will be won by making it safe to leave

The Foundry Price ended with seven questions a buyer should ask before signing a renewal. This essay is the operational sequel. You asked the questions, the answers were poor, and now there is a harder problem: how does a business move?

The claim of this essay: when software becomes cheap to build, the binding constraint is no longer building the replacement. It is moving a living business into it without breaking anything — and whoever industrialises that will hold a more valuable position than whoever ships the better feature.

1

The Moat That Is Not a Feature

Ask a finance director why an expensive contract was renewed and the answer is rarely that the product is excellent. It is some version of: we looked, and leaving seemed worse.

Unpack that sentence and it contains a specific set of fears, each reasonable. The data may not come out complete. Configuration built over six years is undocumented, and nobody remembers why the third rule exists. Integrations will break in ways that surface a fortnight later. Workflows nobody described will turn out to have been load-bearing. Historical records may not transfer, and somebody will need them during an audit. Staff will resist a system that does the same job differently. And underneath all of it, one person will be responsible if the transition goes badly — and that person is usually the one recommending the change.

None of those fears is about the incumbent’s product. They are all about the transition. Which means the incumbent is protected by something it did not build and cannot lose: the customer’s reasonable assessment that the disruption exceeds the saving.

This has always been true. What changes it is the collapse in production cost. When building a competitive replacement took three years and forty engineers, migration difficulty was one obstacle among many and not the largest. When the replacement can be produced in months, migration becomes the only obstacle left standing — and therefore the one worth industrialising.

There is a useful precedent. Mobile telephone numbers were portable long before most people used the right to switch; the point was never that everybody moved, but that staying became a decision rather than a default. A migration factory does the same thing to a software renewal. Its value is not measured only in customers who arrive. It is measured in the moment a buyer stops treating the incumbent as fixed.

2

Implementation Is Not Migration

The software industry has a word for the work of getting a customer onto a new system, and the word is wrong. Implementation describes configuring a product to meet a customer’s requirements. It begins with the new system’s needs: which fields must be populated, which workflows configured, which integrations connected, which users trained. It is a project, it is scoped by the vendor, and it is priced.

Migration begins from the other end. It begins with what the customer cannot afford to lose — and the customer usually cannot say what that is, because most of it was never written down. The undocumented exception. The suppression rule added after an incident in 2022. The report one department depends on that nobody else knows exists. Implementation asks what the new system needs. Migration asks what the old one was doing.

Implementation Migration
Begins with the new product’s configuration Begins with the customer’s current operating reality
Collects the fields and settings required Discovers undocumented workflows, exceptions and dependencies
Optimises for go-live Optimises for continuity, evidence and reversal
Treats disagreement as a configuration issue Treats disagreement as information about policy or history
Ends when the product is live Ends when the business is stable and the way back is known

Table 1. Implementation configures a product. Migration carries a living organisation.

The distinction is not academic, because the two produce different failure modes. A go-live can succeed while the migration fails. The screens work, the records exist, the project is signed off — and a consent rule was simplified, a refund exception vanished, a report no longer reconciles, and one connector silently stopped sending a class of event. The new system is live. The old operating truth did not arrive.

A failed implementation produces a system configured wrongly, which is visible and fixable. A failed migration produces a system that works perfectly and quietly does something the business did not intend — the silent kind of failure that The Expensive Last 10% described, discovered later, by its consequences.

This also explains why migration has never been a product. It has been a professional-services engagement: senior people, spreadsheets, discovery workshops, a project plan, a risk register, and a price that often exceeds the software’s annual cost. That was rational when discovery could only be done by expensive humans interviewing other humans. It is what makes migration the obvious thing to industrialise now, because most of that discovery is reading systems, not reading minds.

3

The Machine

A migration factory is a repeatable production line, not a bespoke project. Eight stages, and the middle two carry the difficulty.

Figure 1. The migration factory. Steps three and four are where the value sits.

Discover. Connect to the incumbent and establish what is running: entities, volumes, configuration, automations, integrations, permissions, retention rules, and — the part that matters — what is running that nobody described. Discovery reads the system rather than interviewing the organisation, which is why it can be done in an afternoon rather than a fortnight.

Extract. Records, configuration, workflow definitions, permissions, templates, history, consent records, suppression state and audit trail, in a form that can be inspected rather than merely loaded. Data export without configuration export moves the nouns and loses the verbs.

Reconstruct. This is the hard one and the whole argument turns on it, so the next section takes it alone.

Reconcile. Discovery always surfaces contradictions: two rules that disagree, a policy documented one way and running another, permissions granted to people who left. Most migration projects quietly resolve these by picking one. A factory does the opposite — it surfaces them and requires somebody with standing to decide, because as The Expensive Last 10% argued, choosing between two contested truths is an institutional act and not a technical one.

Shadow. The new system runs alongside the old, receiving the same inputs, acting on nothing. This is the step that removes the fear, and it is the reason the sequence can be self-serve: nobody has to trust the replacement before watching it be right.

Certify. Equivalence has to be defined before it is tested, or the comparison becomes an argument. Which outputs must match, within what tolerance, over what period, and what constitutes an acceptable difference rather than a defect. A migration that cannot state in advance what would count as success is not a migration; it is a hope.

Switch. By segment rather than by weekend. One customer group, one workflow, one region at a time, with the old system live behind it.

Reverse. A route back that stays open, and is tested rather than assumed. Reversal is not the inverse of switching; it is a designed capability, and it has to exist before anybody needs it.

4

Reconstructing a Specification That Never Existed

Everything above is engineering except the third step, which is closer to archaeology.

The incumbent system does not contain a specification. It contains behaviour — configuration, rules, automations, exceptions and accumulated workarounds that together encode how the business operates, without anywhere stating it. Nobody wrote it down because nobody had to; the system was the statement.

Worse, there is no single account to work from. Four of them exist and they disagree: the written policy, the incumbent’s configuration, what the system observably does, and what the experienced people say happens when the normal rule does not fit.

Figure 2. The reconstruction step resolves four disagreeing accounts into one.

So the reconstruction step converts behaviour into an executable account: which rules are in force, what they do, which are legal requirements and which are preferences, which exceptions have been accepted, who may override, what happens when two rules conflict. *This is the same artefact Software After Code called the specification, produced backwards* — derived from a running system rather than authored ahead of one.

Two properties make it difficult, and both were named in the previous essay. Provenance is missing: the system knows what to do and not why, so a rule cannot be safely changed by anyone who does not already know its origin. And habit is indistinguishable from policy in the artefact itself — a workaround added for a supplier who no longer exists looks exactly like a compliance requirement. Migrate without separating them and the new system inherits the folklore with the same authority as the law.

Which is why the reconstruction step produces two outputs rather than one: the specification, and a list of things the organisation must now decide. The second output is often more valuable than the migration. A business that has never seen its own operating rules written down usually discovers, at this point, several it would not have chosen.

And this is the durable asset. The data lands once and is then just data. The reconstructed specification is what the customer keeps, what the system is regenerated from afterwards, and what makes the next change safe — which is the recompile property from Software After Code, and it cannot exist without this step.

5

Disagreement Is the Product

Shadow operation is the bridge between a plausible replacement and an authorised one, and its output is not a pass mark. It is a list of differences — every case where the two systems would have done something different under identical conditions.

The temptation is to call all of them defects and fix them until the systems agree. That destroys precisely the information the exercise exists to surface, because the differences fall into five categories and only one of them is a defect.

Kind of disagreement What it requires
The replacement is wrong Correct the specification and add a permanent test
The incumbent is wrong Confirm the old behaviour should not be preserved out of familiarity
The data differs Reconcile source, timing, identity or historical state
The policy is ambiguous An authorised person decides, and the decision is recorded with its provenance
Both are acceptable Define the boundary within which either action is allowed

Table 2. Shadow operation converts uncertainty into a finite set of decisions.

The second row is the one customers find hardest and value most. A difference is not automatically an error in the new system; sometimes the incumbent has been doing something wrong for four years and nobody noticed because nothing existed to compare it against.

Certification, then, does not mean the two systems agree. It means the organisation understands why they differ and has authorised the new behaviour — across the ordinary path, the rare and consequential exceptions, the permission model, silent-failure detection and the route back.

Which connects migration to the authority ladder in Software After Code. The replacement begins by observing, then recommending, then preparing, then acting after approval, and only later operating within policy. A migration is not complete when the data arrives. It is complete when authority has been earned.

6

Switching Without the Weekend

Large migrations are still narrated as heroic weekends: freeze on Friday, move everything, test through the night, declare victory on Monday. The story persists because projects are organised around dates. Businesses are organised around continuity.

Progressive switching aligns with the second. Move one capability whose boundaries are clear, or one location, one cohort, one workflow, keeping the incumbent as the fallback for everything else. Compare operating results, support load and exceptions. Expand only where the evidence holds.

This changes the economics as well as the risk. The customer receives value before the whole estate moves. The producer discovers where its own machinery is weak while the blast radius is small. And the incumbent contract can be reduced in stages rather than terminated through one all-or-nothing negotiation — which matters, because that negotiation is often the real reason a switch never starts.

It also permits an ending the industry rarely admits to. The end state need not be total replacement. Some customers will keep the incumbent as a system of record and move workflows, intelligence or selected capabilities out around it. Migration is not ideological purity; it is the disciplined movement of utility to the architecture and the price that serve it best.

7

The Argument Only Works If It Points Both Ways

There is an obvious objection, and any reader will have arrived at it. A company that industrialises leaving expensive software has built a weapon it will eventually want to point away from itself. The migration factory attacks the incumbent’s lock-in on Monday and becomes the new lock-in by Friday.

The only answer that survives is to make exit a shipped feature: full data export, the reconstructed specification exported with it, a documented list of what will stop working, connector inventory, the authority granted to each agent, and a transition window during which both systems can run. Not on request. Not as a retention conversation. As a documented capability with a page describing it.

This sounds like commercial self-harm and is closer to the opposite, for a reason that fits the rest of this series. The company has argued that its price is honest, its obligation is inspectable and its value is continuing. A customer who cannot leave never tests any of those claims, which means the company never learns whether they are true. Retention through lock-in is a mechanism for not finding out.

There is also a plainer argument. A buyer deciding whether to enter a system is making a bet on how hard it will be to reverse. Lowering the cost of leaving lowers the cost of arriving — which, for a company with no salespeople to overcome hesitation, is not a philosophical position. It is the acquisition mechanism.

8

What This Changes

Three consequences, and they run in increasing order of size.

For the buyer, the renewal conversation changes shape. The Foundry Price argued that the moment leaving has a price, staying has a negotiation. Migration machinery is what converts that from advice into a number — and the number is useful whether or not anybody switches.

For the producer, migration stops being a cost of sale and becomes the product’s front half. It is the first thing a customer experiences, it is where the operating truth is captured, and it is the step that decides whether everything afterwards works. A company that treats migration as onboarding has misunderstood which part of its product is load-bearing.

And for the market: incumbency stops being a position and becomes a performance. Software companies have long enjoyed a protection they did not earn and could not lose — the accumulated difficulty of leaving. When that difficulty is industrialised away, the only remaining reasons to stay are the ones a vendor has to keep earning: the product is good, the price is honest, and somebody is answerable when it fails.

One discipline keeps the whole thing from collapsing back into what it replaced, and it should be stated because the pressure is predictable. A difficult migration will invite bespoke services. A valuable customer will be offered a project team. The revenue will look good, and the migration factory will quietly become an implementation practice again. The rule is the same one that governs everything else here: custom work is permitted only when it produces a reusable mapping, connector, specification or component. Measure it accordingly — human hours per migration, proportion of configuration reconstructed without help, disagreements resolved before switching, incidents after it.

The most valuable thing to build in software may not be the replacement. It may be the road out.

Thinks 2079

Mark Zuckerberg: “Invention, not automation, will be the greatest contribution of superintelligence. Early AI could answer questions and do routine work. Soon it will increasingly help discover new knowledge — ranging from discovering new drugs to cure a family member’s disease to finding new ways to improve your business. While the number of questions a person can ask in a day is limited, the number of valuable things superintelligence can invent to help achieve your goals is unlimited. As intelligence becomes abundant, the most important question will be how we direct it. Some argue that superintelligence itself or a small set of experts who control it should decide what is best for humanity. We disagree. The history of democracy and economics has shown that there is no single objective answer to how people define the best life, and therefore the best approach is letting people decide what matters in their own lives.”

FT: “One of the most gloriously oddball texts in contemporary philosophy is Bernard Suits’ The Grasshopper. Remember the parable of the ant and the grasshopper, where the ant works all summer and the grasshopper lazes about and eventually starves? In the traditional version, the ant is the hero — the ceaseless hard worker — and the grasshopper is the object lesson in the perils of laziness. Suits flips the story. In his book, the grasshopper is the hero; he is the embodiment of play…Imagine, says Suits, utopia — some future paradise where technology has solved all our practical problems. We have perfect medicine, unlimited energy and boundless resources. What would we do with our time? We would play games, says Suits, or we would be bored out of our minds. And if playing games is all we do in utopia, then games must be the meaning of life.”

Debashis Basu: “Economic theories explain growth through capital, technology and institutions, but leadership may be the missing force that turns sound policies into sustained prosperity…What are the attributes of a good leader? Mr Goh says leaders should be visionary and diligent, and selflessly devoted to national interests rather than to party or personal ones. For credibility, leaders must have integrity and be incorruptible (or have the incentives to remain so). Clearly, shifting an economy on to a path of sustained high growth demands radical choices: Creating better and better human capital, integrating into global markets, and continuously absorbing technology and fostering efficiency to keep infrastructure costs low. But these cannot happen on their own. It is leadership (honest, wise, visionary and committed to course correction) that makes it all happen. Empirical evidence says so. It is time for growth theories to catch up.” 

NYTimes: “You could call it “hobbyamory”…A growing number of singles [are] emphasizing multiple hobbies over dating. Rather than spending hours swiping and messaging on apps, these singles are investing their time, energy and disposable income in passions like rock climbing, cake decorating and cyanotype printmaking. They say their social calendars are packed, their friend groups are expanding and their lives feel rich, with or without a romantic partner.”

The Company the Price Builds (Software Foundry Series #7)

What one-tenth pricing forces a software company to become

The earlier essays in this series described a production system. The Software Foundry named the collapse in production cost, The App-Stack Tax the waste it exposes, The Foundry Price the number that follows, and Inside the Foundry the machinery that makes the number sustainable. Software After Code and The Expensive Last 10% then widened the argument beyond any one company.

None of them described the company. That is this essay’s subject, and it has an unusual starting point.

The claim of this essay: at one-tenth the incumbent price, most of a conventional software company is not a choice. It is arithmetically unaffordable — and the interesting question is which of its functions were solving a problem that has now gone, and which were solving one that has not.

One disclosure belongs at the front rather than in a footnote. I run a company built on the model this essay describes as obsolete: enterprise selling, implementation, customer success, the whole apparatus. That is not an argument against the analysis, but it is the reason the analysis has a conclusion I would rather avoid — a company of this kind cannot be built as a product line inside a company of the other kind. The functions being removed are not costs to be trimmed. They are the organisation.

1

SaaS Industrialised Distribution, Not Production

The cloud transition is worth revisiting, because it is routinely described as the moment software was transformed and it was nothing of the sort. It changed how software was delivered and paid for: browser instead of installation, subscription instead of licence, shared infrastructure instead of separate deployments, continuous release instead of the eighteen-month upgrade. Real changes, and they built an industry.

But look at what the company behind the software still contained. Engineers writing and maintaining code by hand. Product managers translating customer needs into specifications nobody enforced. Salespeople finding buyers. Consultants making the product work after it was bought. Customer success teams driving adoption that should have happened unaided. Support desks answering questions the product could have answered. Finance constructing contracts complicated enough to require explanation. And managers coordinating all of them.

SaaS put software in the cloud. It left the company that made the software almost exactly where it was. Production stayed artisanal, and the organisation stayed the shape that artisanal production requires.

That is the structure now under pressure — and the pressure does not come from AI directly. It comes from the price. Which is worth stating as a reversal of the usual order, because most software companies discover their price at the end: build the product, hire the organisation needed to sell and support it, add the expected margin, and arrive at the number the customer pays. The invoice is the last layer of the company. Begin instead from the buyer’s bill and work backwards, and the question becomes which company can exist inside that number.

The price is not the last decision a software company makes. It is the first, and it writes the organisation chart.

2

What the Price Forbids

Take the arithmetic seriously for a moment. A product replacing an incumbent bill of a thousand dollars a month cannot carry a salesperson: the commission alone exceeds the revenue. It cannot carry an implementation consultant, a success manager or a support desk staffed to answer routine questions. These are not decisions about company culture. They are line items the price has already ruled out.

Which turns out to be the useful way into the question, because each of those functions was invented to compensate for a specific constraint. Name the constraint and the honest question becomes: is it still there?

Function The constraint it solved What replaces it
Sales team The product could not be found or understood unaided Borrowed distribution and a self-serve diagnostic
Implementation The product could not configure itself Migration machinery, run by the customer
Customer success Adoption failed by default The product observes its own adoption and intervenes
Support desk The product could not explain or repair itself Telemetry, diagnosis, guided repair, reversible change
Custom roadmap Change was expensive, so it was rationed and traded Reusable specifications and shared machinery
Annual contract Revenue had to be secured against churn the product could not prevent Monthly terms, and value that survives the option to leave
Per-seat pricing Value was assumed to track headcount Flat capability pricing with visible, capped metering
Lock-in Retention could not be earned every month Portability, and a documented way out

Table 1. Each function answered a real problem. The question is which problems survived.

Read the middle column as a group and a pattern appears. Almost every expensive function of a software company existed to compensate for a product that could not do something. It could not be found. It could not configure itself. It could not explain its own failure. It could not be left safely, so it had to be locked. The organisation was scaffolding around the product’s incapacities.

That is why the removals are not austerity. Cutting the people while leaving the work behind produces an understaffed SaaS company, not a different kind of company — and the failure is not immediately visible, which is what makes it dangerous. A company that deletes the sales team without solving discovery has not become efficient; it has become invisible.

The substitution is also stricter than putting AI in front of each department, and the distinction is worth being blunt about. A support bot in front of a product that cannot observe itself merely automates the apology. A coding agent inside a company that still builds a separate architecture for every product produces inconsistency faster. A self-serve trial attached to a six-week migration is not product-led growth; it is a sales funnel with the first meeting removed.

Which gives the operating standard: no routine task should require another permanent department. When recurring human work appears, the first question is not whom to hire. It is what the product or the production system failed to absorb.

3

What the Price Cannot Delete

Now the turn, and it is the part of this argument that matters most, because the previous section read alone describes a company that has quietly transferred its risk to its customers.

Three things do not dissolve. Accountability — somebody must be answerable when the software acts wrongly, and no volume of automation supplies that. Migration — a customer arriving from an incumbent is moving a living business, which is the subject of the next essay. Incident ownership — when a connector fails silently or an action goes out that should not have, a named person must own it at an inconvenient hour.

The Expensive Last 10% argued that these obligations are what a producer sells once code becomes abundant. It follows that a low price achieved by removing them is not a low price. It is unpriced risk moved onto the buyer’s side of the arrangement, where it will be discovered later by somebody who did not know they had bought it.

So the discipline is precise, and it is easy to state and hard to hold: industrialise the obligation; never delete it. Deflect the routine question through the product. Keep a named human path for consequential failure. And make the obligation visible rather than asserted — which specifications are current, which dependencies are monitored, what authority the system holds, how an action is reconstructed and reversed. At a low price this is not a nicety. A low price is exactly what an unpriced risk would look like from outside, and inspectable evidence is the only thing that distinguishes the two.

4

A Small Institution Governing a Large Machine

What remains is a company organised around a much shorter list of human responsibilities: deciding which problems deserve solving, writing the operating specifications, designing the architecture and the boundaries of authority, judging whether the system is correct, governing exceptions, holding customer trust, and allocating capital.

Agents do the translation between those decisions — specification to implementation, implementation to tests, behaviour to documentation, telemetry to diagnosis, question to answer, usage to intervention. This is not a conventional company with AI assistance. It is a small human institution governing a much larger machine workforce, and the distinction shows up in what the company does when demand grows.

A conventional software company answers growth by hiring: more account executives, more implementation, more success, more support. Each increment of revenue costs an increment of people, which is why software companies get less efficient as they get larger and why their prices must rise to fund the organisation their prices created.

The alternative is to answer growth by compounding machinery. A connector built once serves every product. A migration mapping improves every future migration. A support resolution becomes an automated playbook. A domain rule becomes a reusable specification. A security fix propagates across the estate. The operating question stops being how many people are needed to reach the next ten million, and becomes which reusable capability would make the next ten million need fewer people than the last.

The phrase “small team” invites a misreading worth heading off. Headcount is not the doctrine, and there may never be a ten-person billion-dollar software company in any market worth entering — that number is theatre. The meaningful curve is different: human intervention per customer, per product and per consequential action, falling as the company learns.

Which sets the rule for hiring, and it is not a headcount target. Add a person when the system faces a class of judgement it does not possess — a new vertical’s domain truth, a new category of risk, a new regulatory context. Never add one because revenue grew.

5

Product-Led from Discovery to Departure

Most companies described as product-led are self-serve at the easiest moment only. A user starts a trial without speaking to anyone, and then meets a human at migration, at configuration, at adoption, at expansion, at support and at exit. That is not a product-led company. It is a company whose lead form has been replaced by a button.

At this price the whole relationship has to be product-led, and it is worth naming the stages because each one is a place where a conventional company inserts a person.

Figure 1. Each stage is somewhere a conventional company inserts a person.

Discovery comes through an existing base, marketplaces, partners and word of mouth — borrowed, because acquisition cost is the one thing AI did not reduce. Diagnosis should happen before purchase: connect the incumbent, show which jobs are used, expose what is being paid for and identify what can move safely. The customer should reach first value before a salesperson would ordinarily have qualified the opportunity.

Migration and shadow operation sit at the centre of this sequence rather than at its edge, and they are the subject of the next essay. The organisational point is enough here: if moving a customer in requires a project manager, a consultant and a weekly call, the price does not survive contact with adoption.

Departure completes the discipline. Data, configuration and the reconstructed operating specification should be exportable without a retention conversation. Easy exit is not commercial naivety. It is what forces the product to hold customers through continuing usefulness rather than captivity — and a company with no salespeople to overcome hesitation needs the entry decision to feel reversible.

6

Pricing Without a Growth Tax

The headline number matters less than the pricing constitution behind it, because a nominally cheap product becomes expensive through seats, contacts, modules, activation fees, compulsory annual terms and opaque usage. The buyer experiences the invoice, not the pricing page.

So: published prices, monthly terms, no quotation for the standard product, no activation fee for ordinary use, and no per-seat or per-contact charge where those units create no real cost. A growing small business should never fear its own success on the bill — which is precisely what per-seat pricing was designed to do, and the reason flat-rate competitors keep winning in categories where headcount and value have come apart.

Where variable cost is real — messages, storage, payments, model calls — metering must be visible, capped by default and paired with budget controls. A low entry price with an unpredictable usage bill is not affordability. It is uncertainty relocated to the customer.

And One Tenth is relative to the incumbent bill being replaced, not a universal figure. A capability at $150 replacing $1,500 of fragmented software is faithful to the idea. A product at $10 that consumes twenty minutes of human support each month is not, whatever its price says.

7

The Economics: The Target Is Not the Margin

It is tempting to describe the opportunity as attacking fat gross margins, and the numbers invite it: monday.com reported an 89% GAAP gross margin for 2025; HubSpot generated roughly $2.62 billion of gross profit on $3.13 billion of revenue; Veeva’s subscription gross margin exceeded 86%.

That reading is wrong, and this series has already argued why. Gross margin at that level funds delivery, security, reliability and the obligation — the things The Expensive Last 10% insists deserve to be paid for. Attacking gross margin means attacking dependability, which is the one thing that cannot be attacked.

The fat is one line lower. HubSpot spent $1.38 billion on sales and marketing in 2025 — about 44 cents of every revenue dollar on being found and believed. monday.com spent roughly $631 million on the same line. That block is not delivering software to anyone. It exists because the product could not be found or understood unaided, which is the constraint the first row of Table 1 describes.

Figure 2. The block being removed is the one beside delivery, not delivery itself.

So One Tenth is not a discount financed by optimism. It is what the price becomes when the go-to-market line is deleted rather than the product. That formulation matters, because it is falsifiable: a company that removes the sales organisation and does not find distribution has not built a cheaper product, it has built an unfound one.

The constraint runs the other way as well, and it is the reason this cannot be assumed. Inference is a cost that scales with usage rather than with customers — the opposite shape from the software industry’s foundational economics — and AI-native gross margins are settling well below the historic norm. The two-plane architecture from Software After Code is the structural answer: intelligence configures the workflow, deterministic execution runs it. Delivery cost is the one line that gets harder in this model, which is why it has to be engineered from product one rather than repaired later.

8

What Management Watches

A company shaped this way cannot be run on the conventional dashboard. Bookings, pipeline, quota attainment and utilisation all measure an organisation this company does not have. What replaces them falls into four groups, and they are worth naming because they are how the model is disproved as well as proved.

Production. Time from specification to deployed capability. Proportion of a new product drawn from existing machinery. The cost of product two relative to product one.

Adoption. Time from connection to first useful output. Migration completion without human help. Retention at eight weeks. Capabilities activated per customer.

Economics. Gross margin after the full cost of keeping the promise — inference, infrastructure, messaging and human minutes. Support minutes per active customer. Acquisition cost by distribution route. Revenue per permanent employee.

Authority. Distribution of actions by authority level. Human interventions per thousand agent actions. Proportion of consequential actions that can be reconstructed. Reversal rate. Time to detect a silent failure.

The last group is the unfamiliar one, and it is the group that keeps this company honest. A conventional dashboard measures how much software was sold. This one measures how much organisational labour the software replaced without weakening the promise — and the two can move in opposite directions, which is precisely why both need watching.

These numbers also guard against a specific kind of decay. A large customer will offer to pay for a branch. A partner will ask for a special margin. Someone will promise a roadmap item to close a quarter. A support queue will look easier to staff than to eliminate. Every one of those decisions is locally rational and collectively fatal, and none of them shows up on a conventional dashboard until the company has already become something else.

Inside the Foundry committed to publishing production numbers once they exist. That commitment stands, it is not yet payable, and the final essay in this series dates it.

9

What This Model Is Bad At

An essay describing a company one intends to build is worth little unless it says where the company should not be chosen, so: this model is a poor answer for any customer that needs bespoke work, institutional procurement, a named human in the room during evaluation, contractual liability at enterprise scale, or a roadmap shaped by their own requirements. Those buyers are well served by the incumbent model, which is why the incumbent model exists.

It also does not generalise across every domain. Where regulation is heavy, where physical operations intrude, where a single error carries severe consequence, or where buyers cannot purchase without institutional selling, the obligation cannot be industrialised by a small team — at least not yet, and not by a company doing it for the first time. The defensible claim is narrower than the exciting one:

AI makes a radically smaller and more automated software company possible in domains where the work is repeatable, the data is reachable, migration can be industrialised, and customers can buy without being sold to. Where a domain fails those tests, the old company shape is still the right one. Which domains pass is the subject of the ninth essay.

The great software companies of the subscription era were built by organising large numbers of people to create, sell, implement and support code. The companies that follow may be built by organising a small number of people to govern specifications, machinery and authority. Their advantage will not be that they employ fewer people. It will be that each additional product, customer and market requires fewer than the last.

Thinks 2078

FT: “The benefits of the AI transformation might well be enormous, even if they are as yet mostly unproven. But in the meantime, the rest of society cannot remain passive passengers in the back of the AGI car, like children on a hot holiday trip, frustratedly asking: are we there yet? The AI companies are intent on driving humanity to a destination that has not been fully discussed, let alone agreed. It is time for the rest of us to acquire more agency over our human future. ”

Khachan: “There’s an Italian saying that all the pieces are equal inside the chess box. The same thing applies to people around the chessboard.” [from an FT article on why chess clubs are sweeping the board]

Akash Prakash: “The introduction of AI models with Mythos-level capability has changed the game for cybersecurity. When attackers can reverse engineer a patch and weaponise it in minutes, speed of response has become critical. At the moment, enterprises are notoriously slow in implementing patches as they try to optimise system uptime and avoid patch-related disruptions. Talking to experts, one gets the sense that India has to urgently upgrade its software, especially in the states and the small and medium enterprise sector. Much of it is very old.” 

NYTimes: “As of June 30, private equity firms had 33,575 unsold companies in their portfolios, according to PitchBook, an industry data firm. That’s up from 32,451 companies at the end of last year and 15,923 companies a decade ago. The growing backlog is a challenge for private equity’s core business model. Typically, such firms aim to buy a company, often add large amounts of debt to its balance sheet, improve its financial performance and then sell it for a profit, usually within five to seven years.”

The Expensive Last 10% (Software Foundry Series #6)

Why generating software is becoming easy and keeping its promise is not

The previous essay, Software After Code, argued that machine generation removes code from the centre of what humans make — and raised an objection it deliberately left open. If producing software has become cheap for vendors, it has become cheap for customers too. Many companies will build what they used to buy, and some of them should.

This essay takes up what happens after that decision — not to argue anyone out of it, but because the part that has become cheap and the part that has not are rarely told apart until somebody is living with the difference.

The claim of this essay: generation produces visible capability. Everything that makes capability dependable sits underneath it, is unaffected by the collapse in production cost, and has become more expensive rather than less — because software now acts.

1

The Application Built in a Weekend

A logistics company with about forty people has a problem that every operations manager will recognise. Delivery exceptions — the failed drop, the refused consignment, the address that does not exist — arrive through three channels and are tracked in a spreadsheet that four people edit and nobody trusts. The commercial software that would handle this properly costs more than the problem appears to be worth, and the last evaluation stalled on a six-week implementation nobody had time for.

So the operations lead, who is technically capable but not a software engineer, spends a Saturday with a coding agent. By Sunday evening there is a working application: exceptions pull in from all three sources, get classified, route to the right person, escalate on a timer, and produce a weekly summary that is more useful than anything the spreadsheet ever gave anyone. It cost a weekend and the price of some inference. On Monday the team starts using it, and within a fortnight nobody would go back.

This story should be granted completely, because it is real and it is going to become ordinary. The application is good. It fits the business better than a purchased product would have, because it was shaped by somebody who understood the work. It cost a rounding error. Any argument that begins by suggesting this could not have happened, or that the result must secretly be bad, is not an argument — it is a vendor protecting a position, and the reader will smell it immediately.

Now move forward twelve months, and watch what the year does.

In month three, one of the three source systems changes its authentication, on its own schedule, with an email nobody read. Exceptions from that channel stop arriving. Because they stop arriving rather than arriving wrongly, nothing looks broken — the queue simply gets shorter, which everybody enjoys for eleven days.

In month five, a customer disputes a charge arising from an automated escalation, and asks the company to show why the system did what it did. There is no record beyond the current state of the database. The reasoning existed at the moment it ran and was never written down.

In month seven, somebody patches the classification rule directly, at eleven at night, to stop a client complaint. It works. It is never reconciled with anything, and nobody is quite sure afterwards whether that patch is still in effect.

In month nine, the operations lead who built it takes a job elsewhere. The application keeps running. Nobody who remains can say with confidence why the timer is set to forty-eight hours, which exception types were deliberately excluded, or which of the several near-identical rules is the one that fires in practice.

In month eleven, an auditor asks a straightforward question about retention of customer data, and the honest answer takes three weeks to assemble.

None of this is a story about a bad application. It is a story about the difference between a thing that works and a thing that can be relied upon — and that difference did not become cheaper when generation did.

The repository took a weekend. The promise began on Monday morning.

Day two Month twelve
A working application A connector has changed, silently
Ordinary cases automated Exceptions have accumulated
A visible workflow replaces a spreadsheet One manual patch is unreconciled
The builder knows every decision The builder has left
The demonstration succeeds Somebody has to keep the promise

Table 1. The same application, twelve months apart. Nothing in the right-hand column is a defect in the left.

Key points

  • The weekend build is real, is often better-fitted than a purchased product, and is going to become ordinary.
  • What arrives in the year afterwards is not a defect in the application. It is the obligation that was never priced.
  • The repository took a weekend. The promise began on Monday morning.

2

The Demo and the Promise

Software has always divided into two parts, and the division was always uncomfortable to talk about because the second part is invisible in every demonstration and consumes most of the effort.

The first part is what the software does when things go as expected. Screens and forms. The ordinary workflow. Reading and writing records. Dashboards. Standard integrations with well-behaved systems. Conventional business logic. This part is now substantially generatable, and getting more so. It is also the entire content of every demonstration, every screenshot and every evaluation — which is why buyers have always over-weighted it, and why the software industry learned to sell to it.

The second part is what the software does when things do not go as expected. The edge case nobody thought of at design time. The exception that is legitimate but looks fraudulent. The operating history that explains why this rule exists. Provenance — who set this, when, on whose authority. Permissions. Consent, and evidence of consent. The connector that changes on somebody else’s schedule. Security. Monitoring that notices absence, not just failure. The ability to reverse an action. Compliance evidence produced on demand. Continuity when the person who understood it leaves. And, underneath all of it, somebody answerable when the software gets it wrong.

Figure 1. Generation reaches the part above the line. The part below it is what a demonstration cannot show.

There is a structural reason the second part gets under-bought, and it is worth naming because it explains a great deal of behaviour on both sides of the transaction. Dependability is largely built from capabilities whose success appears as the absence of an event. Nobody celebrates a suppression rule that correctly stopped a prohibited message; the message simply never arrives, and no one knows it nearly did. A clean audit trail goes unnoticed until the complaint, the dispute or the investigation. A reversal mechanism is unused on almost every day it exists and indispensable on the one day it is needed. Software that displays a feature accrues credit continuously. Software that prevents a disaster accrues none, until the disaster.

A product is not defined by what it does when everything is normal. It is defined by what it does when reality departs from the demonstration.

A demonstration proves that software can work. This second part determines whether it can be trusted.

The split is not new. What is new is the ratio, and it moved twice.

It moved once because generation collapsed the cost of the first part and left the second part exactly where it was. A visible capability that used to take a team three months now takes a weekend; the obligation underneath it takes the same year it always took. Anything that becomes cheap while its complement does not has raised the relative price of the complement — which is the whole economic content of this essay, and it applies to the vendor exactly as much as to the customer.

And it moved a second time, for a reason that has nothing to do with production cost. Software has started to act.

The previous essay set out the ladder along which authority is granted — from observing, through recommending and preparing, to executing, operating and eventually owning a goal. What matters here is only the consequence at the bottom of the ladder. When software displayed information and waited, a missing edge case produced a wrong screen. Somebody looked at it, frowned, and nothing had happened in the world. When software acts, the identical gap produces an action: the message sent to somebody who withdrew consent, the refund issued that should have been reviewed, the escalation triggered against a customer who did nothing wrong, the record altered in a system that other systems believe.

Figure 2. The same specification gap, before and after delegation.

A deterministic system fails visibly and locally. A system with authority and a weak specification fails confidently and at scale — and it fails silently, because there is no screen for anybody to frown at. It is found by its consequences, which arrive later, through a customer, a regulator or an auditor.

So the expensive part is not a legacy concern that better generation will eventually reach. Delegation is what made it expensive. The more capable the software becomes, the more the value sits in the part that generation does not touch.

Key points

  • The first part is what software does when things go as expected. It is now substantially generatable, and it is the whole of every demonstration.
  • The second part is what it does when they do not. It did not become cheaper.
  • Dependability is built from capabilities whose success appears as the absence of an event — which is why it is systematically under-bought.
  • Anything that becomes cheap while its complement does not has raised the relative price of the complement.
  • Delegation moved the ratio a second time: the same gap that once produced a wrong screen now produces an action.

3

What the Customer’s Agent Cannot Finish

There is a comfortable version of this argument that should be refused, because it is wrong and it will not survive a serious reader.

The comfortable version says that models lack domain knowledge — that a coding agent does not understand logistics, or lending, or clinical operations, and that this is what a specialist vendor sells. That claim is already shaky and gets weaker every year. Models know a great deal about how a returns process, a credit assessment, a reconciliation or a triage protocol normally works. They will know more next year. Any thesis that depends on models staying ignorant is a thesis with an expiry date.

The sharper claim survives:

Generic domain knowledge is becoming abundant. Current, organisation-specific operating truth is not.

Consider what an agent generating that exceptions application would have needed to know, none of which is generic. Which classification does this company use, and how does it differ from the industry norm, and why — because it does differ, and the reason is usually an incident that happened once. Which of these rules is a legal requirement and which is a preference that somebody could change if asked? Where did the forty-eight hour timer come from: a contract, a regulation, or a decision made in a meeting in 2021 by somebody who has left? Which exceptions have been deliberately accepted, by whom, and on what grounds? Who is permitted to override, and what happens when two rules apply and disagree? What evidence would demonstrate, to somebody hostile, that the policy was followed? And — the question that quietly governs all the others — is any of this still true?

None of that is in a model, and none of it is in the request the operations lead typed on a Saturday. It exists in fragments. Some sits in policy documents nobody has read since they were written. Some sits in the configuration of an existing system, where it was recorded as a setting rather than a decision. Some lives in spreadsheets, email threads and support macros. And a great deal survives only in the memory of five or six experienced people who know which written rule is not the whole rule. Turning those fragments into an executable account of how the business operates is not clerical work. It is organisational design — and it is difficult in a way that has nothing to do with software.

It also involves a set of distinctions that only an organisation can draw. Policy has to be separated from habit. The current rule has to be separated from the historical workaround that outlived its cause. Authority has to be separated from convenience. And stated intention has to be separated from what the business is prepared to enforce when enforcing it costs something. Without those separations, generated software will reproduce an organisation’s folklore as confidently as its policy, and with the same authority to act on it.

Which sharpens what the customer’s own agent can and cannot do, and the concession matters. It can do more of this than is comfortable to admit. It can interview, read the existing configuration, draft a specification, and — valuably — identify contradictions that nobody had noticed, because it does not share the assumptions that made them invisible. What it cannot do is decide which of two contested truths the organisation should adopt, or determine who is entitled to accept the risk of getting that wrong. That is not a knowledge problem. It is an institutional act, and it requires somebody with standing to perform it and to be answerable afterwards.

This is why the weekend application is so often excellent and so rarely complete. It captures the operating reality as one well-informed person understood it on one day. That is a valuable artefact. It is not the same as an account of how the business works that stays accurate as the business changes — and the gap between those two things is invisible on Monday and expensive in month nine.

Two properties make organisational truth hard to hold, and both are worth naming precisely because neither is a technical problem.

The first is provenance. A rule without a source cannot be safely changed. Anyone who has worked in an established business has met the constraint nobody can explain and nobody dares remove, and the reason the fear is rational is that the rule might be load-bearing. When rules are captured only as behaviour — in code, or in a generated implementation — provenance is exactly what is lost. The system knows what to do and not why, which is fine until the day somebody needs to change it.

The second is currency. Operating truth decays. A jurisdiction adds a requirement. A contract is renegotiated. A category behaves differently after a bad quarter and the rule that was written for the old behaviour quietly stops fitting. Nothing about generation addresses this. A system generated from a perfect specification in January is running on a January account of the world in November, unless somebody has been maintaining it — which is a job, and it is the job that most self-build arrangements have not assigned to anyone.

So the scarce input in software is not code, and it is not domain knowledge in the abstract. It is a current, sourced, tested account of how one particular organisation operates — held by somebody whose job is to keep it true.

Key points

  • Refuse the comfortable claim that models lack domain knowledge. It is weak and it expires.
  • Generic domain knowledge is becoming abundant. Current, organisation-specific operating truth is not.
  • Turning fragments into an executable account of how a business operates is not clerical work. It is organisational design.
  • Without separating policy from habit, generated software reproduces an organisation’s folklore as confidently as its policy.
  • The agent can draft the specification and find the contradictions. Deciding which contested truth to adopt is an institutional act.
  • Provenance and currency are the two properties that make organisational truth hard to hold — and neither is a technical problem.

4

The Maintenance Obligation

The previous essay described the machinery that reduces this burden — change a shared truth once, and every capability that depends on it is rebuilt and revalidated. That machinery is real, and it is the strongest argument a producer has. It also has to be built, staffed and run by somebody, which is the point of this section.

Here is the obligation, stated plainly, because it is usually discussed as an abstraction and it is not an abstraction — it is a set of jobs that appear on somebody’s calendar.

Connectors. Every integration is a dependency on a third party’s roadmap. Interfaces change, deprecate, tighten their authentication, alter their rate limits and occasionally disappear. The work arrives unscheduled, in a quarter that had other plans, and the failure mode that matters is not the loud one. A connector that breaks noisily gets fixed on the day. A connector that silently stops delivering a subset of records is discovered weeks later, and the interval between those two events is where the real damage sits.

Releases. Something has to establish that the current version still does what the last version did. In a generated system this is more important rather than less, because the implementation may have been rebuilt entirely between one week and the next.

Incidents. Somebody answers when it fails, and the answer is required at inconvenient hours by definition. This is the single most under-costed item in every self-build decision, because it is not a task but an availability commitment — which is a very different thing to ask of a person who already has a job.

Security. Vulnerabilities are disclosed on a schedule set by researchers, not by the business. Somebody must be watching, must know which disclosure applies, and must be able to act quickly.

Regulation. Rules change in one market and have to be threaded through a system built for the world as it was. Somebody must notice, interpret, and know which behaviour is affected.

Permissions. People join, change roles and leave. Access granted in a hurry is rarely reviewed, and permission drift is the quiet precondition of most serious incidents.

Support. Somebody answers the colleague who cannot make it work. At small scale this is invisible; it is also how most internal tools die in practice — not through failure, but through nobody being available to explain them until people stop bothering.

Continuity. The knowledge has to survive the person. This is the obligation that self-build arrangements fail most reliably, because it is the only one that produces no symptom at all until the day it produces every symptom at once.

Obligation inherited What happens if it is neglected
Specification maintenance The system enforces yesterday’s truth, confidently
Connector monitoring Silent breakage, discovered by its consequences
Permissions Actions occur without valid authority behind them
Release testing A local fix quietly breaks a path nobody was watching
Incident response The same harm repeats, because the cause was never found
Regulatory tracking The implementation becomes non-compliant without changing
Security A useful tool becomes an exposure
Support and continuity The tool becomes an orphan and is quietly abandoned
Accountability Responsibility is discovered to be nobody’s, after the event

Table 2. Self-building removes an invoice. It does not remove this column.

Figure 3. The two costs are not the same shape, which is why comparing them once is misleading.

That shape is the argument. The cost of creation is a point. The cost of the promise is a line — and the line runs for as long as anybody depends on the software. A comparison made at the moment of creation is not comparing like with like, which is why the weekend build looks so overwhelmingly favourable in the week it is made and merely reasonable three years later.

Now the honest boundary, because this section could easily be read as an argument that self-building is a mistake, and it is not one.

A large company with real engineering depth is already carrying every item on that list for other reasons. It has an on-call rota, a security function, a release process, an access review, a compliance team. For that company the marginal cost of adding one more internally built system is small, and the fit advantage is large. Such a company should probably build more than it does.

The forty-person logistics business is being offered something different. It is being offered the chance to become, in a small way, a software operation — and most companies do not want to be one, in the same way that most companies do not run their own payroll engine or their own payment infrastructure, though both are technically within reach and neither is conceptually hard. The question is not whether the obligation can be carried. It is whether carrying it is what this company wants to spend itself on.

Key points

  • The obligation is not an abstraction. It is connectors, releases, incidents, security, regulation, permissions, support and continuity.
  • The cost of creation is a point. The cost of the promise is a line.
  • A company already carrying that list for other reasons should probably build more than it does.
  • For everyone else the question is not whether the obligation can be carried, but whether carrying it is what the company wants to spend itself on.

5

Accountability for the Act

Everything so far has been about effort — work that must be done by somebody, priced or unpriced. The deepest layer is not about effort at all. It is about things that cannot be produced on demand at any price, because they are accumulated rather than made.

Standing with third parties. A great deal of consequential software depends on how other systems and institutions regard the sender. Whether messages are delivered or filtered. Whether an interface grants elevated limits. Whether a payment processor treats the traffic as ordinary or as risk. Whether a platform certifies the integration. None of that can be generated. It is built over years through consistent behaviour, and it can be destroyed in an afternoon by a system acting confidently on a bad rule. A new application starts with none of it, and that starting position is invisible until the first thing goes wrong.

Consent with provenance. Not a flag in a database recording that permission exists, but a record of how it was obtained, when, under which wording, in which jurisdiction, and what the person was told at the time — durable enough to be produced in a dispute months later, by which point the interface that captured it may not exist in that form. This is straightforward to build in principle. It is almost never built by somebody assembling an application in a weekend, because it is invisible unless somebody asks, and nobody asks until it matters.

Suppression that must never fail. Every system that acts on people needs a set of rules that hold under all circumstances: this person must not be contacted, this account must not be charged, this record must not be exported. What makes these hard is not the logic — the logic is trivial. It is that they must survive regeneration, refactoring, a new developer, a migration and an urgent change made under pressure. This is the clearest case in software where being right on average is worth nothing: there is no persuasive account of a suppression failure that begins with the system having been intelligent most of the time. A rule that holds most of the time is not a suppression rule. It is a preference.

A record that reconstructs the decision. Not logs of what happened, which most systems have, but an account of why: what the system knew, which policy was in force, who set it, what boundary applied, what it did, and what state it left behind. This is the difference between being able to say the system sent the message and being able to say why the system was entitled to send it. A regulator, a court and a serious customer all ask the second question. Intelligence without a reconstructable record is not autonomy. It is opacity — and no organisation should grant authority to something it cannot afterwards examine.

Reversal. The ability to undo an action, including its consequences in other systems that have already acted on it. Reversal is not the inverse operation, and the hardest version is not the clean failure but the partial one: the refund that stops halfway, leaving the order system and the payment system holding different accounts of reality. Nothing has failed loudly; the two systems simply disagree, and each will go on acting on its own version. Reversal is a designed capability, and it has to be designed before it is needed — which means before anybody has evidence that it will be. Where an action cannot be reversed, the authority threshold in front of it should be higher, which is a design rule and rarely treated as one.

A named party who is answerable. When software acts wrongly and somebody is harmed, the question is not only what failed. It is who is responsible. In a purchased arrangement the answer is written down, sits with an organisation that carries insurance and has something to lose, and is enforceable. In a self-built arrangement the answer is the company itself — which may be perfectly acceptable, and is a materially different position that ought to be entered deliberately rather than discovered afterwards.

These items share a property that separates them from everything in the previous section. Effort can be bought late. Standing, provenance and reversibility cannot. A company that decides in month twelve to take its obligations seriously can hire, staff a rota and write the runbooks. It cannot retroactively acquire three years of good conduct with a mail provider, or produce a consent record for a permission captured by an interface that no longer exists, or reverse an action through a design that was never built.

Which is why this layer, rather than the maintenance layer, is where the deflation stops. AI makes software creation cheaper. It makes trusted action more consequential — and the second effect is larger than the first for any system permitted to do anything that matters.

Key points

  • This layer is accumulated rather than made: standing with third parties, consent provenance, suppression, decision records, reversal, and a named answerable party.
  • A rule that holds most of the time is not a suppression rule. It is a preference.
  • Effort can be bought late. Standing, provenance and reversibility cannot.
  • AI makes software creation cheaper. It makes trusted action more consequential.

6

Why the Vendor Still Exists

Return to the objection that opened this essay, now that both sides of it are visible.

The forty-person logistics company was never choosing between having software and not having it. It was choosing where the obligation would sit. Self-building does not remove the software producer from the arrangement. It relocates the producer inside the customer, and the relocation is invisible on the weekend it happens because the obligation has not started yet.

Stated that way, the decision becomes a reasonable one to make rather than a trap. Some companies will look at the obligation, recognise that they are already carrying most of it, and build — correctly. Others will look at the same list and conclude that they would rather buy the promise than become the party who keeps it. Between those positions sits a range that will grow: managed specifications, shared runtimes, certified components, arrangements where the customer holds its distinctive rules and a producer maintains the machinery underneath. The decline of code scarcity produces more ways to source software, not one right answer — and the useful question in every one of them is the same. Who is carrying the obligation, and do they know they are carrying it?

For the producer, this changes what a software company is selling, and the change is not cosmetic. It can no longer sell access to a capability, because capability is becoming abundant. What remains sellable is everything this essay has described: an account of the customer’s operating truth that somebody keeps current; the machinery that carries a change safely across everything depending on it; the migration that makes arriving and leaving survivable; the accumulated standing that lets the software be trusted by third parties; and a named party who is answerable when it acts.

That is a harder business than selling software was. It is also a more defensible one, because none of it can be generated on a Saturday.

One misreading has to be closed off before it takes hold, because this argument can be turned into something it is not. Nothing here is a defence of expensive software. A great deal of what the industry charges for deserves to collapse: features nobody uses, foundations rebuilt for every product, integrations assembled by hand, implementation projects that exist because the product could not be adopted without them, and the organisational overhead layered on top of all of it. The obligation described in these six sections is not what made traditional software expensive. It was one line item among several, and often not the largest.

So the correct conclusion is narrow and it should be stated precisely. The waste should be removed. The promise should not. Affordability that is achieved by deleting the obligation is not affordability — it is unpriced risk sitting quietly on the buyer’s side of the arrangement, and it will be discovered eventually, usually by somebody who did not know they had bought it. The obligation has to be industrialised instead: shared foundations, reusable tests, monitoring that runs without being staffed, deterministic execution for repeated work, and regeneration when the governing truth changes. That is the difference between software that is cheap and software that is affordable, and the distinction becomes more important as the price falls, not less.

Which suggests something a producer can do rather than merely claim. Make the obligation inspectable. Show which specifications are current and when each was last reviewed. Show which dependencies are being watched. Show what authority the system currently holds and who granted it. Show how a decision can be reconstructed, and how quickly an action can be reversed. Every one of those is checkable, and none of them appears on a feature list. Trust moves from a brand promise to inspectable operating evidence — and at a low price that shift is not optional, because a low price is exactly what a buyer would expect an unpriced risk to look like.

The competitive question is no longer who can build the capability. It is who can keep it correct, current, connected and accountable as the world changes.

And the answer to the objection can now be put in a single line. The customer is not being asked to pay for software; the customer can produce software. The customer is being offered the chance not to become the party who is responsible for it — which is worth something, and worth more every time the software is permitted to act on its own.

The future software company will not win because it writes more code. It will win because customers trust it with more authority.

Thinks 2077

Business Standard: “India’s next scientific revolution needs patient philanthropic capital…Private wealth built the institutions that gave India its scientific capability long before the nation had economic strength or independence. There is a need to do it again.”

NYTimes: “To optimists, robots could rescue U.S. manufacturing by increasing productivity, solving shortages of skilled workers and giving Western carmakers a fighting chance at competing with Chinese rivals that enjoy lower costs. Boring but important jobs like sorting parts would be done by robots, freeing humans for more interesting and specialized work. But many experts who have studied the use of robots and other automation tools caution that there is also a gloomier scenario. “With any kind of automation, you generally reduce the amount of labor you need per widget that you make, and that’s generally a gain for society,” said Susan Helper, a professor of economics at Case Western Reserve University who testified on robotics before a congressional committee in April. “But,” she added, “how do you use that extra time that’s freed up? You can use it to just take away the interesting work and leave the human workers with the boring work. And have fewer workers that you pay less.””

Mike Grossman: “Build a great cultural environment because if people really love working for the company, when adversity strikes—and it always does—people will stay…I think of business as just an endless sequence of problem-solving exercises. You’re trying to analyze data in a structured way and then figure out the answer.”

Spyglass: “Open weight models may now be a viable alternative to the closed variety and, relatedly, distillation of such models may actually spur innovation in the sector.”

Software After Code (Software Foundry Series #5)

How machine-generated software changes what software is, how it ages, what it may be trusted to do, and who still needs to sell it

This series has argued four consequences of a single change. The Software Foundry named the change: AI alters the production function of software. The App-Stack Tax traced the waste that change exposes on the buyer’s bill. The Foundry Price named the number that follows. Inside the Foundry walked through the production system that makes the number sustainable. Each of those essays took the change as given and followed it somewhere.

This essay goes underneath them and asks what the change itself is — because the answer is larger than a cheaper way to make the same software.

The claim of this essay: AI will not merely help humans write software faster. It removes code from the centre of what humans make — and once that happens, software changes what it is, how it ages, what it may be trusted to do, and why a company should buy it rather than build it.

1

The Last Generation of Handwritten Software

For seventy years, software has been made in a workshop. The workshop grew spectacularly more powerful. Machine code gave way to higher-level languages. Memory management became automatic. Libraries removed the need to rebuild common functions. Version control made large collaboration possible. Open source let each generation begin higher than the last, and the cloud removed the burden of shipping. Yet the central act stayed recognisable: a person understands a need, translates it into instructions a machine will execute, and hands the result to other people who will maintain it.

The model remained artisanal even inside the largest software companies. A modern engineering organisation may employ thousands of people and automate its testing, deployment and monitoring, and its output still rests on human beings making millions of implementation decisions. Software scaled by applying more engineers, better abstractions and larger organisations to the same act. The craft became industrial in size without becoming industrial in method — which is why the economics never moved. The craftsman acquired better tools and remained the bottleneck, and everything the industry charged was built on top of the bottleneck.

This is a familiar shape, and the industry that lived through it before knows how the story ends. The first machines in the textile trade were sold to weavers as better tools: a faster shuttle, a mechanical spinner, a device that let one artisan do the work of three. For a while that description was accurate. Then it broke — because the significant change was never that a weaver became faster. Production left the workshop. The unit of output changed, the organisation of labour changed, the consistency of output changed, the economics of cloth changed, and the number of people who could afford to wear it changed. Skill moved out of the individual hand and into machines, processes, gauges and standards. The tool story was true and small. The system story was the one that mattered.

AI coding is currently being told as a tool story. Engineers finish features faster. Boilerplate disappears. A small team does the work of a larger one. A backlog shrinks. All of it is likely, and all of it stays in the artisan register.

The consequence that matters is not that programmers will type less. It is that code will stop being the primary thing humans produce.

The profession will be redrawn rather than removed — engineering did not disappear when factories arrived. But its centre of gravity moves. Fewer people will spend their days translating already-understood requirements into routine implementation. More will define architectures, resolve the ambiguous domain question nobody wrote down, design the evaluations that decide what correctness means, investigate failures, govern what the system is allowed to do, and improve the production machinery itself. The craft is not deleted. It is embodied at a higher level and multiplied through machinery — which is what happened to every craft that industrialised, and it was never the smaller job.

Three shifts follow, and this essay is organised around them. Software moves from handwritten to generated. It moves from maintained to regenerated. And — the shift most often announced and least often stated correctly — it moves from deterministic to agentic. Each changes something different: the first changes how software is made, the second how it ages, the third what it is permitted to do. Taken together, they change what a software company is for.

Figure 1. Three shifts, and what each one changes.

Key points

  • For seventy years the production model stayed artisanal. The craft became industrial in size without becoming industrial in method.
  • The Industrial Revolution’s tool story was true and small. The system story was the one that mattered.
  • Three shifts follow: handwritten to generated, maintained to regenerated, deterministic to agentic.

2

Code Becomes an Artefact

Start with what a software product has been. A codebase is the accumulated record of every decision a company ever made about how its product behaves: business rules, edge cases, the fix from the incident three years ago, the workaround for a customer who has since left. Some of that is written down elsewhere — in requirements documents, architecture diagrams, tests. But when those disagree with the implementation, the implementation wins. For most software companies the code is not the record of the truth; it is the truth, and reading it is the only reliable way to find out what the product does. This is why software is expensive to change, why the people who wrote it are hard to replace, and why the sentence “we cannot touch that module” appears in every engineering organisation past a certain age.

Machine generation inverts the arrangement. The chain that ran from human intent to human-written code to running software now runs from human intent to a specification, from specification to a machine-generated implementation, and from implementation to continuous validation against the specification that produced it. The implementation remains necessary. It stops being the primary human-authored asset.

Figure 2. The source of truth moves — and regeneration returns to it, not to the code.

The word specification needs rescuing here, because a generation of requirements documents that nobody read and nothing enforced has degraded it. A prompt requests a result. A specification establishes what must remain true. It states the job to be done, the entities and data definitions, the rules of the domain and where they came from, the actions permitted and the actions prohibited under any circumstances, the evidence required before an action, the expected behaviour in ordinary cases, the treatment of exceptions, the conditions for escalation and the path for rollback. And it contains acceptance tests precise enough for another machine to judge whether the generated implementation satisfies the intent. It is closer to a contract than to a brief.

This changes who can participate in making software. The expert on returns, lending, logistics, payroll or clinical operations need not become a programmer to shape the system. Their valuable knowledge was never the syntax of a language. It is the operating reality: which rule applies, which exception matters, which action is irreversible, which failure is unacceptable. The specification is where that judgement becomes executable.

Code tells a machine how the system was implemented. A specification tells the organisation what must remain true.

From that follows the operating discipline of the whole era, and it fits in one sentence: humans edit the specification; machines regenerate the implementation.

The sentence sounds procedural. It is not. It carries a consequence with teeth, and any organisation adopting generated software will meet that consequence the hard way if it does not adopt the rule deliberately. An emergency patch that is not translated back into the specification is silently deleted by the next regeneration. Somebody fixes something at two in the morning, directly in the implementation, because a customer is down and the fix is obvious. The fix works. Weeks later, a regeneration runs for an unrelated reason and the fix is gone — because it was never part of what the system had been told must remain true. This is not a defect in the machinery. It is the discipline that prevents a new kind of legacy estate: vast quantities of machine-written code whose governing decisions nobody can reconstruct. But it only works as law, and it is brutal as a surprise.

Code does not become irrelevant in this arrangement. It becomes inspectable evidence of an implementation at a moment in time. Engineers will still read it when performance, security or an unusual failure demands it. What moves upstream is the enduring asset.

Nor does this imply one monolithic document governing everything. Specifications will themselves be layered — organisational policy, vertical rules, product contracts, interface definitions, component guarantees, tests — and the relationships between the layers will matter as much as their contents. The central difficulty of this era will be keeping those layers explicit enough for machines to compose and legible enough for people to govern. That is a hard problem and it is not solved. It is, however, a better problem than the one it replaces.

Key points

  • The codebase used to be the source of truth. The specification becomes it.
  • A prompt requests a result. A specification establishes what must remain true.
  • Humans edit the specification; machines regenerate the implementation.
  • An emergency patch not translated back into the specification is deleted by the next regeneration.
  • Specifications will be layered, and keeping the layers both composable and governable is the era’s central difficulty.

3

Software That Does Not Have to Grow Old

Software ages in a way physical products do not. The program may still do exactly what it was written to do while the environment around it keeps moving. A framework loses support. An interface changes its authentication rules on somebody else’s schedule. A library develops a vulnerability. A regulator changes what may be stored or communicated. A browser alters a security policy. The engineer who understood the pricing module leaves, a temporary patch becomes permanent, and an undocumented exception hardens into a dependency nobody dares remove.

Some of that accumulation is poor work. Most of it is the price of time. Every implementation records assumptions about the world at the moment it was written, and as those assumptions decay the organisation layers repairs on top of them. Eventually, changing a small part of the system requires understanding a history that no document fully contains. This is why software companies spend a rising share of their engineering capacity standing still — and why they eventually price that capacity into the invoice.

Specification-driven production offers a different relationship with time. When a shared connector, policy, model or regulatory rule changes, the production system updates the relevant specification or foundation, identifies the products affected, regenerates their implementations, reruns the required evaluations and deploys the validated versions. Traditional software organisations maintain applications one at a time. A specification-driven system maintains the truths from which applications are rebuilt. For the customer, that resolves into a promise the software industry has never been able to make:

Software that does not have to grow old.

The gain is not only speed. It is consistency: a security correction to identity, a repaired connector or a new jurisdictional rule propagates across every capability that depends on it, and each one inherits the same correction, the same test and the same evidence that it was applied. In the handwritten world that same change was a separate project in every application, done to a different standard in each, with the evidence assembled afterwards by hand.

It also changes what a release is. In the handwritten era a release bundled months of implementation work into a version the customer had to accept, and version numbers carried real anxiety. In a regenerated system, change becomes smaller, more continuous and more targeted — a connector rebuilt, a policy updated, a control strengthened, only the affected capabilities revalidated. The release stops being a shipment of accumulated code and becomes a certified statement that the current implementation still satisfies the governing truths.

Now the honesty condition, which belongs in the middle of this argument rather than a footnote, because without it the promise is marketing.

The new technical debt is specification debt. AI does not abolish ambiguity; it moves fragility upstream. The failures of this era will not come from ageing implementation code. They will come from incomplete intent, missing edge cases, policies that contradict each other in a situation nobody imagined, acceptance tests that pass without proving anything, authority boundaries never written down, and domain assumptions that were true when captured and are not true now.

A poor specification can be worse than poor handwritten code, because a human team introduces inconsistency slowly and a production system can spread a flawed rule across every product in minutes. That is not a smaller problem than technical debt. In one respect it is a larger one, because the fault is now systematic rather than local — and it demands far stronger discipline about provenance, review, testing and ownership of the specification itself.

Figure 3. Fragility does not disappear. It moves upstream — and changes character.

So the recompile property is not a free gift of the technology. It is the reward for specification discipline. Software does not stay young because a machine can rewrite it. It stays current because the organisation has kept a precise, tested account of what must remain true while letting the implementation change underneath it.

Key points

  • Traditional software maintains applications one at a time. Specification-driven systems maintain the truths from which applications are rebuilt.
  • The gain is consistency as much as speed: one correction, one test, one piece of evidence, everywhere it applies.
  • A release becomes a certified statement that the implementation still satisfies the governing truths.
  • The honesty condition: implementation debt gives way to specification debt, and a poor specification propagates faster than a human team could.

4

From Deterministic Applications to Agentic Actors

The third shift is the one most discussed and most often mis-stated — and the mis-statement matters, because it leads organisations to build the wrong thing.

Traditional software is predetermined, and the word is precise. A human anticipated the situations the software would meet, encoded a response to each, and shipped the result. Given this input, produce that output. If this happens, run that. The application waits; a person operates it; a situation outside the designed path either fails or is handed back to a human. The intelligence in a conventional application is entirely the intelligence its designers had in advance — which is why every enterprise implementation contains a discovery phase whose purpose is to find out, before go-live, everything that might ever happen.

Agentic software works from the other end. It receives a goal, observes the situation as it is, chooses among available actions and adapts as conditions change. It can interpret intent rather than only accept structured input; decide which workflow applies; construct a path when no predefined path fits; use several systems to do it; ask when the request is ambiguous; monitor the result and change course; and stop when its confidence or its authority runs out. The governing instruction is no longer if this, then that. It is: given this goal, this context and these boundaries, decide what to do next. The software stops being a tool in the user’s hand and becomes a delegated actor inside the organisation.

This is a deeper change than the one usually reported. The popular version says the interface moves from clicks to chat, and the popular version is a distraction. Chat may be convenient for expressing an ambiguous goal; a dashboard may be better for inspection; a form may remain the safest way to approve a precise transaction. Those are design decisions taken moment by moment. Software is not moving from being clicked to being talked to. It is moving from being operated to being authorised.

Delegation does not arrive in a single leap. It forms a ladder — and it is worth stating as a ladder, because no product arrives at the top of it and none should be sold as though it had.

Level What the software is trusted to do
Observe Read the situation and explain what is happening
Recommend Propose an action, and leave it there
Prepare Construct the action in full, and wait
Execute Act, after explicit approval
Operate Act without asking, inside agreed policies
Own a goal Choose its own actions towards an outcome, and escalate the exceptions

Table 1. The authority ladder. Interface is a design decision at every rung; authority is a commercial one.

Two products of similar intelligence may therefore have very different value: one can advise, and the other has earned permission to act.

Which is why authority is earned, not asserted. A system does not become trustworthy because the model behind it is capable, or because it sounds confident. It becomes trustworthy through evidence that its actions were right before, permissions that constrain what it can reach, an audit record that reconstructs why it did what it did, and the ability to reverse an action that should not have happened. A system trusted to prepare a refund, approve a supplier, send a communication or alter a price must be able to show what it knew, which policy applied, which boundary it respected and how the action can be undone. Capability establishes what a system could do. Only that apparatus establishes what it should be allowed to do — and the second is what the buyer is paying for.

There is a second consequence, and it is the reason specification debt is not an abstract worry. Delegation converts specification gaps into consequences. An unclear suppression rule sends the message. An incomplete permission model exposes the data. A missing escalation path approves the exception. When software displayed and waited, a missing edge case produced a wrong screen and somebody caught it before anything happened in the world. A deterministic system fails visibly and locally. An agentic system governed by a weak specification fails confidently and at scale.

So the quality question for future software is not only whether the agent is intelligent. It is whether the institution around it has defined authority precisely enough to let intelligence act safely.

Which is also why the three shifts are not independent. Generated software makes agentic behaviour practical, because workflows no longer have to be anticipated and encoded in advance. Agentic behaviour then makes specification discipline mandatory, because the stated boundaries are the only thing standing between a capable system and an action nobody authorised. And that gives the specification a second job it never had:

In the old world, code defined what an application could do. In the new world, the specification defines both what the system should become and what it is permitted to do.

Key points

  • Deterministic software encodes responses a human anticipated. Agentic software pursues a goal within boundaries.
  • The shift is not from clicked to talked to. It is from operated to authorised.
  • Two products of similar intelligence differ in value: one advises, the other has earned permission to act.
  • Delegation converts specification gaps into consequences.

5

The Two Planes

Everything above invites a conclusion that would be expensive to act on: that software is moving from deterministic to agentic, and that the destination is a system reasoning its way through every task. That architecture would be costly, unpredictable and hard to audit. Correcting the conclusion is the difference between something that can be afforded and something that cannot.

The accurate statement has two halves. Decision-making becomes agentic. Execution stays deterministic — permanently, and by design.

The agentic control plane does the work that benefits from judgement: interpreting what is wanted, configuring the workflow, setting the policy, resolving ambiguity, diagnosing an exception, deciding that something has changed and the approach should change with it. The deterministic execution plane does the work that benefits from being identical every time: processing events, enforcing permissions, checking eligibility, transforming data, calling stable interfaces, running to schedule, writing the audit record.

The agent decides what should happen. The deterministic system ensures it happens correctly, repeatably and cheaply.

Use intelligence to decide what should run. Do not pay for intelligence every time it runs.

An agent might read a new returns policy, translate it into a workflow, and identify the exceptional cases needing human review. Once that is approved, the ordinary path should run as deterministic software until new information gives the control plane a reason to reconsider it.

Two independent arguments arrive at this same architecture, which is usually a sign that it is right.

The first is economic. Inference is not free, and does not become free merely because model prices fall. It is a cost that scales with usage rather than with customers — the opposite of the shape the software industry is built on. A product that reasons afresh on every execution has a cost line that grows precisely where its revenue does not, and the gross margins reported by AI-native products so far sit well below the seventy-five to eighty-five per cent that software companies organised themselves around. Compiling repeated work into deterministic execution is therefore not an economy measure bolted on to protect margin. It is the reason an affordable agentic product can exist at all.

The second argument is about trust, and it is the more important of the two. A system that reasons its way to each individual action cannot fully explain why it acted, because the reasoning was assembled in the moment and is not reliably reproducible. A payment calculation, a consent check or an eligibility rule should not vary because a model sampled a different chain of reasoning. A system that reasons about configuration and then executes deterministically can be explained completely: this is the policy, this is when it was set, this is who approved it, this is the rule that fired, this is what it did, and this is how to undo it. An agentic system that cannot be reconstructed cannot be given authority — which returns the argument to the ladder. The two planes are not merely how agentic software becomes affordable. They are how it becomes accountable enough to be trusted with anything that matters.

The boundary between the planes is itself a design decision, and the most interesting one in the architecture. A novel exception arrives in the control plane and is handled there, with close human review, because nobody has seen it before. Once the organisation understands the pattern, it can be written as policy, tested, and compiled downwards into the execution plane. The system therefore learns in two directions: the agent gets better at handling novelty, and the deterministic machinery absorbs whatever has stopped being novel. Intelligence is spent where uncertainty remains, and withdrawn from where repetition has already removed it.

Figure 4. Agentic control over deterministic execution — and the boundary that moves as the organisation learns.

Agentic where judgement is valuable. Deterministic where reliability is essential.

Key points

  • The shift is not deterministic to agentic. It is agentic control over deterministic execution.
  • Use intelligence to decide what should run; do not pay for intelligence every time it runs.
  • The economic argument: inference scales with usage, not with customers.
  • The trust argument: a system that cannot be reconstructed cannot be given authority.
  • The boundary moves: novelty enters at the top and is compiled downwards once it is understood.

6

The Objection: Why Buy Software At All

Every argument in this essay carries an obvious reply, and any reader who has followed it this far has already formed it. If AI has collapsed the cost of producing software for vendors, it has collapsed it for customers too. A business can describe a workflow to a coding agent, connect a database and have a useful internal tool in days. Some companies are already choosing that route for selected functions rather than renewing expensive general-purpose software.

The argument has to meet this, because a thesis that cannot survive its own logic is not a thesis. And it should meet it without defensiveness, because the objection is correct as far as it goes. The answer cannot be that customers are incapable of building. Many are, and some should. A company with unusual workflows, strong technical leadership and a willingness to own the system may gain speed, fit and substantial savings by generating its own. The software market will contain far more self-built capability than it does today.

But look closely at what the decision is. Self-building does not remove the software producer from the arrangement. It relocates the software producer inside the customer.

The act of generation is only the beginning. Somebody now owns the specification and must keep it current as the business changes. Somebody monitors the connectors and repairs them when a third party ships a breaking change on their schedule rather than yours. Somebody manages permissions as people join and leave, tests every release, responds at midnight when something fails, tracks the regulation that changed in one market, patches the vulnerability disclosed on a Friday, answers the colleague who cannot make it work, and holds the institutional memory when the person who built it moves on. And somebody is answerable when the software takes an action it should not have taken — which, by the previous two sections, is now a category of failure that did not exist when software waited to be operated.

The distinction is easy to miss because software is first encountered as a repository and an interface, and those can now be produced astonishingly quickly. The continuing obligation is the invisible part.

A repository can be generated in an afternoon. A dependable product is a continuing institutional promise — that the system will work not only on the day it is demonstrated, but after an interface changes, a key employee leaves, an exception appears, a regulator asks for evidence, or an automated action has to be reconstructed and reversed.

This is a spectrum, not a binary. At one end a company buys a finished capability and delegates most of the continuing obligation. At the other it generates and operates everything itself. Between them sit managed specifications, shared runtimes, certified components and co-produced systems in which the customer controls its distinctive rules while a provider maintains the production machinery. The decline of code scarcity will create more ways to source software, not one universal model — and the important question in each of them is the same: who is carrying the obligation, and do they know it?

Figure 5. Software still gets built either way. What differs is where the continuing obligation sits.

So the role of the software company changes rather than disappears. It can no longer rest on the scarcity of code. Its reason to exist becomes the production and operating system around the code: maintained specifications, tested components, migration, security, support, permission infrastructure, monitoring, updates, and a named party who remains answerable.

The competitive question in software therefore changes shape. It stops being who can build the capability, because increasingly the answer is everyone. It becomes:

Who can keep the capability correct, current, connected and accountable as the world changes?

That question has a different answer — and it is not the one with the most engineers. The next essay in this series, The Expensive Last 10%, takes up the continuing obligation in detail.

Key points

  • The objection is correct as far as it goes: many companies will build what they used to buy, and some should.
  • Self-building does not remove the software producer. It relocates it inside the customer.
  • The decline of code scarcity creates more ways to source software, not one universal model.
  • The competitive question moves from who can build it to who can keep it true.

7

What Becomes Scarce

Every technological abundance shifts value towards a new scarcity. When computation was scarce, access to machines mattered. When distribution moved to the cloud, product focus and customer acquisition mattered more. As AI makes code abundant, advantage migrates to whatever code generation does not supply — and the list is worth walking, because the list is where a reader decides what to do on Monday.

Verified operating specifications. The tempting claim is that models lack domain knowledge, and it is the wrong claim. Models increasingly know how a returns process, an approval, a replenishment cycle or a reconciliation normally works, and they will know it better every year. Generic domain knowledge is becoming abundant. Current, organisation-specific operating truth is not. Which policy did this business choose, and how does it differ across its categories and jurisdictions? Where did the rule come from, and who owns it? Which exceptions have been accepted, by whom, and on what grounds? Who may override, and what happens when two rules conflict? What evidence would demonstrate compliance if somebody asked next week? None of that lives in a model. It lives in an organisation, mostly undocumented — and converting it into executable form is the scarce act.

Migration. Important software is rarely installed into an empty company. It enters a living organisation full of partial data, undocumented workarounds, stale permissions and processes that exist because somebody once met an exception. Discovering that reality, mapping it, reconciling it and switching without a bad week is not solved by producing a fresh application. The incumbent’s deepest protection was never its feature list; it was the customer’s fear of leaving — which makes this capability worth more now, not less, precisely because everything else about switching has become easier.

Accountability. Identity, permission, consent with legal provenance, suppression that must never fail, the audit record that reconstructs a decision, escalation, rollback, and a named party answerable for an action. As intelligence becomes common, permission to act becomes the differentiator. The valuable system is not the one capable of making the decision. It is the one the organisation is prepared to authorise.

Distribution and trust. AI collapsed the cost of producing a narrow capability. It did nothing to the cost of being discovered, believed, adopted and retained. Buyers still need evidence, references, continuity and confidence that the provider will be there after the first release. Production may be abundant while attention and belief remain scarce — and every argument that begins “software is nearly free to make now” and ends “so the market goes to whoever makes it” has skipped the only step that did not become cheaper.

The recompile system. The machinery of section three: identify which capabilities depend on a changed rule, regenerate them, test them, deploy them safely. A model can produce code. An institution must preserve the dependencies, standards and evidence that make portfolio-wide change reliable. This is also the answer to the self-build objection that no price argument can supply. The buyer is not purchasing software that works today. The buyer is purchasing the machinery that keeps it true tomorrow.

Domain judgement in executable form. As the cost of implementation falls, the person who can state the right rule, the exception, the test and the escalation path becomes more valuable than the person who can implement it. Over time the industry may develop production systems through which such experts contribute verified knowledge without having to found and operate software companies — a lawyer contributing a jurisdictional rule, an accountant a reconciliation policy, a clinician a triage protocol, an operator the exception logic learned over fifteen years, each contribution versioned, tested and governed rather than buried in an advisory document or translated imperfectly by a distant implementation team. The important unit would be the trusted specification rather than another isolated application. That is where this is going as an industry. It is not where anyone is yet, and a producer who announces it before earning it will be asked, reasonably, to show the thing.

AI makes this cheaper AI does not supply this
Writing code Distribution
Producing variants Trust
Modifying software Migration
Building narrow capabilities Accountability
Testing routine behaviour Domain judgement
Serving long-tail cases Permission to act

Table 2. Value moves to the right-hand column.

AI commoditises production and raises the value of everything production cannot commoditise.

This will rearrange the industry rather than merely discount it. Some categories will be absorbed into general-purpose agent environments. Some companies will build more for themselves. Some vendors will become infrastructure providers, custodians of specifications, or operators of trusted vertical systems. The margin once earned from the scarcity of code will have to be re-earned through responsibility, distribution, migration and continuous correctness.

Key points

  • Generic domain knowledge is becoming abundant. Current, organisation-specific operating truth is not.
  • Migration matters more, not less, as everything else about switching becomes easier.
  • The valuable system is not the one capable of deciding. It is the one the organisation is prepared to authorise.
  • The buyer purchases not software that works today, but the machinery that keeps it true tomorrow.
  • The margin once earned from code scarcity must be re-earned through responsibility, distribution, migration and correctness.

8

The Software Foundry

If the factory was the organisational answer to machines producing physical goods, something equivalent is required for machines producing software — and it is not a faster version of a software company. It is a different institution.

This series has described it in detail elsewhere, so name it rather than rebuild it. A software foundry treats specifications as its control surfaces and generated implementations as its output. It validates behaviour through inspection built into production rather than bolted on after it. It builds shared foundations once and lets each product be only the job it exists to perform. It pairs agentic control with deterministic execution. It grants authority through permissions, evidence and reversibility rather than through confidence. And it regenerates its estate when the world underneath it changes. The economics of that arrangement, the price it makes possible and the production system that sustains it are the subjects of the earlier essays.

This essay has stayed above that floor deliberately, because the larger change is not one company or one product category. It is a change in the nature of software itself, and it will be true whether or not any particular institution is built to meet it.

What this closing movement can add is only this. A foundry is not a way to make software cheaply. It is a way to be accountable for software at a price a far larger number of businesses can pay — and those are different ambitions, which is why cheapness alone has never won a software market and will not win this one.

The first age of software taught humans to translate intent into code. The next teaches machines to turn intent into software. Code does not disappear; it recedes from the centre of human attention, in the way machine instructions receded when compilers arrived and almost nobody mourned them. What moves into the centre is harder and more interesting: the specifications that state what must remain true, the systems that can rebuild software safely when the world shifts underneath it, and the institutions trusted enough to let software act.

The advantage will not belong to whoever writes the most code. It will belong to whoever can turn human judgement into dependable capability — and keep it true.

Thinks 2076

Sanjeev Sanyal: “In most countries, educated youth in the early to mid-twenties tend to have somewhat higher unemployment rates than others. It evens out, by the way, by the time they are 30 (years of age). But there is a phase when they tend to be somewhat more unemployed. In the specific case of India… there are many reasons this happens. One of them, of course, is that this is a phase where many, many educated people take time out to write various kinds of government and other exams like UPSC, etc.”

David Dunning: “One of the reasons why people are overconfident in general is that they tend to look for reasons supporting their ideas. In fact, one of the pieces of advice we always give is: If the decision is important, stop and ask why you might be wrong. In planning, one of the procedures that’s recommended for a group is to project yourself into the future and imagine your initiative failed spectacularly. What happened to make it fail? And say, “OK, what can we do to prevent those stories from happening?” That’s called a “pre-mortem.” Doctors do this as a matter of course. They have a different name for it. They don’t diagnose you; they do a “differential diagnosis.” They may have an idea, but they have to think about what else it can be. OK, you might have Lyme disease, but what else are your symptoms consistent with? We’re going to test you to rule out everything that we can. That’s core to medicine, thinking of alternatives, thinking how you could be wrong.”

: “We should try to get our incentives right, but also—incentives be damned, do the right thing. If we seek an end to our permanent problems, we must cultivate in ourselves the same three things Charles Sumner said we should look for in a leader: “The first is backbone, the second is backbone and the third is backbone.””

FT: “IT services have played a crucial role in driving India’s economic growth and expanding its middle class. But many of the formulaic tasks that underpin the industry, which employs 6mn people and contributes about 7 per cent of GDP, can now be performed by generative AI. If the sector cannot pivot to higher-value work, that could spell more trouble for a government already struggling to create enough jobs for India’s huge workforce. After years of headcount growth the sector has begun to shed jobs; a July report from rating agency S&P stated that by March this year Infosys and Wipro had reduced staffing by 5-6 per cent from 2023 levels, with similar trends evident at Tata Consultancy Services (TCS) and Tech Mahindra Ltd. Investors are also turning bearish on the sector’s prospects.”