When there is suddenly too much output
Many times more output, and nothing moves faster: checking now costs more than creating
Organisations that language models have reached tend to see the same picture: producing a working artifact — a contract, a policy, a report, an operating manual, a deck, a mockup — has become many times cheaper, while the organisational cycle has shortened nowhere near proportionally. For an individual the gain is obvious and measurable. At the level of the organisation it gets lost somewhere along the way.
The stock explanation AI apologists offer is that people haven't mastered the tool yet — give them a year. It's partly true: every technology has its J-curve, where the payoff arrives only after the processes have been rebuilt around it. But in my view a second reason is at work, and training doesn't cure it. Production and the grounds to trust the result used to be glued together in one person and one stretch of time: whoever wrote the document earned the right to be trusted by the same labour. He thought, cross-checked, tried and discarded options, and the finished document carried the traces of that work. Generative models unglued these two things: the artifact is produced in minutes, and the proof of why it deserves belief is nowhere supplied.
What follows are my musings on how that proof used to assemble itself, why it stopped assembling, and what to do now.
Expensive production doubled as proof
Volume was an honest signal of time invested — right up until faking it cost as much as doing the work
There's a habit few people ever put into words, but the management of knowledge work rested on it: volume was an honest signal of time invested. A forty-page memo meant somebody had spent weeks, and deserved attention for that alone. We read not only the content — we read the evidence of labour sewn into the sheer thickness of the document. And the signal didn't lie, because faking it cost roughly as much as doing the work for real.
Design review, state expert appraisal, incoming inspection, sampling-based acceptance solve different problems, but they all grew up under the same constraint: checking capacity is limited, so it has to be rationed.
Accountability rests on the same assumption. The formula «the responsible officer reviews and signs» quietly presumes that the volume is physically readable, and that the signature certifies a person actually looked and agrees with what's written.
But something changed. The thickness of a document no longer means anything — do we still believe it? At the Monday meeting the department head reports that this week the team delivered a policy, two reports and a forty-page requirements spec, and everyone nods. Does anyone ask whether those forty pages were read by a single person capable of confirming what's written there is true? At a factory that pile of output is called work in progress, and there everyone can see it: it takes up floor space, gets in the way, annoys the section chief. On the shared drive it looks like a productive week.
The intent went missing from the result
What gets handed over is the result, not the intent — and the receiver rebuilds it at his own expense
Between people, the result is what gets passed on — not the intent. A text, a drawing, a mockup or a spreadsheet contains what came out, and not why it was done, which options the author tried and discarded, where he had doubts and where he knew for sure. The receiver has to reconstruct the intent, and that reconstruction has a cost that your own production doesn't.
That doesn't mean checking is always dearer than doing: a short calculation is faster to read than to perform, a boilerplate contract faster to proofread than to draft from scratch. But acceptance carries a fixed surcharge for rebuilding context, and while production was expensive, the surcharge got lost against that background. The moment generation became cheap, the surcharge turned into the main cost line.
Two more things used to absorb the difference. Volume: a person wrote as much as he could manage, so the output matched the attention span of one reader. And the presence of an author: you could ask him what he meant, and half of the acceptance happened in a short conversation rather than a line-by-line read.
Now volume is no longer bounded, and there's often no intent behind the text — not because the author hid it, but because he never formed it: the artifact was produced by a model from a single-cell prompt, and the person who brought it sometimes can't explain himself why it says what it says.
There's no one left to ask.
Hence a consequence that seems to me more important than the overproduction itself. Behind forty unverifiable pages, the culprit is more often a missing source of truth than excessive volume. The old product of human diligence was capital: people wrote ten pages about what they firmly knew. The model writes forty and fills with plausible text exactly those spots where the process has no accessible, authoritative source of truth. The knowledge may well exist in the organisation — a piece each with three different people, in unwritten form, in the next department, in the product itself or in a physical process — but the model has no access to it, and nobody owns what counts as the right answer here. What's written can't be verified not because it takes long, but because there's nothing to check it against.
Software noticed later than everyone — but louder
As long as verification scales with production, overproduction is acceptable
Curiously, the loudest talk about overproduction today comes from software development, though by my observation it surfaced there later than in legal, engineering-design and management processes. For lawyers and design engineers it showed up almost immediately: the assistant assembled a contract in fifteen minutes, and the contract went into the queue for the lawyer; a project manager produced forty pages of an operating manual, and they joined the line to the only person able to confirm the text matches the real product.
Development held out longer for an understandable reason. Engineers invented testing long ago — the part of the acceptance criterion they managed to express in machine-executable form. The trade got used to the proof of fitness living next to the result as a separate artifact and running automatically. As long as verification scales with production, overproduction is acceptable, and code took the multi-fold growth of generation calmly.
It hit the wall here too — exactly where checking became human again: reading other people's changes, judging architectural decisions, asking «are we even building the right thing». So development is walking the same road, just with a head start of accumulated discipline.
The conclusion comes out simple: overproduction is acceptable where the result has something to be checked against — a test, a calculation, a measurement, a fact. And destructive where you have to check with your eyes against nothing.
People signed without reading before, too
They did — but behind every paper stood an author you could hold to account
You'll say: come on, nobody really read anything before either. Policies were signed unread, volumes of design documentation were skimmed diagonally, contracts were checked on three clauses out of forty, and the board of directors got its papers a day before the meeting and opened them in the car on the way. Organisations ran on trust in people, not on reading documents, and the unread page was a fiction long before language models.
I partly agree with the objection — but there was a safety catch, and it wasn't reading. It was the author.
Behind the document stood a person you could question, and answer to later. He knew he would be asked, and that knowledge worked harder than any proofreading: writing obvious nonsense is a bad deal when a month later you'll have to explain it out loud at a meeting. The real check happened not at the moment of signing but in the author's head, in advance, and the acceptor's signature was less an act of control than a transfer of reputational risk onto a specific surname.
That is exactly what broke. The person who brings forty model-made pages didn't pre-check them in his own head — he didn't write them. You can still hold him to account, but the substance of his answer will be «that's what it generated».
The usual argument is about what happens next: either we hit the wall of regulation and slow down, or we wave it off and lose quality. To me the argument is empty — if nothing changes, the numbers will stay great, the on-paper controls will keep working, and the accumulated defect will surface all at once and at the worst possible moment. And nobody will be to blame: when the document is longer than a human can read, the signature turns into a checkbox automatically. It won't glue itself back together — verification will have to be rebuilt, by hand and on purpose.
Proof is now a separate line item
The cost of checking moved from man-hours into compute — and compute can be bought
Checking by reading costs exactly as many pages as there are: twice the text — twice the man-hours. Reasoning models with a long chain of thought break that arithmetic. Where they're given the source, a standard, a checklist, a fact base or an external tool, they take over most of the routine: they check the document against its source, catch internal contradictions, find the gaps, verify the form.
Hence a rule: it is reasonable to spend more on checking than on production. Today it's usually the other way round — one pass of generation and, at best, one proofread. How much to spend follows from risk: the scarier the missed error and the weaker your available sources of truth, the more checking.
Sampling doesn't die from this — it moves up a level. When every document goes through a full automatic check, what you sample is no longer the documents but the loop itself: what it systematically misses and whether it has drifted over time. And where a test is expensive or destroys the product, sampling isn't going anywhere.
The whole construction has one condition: the checkers' errors must be independent. Two models of the same family checking each other are two pressure gauges from the same factory with the same zero offset: the readings will agree even where both are lying. Same architecture, same training recipe, overlapping sources, identical task framing, one shared pull towards the plausible — such a pair catches random noise beautifully and doesn't see the systematic error. What exactly was in the training data nobody outside knows, but the correlation of errors between models shows up reliably, and different families reduce it rather than remove it.
So dissimilarity has to be built on purpose: different model families and different roles — one checks against the source, another plays the opponent and is obliged to find a defect, a third walks the standard formally. And the dissimilarity of evidence matters more than the dissimilarity of models: a model plus a calculation plus a registry extract plus a measurement on site is stronger than just five different models.
Coherent is not the same as true
Inside a text you can verify everything except one thing: whether it's true
The temptation after all this is obvious: build a conveyor where the machine produces and checks, and the human appears at the end to sign.
It won't work, and not because of model quality. A model will verify the document is coherent: it doesn't contradict itself, its inputs, the standard or the previous version. But that the inputs are wrong, the goal is off, and the machine on the shop floor isn't the one named in the spec — that it will not see: none of it is in the text, all of it is in reality, and the model only sees someone's description of that reality.
And the boundary here doesn't run between machine and human. What leads out to reality isn't necessarily a person: a sensor, a test rig, a lab assay, a state registry, a bank transaction, a geodetic survey — all of these are sources of truth, and every one of them beats the most attentive reading. While a person reading the same wrong document may not touch reality at all.
Two things remain with the human. Saying what counts as good here — which piece of reality is taken as true and what the checking runs against. And making the decision together with its consequences. Disputed cases land there too: when two sources disagree, when the defect is of an unfamiliar kind, when the goal itself has changed, or when safety has to be traded against money and deadlines — where there's no criterion, a human decides. Everything else — working on text by means of text — goes to the machine.
And a check built around a wrong source of truth is more dangerous than no check at all, because it manufactures unearned confidence. A crooked input used to be caught by the first live person who read it and winced. Now it passes five green checks and arrives at the signer with a report of full compliance.
What lands on the signature changes too. Take the intermediate people out of the chain and the whole volume rides through to the final signature; the signer becomes the bottleneck. So what goes on his desk should be not forty pages but a digest: what was checked and with what, what didn't add up, what was fixed, and which questions the machine couldn't settle. He still signs the whole document — the digest shortens the reading, not the responsibility. And the signature itself, strictly speaking, doesn't prove the person read anything: signing rights get delegated, responsibility gets distributed, the signature only fixes who answers for it. All the more reason for the evidence pack to travel with it.
What the new rulebook looks like
The verification loop is a product too, and it needs testing itself
The old rulebook answers one question — who checks whom: positions, signatures, approval deadlines. The new one will have to answer different questions: what each document is checked against, how dissimilar the people and models checking it are, where the conveyor is obliged to stop, and at what point a human steps in.
Testing regulations isn't new — internal audit does exactly this: it looks at whether the control catches what it was set up to catch. What's new is that the model loop can be measured in numbers: regularly run documents with pre-seeded defects through it and count how many get caught.
The share of caught defects alone isn't enough: what matters is how many false alarms, which defect types slip through more often, and whether quality drifts over time. Otherwise it's easy to build a checker that finds a defect in every paragraph and looks flawless on paper. As far as I can see, this is the cure for box-ticking: trust in the check becomes measured rather than declared.
The second thing to build is memory. In document approval a remark usually stays in the email thread or in the expert's head and never becomes a rule the machine can apply on its own. Without that translation the loop checks every document from scratch, and the expert makes the same remark for the hundredth time — only faster now, and in greater volume. A remark that never became a line in the reference standard is a person's time and attention spent for nothing.
What will set organisations apart
What will differ is not the ability to generate, but the ability to move the criterion outside
My expectation is that the gap in the ability to generate will close faster than the strategy about it can be written: the models are shared, access is roughly equal, and people already use them even where they're forbidden to.
The skill to build is taking a person's taste, experience and gut feel out of the head and into a form the machine can apply on its own. And the skill of keeping the result connected to reality rather than to a description of reality: going back to the machine, the site, the client, and checking that the world is still the way it was once written down. The first gives speed; the second keeps that speed from driving off in the wrong direction.
Neither can be bought or handed to a contractor — it takes the time of the person whose competence we'll be drawing on in this work. And at the start it takes a lot of it.
A lot — I'm speaking from practice here: I started building such HITL systems (human-in-the-loop) with myself, and today the gap between what that combination produces and what people who have only just got their hands on LLMs produce is hard even to measure.
If the article turned out useful and you feel like discussing something — don't hesitate.
