UNDERTOW 021: Verdict Theory — Who's Standing at the Pass?
Machines can make anything now except somebody to answer for it. Which is why the person willing to put their name on the work is about to be worth more, not less.
The one thing AI cannot generate is someone to blame.
It can generate everything else. Code, images, videos, campaigns, entire plausible bodies of work, in volume, on demand, at close to no cost. What it cannot produce, at any price, is a specific person willing to look at something and say: this one is good, this one ships, and if it’s wrong, that’s on me. Call that act the verdict. Not the opinion, which is free and everywhere. The named, liable act of standing behind a thing.
The verdict behaves the way it does because of what happened to the rest of judgment. The half of judging that can be measured, scored, or turned into a checklist gets absorbed into the tools almost immediately; whatever becomes a metric, the machines learn to hit, and Goodhart wrote the epitaph for that half decades ago. What remains is the part a score does not capture: a choice, made by someone reachable, who bears the consequences if it is wrong. Every star, stamp, rating, and certification the public trusts is one of these, held by somebody. And here is the strange part: nobody guards them. Institutions give theirs away in public, constantly, and the way you find out it was the most valuable thing you owned is that people stop believing you.
There are four things an institution can do with its verdict.
It can hold its verdict: keep the judgment inside, run by procedure, sold to nobody being judged. Michelin’s restaurant stars have done this for a century. Anonymous inspectors booking under false names, paying for their own meals, each visiting only once so faces do not become familiar, because no star is ever awarded or taken away on a single meal. The procedure does not care who you are. That indifference is the product.
It can name its verdict: put one person’s name on the judgment so the deference has an address. The New York Times restaurant critic moves reservations in a way ten thousand anonymous star ratings never will, because the critic can be wrong in public and the swarm cannot. When Pete Wells took Per Se from four stars to two, reversing two critics who had come before him, Thomas Keller answered him by name inside two weeks, the whole city glued to the argument. A restaurant sitting at 3.2 stars on a review site has no one to answer. The number is not anyone’s opinion; it is the residue of the faceless reviews, and you cannot argue with a residue. There is nobody there.
It can sell its verdict: take money from the things being judged. All through its early years, the top of an Amazon search was a verdict, an algorithm’s honest best answer to what you asked for. Then the slots went up for sale. Today the first products you see are usually ones that paid to be there, placements bought by the products being judged, set in the same font as the ranking they replaced. The verdict is still on the page. It just works for the other side now.
Or it can rent its verdict: borrow someone else’s held verdict for a fee. A movie star who has never held the token looks into a camera and vouches for a crypto exchange, spending in ninety seconds a verdict earned across decades of their other work. When the exchange collapses, and that genre of exchange collapses reliably on schedule, the star’s lawyers explain that no reasonable person could have taken the endorsement as advice. Hold that defense up to the light. A rented verdict, at the one moment it gets tested, is worth nothing by its own account.
Hold, name, sell, rent. For a while that looked like the whole theory: an institution’s future was predictable in which of the four it chose, and you could read the choice years before the consequences arrived. Then the pattern got tested against institutions it was never built from: kosher certification, Champagne, medical boards, app stores. The theory survived the test, but not in its original shape. The four moves turned out to be the surface. They describe what an institution does with its verdict, and nothing more. Underneath them sat two conditions the moves had been pointing at all along, and it is the conditions, not the choice of move, that decide whether an institution’s authority survives.
A rented verdict, at the one moment it gets tested, is worth nothing by its own account.
The first condition is the firewall: money from the people being judged must never reach the judging itself. Taking their money is fine. Letting their money change the answer is fatal.
Consumer Reports buys every product it tests, at retail, anonymously, and has refused advertising for ninety years, which is why a corporation with a billion-dollar launch cannot buy one sentence of its verdict.
The arrangement is old wherever it survives: the Orthodox Union charges food companies for kosher certification and has held authority for a hundred years because the fee buys the audit and never the answer, and Champagne’s governing body is funded by the very producers it polices and still tells them no. Michelin’s own guides outside France work the same way: tourism boards pay to bring the inspectors to their state, and no restaurant pays a cent of it.
Every institution that lasts pays the same price, which is turning down money, constantly, forever. That is the firewall premium. And the grade for failing is always the same. The Better Business Bureau once gave an A minus to a company that did not exist, invented by pranksters and registered under the name Hamas, because the membership fee had cleared. One grade, and a century of accumulated trust became a punchline. The moment money can change the judgment, the stamp is dead; the believing just takes a while to catch up.
The second condition is the exit: the audience has to be free to walk away. Real authority is deference people give when they could go elsewhere; that freedom is what forces a verdict to stay honest. When the audience is trapped, a bad verdict can boss people around for years. Consider any developer who has shipped on the phone platforms since the beginning. She does not trust the app review process. She has watched it approve clones of her product and then reject her update over a button, and she can tell you, fight by fight, that the procedure feels arbitrary. It does not matter. There is only one door to two billion phones, so she rebuilds her interface around the reviewer’s tastes anyway, and so does every developer she knows, billions of dollars of effort steered by a verdict nobody involved believes. From a distance that looks like authority. Get up close and it is obedience, and the difference only shows if there comes a day a second door opens.
A warning before we use any of this. If money can touch the judgment and the audience can leave, the authority will die. That prediction is nearly a law of nature. There is no matching law in the other direction. Running everything clean guarantees nothing, and American medicine is the proof.
The specialty boards that certify physicians do not certify them once; they require doctors to keep re-proving themselves across their careers, through recurring exams, coursework, and fees. And the boards run that machinery cleanly, by any test this essay has: no hospital money in the judging, no way to buy a pass. The doctors revolted anyway. Twenty-two thousand physicians signed a petition against the system, and whole specialties have backed rival boards to escape it, not because the verdict is for sale but because nobody has shown that any of it actually makes anyone a better doctor. The complaint is not corruption. The complaint is that the verdict costs a career’s worth of hours and buys nothing. So the two conditions tell us which verdicts are safe to trust. They cannot tell us which ones are worth obeying.
So the two conditions tell us which verdicts are safe to trust. They cannot tell us which ones are worth obeying.
By now the verdict theory should feel finished. Four moves on the surface, two conditions underneath, one warning about direction, and every case in this essay fits it. That is exactly when to distrust a new theory. Agreement is cheap. The only test that counts is a case that refuses to fit. This July, one arrived, and it came from inside the theory’s own founding exhibit.
Michael Fridjhon is a wine judge, and not a romantic about it. He has sat tasting with Eben Sadie, one of the great winemakers alive, in near silence. He has also opened a bottle of Sadie’s Mev Kirsten and found it dead in the glass, killed by its own closure. The cork had cost forty rand apiece, the best that money buys, and it failed anyway. That is what a lifetime of judging wine teaches. Dead bottles happen to the greatest producers alive, at random, and no reputation can tell you which bottle, because the only way to know what is in a bottle is to pull the cork. So when Michelin decided this summer to rate wine producers, Fridjhon was exactly the wrong man to do it in front of.
This July, in Dijon, Michelin started rating wine producers, something it had never formally done: ninety-four Burgundy estates, ranked on a table of merit, one, two, or three grapes, the star system with the nouns changed. No blind panel was ever convened. No vintages were tasted side by side against the ninety-four names on the table. They were judged, in Fridjhon’s rendering of the blurb, holistically: on identity, balance, consistency, technical mastery, commitment to long-term excellence. Balance is a word you can only reach by drinking. Read that list of criteria again, slowly. It is a description of reputation. Michelin ranked the producers on their reputations, the one thing its restaurant system was built to overrule. Within a week, one of the ninety-four, Domaine Arnoux-Lachaux, publicly asked to be taken off the list. The honor was declined by the honored, on the grounds that nobody had asked them, and nobody had tasted.
Michelin ranked the producers on their reputations, the one thing its restaurant system was built to overrule.
Michelin owns Robert Parker’s Wine Advocate, the empire built by the most influential wine critic who ever lived. It bought the operation in 2019, and with it one of the most respected wine-tasting apparatuses in the world: trained tasters, blind protocols, decades of scoring method. Some of those tasters now work for Michelin directly, employed to assess estates for the new guide. So the palates crossed over, and the protocol did not. Michelin brought across the people and left behind the procedure that made the people trustworthy: the anonymous visit, the repeat check, the blind comparison. That is the detail that separates a lapse from a choice.
Fridjhon’s complaint, in Business Day, was about method, not merit. He was not arguing the ninety-four names are wrong. His complaint was simpler and bigger than that.
Michelin’s authority was built by testing and by nothing else: a century of anonymous visits, paid meals, repeated checks. The name on the cover is a promise that the checking happened. And the company put that name on a wine list where no checking happened.
His question back was harder than his complaint: fine, no wine was tasted blind. But what would tasting even look like? Ninety-four producers, all their wines, several vintages each. Nobody can drink their way to that ranking and survive. Even the experts mostly agree on the famous names without tasting them, and almost never blind. Once a reputation is made, nobody re-tests it; the greatness just gets repeated, and the repeating is what keeps it great. Fridjhon does not offer one, and neither does this essay yet. We will come back to it.
Test the wine list against the verdict theory, and watch it break the theory twice. Start with the four moves. Michelin did not hold its verdict, name it, sell it, or rent it. It did something the list has no word for: it took a name that a century of testing had made trustworthy and spent it in a category where no testing happened.
Then the two conditions. Money first: nobody paid to be on the wine list, so the firewall holds. Exit second: nobody is trapped into trusting it, so the door is open. The list passes both, and it is still wrong. That is because both conditions are questions about a judgment somebody has already made. Michelin never made one. There is no judgment inside the wine list to test. No vintages were tasted blind, no checking was done, and what separates the list from a magazine’s guess is only the name on the cover, a name built by a century of refusing to work exactly this way.
Taste, in 2026, has had its run as the word of the moment. Among the people who set the discourse it is already an embarrassment, a bell ringing to say you arrived late. But it is still everywhere below that line, the strategy decks, the LinkedIn posts, the tech tweets, because a word takes years to travel from the people who coin it to the people who repeat it, and taste is deep into the repeating years.
Ask why the whole discourse landed on that word and not another. Taste is the one professional asset that never has to survive a test. A maker can be tested on the spot: hand a designer a brief Monday and by Tuesday there is a comp on the wall that is either good or not. A judgment gets scored by the world, on a delay: somebody signed off on the 2024 Jaguar rebrand, and the world graded it twice: the internet inside a day, the sales chart over the year that followed.
Taste commits to nothing, signs nothing, leaves no record to check. Your taste can age badly, and plenty has. But when it does, you pay nothing. You wince, you update, you stop mentioning the band. The bill only arrives when taste files a claim: the moment somebody signs the act or kills the campaign, it has stopped being taste and become a verdict, and verdicts have addresses. Embarrassment is the worst thing taste can ever cost.
The discourse is so overgrown that culture has already turned on it. Tasteslop, Emily Segal called it this spring: the visible signs of taste extracted from the social world that made them mean anything, redeployed generically, everywhere. Taste, in her account, was never a property of objects at all; it is a relation between objects and people and scenes and timing, and it does not exist unless somebody else can see it. “There is no such thing as taste,” she wrote, “if it falls in a forest.”
And she filed one more thing under the new word, the sharpest thing in the definition: tech’s recurring discourse about the significance of taste. The moat talk is slop about slop, which follows from her definition whether or not she would put it that way. You know the man saying it. So do I; there are a dozen of him and I have read their LinkedIn posts. Front row of the AI panel, lanyard still on at dinner, and when he says taste is the moat he says it the way he said design thinking in 2016 and storytelling in 2011, which is to say beautifully, with total conviction, in the tone of a man who has never once been asked to prove it. He will say the next word in that tone too. The tone is the job.
The panic underneath is real. He is rich in the kind of capital that can be counted and anxious about the kind that cannot be bought, and taste is the name he gave the difference so it would fit on a slide. That is what the moat was ever guarding. But follow Segal’s own definition one step further than she takes it. If taste only exists when somebody else can see it, then in a world where machines generate the visible signs of taste in unlimited supply, being seen is no longer the scarce condition either. Signed is. The scarce thing was never the eye, and it is not even the audience. It is the name underneath, the one part of the arrangement that pays when the work is wrong. The eye was never the moat. The neck was.
The industry that made making cheap is running the cleanest test of all this on itself, right now. AI labs live and die by benchmarks, standardized exams that score how capable a model is, and those scores are not a sideshow; they are the sales material. They move lucrative enterprise contracts, valuations, and headlines. So follow the money through one release cycle. The lab builds the model. The lab, in some cases, funds the exam, or commissions it. The lab administers the exam to its own model, in-house, and publishes the number. The number sells the product. The lab is the student, the proctor, and the press office all at once. Money here is not merely near the judgment. Money is the judgment’s employer.
The clearest case on the public record is a benchmark called FrontierMath: research-grade math problems hard enough that frontier models were expected to fail them, built by an outside group precisely so that no lab could study for the test. The one exam that could not be gamed. Then a leading lab announced record scores on it. The funding surfaced the same day as the score, in a line almost nobody read for weeks: the lab had paid for the benchmark’s creation and held access to most of the problem set. Contributing mathematicians said they might not have participated had they known. And when the arrangement surfaced, the safeguard revealed was this: a verbal agreement that the lab would not train on the questions it had paid for. The details are still argued over; the structure is not. Write the sentence out plainly and it refutes itself. The un-gameable exam was funded by a student holding the answer key, on his word.
The lab is the student, the proctor, and the press office all at once.
The fix is already being built, and its shape is familiar. Independent evaluation groups now exist whose entire purpose is to test models while taking no money from the labs they test, and governments are standing up institutes to do the same from inside the state. This is the crash test arriving in software: the Insurance Institute for Highway Safety is paid by insurers, who profit when cars kill fewer people, not by automakers, which is why its ratings can end a car’s reputation in one release and why automakers redesign around them. An arrangement with the same logic is old enough to have once been stamped into silver. Whether the new evaluators can hold the position is another matter. They are small; the labs they judge are becoming the richest companies on earth; and the pull of lab money, access, and jobs is constant. One of the most principled evaluators in the field takes no lab money and still runs on donated lab compute, which it discloses, and that is a measure of how hard this firewall is to build: even the cleanest referee’s electricity comes from the players. The verdict theory does not care about anyone’s stated mission. It says one thing. The day that lab money can touch the evaluation, the stamp dies, and every score after that is marketing.
The other condition, the exit, is being decided in procurement departments right now, and for once the audience has real doors. The same enterprise work increasingly runs on rival systems, including Chinese open-weight models that undercut the American frontier labs by most of the price; by one routing platform’s count, a meaningful share of American AI traffic now flows to them, and companies that switched describe watching their cost curve fall off a cliff at near equal quality. That freedom is doing quiet enforcement. It is why serious buyers have stopped taking launch benchmarks at face value and started running private exams on their own tasks: when you can leave, you can afford to check.
But look at what the labs are selling next. Agents welded into workflows. Proprietary data pooling inside one vendor. Multi-year contracts. Every one of those is a door closing, and the vendors know it; the doors are the product. Remember what the open door was doing a moment ago: because buyers could leave, they could afford to run their own exams instead of taking the launch numbers on faith. Slam the doors shut and that arithmetic runs backward. A company three years into a contract, its data pooled in one vendor’s systems, its workflows welded to that vendor’s agents, is not going to run a private exam whose bad result it cannot act on. Checking only pays when leaving is possible. So the checking stops, the scores go back to being taken at face value, and the captive verdict arrives in enterprise software the way it arrived on the phone.
If AI keeps doing what it appears to be doing, none of this stays contained. More generated work means more things nobody can personally check, which means more trusting of stamps, scores, and lists. Every quarter, the verdict layer carries more of the economy’s weight, and it is resting on institutions that have mostly never been stress-tested. I am watching this from inside as it starts, so I will call that a bet rather than a law. The two conditions, the firewall and the exit, are the law; they were true before AI and they are proven on a century of cases. The bet is only about scale: that everything is about to run through them, all at once.
One more industry, and I should be honest about why it comes last: it is mine, and I’ve been putting it off. I run creative for a living. My actual job, stripped of the title, is that work leaves the building when I say it leaves, and if it’s wrong, the client doesn’t call the person who made it. They call me. So the job has always had two halves: the making, which my teams do, and the answering, which is mine. And when an anonymous holding-company chief executive interviewed by a research firm says that the industry will double profits and halve its people by 2028, I read that as a spreadsheet sorting decision, already made. You can’t cut half an industry without first deciding who goes and who stays, and his math tells you the decision: the makers go, because machines can make now, and the people who answer for the work stay, because someone still has to sign it and take the call when it fails. Fewer people, same signatures. He may be right. And that’s why I can’t read his quote as a weather forecast. It’s a memo about which half of my job he believes I am in.
The month the industry’s biggest merger closed, four thousand people went, and DDB, the name on the agency that taught this business what an idea was, was retired as a cost synergy. That name was Bernbach’s agency, the shop that invented the modern creative idea, decades of the best making this industry ever produced, stored in three letters. If the making were the product, deleting those letters should have cost something. It cost nothing. Bernbach’s name came off the door and the market barely looked up, because investors had already concluded what this essay has been circling: the making was never the product. The product was always somebody willing to be called.
Watch the industry confess it under pressure, twice in one season. Coca-Cola generated its Christmas ad with AI, heard the word soulless from all directions, then shipped another one the following year, proud that the production timeline shrank from twelve months to one; McDonald’s tried the same and pulled its ad within days. And Cannes, the industry’s own verdict factory, responded to a cheating scandal with a package of new rules, at the center of it a requirement that every entry carry the personal endorsement of the agency’s chief executive and the client’s chief marketer. Entries promptly fell by a quarter. The award’s answer to a credibility crisis was to demand a named human signature on every piece of work, and a quarter of the work did not come back. Agencies that understand what just happened will start selling the person standing behind the work, by name. The rest will keep selling the one thing AI just made free, and wonder where the margin went.
The pass, in a restaurant kitchen, is the counter where every plate stops before it reaches the dining room. Stand there on a Saturday night. The rail is overflowing with tickets, the heat coming off the range in sheets, plates landing every few seconds, and one person’s eyes are on every single one of them before it goes. Tired eyes. Liable eyes. The kitchen can run at that speed only because every cook in the place knows exactly whose eyes they are, and the dining room, without ever thinking about it, is eating that fact. That’s the whole affirmative content of the machine. A verdict is not a score, a stamp, or a list. A verdict is somebody standing at the pass.
If your work carries a sign-off, if you are the creative director, the editor, the chief marketer whose approval is on the campaign, the one at the lab who signs the model card before a launch, then you already stand at the pass. Here’s the uncomfortable half, and I’ll count mine first, because I thought my count was clean. I don’t trust reviews anymore; I stopped years ago. I don’t book Resys off stars. I have never once hired anyone off a certification or a client-logo slide; I trust research and the vouch of people I have actually worked with, which means I had already fled to named verdicts before I could have told you why.
But watch the residue. When a screenshot of a lab’s benchmark results scrolls past on my Twitter feed, the number still lands first. The opinion forms, and then I make it earn its way back down: I interrogate it, I read the people who retest these things, I do the work. The stamp still gets its half second before the audit arrives.
And Michelin: until I wrote this essay, everything I knew about the most trusted rating system on earth came from screens. Ratatouille, where the great chef Gusteau loses a star and dies of it, the grief played straight in a children’s cartoon. Bradley Cooper hunting a third star through two hours of kitchen paranoia. Carmy white-knuckling toward one on The Bear. Set dressing. Thirty years of deference built entirely out of fictional scenes.
Even my fix confesses: when I want the truth about anything now, I ask a model very specific questions and read what comes back, which is to say I replaced the verdicts I couldn’t audit with an artificial judge whose deliberations I can’t see at all. So when I tell you to turn around and count your own, the certifications in your footer, the scores you repeat because the meeting expects them, I’m not preaching from outside the tent. From the outside, a held verdict and a coasting one look identical. I know. I extended credit to both for decades, and the only reason I checked was to write this paragraph.
Which brings back the question this essay promised to return to. Fridjhon never claimed the ninety-four names were wrong; he asked what checking them would even look like, and the honest answer is that it wouldn’t. Walk his dead end step by step. Nobody can re-taste the world: there are more wines, restaurants, models, and campaigns than any working life could ever check firsthand. Because nobody can, every judge starts from standing they inherited, reputations built by other people, in other decades, that they never tested either. And a reputation, once made, feeds itself: the famous names get repeated because they are famous, and the repeating is what keeps them famous. Follow those three steps to the end and you arrive where Fridjhon did. If the question is who truly deserves their authority, nobody knows. Nobody has ever known. There is no procedure that will tell you.
But stand in that dead end for a minute, because it blocks less than it seems to. What it blocks is knowing whether a verdict is right. What it leaves open is knowing whether the verdict was ever exposed to being wrong: whether it was made under conditions where a wrong answer was possible and would have cost something. Michelin’s restaurant stars were built that way; its wine list was not, and you did not need to taste a single bottle to know it. That is the move. You stop asking about the wine and start asking about the arrangement, and the arrangement can be read in two questions. Can money from the judged touch the judgment? Can the audience leave? You cannot verify every verdict you rely on. But you can check whether the check gets run.'
You cannot verify every verdict you rely on. But you can check whether the check gets run.
A theory earns its keep only if it can be wrong, and an essay about named liability that ends with its author risking nothing would be its own counterexample. So here is my neck. Three predictions, my name on them, scored in public as they land.
First, and dated. Michelin has until September 2027, when the second edition of the wine ratings is due. By then, one of two things happens. Either Michelin adds a real verification procedure to the wine list: tasting panels, blind protocols, something an outsider can point to and say “that is the check.” Or the wine trade stops treating the list as a rating and starts treating it as marketing. If neither happens, if the list stays untested and people keep deferring to it anyway, then this theory is wrong, and wrong at its center, because it would mean a trusted name can carry authority indefinitely with no checking behind it. I’m stating that in advance so the result can be scored.
Second. AI-generated fake reviews are about to overwhelm every rating system built on anonymous votes. The platforms will try filters first. Filters will fail, because the fakes are now as plausible as the real thing and there are millions of them. The fix that actually works will be names: reviews tied to real, accountable people with track records. Here’s what that means in practice. If your business has spent years accumulating five-star ratings from anonymous reviewers, that asset is about to lose its value. If you have been attaching your own name to your judgments, in public, with a record anyone can check, that asset is about to gain value. Most people and most businesses are holding the first kind and not the second.
Third, and this is the one my own next act depends on. Within a few years, “who signed off on this” becomes something companies advertise. AI labs will market which people stood behind their test results. Agencies will market which people stood behind the work. Same move, both industries. And once the sign-off is the product, the careers follow: seniority starts going to the people willing to put their name where the blame lands. If this is wrong, then I’ve misread the industry I have spent my life in, and the public scoreboard will say so, under my name. That is the deal this essay ends on: it argued that a verdict only counts when somebody can be held to it, so it closes with three that I can be held to.
In the meantime the theory folds down to something anyone can run in a minute, on anything. Take any verdict you rely on: the certifier, the benchmark, the guide, the critic, the board, the app store, a list of ninety-four Burgundy producers. Ask whether money from the judged can touch the judgment. Ask whether you could leave. Then ask the only question left, the one Michelin forgot to ask itself this summer.
Who’s standing at the pass?


