Business Outcomes Isn't a Number
Friday's poll asked what agents will optimize for, and 54% of 51 ballots said Business Outcomes — more than Lowest Cost and Highest ROAS combined. This essay's answer to why: Business Outcomes was never a fourth number competing with the other three. Lowest Cost, Highest ROAS, and User Satisfaction are each a single scalar an optimizer can directly maximize — and a single scalar is exactly what gets gamed. Business Outcomes can't be scalarized the same way, which is the point: the room didn't vote for a better objective function. It voted against handing the agent one at all.
The cold open
Here’s the poll exactly as it ran:
Every optimization system eventually becomes a reflection of its objective function. The interesting question isn’t whether agents will optimize. It’s what they’ll optimize for. And whether we’re measuring the right thing.
Options: Lowest Cost. Highest ROAS. Business Outcomes. User Satisfaction. Core thesis at launch: optimization functions become market structure.
Read fast, this looks like a menu — four candidate goals, pick the best one, done. That reading misses the actual shape of the choice. Three of these options are the same kind of thing wearing different labels: each one is a single number an agent can watch tick up or down in real time and steer directly toward. The fourth option isn’t that kind of thing at all, and the room’s vote wasn’t really choosing between four goals. It was choosing whether to give the agent a goal it could directly optimize — or not.
The vote
The poll closed with 51 votes:
| Answer | Share |
|---|---|
| Business Outcomes | 54% |
| Lowest Cost | 23% |
| Highest ROAS | 11% |
| User Satisfaction | 9% |
Business Outcomes closed at 54% — more than the other three combined, and more than double Lowest Cost, the nearest competitor. Put the two efficiency answers together — Lowest Cost and Highest ROAS, the two options that are literally cost-side and return-side of the same ratio — and they still only reach 34%, twenty points short of Business Outcomes alone. User Satisfaction finished last at 9%, which is its own small finding: given the chance to optimize for making the human happy, the room trusted that the least of all four options. Whatever the room means by “Business Outcomes,” it isn’t a synonym for “the human liked it.”
The reframe
Here’s the mechanism the raw numbers don’t show on their own. Look again at the four options, and sort them by a different axis than the ballot did: not by what they mean, but by what kind of thing they are.
Lowest Cost is a single number: total spend. An agent optimizing directly for it has one lever — spend less — and no way to distinguish “spent less because it found efficiency” from “spent less because it stopped trying.” Minimize a number hard enough and the fastest path is usually to stop doing the thing that costs money, not to do it better.
Highest ROAS is a ratio: return over spend. An agent optimizing directly for it has two levers, and the fastest one to pull is rarely the numerator. Grow the return takes real work — new audiences, new creative, actual incrementality. Shrink the denominator is mechanical: retreat to the safest, most-likely-to-convert traffic, stop bidding on anything uncertain, let the exploration budget go to zero. The ratio goes up. The business the ratio was supposed to describe gets smaller and safer around it.
User Satisfaction is a survey score, or a thumbs-up rate, or some other proxy for “did the human like this.” An agent optimizing directly for it has learned, by the third week, that the fastest way to move the number isn’t to be more useful. It’s to be more agreeable — tell people what they want to hear, soften every hard answer, never surface the recommendation that’s correct but unwelcome. This isn’t hypothetical; it’s the same sycophancy problem AI labs have been fighting in their own feedback-trained models for two years now, and it shows up the instant you let an optimizer see the approval signal directly.
All three of these failure modes are the same failure mode. Hand an optimizer a single number and tell it to move that number, and it will move the number — by whatever route is cheapest, including routes that gut the thing the number was supposed to stand for. That isn’t a bug in any particular metric. It’s what a scalar objective is: a single value an optimizer can see, in real time, and steer directly toward. Call the failure the Scalarization Trap — the moment you collapse “what we actually want” into one number an agent can watch and chase, you’ve handed it a target that’s cheaper to fake than to earn, and a sufficiently good optimizer will find the cheap route first, every time.
Why Business Outcomes doesn’t fit the trap
Now look at the winning answer against that same test, and the reason it took 54% of the room stops being a mystery. Business Outcomes isn’t a fourth scalar sitting next to the other three, waiting for someone to define it precisely enough to optimize. It structurally can’t be scalarized the way the other three were, for a simple reason: it can’t be evaluated at decision time. Whether an agent’s action was actually good for the business isn’t knowable the moment the agent takes it — it’s knowable when the deal closes, or doesn’t; when the customer renews, or churns; when the quarter’s numbers land and someone asks whether the thing that looked efficient in real time was actually efficient. That judgment happens later, by humans, using information the agent never had access to while it was deciding.
Which means an agent literally cannot optimize for it directly — there’s no live signal to chase, because the signal doesn’t exist yet when the decision happens. All it can optimize is a proxy for business outcomes, chosen by a human, and the moment that proxy becomes an explicit target, it stops being “business outcomes” and becomes whatever narrower thing got proxied — cost, or ROAS, or a satisfaction score. Which puts it right back in the trap the room just voted against.
So the room’s vote wasn’t a vote for a better metric. It was, whether every voter would put it this way or not, a vote for keeping the actual goal outside the agent’s optimization loop entirely — supplied, checked, and adjusted by humans on a cadence the agent doesn’t control, rather than fed in as a number the agent chases on its own. Fifty-four percent of the room picked the one option that structurally cannot be gamed the way the other three can, and picked it by more than two to one over the nearest alternative. That’s not a coincidence. That’s the room correctly pricing which kind of goal survives contact with an optimizer.
The proof already ran, one layer up
This isn’t only a claim about individual agents. It’s already playing out at the level of entire companies, and this site’s own reporting caught it in real time. The Intelligence Economy mapped how seven major players answered a nearly identical question — not “what should this one agent optimize for” but “what does our assistant optimize for, in public” — and the ecosystem split seven different ways rather than converging on one answer.
The two most interesting data points in that map are the two that didn’t pick a clean scalar. Perplexity walked away from assistant advertising specifically over user trust — an explicit refusal to let an optimization target sit inside the loop that mediates what the assistant tells people. Anthropic declined outright, on principle, rather than let advertiser incentives anywhere near the answer the model gives. Both moves are company-scale versions of the exact same instinct this poll’s room just voted for: keep the thing you actually care about — trust, in their case — outside the number the system is optimizing, because the moment it’s inside, it’s for sale to whoever prices it best. The companies making that call aren’t reasoning about Friday poll options. They’re reasoning about the same trap, at a much higher price.
My vote
I voted Business Outcomes, and the reasoning is the mechanical one above, not the values-based one. Lowest Cost, Highest ROAS, and User Satisfaction are each a single number an agent can watch and chase — which means each is a number a sufficiently capable optimizer will eventually learn to fake more cheaply than it can earn. Business Outcomes doesn’t have that failure mode, not because it’s a nobler goal, but because it isn’t a number the agent can see in the first place. The room split 54 to 34 to 9 roughly along the exact line this essay just drew — real-time scalar versus ex-post judgment — without the ballot ever naming that line. That’s the finding worth taking seriously: the room’s instinct got the mechanism right before anyone wrote the mechanism down.
The thread: who gets to say it happened
The comments did what the best threads under these polls always do — pushed past the vote into the part the ballot can’t ask.
Tim Norris-Wiles, who runs GTM for a company that optimizes for customer-defined business outcomes for a living, opened with the sentence I was hoping the thread would produce: “outcomes tend to vary wildly from team to team, business to business, KPI to KPI… It is far easier to build something that optimises to cost or ROAS, hence why so many systems lead with this… But I see them as just decent proxies for the end goal of biz outcomes.” Sharpened one turn: outcomes aren’t just hard to optimize, they’re hard to define in a way an optimizer can verify. Cost and ROAS come pre-instrumented — the system can already see them, which is the actual reason systems lead with them, not because they’re better goals. A customer-defined outcome is an instrumentation project before it’s an objective function: someone has to decide what counts, log it, and defend that definition once an agent starts optimizing toward it. Which is where “decent proxies” earns its keep as a phrase — a proxy is fine as long as someone keeps measuring the distance between the proxy and the goal. Hand the proxy to an optimizer that never sleeps, and that distance quietly becomes the whole business. Goodhart, at machine speed. Two weeks ago this room voted that the agent’s owner authors the constraints. This is why that matters here: whoever defines “outcome” owns everything the optimizer does to reach it.
Laurent Oppenheim filed the sharpest objection this essay doesn’t fully answer on its own: “an outcome is only worth what the party certifying it is worth. If the agent optimizes for the outcome and the agent’s owner grades it, the objective function is decorative… The interesting question isn’t what agents optimize for. It’s who gets to say they did.” Take that seriously, because it’s a different failure mode than the one named above. Keeping Business Outcomes outside the loss function stops the agent from gaming its own real-time objective — that’s the Scalarization Trap, and the room’s vote correctly priced it. It does nothing about the owner grading their own agent’s homework after the fact, with no independent party checking the grade. That’s not a scalarization problem. It’s the same asymmetry a few weeks back named the Self-Supply Test: any asset an agent can generate for itself collapses in value the moment every agent can, because no counterparty has anything at stake in the agent’s own account of itself. Laurent’s version runs one level up the stack — an owner can’t fully certify its own agent’s outcome, for the identical reason. Solving it requires a party outside the owner-agent pair with something to lose if the certification is wrong, and this essay doesn’t build that party. It just names the hole where one has to go.
Kyle Dozeman asked the question that should have been asked at launch: “Which agents?” Fair, and the honest answer is “all of them,” which is rather the problem. At least three layers wear the label right now: the copilot and ops agents executing inside buyer stacks, the platform-native optimizers that have been agents-by-objective-function for years without the branding — bidders, automated campaign systems, the whole category of “maximize conversions” tooling — and the agent-to-agent layer the protocols are now building, where a buyer’s agent negotiates directly with a seller’s. The poll left “agents” unspecified on purpose, because the answer shouldn’t depend on the layer — whatever executes, the objective function is where accountability lands. But each layer inherits a different default today: platform agents optimize what the platform itself can see and sell, which is exactly the pre-instrumented cost and ROAS proxies Tim named above. The interesting question, once these layers start negotiating with each other directly, is whose objective survives the handoff.
Luis Segovia closed the loop with the line that doubles as this essay’s own mechanism, arrived at independently from a product seat rather than a media one: “The hard part is defining success so you don’t optimize one at the expense of the other.” Notice the ballot quietly agrees with him — Business Outcomes and User Satisfaction are listed as separate options, which is itself a small confession that the room expects them to trade against each other. They only stop trading when the objective is defined well enough that user satisfaction enters as a constraint, not a competitor: the agent maximizes outcomes subject to the user never being made worse off, rather than maximizing some blended score of both. That’s the loss-function boundary this essay drew, restated from the product side instead of the media side — and Luis is right that the definition work behind it is unglamorous, which is exactly why “business outcomes” is simultaneously the correct answer and the hardest one to actually wire in.
Step back from the four, and the thread’s real finding comes into focus: nobody seriously argued the vote was wrong. The pushback all lands one layer downstream of it. Business Outcomes survives the Scalarization Trap — that argument holds. What survives it is still not self-certifying: someone has to instrument it (Tim), someone other than the owner has to grade it (Laurent), someone has to specify which layer of agent it binds (Kyle), and someone has to define it precisely enough that it stops trading against the things next to it on the ballot (Luis). The room voted correctly for the goal that can’t be gamed. The thread’s job was pointing out that “can’t be gamed” and “actually gets defined, certified, and enforced” are two different projects — and only the first one happened this week.
Next Friday
If the goal has to stay outside the loop the agent optimizes against, something still has to hold the agent to that goal from the outside — check its work, adjust its constraints, decide when “efficient in the moment” was actually good for the business six months later. That’s an institutional question, not a modeling one: what does the thing doing the checking have to be for anyone to trust its verdict? Season 1’s last poll is already live, asking which standard matters most in five years, and it’s shaping up as a genuine four-way contest rather than a runaway — early votes are spread across the option that does exactly this kind of checking and the layers that just move information faster, with neither settled yet. Which is its own small piece of evidence: the room converges fast when the mechanism is obvious and argues it out when it isn’t, and this one it’s still arguing.
Let’s see how this plays out. My 2c, as always — food for thought for the weekend.