New Coke: they verified the wrong thing, flawlessly
New Coke is the opposite trap from Quibi. Quibi skipped verification. New Coke ran one of the most rigorous verification programs in the history of consumer products — and still walked off a cliff.
New Coke is the opposite trap from Quibi. Quibi skipped verification. New Coke ran one of the most rigorous verification programs in the history of consumer products — and still walked off a cliff.
Coca-Cola reformulated its flagship product on the back of nearly 200,000 taste tests. In blind tests, the new formula beat both the old Coke and Pepsi. The data was clean, the sample was enormous, the result was unambiguous. By the standard the company set for itself, the project was a triumph.
It launched on April 23, 1985. Seventy-nine days later, on July 11, 1985, the company reversed itself and brought the original formula back as Coca-Cola Classic.
This is not a story about lazy research. It's a story about a team that verified the wrong thing — perfectly. They had a flawless answer to a question that wasn't the one the company was actually betting on.
What 'flawless execution' actually looked like
Strip out hindsight and read the project the way its own status report would have read in early 1985:
- Research volume: roughly 200,000 taste tests. Risk: retired.
- Result: the new formula won, head-to-head, against Coke and Pepsi. Risk: retired.
- Strategic rationale: Pepsi was winning on taste in public challenges; reformulating answered the competitive threat directly. Risk: retired.
- Rollout: national launch, on schedule, with full marketing weight behind it. Risk: retired.
If you were grading the execution of this project, it's close to an A. The research was more thorough than almost anything a competitor was doing. The dashboard was green on every line. And that is exactly the trap — because the entire research program measured one variable, and the company was betting on a different one.
The question they answered vs. the question that mattered
Every taste test asked, in effect, one thing:
Which liquid do you prefer when you sip it without knowing what it is?
That is a real question. It has a clean, measurable answer. New Coke won it.
But the company wasn't actually betting on which liquid people preferred in a blind sip. It was betting on something the tests never put on the table:
Will the people who have built sixty years of identity, memory, and loyalty around this specific brand accept us replacing it — even with something that tastes better in a cup?
The blind test, by design, stripped out the one thing that made Coke valuable: the brand. A sip with no label is a sip with no meaning. The research optimized away the variable the whole business ran on, and nobody noticed because the number it produced was so large and so clean.
The signals were knowable in advance, not in hindsight:
- People don't drink Coke blind. They drink the idea of Coke — the can, the script logo, the childhood, the ritual. Removing the label removed the product.
- A preference of 'tastes slightly better' is weak; an attachment of 'this is mine, don't touch it' is strong. The company measured the weak force and never instrumented the strong one.
- The protests, the hotlines, the hoarding — that emotional intensity was latent in the customer base before launch. A different question, asked of the same people ('how would you feel if we replaced the original forever?'), would have surfaced it for a fraction of what the reversal cost.
None of this required foresight. It required someone, before the reformulation was committed, to ask whether taste preference was the right outcome to verify at all — or just the easiest one to measure.
Through the Orchestration Loop
The Outcome Orchestration discipline frames any delegated effort as a loop: Frame, Delegate, Verify, Steer. Quibi failed by leaving Verify empty. New Coke failed one step earlier — at Frame — and the flawless Verify that followed only made the error more confident.
Frame. The failure lives here. The team framed the problem as 'we are losing a taste war, so produce a better-tasting cola.' That framing silently redefined the outcome from protect the brand to win the sip test. Once the goal was mis-set, everything downstream inherited the error.
Delegate. Done superbly. The taste-testing operation was world-class — scale, rigor, controls. This is what a great execution culture buys you, and it was aimed at the wrong target.
Verify. Done — flawlessly, against the wrong specification. This is the part worth sitting with. Verification passed. The data was sound. Verifying hard against a mis-framed outcome doesn't catch the error; it certifies it. A green check on the wrong question reads exactly like a green check on the right one.
Steer. The one thing that saved them. When the real outcome arrived — the backlash — the company could still reverse, because a formula change is cheap to undo. They steered in 79 days. (Quibi, by contrast, had spent its room to steer before the verdict came in.) The recovery is the only reason this is a teardown and not an obituary.
Why this is the defining failure mode of the AI era
Here's why a forty-year-old soda reversal matters to a project manager in 2026.
When you delegate work to AI, the system will optimize — relentlessly, tirelessly — against the metric you give it. It will run the equivalent of 200,000 taste tests before lunch and hand you a clean, confident, beautifully-charted answer. And if the metric is a stand-in for the real outcome rather than the outcome itself — completion rate for 'did it work,' resolution for 'was it resolved well,' taste preference for 'will they accept this' — the machine will hit that target perfectly and walk you off the same cliff, faster and with more conviction than any 1985 research department could manage.
The danger of a capable autonomous system isn't that it fails. It's hyper-competence aimed at a flawed metric. It will not save you from a badly framed outcome. It will execute that outcome with terrifying efficiency and produce a dashboard that says you succeeded.
Which means the scarce skill is no longer measuring things well — the machine measures better than you. It's deciding which outcome is worth measuring, and noticing when your clean green metric has quietly become a proxy for something it no longer represents. That's not a tooling problem. It's an orchestration problem — the Frame step — and it's the first thing a 'flawless execution' culture skips, because framing feels like it's already done by the time the real work starts.
New Coke optimized the part it could measure. It never checked whether the part it could measure was the part that mattered.
What a strategic PM — with AI used well — would have caught
New Coke is the harder, more instructive case, because the team's execution was flawless — nearly 200,000 taste tests. What was missing wasn't rigor; it was the strategic chain — understanding, judgment, decision — applied to the right question before that rigor was unleashed. The prevention isn't 'verify more'; it's a PM using AI to understand, judge, and decide better about what to measure. (The chain is the capability AI amplifies; the Orchestration Loop is the operating rhythm it runs through — both essential, the chain the decisive one here.)
Understand what the test measures. A strategic PM uses AI to surface what a culture of execution can't see from inside its own success: here's our objective — protect the brand — and here's our test, a blind taste preference. What does it not measure? In minutes it returns the gap: the blind test strips the label, and the label — sixty years of identity, memory, and loyalty — is the entire source of the brand's value.
Judge whether it's the right outcome. The silent failure was a judgment failure: 'win the sip test' was quietly substituted for 'protect the brand,' and no one weighed the swap. AI helps interrogate that judgment and model the outcome that actually mattered — how would loyal customers react to replacing the original forever? A different question than 'which do you prefer in a cup,' with a different, knowable answer.
Decide what to verify, and when. Understanding and judgment converge on a decision: test the real bet cheaply, before committing. Loss-aversion and identity-attachment analysis, sentiment modeling, analogs of brand-replacement backlashes — any of them would have surfaced the revolt for a fraction of what the 79-day reversal cost.
Why this is strategic AI, not mechanical: AI optimizes relentlessly against the metric you give it — 200,000 taste tests before lunch, a clean and confident wrong answer. Mechanical AI makes you faster at hitting the target you set; strategic AI sharpens the understanding, judgment, and decision that set the target. New Coke had rigor to spare; what it needed was the chain — used to ask, before the work, whether the metric was the outcome it was betting on.
Sources: Coca-Cola company history and contemporaneous reporting on the 1985 reformulation and reversal. Analysis is the author's.
Get the next teardown when it drops.
An open discipline for governing outcomes in the age of AI. Teardowns, frameworks and playbooks, published openly under CC BY-ND 4.0.