Abstract turquoise architectural fragments glitching apart against black, evoking overbuilt, disconnected structure
All Insights

INSIGHT · PRODUCT DELIVERY ECONOMICS™

Not Every Feature Deserves to Be Built.

Why product investment decisions need more economic rigor before delivery capacity is committed.

5–7 min read · August 2026

Not Every Feature Deserves to Be Built.

Jon Encarnacion
August 20265–7 min read

In Brief

  • At Microsoft, Bing, and Airbnb, 66–92% of tested product ideas failed to move the metric they were built to improve (Kohavi et al.).
  • These are among the most data-driven product organizations in the industry—the pattern isn't a sign of weak instincts, it's the base rate.
  • "It seemed like a good idea" isn't a business case; intuition alone was never a reliable filter for which ideas will work.
  • The fix: fund a portfolio of small, validated bets before committing full delivery capacity to any one of them.

The most consequential decision in product delivery isn't how well something gets built. It's whether it should have been built at all—and that decision usually gets made with far less scrutiny than the build itself receives.

What Happens When You Actually Check

Most organizations assume their product ideas are good ideas, more or less by default. The evidence says that assumption is wrong more often than it's right.

Ronny Kohavi, who built and ran large-scale controlled-experimentation platforms at Microsoft, Bing, and later Airbnb, has published some of the most extensively reviewed data available on this question—what actually happens when a shipped idea is tested against a real, randomized control group instead of just being assumed to work. The pattern holds with remarkable consistency across very different companies: at Microsoft, roughly two-thirds of tested ideas failed to improve the metric they were built to improve. At Bing, the failure rate ran higher still, around 85%. At Airbnb, roughly 92%. Booking.com, running more concurrent experiments than almost any company in the world, has reported a similar result—the large majority of ideas its own product teams believed would help did not, when actually measured.

Exhibit 1

Share of tested product ideas that failed to move the metric they were built for

Microsoft66%
Bing85%
Airbnb92%

Measured against a real, randomized control group — at three of the most data-driven product organizations in the industry.

Source: Kohavi, R. et al., Trustworthy Online Controlled Experiments; exp-platform.com.

Exhibit 2

An idea seems obviously good

Full delivery capacity gets committed

It ships as planned, on schedule

It's tested against real usage—or not tested at all

Most of the time, it doesn't move the metric it was built for

Why shipping on schedule and shipping something that works are two different outcomes.

These are not companies with weak product instincts. They are among the most sophisticated, data-driven product organizations in the industry, and their own numbers say the same thing: most ideas that look good enough to build turn out not to earn their cost once someone checks.

The ideas that survive contact with a controlled test are the exception, not the rule—for everyone, not just for teams that are getting it wrong.

"It Seemed Like a Good Idea" Isn't a Business Case

If even the best product organizations in the world are wrong most of the time about which ideas will work, the honest conclusion isn't that those organizations are bad at their jobs. It's that intuition alone was never going to be a reliable filter—and every organization that skips validation and commits delivery capacity straight from "this seems like a good idea" is making the same bet those companies' own data shows usually doesn't pay off.

Melissa Perri calls the organizational pattern that results from skipping this step the build trap: equating more shipped output with more success, and losing track of whether any particular thing that got shipped actually created value. Her proposed fix reframes the whole problem as a capital allocation question—fund product work the way a venture investor funds a portfolio, putting a small amount of capacity against many unproven ideas, and only committing serious capacity once an idea has evidence behind it.

That's also the core discipline behind Eric Ries's build-measure-learn loop: treat what you ship as a test of an assumption, not a finished commitment, until the data says otherwise. The goal isn't to move slower. It's to spend the smallest amount of capacity necessary to find out whether an idea is one of the roughly one-in-three that works—before spending the much larger amount of capacity it takes to fully build, harden, and maintain it.

The Economics of Checking First

This connects directly to the same economic logic that should govern any prioritization decision: capacity is finite, and every dollar of it committed to an unvalidated feature is a dollar not available for the smaller share of ideas that would have actually earned their investment.

The fix costs far less than the mistake. A validation step—a small experiment, a narrow release, a real test against real usage—costs a fraction of what building, shipping, and then indefinitely maintaining the wrong thing costs. The organizations with the best data on this question aren't the ones that guess less often. They're the ones that built the discipline to find out before they commit.

The Real Question, Asked Earlier

Not every feature deserves to be built—not because most product ideas are bad ones, but because most ideas, even from strong teams, don't turn out to be worth what they'd cost until someone actually checks.

ASK YOURSELF

Are we funding ideas, or funding evidence?

01

How many of our last ten shipped features were validated before full build, not just after?

02

What's our actual hit rate once something ships—do we even measure it?

03

How much capacity would a small validation step cost compared to building the wrong thing fully?

04

Do strong opinions or real evidence carry more weight in our prioritization conversations?

05

If we're honest, what's the base rate we should expect our own ideas to beat?

THE COHERENZ PERSPECTIVE

Capacity committed to an unvalidated idea isn't confidence. Given what the data shows across company after company, it's a coin flip weighted against you—with your most expensive resource on the table.

VALUE → CAPACITY → FLOW → OUTCOME

What changes?

Instead of asking: "Does this seem like a good idea?"

Ask: "What's the smallest, cheapest way to find out if this idea is one of the ones that works?"

01

Fund a portfolio, not a single bet

Put small capacity against many ideas before committing large capacity to one.

02

Validate before you build fully

A narrow test costs a fraction of what shipping, hardening, and maintaining the wrong thing costs.

03

Treat capacity like capital

Every unit committed to an unvalidated idea is a unit not available to the ones that would have earned it.

The organizations that treat validation as optional aren't moving faster. They're spending delivery capacity on the same odds everyone else is working against, just without bothering to look at them first.

From Insight to Action

Want to know how much of your current roadmap has actually been validated—and how much is running on conviction?

Coherenz helps product and technology leaders bring economic rigor to investment decisions before capacity gets committed.

Sources

  1. Kohavi, R., Tang, D., Xu, Y. — Trustworthy Online Controlled Experiments (Cambridge University Press, 2020); Microsoft/Bing experimentation-platform findings, exp-platform.com.
  2. Reported experimentation failure rates at Airbnb and Booking.com — industry interviews/case discussion (e.g., abtasty.com "1,000 Experiments Club" interview with Ronny Kohavi).
  3. Perri, M. — Escaping the Build Trap: How Effective Product Management Creates Real Value (O’Reilly, 2018).
  4. Ries, E. — The Lean Startup (Crown Business, 2011) — validated learning, build-measure-learn.