7 products live across Labs
Product Strategy

Product Strategy Frameworks: From Idea to Decision-First Roadmap

The real, correctly-sourced mechanics behind Jobs to Be Done, RICE, Opportunity Solution Trees, OKRs, and Working Backwards — and how to combine them into a roadmap built on decisions instead of a list of feature requests.

By Loomstrat Studio TeamPublished September 5, 2026Updated September 5, 202626 min read

Most product roadmaps are lists of features with dates next to them, assembled from whoever asked loudest — a sales team's biggest deal, an executive's pet idea, the last customer who complained. None of that is a strategy. It's a queue. And the honest data on how product teams actually operate backs this up: Pendo's 2022 State of Product Leadership survey found that even among self-described “data-driven” companies, a third of product leaders described their decisions as mainly or fully instinct-driven, and a separate 2026 industry survey found that 64% of internal stakeholders admit they only sometimes or rarely read the roadmaps their own product teams publish.

This guide covers the real, correctly-sourced frameworks serious product teams actually use to replace that guesswork — what each one actually does, who actually created it (several of the origin stories repeated online are garbled or embellished, and this guide corrects them), and how to combine them into a single, coherent process for deciding what to build, rather than treating them as interchangeable buzzwords to bolt onto a feature list after the fact.

This guide is written for a founder or product leader with a validated product deciding what to build next — the natural continuation of the question covered in our multi-product SaaS portfolio guide, which covers when and how to expand a portfolio, but not the day-to-day discipline of deciding what any single product should build next. The frameworks below apply whether you're running one product or several: the specific tool changes with company size and stage, but the underlying discipline — decisions backed by evidence, not the loudest voice in the room — doesn't.

Why Most Roadmaps Fail

Why do most product roadmaps fail?

Most roadmaps fail because they list outputs (features to ship) instead of outcomes (results to achieve), which means a team can hit every date on the roadmap and still fail the business. Marty Cagan of the Silicon Valley Product Group calls this the “feature factory” problem: teams “measured by output and not outcome” who, when a shipped feature doesn't work, can point out that generating results “was not their job.”

Cagan's distinction is specific and worth stating precisely rather than paraphrasing loosely: feature teams are “given a list of features to build” and are not accountable for whether those features actually work, while empowered product teams are “given problems to solve” and are measured by the business outcome, not the shipping of any particular feature (Marty Cagan, SVPG, “Product vs Feature Teams”). Melissa Perri makes a closely related argument in her book Escaping the Build Trap (O'Reilly, 2018), building on a blog post she published years earlier: companies fall into what she calls the “build trap” when they measure success by output — features shipped, story points closed — rather than by the value or outcome those features actually created (Melissa Perri, “The Build Trap,” 2014).

This is the organizing idea behind every framework in this guide. None of them is a silver bullet on its own, and several solve genuinely different parts of the problem — discovering the right problem, scoring competing ideas, communicating a roadmap, or forcing clarity before building. Used correctly, together, they form a single pipeline: from a fuzzy idea to a specific, defensible decision about what to build next.

Reading about a framework and actually running it are different skills, and the gap between them is exactly where most adoption attempts stall. A team can read Cagan's essays, nod along, and still hand its engineers a features list the following Monday, simply because changing how decisions actually get made is harder than agreeing with an argument about how they should be made. The rest of this guide is written with that gap in mind — not just describing each framework, but naming specifically what changes in a team's daily behavior when it's actually adopted, not merely cited.

The Frameworks at a Glance

The 8 frameworks in this guide, by what each one is actually for
FrameworkCreated BySolves
Jobs to Be DoneClayton Christensen, Bob MoestaUnderstanding why customers actually "hire" a product
Opportunity Solution TreesTeresa TorresMapping customer needs to candidate solutions before committing
RICE / ICE scoringIntercom (RICE); widely attributed to Sean Ellis (ICE)Comparing competing ideas objectively
North Star MetricSean Ellis (concept); Amplitude/John Cutler (framework)Aligning teams around one measurable outcome
Kano ModelNoriaki Kano, 1984Classifying features by their real effect on satisfaction
OKRsAndy Grove (Intel); popularized at Google by John DoerrSetting ambitious goals with measurable results
Now-Next-LaterJanna Bastow, ProdPadCommunicating a roadmap without false-precision dates
Working BackwardsAmazonForcing customer-value clarity before building anything

Jobs to Be Done

What is the Jobs to Be Done framework?

Jobs to Be Done (JTBD) holds that customers don't buy products based on demographics or category — they “hire” a product to make progress on a specific job in a specific circumstance. The framework was developed by Clayton Christensen with practitioners Bob Moesta and Rick Pedi, and is most fully explained in Christensen's book Competing Against Luck (2016).

Christensen first laid out the core idea in a 2005 Harvard Business Review article, “Marketing Malpractice: The Cause and the Cure” (HBR, December 2005), arguing that most product failures trace back to segmenting customers by attributes — age, income, company size — rather than by the job they're trying to get done. The framework's fullest treatment came a decade later in Competing Against Luck, co-written with Taddy Hall, Karen Dillon, and David Duncan.

The framework's most famous illustration — the “milkshake study” — is worth getting right, since the version that circulates online has been embellished well past what its own practitioner actually documented. Bob Moesta's own account, published by his firm The Re-Wired Group, describes the client only as “a leading fast food restaurant chain” — not McDonald's by name, contrary to how the story is usually retold (The Re-Wired Group, “Milkshakes in the Morning”). The actual methodology: researchers spent extended time observing purchases in restaurants, then interviewed the specific customers who bought milkshakes about the circumstances of that purchase — not surveys, not focus groups. The finding, per Moesta's own account: roughly 40% of milkshakes were bought in the morning by solo commuters, who were “hiring” the milkshake to make a boring commute more interesting and to stay full until lunch — which meant the milkshake wasn't really competing against other milkshakes at all. It was competing against bagels, bananas, and coffee. Popular retellings that add a specific sales-increase multiple or a switch to yogurt don't appear in Moesta's own telling and shouldn't be repeated as fact.

The distinction matters for more than pedantry. The embellished version of the story implies a clean, dramatic before-and-after that a single research insight produced almost automatically. Moesta's actual account is less tidy and more useful: the insight came from noticing an unexpected pattern in when and by whom the product was purchased, not from a single eureka moment, and turning that pattern into a validated business decision still required real follow-up work. Treating the tidied-up version as the standard for what a JTBD insight should look like sets founders up to expect a single interview to produce an obvious, dramatic answer, when in practice the real value usually comes from noticing a pattern across several conversations rather than one striking anecdote.

Customers don't buy products. They pull them into their life to make progress.

Clayton Christensen, Competing Against Luck (2016)

The practical value of JTBD for a roadmap process is that it changes the unit of analysis. Instead of asking “what feature should we build,” a JTBD-oriented team asks “what job is the customer trying to get done, and what's currently getting in the way of it” — a question that often surfaces solutions well outside the obvious feature request a customer originally asked for.

In practice, JTBD interviews follow the same discipline Moesta's own methodology used: they focus on a specific, recent purchase or switch decision, not a hypothetical future preference. Asking a customer “would you use a feature that does X” produces unreliable, overly agreeable answers, because there's no real cost to saying yes. Asking a customer to walk through the actual circumstances of the last time they switched tools, signed up, or gave up on a competing product surfaces the real forces at play — the push away from what they were using, the pull toward what they chose, and the anxieties that almost stopped them. This is a fundamentally different research method than the surveys and feature-request tallies most roadmap processes rely on, and it's the reason JTBD interviews take real time to run well rather than something a team can shortcut with a quick customer-feedback form.

Opportunity Solution Trees

What is an Opportunity Solution Tree?

An Opportunity Solution Tree is a visual mapping technique, created by Teresa Torres of Product Talk, that connects a single desired business outcome to the customer opportunities (needs, pain points, desires) that could drive it, the range of candidate solutions for each opportunity, and the specific assumption tests used to validate a solution before committing to build it.

Torres introduced the concept in 2016 and gave it its fullest treatment in her book Continuous Discovery Habits (2021). The tree has four layers, read top to bottom: the outcome sits at the root; beneath it branches the opportunity space — the actual needs and pain points uncovered through customer interviews; beneath each opportunity branches a range of candidate solutions; and beneath each solution sit the specific, small experiments a team runs to test whether it actually works before committing real engineering time to it (Teresa Torres, Product Talk).

What makes this framework distinct from a typical roadmap exercise is where it starts: not with a list of solutions already in mind, but with a single outcome and an open exploration of the opportunity space underneath it. Torres's own synthesis argument — directly relevant to how this guide recommends combining frameworks later on — is that most teams jump straight to comparing solutions (the RICE/ICE scoring covered next) before they've actually mapped the opportunity space widely enough to know whether they're comparing the right solutions at all.

The practice Torres recommends alongside the tree itself is weekly customer touchpoints — continuous, small-scale interviews rather than periodic, large research projects — specifically so the opportunity space stays current as customer needs shift, rather than being mapped once and treated as permanently settled. A tree built from interviews conducted a year ago is a snapshot of an opportunity space that has likely already moved; the framework's real value comes from treating discovery as an ongoing habit rather than a project with a start and end date, which is the specific argument behind the title of Torres's book.

RICE, ICE, and Impact/Effort Scoring

Once a team has a set of candidate ideas — whether surfaced through an Opportunity Solution Tree, customer requests, or internal proposals — a scoring framework gives a shared, numeric way to compare them instead of relying on whoever argues most persuasively in the room.

RICE

RICE was created inside Intercom and documented publicly by Intercom PM Sean McBride in a January 2018 company blog post (Intercom, “RICE: Simple Prioritization for Product Managers,” Jan. 2018) — some secondary sources attribute it to Intercom co-founder Des Traynor in 2014, but McBride's post is the earliest confirmed, dated primary source. The framework scores each idea on four factors and combines them into a single number:

  • Reach — how many people or events the idea affects per time period.
  • Impact — scored on a fixed scale (3 = massive, 2 = high, 1 = medium, 0.5 = low, 0.25 = minimal).
  • Confidence — how sure the team is about its Reach and Impact estimates (100% = high, 80% = medium, 50% = low).
  • Effort — estimated in person-months.

The final score is (Reach × Impact × Confidence) ÷ Effort. McBride's own stated reason for building it was to counter three specific biases Intercom's product team kept falling into: favoring “pet ideas” regardless of real impact, over-valuing clever-sounding but low-impact projects, and underestimating how much effort a project would actually take.

ICE

ICE — Impact, Confidence, Ease — is widely and consistently attributed to Sean Ellis, the growth marketer credited with coining the term “growth hacking,” for prioritizing fast-moving growth experiments. Unlike RICE, no single dated primary document defines ICE's exact origin, so treat the attribution as well-established through consistent secondary sourcing rather than a citable original post. Each factor is scored on a simple 1–10 scale and combined (typically averaged or multiplied, depending on the team's implementation) into one score — a deliberately lighter-weight version of RICE, suited to teams running many small experiments quickly rather than scoping a smaller number of larger initiatives.

Low effortHigh effortHigh impactLow impactQuick winsMajor betsFill-insTime sinks
An illustrative framework, not a data chart: how impact-vs-effort scoring (the logic behind RICE and ICE) sorts initiatives into four zones.

Both frameworks are formalized versions of the same underlying logic every impact-vs-effort quadrant represents: quick wins (high impact, low effort) generally deserve to be built first, major bets (high impact, high effort) deserve deliberate, resourced commitment rather than being squeezed in alongside other work, and time sinks (low impact, high effort) deserve to be cut regardless of how interesting they seem.

Both frameworks share the same real limitation worth naming plainly: they produce an objective-looking number out of inputs that are still, ultimately, subjective estimates. A Confidence score of 80% is someone's judgment call, not a measured fact, and a team that treats the resulting RICE score as more precise than the guesses underneath it is fooling itself with false objectivity. The honest use of these frameworks is to structure a debate and surface disagreement about specific inputs — if two people score the same idea's Effort wildly differently, that disagreement is more valuable than either number alone, since it usually reveals a hidden assumption worth resolving before the idea gets built either way.

The North Star Metric

What is a North Star Metric?

A North Star Metric is the single measure a product team rallies around because it best captures the core value the product delivers to customers, chosen so that improving it reliably drives the business's actual outcomes. Its origin is genuinely shared rather than attributable to one person: Sean Ellis promoted the closely related concept of a “One Metric That Matters,” while Amplitude — largely through John Cutler, its former Head of Product Research & Education — developed and popularized the specific “North Star Metric” framework and terminology from around 2017 onward.

The mechanism, per Amplitude's own North Star Playbook, connects three layers: input metrics (the specific actions a team can directly influence), the North Star metric itself (a measure of value delivered to customers), and business outcomes (revenue, retention) that the North Star is chosen specifically to predict (Amplitude, North Star Playbook). The framework's value is forcing a team to pick one thing to actually optimize for, rather than reporting a scoreboard of a dozen metrics that can each be cited selectively to justify whatever the team already wanted to build.

It's worth naming the framework's real limitation alongside its appeal: former Pinterest growth leader Casey Winters has published credible criticism of over-relying on a single North Star metric, since a single number can be gamed, can obscure important secondary effects, and can encourage a team to over-optimize one dimension of the product at the expense of others a single metric can't see. A North Star metric is a genuinely useful alignment tool, not a complete measurement system on its own.

The Kano Model

What is the Kano Model?

The Kano Model, created by Noriaki Kano and co-authors in a 1984 paper, classifies product attributes by their non-linear relationship to customer satisfaction — distinguishing features whose absence causes dissatisfaction (Must-be) from features that improve satisfaction proportionally (One-dimensional/Performance) from features that delight without being expected (Attractive), among other categories.

The original paper, “Attractive Quality and Must-Be Quality,” published in the Japanese Society for Quality Control's journal Quality in 1984, defines six categories in total — Must-be, One-dimensional, Attractive, Indifferent, Reverse, and Questionable — though most popular summaries simplify this to just the first three (Kano, Seraku, Takahashi & Tsuji, “Attractive Quality and Must-Be Quality,” 1984). The model was influenced by Frederick Herzberg's Two-Factor Theory of workplace motivation, applied to product attributes instead of job satisfaction.

The practical value for a roadmap is separating two very different kinds of work that a flat feature list treats identically: Must-be attributes (reliable uptime, basic security) don't win you customers when present, but lose you customers fast when absent, while Attractive attributes are what actually differentiate a product — but only up to a point, since yesterday's delighter tends to become tomorrow's Must-be as customer expectations rise across the whole category.

OKRs

What are OKRs and where did they come from?

OKRs (Objectives and Key Results) pair an ambitious, qualitative Objective with a small number of specific, measurable Key Results that indicate whether the objective was achieved. The system was developed by Andy Grove at Intel in the 1970s, documented in his own book “High Output Management” (1983), and brought to Google in 1999 by John Doerr — a venture capitalist who had learned the system at Intel — as documented in Doerr's own book “Measure What Matters” (2018).

Grove's original system was influenced by Peter Drucker's earlier “Management by Objectives” concept but added the specific, measurable Key Results component. Doerr introduced the system to Larry Page and Sergey Brin while Google was still a very small company, and Google has since published its own official guide to the practice through its re:Work initiative (Google re:Work, “Set Goals with OKRs”), which specifies that Objectives should be ambitious enough to feel somewhat uncomfortable, while Key Results should be concrete and easily gradable.

One nuance worth including for accuracy: Rick Klau, a Google Ventures partner whose 2012–2013 talk on “How Google Sets Goals” became one of the most widely cited public explanations of Google's OKR practice, later published a follow-up acknowledging that parts of his original framing had been oversimplified or wrong — a useful reminder that even the most-cited practitioner accounts of a framework can need correction over time, which is part of why this guide prioritizes primary, dated sources over popular retellings wherever possible.

OKRs solve a different problem than RICE or Kano: they're not a prioritization tool for comparing ideas, they're a goal-setting structure that sits above the roadmap, defining what success actually means for a given quarter before any specific feature gets scored or scheduled.

The most common failure mode with OKRs, well-documented across practitioner writing including Klau's own later corrections, is writing Key Results that are actually just a disguised list of features — “ship Feature X” instead of a measurable result like “increase weekly active usage of the core workflow by 15%.” A feature-shaped Key Result quietly reintroduces the exact output-over-outcome problem Cagan and Perri describe, just wearing an OKR label. The discipline of OKRs only holds if every Key Result is something the team could, in principle, achieve through several different features — forcing the actual choice of which feature to build back into the RICE/ICE scoring step, rather than baking a specific solution into the goal itself.

Now-Next-Later Roadmaps

What is a Now-Next-Later roadmap?

A Now-Next-Later roadmap organizes work into three columns by confidence level rather than calendar dates — Now (actively being built), Next (validated problems being scoped), and Later (strategic bets not yet committed) — avoiding the false precision of a roadmap with specific ship dates for work that hasn't been scoped yet. It was created by Janna Bastow and Simon Cast, co-founders of ProdPad.

Bastow and Cast sketched the format in 2012, originally under the labels “Current / Near Term / Future,” not “Now/Next/Later” — by Bastow's own account, the now-standard name was actually coined afterward by one of ProdPad's early customers, not by Bastow herself (Janna Bastow, ProdPad, “Why I Invented the Now-Next-Later Roadmap”). The format directly addresses a specific dysfunction: a roadmap with hard dates six months out implicitly promises a level of certainty that doesn't exist, and once stakeholders see a date, they treat it as a commitment regardless of how much has actually been validated. Organizing by confidence level instead of calendar time keeps the roadmap honest about what's actually known.

Working Backwards

What is Amazon's Working Backwards process?

Working Backwards is Amazon's internal product-development process: before any significant engineering work begins, the team writes an internal press release, as if the product has already shipped, plus an FAQ document addressing customer and business questions — forcing clarity on customer value before committing resources. It's documented most fully in Colin Bryar and Bill Carr's book “Working Backwards” (2021), both former long-tenured Amazon executives.

The earliest well-documented public account came from Ian McAllister, then an Amazon director, in a 2012 Quora answer describing the practice. The press release is written from the customer's point of view first — what problem does this solve for them, and why would they care — and the FAQ, typically capped at around five pages, is where the team works through the hard internal questions a customer-facing press release wouldn't normally surface: cost, technical feasibility, and what could go wrong (Working Backwards, PR-FAQ Process).

What makes this process distinct from the scoring frameworks above is timing: it happens before a team has committed to build anything, specifically to catch a project whose customer value doesn't actually hold up once someone has to write it down in plain, specific language a real customer would read. A press release full of vague claims (“an innovative new way to manage your workflow”) is itself a signal that the underlying idea hasn't been thought through as clearly as it needs to be before engineering time gets committed.

Shape Up

Basecamp's own product methodology, documented by longtime product lead Ryan Singer in the freely published book Shape Up (37signals, 2019), is worth including as a real, working alternative to a traditional backlog-and-sprint model. Work happens in six-week cycles followed by a roughly two-week cooldown period. At a “betting table,” stakeholders choose which “pitches” — scoped problem-and-approach proposals — get bet on for the next cycle, rather than pulling from an ever-growing backlog. Each pitch carries a fixed “appetite” — a time budget the team commits to — instead of an open-ended effort estimate, and critically, if a project isn't finished when its cycle ends, it does not automatically get extended; the team has to make a deliberate decision to rebet on it.

Shape Up is a useful case study specifically because it replaces several of the frameworks above with one integrated system, rather than layering all of them on top of each other — a reminder that the frameworks in this guide are tools to combine deliberately, not a checklist every team needs to run simultaneously.

Choosing the Right Framework

No single authoritative source publishes a clean “use this framework at this company stage” chart, and treating one as if it existed would misrepresent every framework covered here. What two of the most credible voices in this space do argue, directly and specifically, is more useful than a stage chart would be:

Marty Cagan's consistent argument across SVPG's writing is that prioritization frameworks matter far less than whether a company has empowered product teams at all — teams given problems to solve and held accountable for outcomes, rather than teams handed a features list. In his framing, adopting RICE inside a feature-factory organization just produces a more precisely ranked feature list; it doesn't fix the underlying accountability problem. Teresa Torres makes a complementary point from the discovery side: her Opportunity Solution Tree framework exists specifically because most teams jump straight to scoring solutions (RICE, ICE) before they've mapped the opportunity space widely enough to know whether they're even comparing the right solutions.

Read together, these two arguments suggest a sequence rather than a menu: fix team accountability first (Cagan's point), map the opportunity space before committing to solutions (Torres's point), and only then reach for a scoring framework to compare the specific solutions that survive that mapping. A company that skips straight to RICE scoring without doing either of the first two steps is optimizing the wrong layer of the problem.

Company size and stage do change which frameworks are practical to run in full, even if the underlying sequence stays the same. A two-person founding team doesn't need a formal Opportunity Solution Tree diagram to benefit from Torres's underlying discipline — talking to customers before committing to a solution is the point, not the specific artifact. A fifty-person product organization with several teams working in parallel benefits much more from the visible, shared structure a formal tree or a written PR-FAQ provides, precisely because coordination between teams gets harder as headcount grows and shared artifacts do real work keeping everyone aligned on the same opportunity space. The frameworks scale in formality with team size; the underlying discipline they encode doesn't change.

When Strategy Frameworks Fail

Frameworks fail in two distinct ways worth separating clearly: being skipped entirely, and being copied without the underlying conditions that made them work elsewhere.

Quibi's 2020 shutdown, only months after launch, is a documented case of the first failure mode. Founder Jeffrey Katzenberg wrote in his own Fortune op-ed that the company's “failure was not for lack of trying” ( Jeffrey Katzenberg, Fortune, Oct. 2020), but the company had bet heavily on proprietary “Turnstyle” technology — seamless switching between portrait and landscape video — that was never validated as something users actually wanted or would pay for, and shipped as a finished, unadjustable product rather than something iterated against real usage. A Working Backwards press release, or even a simple Opportunity Solution Tree asking what job that switching technology was actually hired to do, would have surfaced this risk before hundreds of millions of dollars were committed to it.

The Spotify “squads and tribes” model illustrates the second failure mode. Spotify's 2012 engineering-culture video went viral and was widely copied by other companies as a template for organizing product teams. Spotify itself has since described it as a snapshot of one particular moment, not a blueprint, and outside accounts — including engineer Jeremiah Lee's detailed post “Spotify's Failed #SquadGoals” — document that the promised team autonomy was largely a myth in practice, and that Spotify itself moved away from the model. The lesson isn't that squads and tribes are inherently bad; it's that copying a structure without the underlying culture, trust, and accountability that made it appear to work elsewhere rarely reproduces the result.

Building a Decision-First Roadmap

Combining the frameworks above into one working process looks roughly like this, in sequence:

  1. 1

    Set one Objective and a small number of Key Results

    Before scoring any individual idea, define what success actually means this quarter using OKRs — an ambitious, qualitative goal with 2-4 measurable results.

  2. 2

    Pick or confirm a North Star Metric

    Choose the single measure that best captures value delivered to customers, and connect it explicitly to the OKR above so both point at the same outcome.

  3. 3

    Map the opportunity space with an Opportunity Solution Tree

    Run customer interviews framed around Jobs to Be Done, and map the real needs and pain points underneath your North Star metric before considering any specific solution.

  4. 4

    Classify candidate features with the Kano Model

    Separate Must-be work (can’t skip, won’t differentiate) from Attractive work (differentiates, but isn’t urgent) so scoring in the next step compares like with like.

  5. 5

    Score the surviving solutions with RICE or ICE

    Only now compare specific solutions numerically — after opportunity mapping has already narrowed the field to genuinely relevant candidates.

  6. 6

    Write a Working Backwards press release for the top-scored bet

    Before committing real engineering time, force the team to articulate customer value in plain language a real customer would actually read.

  7. 7

    Publish the result as a Now-Next-Later roadmap

    Communicate the decision by confidence level, not calendar dates, so stakeholders don’t mistake an early-stage bet for a committed deadline.

Not every team needs every layer of this on every decision — a small, well-scoped bug fix doesn't need a full Opportunity Solution Tree. The sequence matters more for larger, more consequential bets, where skipping a layer is exactly how a company ends up building its own version of Quibi's Turnstyle: a well-executed feature nobody actually needed.

Mistakes to Avoid

  • Adopting a scoring framework to fix an accountability problem. Cagan's core argument applies directly: RICE scoring inside a feature-factory organization just produces a more precisely ranked features list, not empowered product teams.
  • Skipping opportunity mapping and scoring solutions you already had in mind. This inverts Torres's core insight — scoring frameworks compare solutions, they don't tell you whether you're considering the right ones.
  • Treating Now-Next-Later dates as committed deadlines anyway. The format only works if stakeholders actually treat “Later” as genuinely uncertain, not as a euphemism for a date nobody wanted to commit to publicly.
  • Copying another company's org structure without its underlying culture. The Spotify model is the clearest documented example — the visible structure was copied far more often than the trust and accountability that (arguably, even at Spotify) made it function.
  • Using a North Star Metric as the only metric that matters. Casey Winters's critique applies here — a single number can be gamed and can hide important secondary effects a broader metric set would catch.
  • Writing Key Results that are really just features in disguise. A Key Result phrased as “ship Feature X” smuggles the output-over-outcome problem back in under an OKR label, and removes the actual decision about which solution to build from the prioritization step where it belongs.
  • Repeating a framework's garbled origin story instead of its actual mechanism. The specific inaccuracies in this guide — the McDonald's milkshake detail, RICE's misattributed 2014 origin — spread because retelling a good story is easier than checking a primary source, and the same laziness that garbles an origin story tends to garble the mechanism too.

Frequently Asked Questions

What is the difference between RICE and ICE scoring?

RICE (Reach, Impact, Confidence, Effort), created at Intercom, is a more detailed scoring model suited to comparing a smaller number of larger initiatives. ICE (Impact, Confidence, Ease), widely attributed to Sean Ellis, is a lighter-weight version suited to prioritizing many fast-moving growth experiments quickly.

What is the real story behind the Jobs to Be Done milkshake study?

Bob Moesta's own account describes the client only as "a leading fast food restaurant chain," not McDonald's by name as commonly retold. Researchers observed purchases and interviewed customers, finding that roughly 40% of milkshakes were bought in the morning by solo commuters using the milkshake to make their commute more interesting — competing against bagels and coffee, not other milkshakes.

Who actually invented OKRs?

Andy Grove developed the system at Intel in the 1970s, documented in his own book "High Output Management" (1983). John Doerr, who learned the system at Intel, introduced it to Google's founders in 1999 while the company was still small, and documented that history in his own book "Measure What Matters" (2018).

Should a startup use OKRs or a North Star Metric?

They solve different problems and work well together. OKRs set an ambitious goal with measurable results for a specific period. A North Star Metric is the ongoing measure a team optimizes toward across periods. Many teams use a North Star Metric to anchor what their quarterly OKRs should actually target.

What is an Opportunity Solution Tree used for?

It maps a single business outcome down through the customer opportunities (needs and pain points) that could drive it, the range of candidate solutions for each opportunity, and the specific tests used to validate a solution before committing to build it. Created by Teresa Torres, it exists specifically to prevent teams from jumping straight to comparing solutions before mapping the underlying opportunity space.

Is the Spotify squad model a good framework to copy?

Not without significant caution. Spotify's own 2012 engineering-culture video went viral and was widely copied, but Spotify itself has described it as a snapshot of one moment rather than a blueprint, and outside accounts document that the promised team autonomy was largely a myth in practice. Copying the visible structure without the underlying trust and accountability rarely reproduces the intended result.

What should come first: a scoring framework like RICE, or team structure?

Team structure and accountability, per Marty Cagan's consistent argument. Adopting a prioritization framework inside an organization that hands teams a features list to build, rather than problems to solve, just produces a more precisely ranked features list — it doesn't fix the underlying lack of outcome accountability.

What is the Kano Model actually used for in roadmap planning?

It classifies features by their real relationship to customer satisfaction: Must-be features (their absence causes dissatisfaction, but presence doesn't increase satisfaction), One-dimensional/Performance features (more is proportionally better), and Attractive features (delighters that differentiate). This prevents a roadmap from treating a basic reliability fix the same way it treats a genuinely novel, differentiating feature.

What is a good Key Result for an OKR?

A good Key Result is a measurable outcome, not a feature to ship — for example, "increase weekly active usage of the core workflow by 15%" rather than "ship Feature X." Phrasing a Key Result as a specific feature smuggles the output-over-outcome problem back into a framework designed specifically to prevent it, and removes the actual prioritization decision from the step where it belongs.

How does Amazon's Working Backwards process actually work?

Before committing engineering resources, the team writes an internal press release as if the product has already shipped, plus an FAQ (typically capped at around five pages) addressing customer and business questions. The press release is written from the customer's point of view first, forcing the team to articulate real customer value in plain language before any significant building begins.

It's worth ending on the same caution Shape Up illustrates: none of this is meant to become bureaucracy for its own sake. A team that runs every one of these frameworks in full, on every decision, regardless of size, has replaced one form of dysfunction (deciding by whoever's loudest) with another (deciding by whoever's most patient with process). The point of a decision-first roadmap is that the rigor scales with the stakes — a small fix gets a quick gut check informed by the same underlying discipline, while a major bet gets the full sequence, because a major bet is exactly where skipping a step is most expensive.

Every framework in this guide answers a different, specific question: what job is the customer hiring your product for, which opportunities are worth pursuing, which solutions deserve engineering time, what does success actually mean this quarter, and how do you communicate the resulting decisions honestly. None of them replaces judgment, and none of them works as a bolt-on to a feature list assembled the old way. Used together, in sequence, they turn a roadmap from a list of promises into a chain of defensible decisions — which is the entire difference between a product strategy and a queue.

Have a build brief already forming in your head?

Loomstrat Studio scopes, builds, and hands over production software in 3–6 weeks — fixed price, 100% repository ownership.