Skip to main content

Behavioral Strategy Synthesis: Why Behavior Selection Is High Leverage #

Jason Hreha· Updated September 5, 2026

This page synthesizes patterns from success and failure cases across consumer, enterprise, and public-sector settings to ask: How does target-behavior selection interact with context and execution? The cases suggest that behavior selection is an early, high-leverage decision. They do not establish a universal causal ranking of success factors.

Evidence note: Fit explanations are retrospective hypotheses, not measurements made by the companies. Use the linked cases and Evidence Ledger for sourced outcomes and their limits. Planning examples do not establish company results or universal thresholds.

The Behavior Fit Assessment is a practitioner decision tool for comparing candidate behaviors across Dispositional Fit, Capability Fit, and Context Fit. It is not a validated measurement instrument. Treat the minimum dimension as a bottleneck and prioritization heuristic; it is not a deterministic probability of behavior.


Executive Summary: A Recurring Pattern #

Across the cases reviewed, including Instagram, Slack, Duolingo, M-PESA, Netflix, Spotify, Zoom, Proposify, Airbnb, and Robinhood, a recurring pattern is:

Cases interpreted as stronger fits pair target-behavior selection with supportive context, motivation, and ability. Execution remains important, and the cases do not isolate the causal contribution of each factor.

Across the failure cases reviewed, including Quibi, Google+, Burbn, selected gamification efforts, and corporate wellness, a recurring retrospective hypothesis is:

Demanding changes to the environment, relatively enduring preferences, or active motivation can make a target behavior harder to sustain. Treat this as a hypothesis to test, not a single-cause explanation of an initiative’s outcome.

This synthesis offers five practitioner interpretations to test when comparing behaviors and designing solutions.


Insight #1: Validate Behavior Selection Before Scaling Execution #

The Pattern in Successes #

Several success narratives include a pivot moment in which teams shifted toward a behavior that appeared to have stronger fit:

  • Instagram: The pivot from Burbn toward photo sharing illustrates selection informed by observed use. Stronger Dispositional Fit is a retrospective interpretation.

  • Slack: An internal tool became a product for persistent, searchable team messaging. The case reports a company-observed association between early message volume and team retention; it does not isolate the effect of behavior selection.

  • YouTube: The early platform broadened toward general video sharing. The role of dating in its origin story is disputed; the case does not establish a clean causal pivot from one prescribed behavior.

  • Duolingo: Short lessons illustrate a behavior that may fit brief periods of available time. The case reports proficiency among selected course completers; it does not show that lesson length caused learning or retention.

  • M-PESA: Enabled phone-based remittance through mobile access and an agent network, supporting an existing need to send money. The case describes the infrastructure and reported outcomes without attributing them to one feature alone.

The Pattern in Failures #

Several failure narratives suggest teams either:

  1. Never validated behavior fit before building, or
  2. Continued optimizing the wrong behavior despite poor fit signals
  • Quibi: Selected passive short-form video consumption (TV behavior) for a mobile context that also supports interactive behaviors. Production quality did not produce a sustainable business; context mismatch is one retrospective hypothesis among several.

  • Google+: Building a social network required users to establish connections and sharing routines alongside existing services. Switching costs and group coordination are candidate constraints; the failure note does not isolate their effects.

  • Burbn: The founders shifted from a product with check-ins and photos toward photo sharing. The case supports examining the required behavior, but does not isolate why the earlier product struggled.

  • Selected gamification failures: Applied point/badge collection to tasks users did not appear to value. Extrinsic rewards may not compensate for low motivation or high friction, so the underlying behavior still requires direct validation.

The Insight #

Behavior selection deserves validation before execution scales. This is a prioritization claim, not a measured ratio between selection and execution quality. Teams should test plausible behavior alternatives early rather than infer fit from implementation effort.


Insight #2: Dispositional Fit Is a Candidate Constraint #

What Is Dispositional Fit? #

Dispositional Fit means: Does the behavior match the population’s relatively enduring tendencies and preferences over the decision-relevant time horizon?

Look for evidence about interests, values, characteristic priorities, tolerances, recurrent motives, typical responses, and established behavior patterns. Self-concept can be one signal in a particular case, but it does not define the dimension. “Identity Fit” is retained only as the legacy alternate name.

Successful Dispositional Fits #

  • Instagram (photo sharing): Existing photo-taking and visual-sharing behavior indicated interest in capturing and expressing moments. The behavior drew on preferences already visible in the target population.

  • Duolingo (micro-learning): Short lessons can fit people with an existing interest in language learning but a preference for brief, structured practice rather than long study blocks.

  • Spotify (Discover Weekly): A ready-made discovery behavior fits listeners who already value music and novelty while removing the effort of searching a large catalog.

  • Airbnb (peer-to-peer booking): Fit differs sharply by segment. Comfort with novelty, tolerance for interpersonal risk, and interest in hosting or local travel are more useful hypotheses than a generic “traveler identity.”

  • M-PESA (mobile remittance): The product enabled an existing, durable priority, supporting family, through a behavior that was faster and safer than prior remittance methods.

Failed Dispositional Fits #

  • Quibi (passive video): Focused ten-minute premium viewing appeared poorly matched to established mobile-use patterns and preferences for interruptible consumption. Context was also a major constraint, so the case should not be reduced to disposition alone.

  • Google+ (social sharing): Participation competed with established communication preferences and existing social graphs. Whether the target population wanted another sharing routine was a question to test.

  • Corporate wellness gamification: Badges and points may be a poor match for employees who dislike public competition, tracking, or patronizing mechanics. Those preferences should be measured rather than inferred from a professional label.

  • Habit-tracking apps: A product can mistake an aspiration to change for a durable preference to record behavior every day. Tracking willingness and routine preferences should be observed directly.

The Insight #

Dispositional Fit can help explain why one candidate behavior appears easier to adopt or sustain than another. Use it as a comparative hypothesis, then verify it through observation and realistic trials rather than assuming it determines persistence.


Insight #3: Context Fit Tests Environmental Support #

What Is Context Fit? #

Context fit means: Does the user’s environment enable this behavior to happen naturally?

Context includes:

  • Physical environment: Where users are, what devices they have, what tools are nearby.
  • Temporal patterns: When users have attention, how fragmented or focused their time is.
  • Social environment: Who else is present, what are group norms, what do peers reinforce.
  • Infrastructure: Access to networks, agents, APIs, payment systems, etc.

Successful Context Fits #

  • Slack (persistent messaging): Distributed teams, available devices, and a need for searchable history can support recurring messaging. This is a context hypothesis; the case does not supply a measured time-to-first-message benchmark.

  • Zoom (video meetings): Remote work increased the need for synchronous collaboration. Joining through a meeting link can reduce access steps, but this synthesis does not establish a measured TTFB or separate the effect of design from the pandemic context.

  • M-PESA (mobile remittance): Feature phones and agent access can support remittance where bank access is limited. Agent availability and other transaction constraints still matter; this case does not establish a universal TTFB benchmark.

  • Duolingo (micro-lessons): Short guided practice may fit settings with brief periods of available time. The cited outcome study does not test context fit or establish that short lessons caused repeated use.

  • Spotify (Discover Weekly): A ready-to-play weekly playlist may reduce selection effort during commuting or household tasks. These are possible use contexts to test; the case does not measure automaticity or establish habitual consumption.

  • Proposify (guided onboarding): The redesigned tutorial guided users through the proposal workflow. The company reported early associations between tutorial completion, proposal sending, and paid conversion, but supplied no usable effect estimate or measured TTFB. BS-0078

Failed Context Fits #

  • Quibi (mobile viewing): Interrupted mobile use is a possible constraint on focused viewing. The pandemic also removed much of the intended commute context, and the internal decision process is partly unknown. The shutdown does not isolate a single context-related cause. BS-0070

  • Google+ (network participation): Joining a service involves more than an individual’s interest when the useful behavior requires other people to participate. Existing networks and coordination costs are context questions, not proof of a single cause of failure.

  • Corporate wellness (hypothetical diagnostic example): An office gym may be difficult for remote employees to use. Test access, schedules, and competing obligations in the named workforce before attributing participation or retention to context.

  • Mint and YNAB (budgeting): Mint automated spending tracking; YNAB centers active allocation before spending. The contrast raises a question about when budgeting decisions occur. Mint did not require batched monthly data entry, and its shutdown does not establish that passive tracking caused behavioral failure.

  • Habit-tracking app (hypothetical diagnostic example): A daily tracking action may be difficult when schedules or locations change. Observe the intended users’ routines and interruptions; this example does not estimate how common stable routines are.

The Insight #

Context Fit asks whether the Social and Physical Environments support the behavior under realistic conditions. Weak support can be a meaningful constraint even when Dispositional Fit appears strong. Test whether matching or redesigning infrastructure, physical space, timing, or social norms changes observed behavior.


Insight #4: Test Nudges Against the Specific Decision #

“Nudge” is often used as a catch-all for good UX, good onboarding, and good product design. On this site we use a narrower meaning: choice-architecture tweaks (defaults, framing, reminders, simplification) that aim to shift behavior without changing the underlying feasibility or value of the behavior.

The compared nudge-unit trials reported a positive average effect of about 1.4 percentage points; some publication-bias-adjusted syntheses estimate an average near zero for the studies they analyze. These findings concern particular samples and analysis choices. They do not forecast every intervention or compare Four-Fit or DRIVE with other methods. See: Behavioral Strategy vs Nudging and Why Nudges Fail. BS-0003 BS-0027

Practitioner recommendation #

The author’s proposal is to investigate behavior fit and enablement before relying on a prompt, default, or framing change. This is a workflow recommendation, not a study-proved advantage over nudging. Test whether:

  • the behavior fits the segment and context,
  • the system enables the action,
  • the value is sufficient and timely.

How to apply the proposal #

  1. Compare candidate behaviors across Dispositional Fit, Capability Fit, and Context Fit, then validate the selected behavior.
  2. Enable the behavior (tools, workflow, infrastructure, incentives/governance where appropriate).
  3. Test candidate choice-architecture changes with clear success criteria and compare their observed outcomes and costs.

When testing a nudge #

Treat it as a falsifiable experiment:

  • define the target behavior, denominator, and window
  • pre-commit to a minimum effect that justifies adoption
  • set rollback criteria, especially for trust/ethics

Insight #5: Observe Behavior Before Designing Solution #

The Validation Pattern in Successes #

Several success narratives include observation of what users were already trying to do before or during design:

  • Instagram: Founder accounts describe a move toward photo sharing based on early use. The case does not establish the founders’ use of BFA or a controlled test of selection against execution.

  • Slack: Observed that teams were trying to organize scattered communication (email, Campfire, AIM). Behavior signal: users wanted searchable, persistent team chat. Slack built exactly that, not a “better email” or generic collaboration suite.

  • YouTube: General video sharing accommodated a wider range of creator interests. Because the dating-origin narrative is disputed, treat this as a behavior-broadening interpretation rather than a settled account of a formal validation process.

  • Duolingo: Short guided practice suggests a hypothesis about available time and required skills. The cited study concerns selected course completers; it does not document the team’s initial research or prove the proposed mechanism.

  • M-PESA: Existing remittances moved through channels such as bus drivers and traveling relatives. Mobile access and cash-in/cash-out agents supported that need through a different transaction system.

  • Zoom: The product existed before the pandemic, which sharply increased demand. Its guest join flow illustrates a design question about access barriers; the growth record does not show that observing pandemic workers led to the original design.

  • Spotify Discover Weekly: Personalized pre-selection suggests a hypothesis about reducing discovery effort. The cited case does not document the team’s historical observations of search time or establish that choice overload drove the design.

The Validation Pattern in Failures #

The public failure records raise validation questions. The diagnostic examples below are hypothetical and do not establish what any team’s internal research found or omitted.

  • Quibi: The mobile-viewing hypothesis faced context and competition problems. The internal research process is partly unknown, so the public record does not justify saying the team never tested it.

  • Google+: The case prompts a test of willingness to establish another social network. The public failure note does not document the team’s research well enough to say it never examined existing networks or switching costs.

  • Gamification (hypothetical diagnostic example): Before adding badges or points, test whether intended users want to perform the underlying behavior and how they respond to competition or tracking. This example makes no claim about the prevalence or research history of failed programs.

  • Habit tracking (hypothetical diagnostic example): Observe when tracking is possible and whether recording behavior is useful to the intended population. Test routine stability and time constraints without assuming that a professional role determines either.

  • Corporate wellness (hypothetical diagnostic example): Ask whether time, care obligations, location, or access limits the proposed action, then observe realistic attempts. These are possible barriers to investigate, not reported findings about an unnamed workforce.

The Validation Process #

These retrospective cases suggest a Problem Market Fit validation workflow to test, rather than proving that every team used it:

  1. Observe what target users are already doing (behaviors, workarounds, informal solutions).
  2. Validate that this observed behavior is a real problem (not an edge case or wishful thinking).
  3. Rank candidate behaviors by frequency, importance, and user motivation (not by what’s novel or exciting to build).
  4. Select the behavior with the strongest evidence across Dispositional Fit, Capability Fit, and Context Fit.
  5. Build the solution around enabling this behavior, not changing it.

The Insight #

Observation of what users are already doing is stronger starting evidence than unsupported forecasts from surveys or focus groups alone. Survey data can reveal beliefs and stated preferences; behavioral data shows what people did, at what frequency, and with what effort. Behavioral Strategy starts with both forms of evidence, then tests candidate behaviors in realistic contexts.


Five Possible Moves in Behavioral Strategy #

The cases suggest five possible moves to consider:

Move 1: Pivot When Behavior Doesn’t Fit #

  • Instagram pivoted from check-ins to photo sharing.
  • Slack pivoted from internal coordination tool to team messaging platform.
  • YouTube broadened toward general video sharing; accounts differ on the role of dating in its origin.

Key principle: When early data misses the relevant behavior criteria (long TTFB, poor retention, low completion), investigate fit and compare alternatives before building further. These measures identify questions; they do not determine the cause.

Signal to watch: If first completion or retention falls below a pre-registered criterion for the behavior, population, and window, investigate the limiting conditions and test alternatives. There is no universal first-completion or D7 cutoff.

Move 2: Simplify Behavior to Increase Capability Fit #

  • Duolingo uses short guided lessons. Whether this format improves completion or learning requires a suitable comparison.
  • Proposify guided users through a tutorial and simulated proposal experience before a real proposal; any effect on first-proposal completion requires a separate test.
  • M-PESA reduced money remittance to USSD menu taps (vs. bank account setup, account minimums, documents).
  • Zoom reduced joining to one-click link (vs. remembering meeting IDs, navigating software).

Key principle: When users struggle with first behavior completion (high TTFB, low completion rate), simplify the behavior itself rather than the interface alone. Reduce steps, reduce decisions, reduce cognitive load.

Signal to watch: If TTFB exceeds the pre-registered domain threshold or first-instance completion falls below its pre-registered criterion, investigate capability friction and other causes. Test whether a simpler or more modular behavior improves observed completion.

Move 3: Change Context to Increase Context Fit #

  • Peloton brought fitness to home context (vs. gym context).
  • Google Photos leveraged cloud storage and automatic backup (matching user’s mental context: “I want my memories safe and accessible,” vs. “I want to organize files”).
  • M-PESA leveraged distributed agent networks (matching user’s physical context: nearby, trusted, accessible).
  • Spotify Discover Weekly leveraged weekly rhythm (matching user’s temporal context: work week, commute routine).

Key principle: When Dispositional Fit and Capability Fit appear strong but adoption remains weak, Context Fit may be blocking. Test environmental changes that enable reliable performance before assuming the candidate behavior itself is wrong.

Signal to watch: If users report that timing, location, or tooling prevents behavior (qualitative), redesign context first before optimizing interface.

Move 4: Target a Different Actor When Direct Approach Fails #

This is less common but powerful:

  • Spain’s organ donation system: Instead of changing individual decision-makers, targeted hospital coordinators as the actor. Coordinators became behavioral enablers for the system.
  • Digital health platform redesign: Instead of expecting individual IT managers to self-train on complex integration, redesigned behavior to involve webinars and support staff (shifting actor from individual to group).
  • Server training vs. patron education: In restaurant/bar settings, it’s often more effective to train servers (what menu items to recommend, how to present) than to educate patrons directly. Change the actor who influences the decision.

Key principle: When target user isn’t adopting, sometimes you need to work backward through the system to find a leverage point. Who influences the target user? Can you design behavior for that influence actor instead?

Signal to watch: If target users consistently fail to adopt despite low friction and high motivation, ask: who else is involved in this decision chain? Can we design for them instead?

Move 5: Measure via Durable Behavior, Not Engagement Metrics #

The following are candidate measures for these products. They are not claims that each company used the listed metric exclusively or published a test of it.

  • Instagram: Photo uploads and repeat posting per defined cohort.
  • Slack: Message completion and repeated team use.
  • Duolingo: Lesson completion paired with repeated skill assessments.
  • Proposify: First proposal sent and time from signup to that event.
  • Spotify: Discover Weekly listening and repeat discovery per exposed user.

Key principle: Engagement metrics (DAU, logins, page views) are cheap to game and tell you little about actual behavior adoption. Target behavior metrics (frequency, completion rate, retention) are harder to game and reveal real fit.

Signal to watch: If DAU is high but target behavior completion is low, or if retention is poor despite engagement spikes, re-examine whether measured behavior is the right behavior.


Hypotheses About Why Failures Happened #

Type 1: Conceptual Failure (Wrong Behavior Selected) #

Root cause: Designers selected a behavior that does not fit the population’s dispositions, capability, or context.

Hypotheses to investigate: Quibi’s mobile-viewing demands, Google+’s participation and switching demands, or a gamification scheme that rewards a behavior the population does not value. These examples do not establish a single cause of a named outcome.

Prevention: Validate behavior market fit before building. Use TTFB, first-completion rates, and retention metrics to test whether behavior is naturally desired.

Type 2: Design Failure (Right Behavior, Too Much Friction) #

Root cause: Behavior is right, but solution adds friction instead of removing it.

Illustrative examples: A habit tracker requiring unnecessary daily entry, or a wellness program offering gym access when the target population needs home-based options. These are diagnostic scenarios, not established causes of a named company’s outcome.

Prevention: Test TTFB and first-completion rates with early prototypes. Investigate friction when either measure misses its pre-registered, domain-specific criterion.

Type 3: Context Failure (Right Behavior, Wrong Environment) #

Root cause: Behavior is right, but user’s actual environment doesn’t enable it.

Examples: Wellness programs for remote workers (assuming office context), financial apps (assuming moment of decision is when you’re planning budget, not when you’re spending).

Prevention: Test in actual user contexts, not controlled labs. Observe where, when, and how users attempt the behavior naturally.

Type 4: Scaling Failure (Works for Early Adopters, Not Mainstream) #

Root cause: Early adopters have high motivation and fit; mainstream users don’t. As adoption spreads, behavior fit declines sharply.

Examples: Peer-coaching apps (works for motivated enthusiasts, fails for casual users), niche communities (high fit for enthusiasts, low fit for mass market).

Prevention: Measure behavior fit separately for early adopters (enthusiasts) and mainstream segments. Plan for declining fit as adoption spreads.

Type 5: Motivation Misconception (Extrinsic Rewards Crowd Out Intrinsic) #

Root cause: Behavior is right, but solution adds incentives that undermine intrinsic motivation.

Examples: Gamification of learning (badges crowd out curiosity), pay-for-behavior programs (payment crowds out altruism), corporate wellness incentives (bonuses crowd out autonomy).

Prevention: Distinguish intrinsic motivation (desire to do behavior) from extrinsic (rewards for doing behavior). When intrinsic motivation exists, extrinsic rewards often backfire.


The Success Framework: A Practical Guide #

Use this framework to evaluate any behavioral strategy initiative:

Phase 1: Behavior Selection (Problem → Behavior) #

Questions to ask:

  • What behavior are users already attempting (with difficulty)?
  • What is the natural frequency of this behavior? (daily, weekly, episodic?)
  • What is users’ current time-to-first-behavior (TTFB)? (seconds, minutes, hours?)
  • What percentage of target users complete the first instance naturally, relative to the pre-committed domain target?
  • What is retention at intervals appropriate to the behavior, relative to baseline and the value-delivery requirement?

Red flags:

  • Users rarely attempt behavior naturally.
  • TTFB exceeds the pre-registered domain threshold.
  • Dispositional mismatch (the behavior runs against characteristic preferences or priorities).
  • Context doesn’t naturally support frequency needed.

Success signal:

  • Users can articulate why they want to do this behavior.
  • TTFB meets the pre-registered domain criterion.
  • Users repeat the behavior at the decision-relevant interval when repetition is required.
  • No obvious bottleneck appears across Dispositional Fit, Capability Fit, and Context Fit.

Phase 2: Solution Design (Behavior → Solution) #

Questions to ask:

  • What are the current friction points in this behavior? (ability, motivation, environment?)
  • What is the minimum viable behavior (the simplest version)?
  • Which observed barriers can we remove while preserving the core behavior?
  • What context or environment changes would enable natural triggering?

Red flags:

  • Solution adds new steps instead of removing them.
  • Relies on extrinsic incentives (badges, points) for motivation.
  • Ignores actual user context or environment.
  • Assumes behavior change will happen without environmental support.

Success signal:

  • TTFB improves relative to the status quo and meets the pre-registered domain criterion.
  • Completion rate improves relative to baseline and meets the pre-registered domain criterion.
  • Users report behavior feels natural, not forced.
  • Retention at decision-relevant intervals meets the pre-registered criterion.

Phase 3: Validation (Solution → Market) #

Questions to ask:

  • Do actual users, beyond the enthusiasts, adopt this behavior at scale?
  • Does behavior sustain beyond the initial trial at the decision-relevant intervals?
  • Are there unexpected context barriers we missed?
  • Does the behavior create network effects or does it decay?

Red flags:

  • High initial adoption but retention below the pre-registered criterion.
  • Behavior adoption varies wildly by segment (some adopt, others don’t).
  • Context barriers emerge at scale that weren’t visible in small tests.
  • Requires ongoing incentives or nudges to maintain adoption.

Success signal:

  • Retention meets the pre-registered, domain-specific criterion.
  • Behavior adoption is consistent across target segments.
  • No new friction barriers emerge at scale.
  • Users adopt without external incentives or nudges.

Common Anti-Patterns to Avoid #

1. The Rational Actor Fallacy #

Pattern: Designing for how people “should” behave (following stated preferences) instead of how they do behave (following actual context and motivation).

Example: Fitness app designed for “dedicated morning runners” when target market is “busy parents with fragmented time.” Behavior doesn’t match actual user context.

Fix: Observe actual behaviors first. Design for reality, not ideals.

2. The More-Is-Better Trap #

Pattern: Adding features, behaviors, or incentives thinking it increases value when it actually increases friction.

Example: Gamification that adds badges, leaderboards, and daily challenges when core behavior (exercising) already has natural rewards.

Fix: Start with the necessary features. Test whether a new feature improves target-behavior completion or another pre-registered outcome enough to justify its cost and added complexity. There is no universal percentage reduction in friction that each feature must achieve.

3. The Early Success Bias #

Pattern: Mistaking early adopter enthusiasm (high intrinsic motivation, matching context) for mainstream viability.

Example: Fitness community app works brilliantly with fitness enthusiasts but fails to generalize to casual exercisers.

Fix: Measure behavior fit separately for early adopters and target mainstream. Plan for declining fit as adoption spreads.

4. The Metric Gaming Problem #

Pattern: Optimizing for easy-to-measure behaviors that don’t connect to real outcomes.

Example: Optimizing for app opens instead of target behavior completion, or for daily active users instead of behavior frequency.

Fix: Measure durable behaviors, not engagement proxies. Report Δ-B (behavior change percentage points) with clear windows and denominators.

5. The Context Blindness Error #

Pattern: Ignoring environmental and social factors that override individual interventions.

Example: Wellness program assumes all employees have gym access and stable routines; misses that remote workers, shift workers, and caregivers face different constraints.

Fix: Test in actual user contexts. Observe environmental barriers qualitatively. Design for context, not for labs.

6. The Motivation Misconception #

Pattern: Assuming extrinsic rewards (incentives, gamification) can drive intrinsically-unmotivated behaviors.

Example: Using points and badges to drive compliance training completion when users see no value in the training.

Fix: Validate intrinsic motivation first. If users don’t naturally want the behavior, no incentive system will sustain it.


The Path Forward: Behavioral Strategy as Discipline #

What This Synthesis Reveals #

  1. Behavior selection is an early, high-leverage decision.
    • Strong execution cannot by itself establish fit for a poorly chosen behavior.
    • The cases favor testing alternatives early and validating before scaling.
  2. Fit across three BFA dimensions structures behavior-selection decisions. BFA is informed by BSM, but it is not a literal one-to-one condensation or reassignment of the eight BSM components.
    • Dispositional Fit: Does the behavior match relatively enduring tendencies and preferences over the decision-relevant horizon?
    • Capability Fit: Does the population have the actual abilities and skills required?
    • Context Fit: Does the external social and physical environment support the behavior?
  3. Treat nudges as marginal experiments, not proof of fit.
    • Choice-architecture changes may add incremental effects after feasibility is established.
    • They should not substitute for testing behavior, capability, or context fit.
    • System changes to context, ability, or infrastructure may be more appropriate when those are the binding constraints.
  4. Observation strengthens behavior-selection decisions.
    • Start by watching what users are already doing.
    • Validate market fit before designing solutions.
    • Measure via durable behavior, not engagement metrics.
  5. Failures teach us where fit breaks down.
    • Conceptual failure: wrong behavior selected.
    • Design failure: right behavior, too much friction.
    • Context failure: right behavior, wrong environment.
    • Scaling failure: works for enthusiasts, not mainstream.
    • Motivation failure: extrinsic rewards undermine intrinsic motivation.

The Competitive Advantage #

In an era where product execution is commoditized (design, engineering, marketing all become table stakes), behavior selection becomes the last defensible advantage.

Teams that:

  • Observe users before designing
  • Validate behavior market fit early
  • Simplify behavior to remove friction
  • Adapt context to enable natural adoption
  • Measure durable outcomes

…will consistently outcompete teams that:

  • Assume which behavior matters
  • Build first, validate later
  • Optimize interfaces instead of behaviors
  • Rely on nudges and incentives
  • Measure engagement proxies

This is Behavioral Strategy.



In this section