BehavioralStrategy.com teaching pack v1.0
Author: Jason Hreha
Source file: methodology/validate-behavior-market-fit.md
License for original text: CC BY-NC-SA 4.0 (existing site prose). External sources retain their rights.
Editable text copy; Markdown headings and tables are retained.

# How to Validate Behavior Market Fit

> **Definition.** Behavior Market Fit is established when observed evidence shows that the target population can and will perform the target behavior in realistic contexts. A validation decision must state the population, behavior, support, observation window, and limits of that evidence.

This guide turns a candidate behavior into a testable operating decision. BFA ratings help identify questions; they are practitioner judgments, not measured probabilities or proof of BMF. There is no universal completion percentage, sample size, or number of contexts that establishes fit.

Use [workbook sections 5 to 8](https://behavioralstrategy.com/education/decision-workbook/#step-5) to record the protocol before gathering data. The [measurement lab](https://behavioralstrategy.com/education/measurement-lab/) lets you practice without recruiting anyone or claiming to have run a study.

## 1. State the decision and the behavior

Write who must do what, on which opportunity, with what support, and by when. Define the useful outcome separately. Include every actor required for completion.

**Invented teaching example.** An outgoing shift lead records each unfinished task's next action and owner; the incoming lead acknowledges the handover before taking over. A saved form is an attempt in this example. Completion requires the recipient's acknowledgement of a usable handover. Whether the next shift follows the information is a separate outcome question.

State what the test can change: a limited rollout, the support model, the target action, or the decision to invest at all.

## 2. Test the intended service in relevant conditions

Recruit from the intended population and record who declines, cannot be reached, or is excluded. Select contexts because they matter to the decision, such as interrupted work or unavailable staff. Do not substitute convenient enthusiasts for a broader population without stating that limit.

List the normal service: tools, training, available help, approvals, and other people's work. Record exceptional assistance separately. A supported trial can establish evidence about a supported service. It cannot by itself forecast self-service completion.

Measure staff time, failures, queueing, and demand peaks. Confirm that the proposed staffing can deliver the support at the intended volume. Affordable ongoing help is a valid design choice; unsupported repetition is not a universal requirement.

## 3. Choose a comparison that answers the question

Observation can expose where an action breaks. A feasibility pilot can test whether the service can deliver it under stated conditions. Neither alone identifies the effect of the service compared with another route.

If the decision requires an effect estimate, define the comparator and use an appropriate design, such as random assignment when workable. Compare equivalent populations, opportunities, and windows. For team behaviors, account for shared conditions and interaction between participants. Repeated records from the same team are not independent people.

Set sample size and follow-up from the needed precision, expected baseline, meaningful difference, grouping, and missingness. A small diagnostic exercise can justify repairing a process without establishing a population-wide success rate.

## 4. Define measures and missing states

| Measure | Numerator and denominator |
|---|---|
| Reach | People with confirmed delivery or exposure / intended eligible people; define which kind of reach |
| Opportunity | Eligible units with an occasion on which the action was applicable / eligible units observed |
| Attempt | Attempts / known applicable opportunities |
| Completion | Verified completed actions / known applicable opportunities |
| Conditional completion | Verified completions / attempts; retain the broader opportunity rate too |
| Repeat completion | First completers who complete again / first completers with a known later applicable opportunity |
| Useful outcome | Units receiving the defined benefit / a population and window specified for that outcome |

Define whether the unit is a person, household, team, or opportunity. Record an attempt boundary, a completion boundary, timestamps, and a quality check. A click, registration, or submission can be a useful intermediate measure without being the final behavior.

Use separate states for **yes**, **no**, **not applicable**, and **unknown**. A missing receipt is not automatically a failed action. A person without another opportunity is not automatically a failed repeat user. Report unknown opportunity counts and unresolved outcomes, and explain any exclusions.

Show how conclusions change if unresolved outcomes are successes or failures. Such bounds describe missing-data uncertainty; they do not replace confidence intervals, address selection bias, or prove causation. When the denominator is zero, report the rate as undefined, not zero.

## 5. Set specific rules before seeing results

Choose a threshold from the service's needs and consequences. State its denominator, window, support condition, precision requirement, and rationale. Specify how missing data and unexpected events affect the decision.

**Illustrative planning rule, not a BMF standard.** A team needs at least 70% of applicable handovers to complete under normal support because its fallback can handle at most the remaining 30%. Before use, it must verify that capacity assumption. Its test plan must also require enough observations to distinguish acceptable performance from unacceptable performance, a quality check, and an explicit rule for unresolved records.

A point estimate above 70% from a few observations does not meet that broader requirement. If the uncertainty crosses the decision boundary, collect the specific missing evidence or report an inconclusive result.

Record the rule with a date before observation. Changes remain possible, but preserve the original rule and explain the change. Do not call an internal dated plan a publicly preregistered study unless it was actually registered.

## 6. Diagnose an unexpected result

Follow the chain before assigning a cause.

| Observation | Live explanations | Useful next evidence |
|---|---|---|
| Few people start | No delivery, no relevant opportunity, low value, competing work | Delivery records and observation of applicable occasions |
| Many start but do not finish | Missing skill, approval, information, time, or service capacity | Location and reason for the break; work done by other actors |
| Completion rises only with exceptional help | Help removes a barrier; coached participants differ; recording differs | A planned support comparison and delivery-cost check |
| Completed action yields little benefit | Quality failure, incorrect outcome link, delay, another missing step | Outcome and quality measures with an appropriate window |
| Repeat completion is unclear | No new occasion, short follow-up, missing records, falling value | Opportunity-aware follow-up and reasons for noncompletion |

These are hypotheses, not diagnoses from a single rate.

The [California EITC exercise](https://behavioralstrategy.com/education/measurement-lab/#eitc) is a public example of why outreach engagement and completed claims answer different questions. Its null result does not by itself establish that the target behavior was wrong. The [Wikipedia exercise](https://behavioralstrategy.com/education/measurement-lab/#wikipedia) separates first action, a quality proxy, and repeat action. Source records and limitations are provided in the lab.

## 7. Record one of five next decisions

| Decision | When it is warranted | What to record next |
|---|---|---|
| Proceed narrowly | Evidence meets the stated operating and uncertainty rules in the tested scope | Population, support, scale limit, monitoring, and reversal condition |
| Change support | A specific delivery barrier has a credible, affordable repair to test | The changed service, cost/capacity assumption, and comparison |
| Reselect | Evidence favors a materially different action or actor over repairing the current route | The alternative, supporting evidence, and next test |
| Stop | No worthwhile route remains under the stated value, cost, or operational constraints | What was ruled out and what evidence could reopen the decision |
| Inconclusive | Missingness, measurement failure, precision, or opportunity coverage prevents a decision | The smallest useful evidence request and a limit on further work |

A disappointing treatment effect can justify stopping that treatment while retaining the objective. It does not automatically require abandoning the selected behavior. Conversely, observed completion alone does not establish acceptable cost, benefit, or product demand.

## What a completed protocol provides

Keep the behavior definition, actor chain, service specification, comparison, measures, thresholds and rationale, raw observation rules, limitations, and next decision together. Another reader should be able to tell exactly what the evidence supports and which claims remain untested.

The [workbook](https://behavioralstrategy.com/education/decision-workbook/) supplies those fields. The [M-PESA worked decision](https://behavioralstrategy.com/education/mpesa-worked-decision/) shows how a complete response can finish with a proposed test rather than an invented historical verdict.
