Skip to main content

Research Agenda for Behavioral Strategy #

Jason Hreha· Updated September 5, 2026

Research status: This page proposes questions and study requirements for evaluating Behavioral Strategy. It does not report completed studies, registered protocols, recruited participants, or an established research partnership. The integrated practitioner approach has not been independently validated.

Research on individual behavioral mechanisms can inform practice without establishing the value of the complete workflow. Retrospective company cases can illustrate an interpretation without showing that the original team used the method. The evidence overview separates those evidence types.

The next research task is to test what the proposed approach adds, where it adds cost, and when it fails. A useful study must be able to find no benefit, identify harm, or support a simpler alternative. Publication, a persistent identifier, or association with a well-known case would not substitute for that test.

1. Do the tools support consistent judgments? #

The Behavior Fit Assessment compares Dispositional Fit, Capability Fit, and Context Fit. Its ordinal ratings, minimum-dimension heuristic, and starting threshold remain practitioner judgments. Research should examine these elements separately before treating a score as a calibrated measure.

Proposed question Evidence a study would need What would challenge the tool
Do raters understand the three constructs in the same way? Independent coding of the same behavior, population, context, and evidence packet, with reasons for each rating Persistent confusion between enduring preferences, actual skills, and external opportunity
Do raters agree enough for the intended decision? Agreement estimates with uncertainty, ratings from different experience levels, and analysis of disagreements Decisions changing mainly with the assigned rater rather than the supplied evidence
Does the minimum rating identify a useful research priority? Compare the suggested bottleneck with barriers found through later observation Other dimensions or a simpler checklist identifying relevant barriers as well or better
Does a proposed threshold improve decisions? Performance across candidate thresholds, with explicit costs of false rejection and false acceptance A cutoff failing to transfer, discarding useful candidates, or adding no value beyond reviewing the evidence

Agreement alone would not establish validity. Raters can agree on a mistaken interpretation. A study should also examine whether ratings add information beyond available observations, prior knowledge, or a simpler comparison. Preserve the version of the definitions and anchors used; do not silently combine legacy Identity Fit ratings with current Dispositional Fit ratings.

2. Does the workflow improve decisions against an actual alternative? #

The integrated approach includes selecting candidate behaviors, examining fit, designing support, and verifying results. A comparative test should specify which parts are being evaluated rather than treating the name of a framework as the intervention.

The comparator could be an organization’s documented current practice, a simpler behavior-selection checklist, or another method implemented according to its own guidance. Choose it because it answers the same decision and is credible in that setting. The framework comparisons help identify possible alternatives; they are not comparative effectiveness studies.

A proposed comparison should specify:

  • The same decision problem, available information, time, staffing, and access to participants across conditions.
  • What training each team receives and how researchers check that each method was actually used.
  • The unit assigned to a method, such as teams or projects, and how shared staff or learning could contaminate the comparison.
  • Decision criteria recorded before outcomes are known, with independent assessment where practical.
  • Costs of using the method, including research time, delayed action, and unnecessary requirements.

Decision quality might include the clarity of the selected behavior, whether critical constraints were found, whether evidence changed the choice, and whether the team avoided an unsupported commitment. These proposed criteria need their own definitions and reliability checks. A panel’s preference for a more detailed memo is not enough to establish a better real-world decision.

3. Do better decisions produce better outcomes? #

A favorable process score does not establish practical value. Follow-up would need to test whether the resulting system changes the target behavior and improves the desired outcome under real conditions.

Define the behavior, eligible or exposed denominator, observation window, baseline, missing-data rules, and uncertainty before comparing results. Match the observation period to the behavior: an episodic action should not inherit the retention schedule of a daily-use product. The measurement standards provide the reporting vocabulary, but do not supply universal success thresholds.

Measure the desired outcome alongside completion, time to first behavior, and repetition. Increased activity could be unhelpful if the activity is a poor route to the goal. Examine whether gains come from shifting effort or cost to someone else, excluding harder-to-serve users, or displacing another useful behavior.

Relevant adverse outcomes should be specified for the domain. These could include burden, unwanted pressure, loss of choice, privacy costs, errors, or unequal access. Product studies also need economic evidence; public programs need evidence of sustainable operations and funding. Any consequential study would require suitable research and domain review before it begins.

4. Where do findings apply, and where do they fail? #

Research should describe the population and setting precisely enough to identify the limits of a result. Volunteers, experienced practitioners, or highly motivated early adopters may not represent the people who will later use the method or perform the behavior.

Prespecify the differences most likely to matter, such as prior skill, recurring preferences, available resources, organizational authority, and whether several actors must coordinate. Report which groups and contexts were absent. Sample size should follow the research question, assignment unit, expected variability, and required precision rather than a universal participant count.

Useful failure conditions include a method producing unstable choices, overlooking an important constraint, increasing decision cost without improving outcomes, or working only when its creator supplies substantial interpretation. A positive result in one domain should motivate a new transfer test rather than establish a general claim across products, healthcare, education, and policy.

What a future study should make reviewable #

Before data collection, a study should identify its question, protocol version, comparator, primary outcome, analysis plan, stopping rules, and relevant conflicts of interest. Any registration should be linked only after it exists. Separate exploratory analyses from the prespecified test.

Afterward, report methods and results under the same editorial rules whether findings are favorable, null, or adverse. Where consent and data rights permit, share materials, de-identified data, and analysis code sufficient for independent scrutiny. Explain deviations, missing observations, uncertainty, and limitations. An author-led demonstration can inform further research; it should not be labeled independent validation.

Contribute a question, critique, or study proposal #

Use the existing contact page to suggest a comparator, challenge a construct, propose a research collaboration, or share a relevant source. A useful proposal states the question, setting, available evidence, comparison, outcome, and what finding would change the claim. Identify the researchers’ roles and any interest in the result.

For a factual error in current public material, use the corrections process. This agenda is a statement of open questions. A future partnership or study would be described separately with its actual status and supporting record.