Skip to main content

Evidence Ledger #

Jason Hreha· Updated September 5, 2026

Each global claim links here. Rows include the evidence class, domain, measurement, effect size, window, case bindings, sources, and confidence level.

How to read a row #

  • Effect size reports the study-specific estimate or bounded outcome in the source’s unit, such as percentage points, standardized effects, counts, proficiency levels, or qualitative bounds
  • Window is explicit (for example D0-D7)
  • Evidence class uses the five classes defined in the Editorial Policy
  • Case bindings identify public cases that cite the canonical ledger record
  • Sources link to a case write-up and raw data when possible
ID Evidence class Claim Domain Measure Effect size Window Confidence Case bindings Sources
BS-0001 Observational case The case library uses behavior-first validation as a planning principle across documented examples. Cross-industry Presence of the planning principle across case interpretations Qualitative synthesis only; no comparative effect estimate Project-specific Working None Case
BS-0002 Illustration The measurement guide illustrates how to test whether shorter time to first target behavior predicts higher medium-term retention. Consumer and SaaS Minutes to first target behavior vs 90-day behavior retention Product- and cohort-specific (validate per segment; no single published effect size) Baseline to 90 days (or product-specific) Working None Method
BS-0003 Comparative test At-scale nudge-unit RCTs show small average effects (~1.4 pp), far below academic journal averages (~8.7 pp). Public sector and enterprise Average percentage-point change in target outcome across large RCTs ~1.4 pp average (vs ~8.7 pp in academic journals) 126 RCTs; 23 million individuals High None Study
BS-0004 Comparative test Opt-out (default) organ donation regimes do not reliably increase transplantation; complementary system investments drive outcomes. Healthcare policy Deceased and living donor rates; transplant rates No reliable increase in deceased donation after opt-out policy shifts; some evidence of decreased living donation (−29%) Cross-national, multi-year comparisons High Case Review, Study, Study
BS-0005 Observational case Instagram pivoted from check-ins to photo sharing after observing how early users behaved, and the change preceded rapid reported user growth. Technology Reported sign-up and user counts; no active-user, photo-post completion, or retention rate established 25,000 sign-ups on launch day (Systrom, CBS, 2015); 7 million users within nine months (Fortune, 2014). No time-to-first-post measurement or causal effect estimate. Launch day through the first nine months Working None Interview, Podcast, Interview
BS-0006 Observational case Slack’s founder reported invitation requests and continued use among teams that had exchanged 2,000 messages; its 2019 S-1 reports aggregate message activity and net dollar retention. Enterprise software Invitation requests, selected-team continued use, aggregate messages and revenue retention; distinct measures with different denominators 8,000 preview invitation requests on day one and 15,000 after two weeks; 93% continued use among teams crossing 2,000 messages; over 1 billion messages in the week ended January 31, 2019 and 143% net dollar retention as of that date August 2013 preview requests; selected-team status at the 2015 interview; January 2019 filing metrics Working Case Founder interview, Company filing
BS-0007 Comparative test Studies of M-PESA in Kenya report risk-sharing and poverty-related estimates under their respective study designs; these do not isolate removal of an environmental bottleneck as the sole mechanism. Finance / Emerging markets Adoption %, transactions, poverty outcomes ~194,000 households (2%) lifted out of poverty in Kenya (study estimate); improved risk-sharing and consumption smoothing Multi-year (Kenya; study periods vary) High Case Study, Study
BS-0008 Observational case Spotify reports sustained, large-scale use of Discover Weekly, a personalized automated discovery feature; aggregate milestones do not estimate incremental lift. Technology / Media Cumulative Discover Weekly listening hours and tracks; weekly new-artist discoveries 2.3 billion listening hours from 2015-2020; more than 100 billion tracks streamed through the feature and 56 million new artist discoveries per week by 2025 (company-reported) 2015-2025 company milestones Working Case Company post, Company post, Company post
BS-0009 Observational case An observational study recruited 225 adult Duolingo Spanish or French beginning-course completers with U.S. IP addresses who self-reported little or no prior proficiency and no concurrent language classes, programs, or apps; 220 supplied reading scores and 220 supplied listening scores, and speaking and writing were not assessed. Education / EdTech ACTFL Reading Proficiency Test and Listening Proficiency Test scores; logged Duolingo learning time Spanish: Intermediate Low reading (n=132) and approaching Novice High listening (n=131). French: approaching Intermediate Mid reading (n=88) and Novice High listening (n=89). No concurrent comparison group or causal estimate. Mean 562 calendar days for Spanish and 634 days for French from first target-language lesson to study participation; median 112 logged learning hours overall Working Case Study
BS-0010 Observational case Zoom reported daily meeting participants rising from 10 million to 300 million from December 2019 to April 2020 as pandemic restrictions coincided with rapid expansion of remote activity; the pattern does not isolate the effect of access friction. Enterprise / Communication Daily participants, join TTFB 10M → 300M daily participants (Dec 2019-Apr 2020); TTFB near-instant Q1-Q2 2020 Working Case Company post
BS-0011 Comparative test Meta-analyses of choice architecture interventions show small-to-medium average effects with substantial heterogeneity and publication bias. Cross-domain meta-analysis Cohen’s d across published and unpublished studies d ≈ 0.43-0.45 (small-to-medium) 214 publications; 455 effect sizes High None Meta-analysis
BS-0012 Comparative test Across five field experiments at one large organization, Study 1 produced a small increase in carpool registration but no increase in active carpooling; Studies 2, 3a, 3b, and 4 found no positive effect on active carpooling, discounted transit-pass purchases, or single-occupancy-vehicle commuting. Transportation / Public policy Carpool registration and activation; discounted transit-pass purchases; self-reported single-occupancy-vehicle commute days Study 1 registration was 0.05% in control versus 0.22% with a letter (d=0.055, P<0.001), but only seven participants became active carpoolers, four in control. Effects in Studies 2, 3a, 3b, and 4 were d=0.00, -0.01, -0.01, and 0.01 and were statistically equivalent to zero under post-hoc bounds set to each study’s 80%-power minimum detectable effect. Study 1: two months; Study 2: four weeks; Study 3a: one-week trial plus one-month follow-up; Study 3b: one month; Study 4: baseline to one month after the final travel-plan session Working None RCT
BS-0013 Comparative test Home energy reports (Opower) generate small, persistent behavior changes; effects are modest and context-dependent. Energy / Residential kWh reduction vs baseline/cohort Small average reductions; persistence with decay (10-20% decay/year after reports stop) 12-24 months High None RCT
BS-0014 Observational case Robinhood’s zero-commission model removed transaction fees, and its filing reports adoption among many first-time investors; the observational record does not isolate fee removal as the cause. Finance New account openings, first-time investors, transaction frequency Large platform growth; industry-wide fee elimination followed 2013-2019 Working Case Industry
BS-0015 Comparative test Reputation and trust systems (Airbnb) enable high-stakes sharing behaviors; conversion improves when trust cues are salient. Technology / Marketplaces Listing conversion, booking behaviors Conversion deltas vary by study; generally positive 12-36 months Working Case Experiment
BS-0016 Illustration Behavioral public strategy connects micro-foundations (individuals, teams, tools) to meso-level performance; simple nudges are brittle without system design. Public policy Conceptual framework (peer-reviewed) Not applicable; conceptual synthesis Not applicable High None Article
BS-0017 Observational case Reviews find that measurement windows and cadence are associated with what longitudinal studies detect and recommend explicit temporal hypotheses. Methods / Measurement Meta-analytic conclusion Not applicable; meta-analysis of temporal aspects Not applicable High None Meta-analysis, Review
BS-0018 Observational case Self-identity explains additional variance in behavioral intentions beyond attitudes, subjective norms, and perceived behavioral control (TPB). Psychology / Theory Meta-analytic correlation (r) and incremental R² r+ = 0.47; ΔR² ≈ 6% (intention) and 9% (identity) beyond TPB components 40 independent tests; N = 11,607 High None Meta-analysis
BS-0019 Illustration Minimum-component (bottleneck) logic: the lowest-scoring determinant constrains behavior; addressing bottleneck components is necessary for reliable change. Methods / Theory Conceptual (model-based) Not applicable Not applicable High None Model
BS-0020 Comparative test Values-affirmation (identity-aligned) interventions yield durable academic benefits versus reminders alone. Education Longitudinal academic outcomes Context-dependent; durable effects reported in field studies Multi-year Working None Field
BS-0021 Comparative test Psychological targeting (personality-tailored messaging) outperforms generic messaging for persuasion and conversion. Marketing / Persuasion Conversion/lift vs. generic controls Positive treatment effects; varies by context Campaign-specific Working None Experiment
BS-0022 Observational case In CB Insights’ startup post-mortem analysis, “no market need” is a top failure reason (#2).              
Product / Strategy Post-mortem analysis 35% (ranked #2; CB Insights) Various; cross-company High None Report      
BS-0023 Illustration The Fogg Behavior Model (B=MAP) established that behavior occurs when Motivation, Ability, and Prompt converge at the same moment. Behavior Design / Theory Published conceptual model Defines Motivation, Ability, and Prompt as simultaneous conditions in the model; does not validate the later Behavior Fit Assessment Since 2009 High None Conference paper
BS-0024 Observational case Reviews of value-based decision making and action control distinguish proposed Pavlovian, habitual, and goal-directed systems with different computational and neural properties. Behavioral Neuroscience Review of computational accounts, neural evidence, and action-control research Supports a qualified distinction among Pavlovian, habitual, and goal-directed control; does not imply that three simple systems universally govern every behavior Established body of research High Case Review, Review
BS-0025 Observational case Reviews and computational models distinguish goal-directed actions as outcome-sensitive and habitual actions as relatively insensitive to reward devaluation. Behavioral Neuroscience Outcome devaluation paradigm Robust dissociation across human and animal studies Established body of research High Case Computational model, Review
BS-0026 Observational case The widely cited 66-day median came from a study of simple behaviors such as drinking water and does not establish a formation time for complex behaviors. Habit Formation Days to 95% automaticity Median 66 days (range 18-254) 84-day study High None Longitudinal study
BS-0027 Comparative test Bias-corrected analyses find an aggregated nudge effect (d = 0.270) that drops to near zero (d = 0.004) after publication-bias adjustment. Nudge Effectiveness Cohen’s d (aggregated vs. bias-corrected) d = 0.270 aggregate; d = 0.004 after publication-bias adjustment 13 articles (14 meta-analyses); 1,638 primary studies; ~30M participants High None Meta-analysis, Second-order meta-analysis
BS-0028 Comparative test Implementation intentions (if-then planning) improve goal achievement with medium-to-large effect size. Behavior Planning Cohen’s d d = 0.65 (medium-large) 94 studies, 8,000+ participants High None Meta-analysis
BS-0029 Comparative test Intentions explain less than one-third of behavior variance; the intention-behavior gap is substantial and systematic. Behavior Theory Variance explained TPB meta-analysis: ~27% of behavior variance explained; intention-change interventions shift intentions (d = 0.66) more than behavior (d = 0.36) Meta-analyses across 185+ studies (TPB) and 47 experimental tests (intention → behavior) High None Meta-analysis, Meta-analysis, Meta-analysis, Review
BS-0030 Illustration Behavior is predicted by attitudes, subjective norms, and perceived behavioral control (Theory of Planned Behavior). Behavior Theory Multi-factor prediction model R² = 0.27-0.39 for intentions; varies for behavior Foundational theory since 1991 High None Theory paper
BS-0031 Observational case A review characterizes habits as responses cued by context and reports greater behavioral stability in stable contexts, chiefly for simple repeated behaviors. Habit Formation Context-behavior associations Strong context effects for simple, repeated behaviors Review synthesis High None Review
BS-0032 Observational case Participants in an automatic-escalation savings program increased their saving rate from 3.5% to 13.6% over 40 months; this record is a before-and-after field pattern. Finance / Behavior Design Savings rate Savings rates increased from 3.5% to 13.6% over 40 months Multi-year longitudinal High Case Field study
BS-0033 Observational case The 2007 review proposed that self-control draws on a limited resource that can be depleted by prior exertion; it did not test whether environmental design sustains behavior better than willpower. Self-Regulation Self-control performance after exertion Significant depletion effects (though effect sizes debated in replications) Meta-analytic synthesis Working None Review
BS-0034 Observational case General cognitive ability (GCA) predicts job performance with corrected validity around r=0.22 in 21st-century samples. Cognitive Ability / Work Corrected validity correlation (r) Observed r = 0.16; corrected r = 0.22 153 samples; N = 40,740 (21st-century) High None Meta-analysis
BS-0035 Observational case Academic performance is jointly predicted by cognitive ability and Big Five traits; cognitive ability accounts for most of the explained variance; conscientiousness is the largest trait contributor. Education / Psychometrics Variance explained; relative importance Combined explains 27.8% variance; cognitive ability 64% of explained variance; conscientiousness 28% 267 samples; N = 413,074; 228 studies High None Meta-analysis
BS-0036 Observational case Big Five traits predict performance across criteria; conscientiousness shows the most consistent positive association; other trait effects vary by performance category. Personality / Performance Meta-analytic correlations (ρ) across performance criteria Overall: conscientiousness ρ = 0.19; extraversion ρ = 0.10; agreeableness ρ = 0.10; neuroticism ρ = −0.12; openness ρ = 0.13. Category-specific: −0.13 to 0.24 54 meta-analyses; k = 2,028; N = 554,778 High None Meta-analysis
BS-0037 Observational case Big Five measures do not fully capture HEXACO scale variance; using Big Five instead of HEXACO entails a large loss of information. Personality / Psychometrics Cross-instrument scale variance capture Large deficiency in capturing HEXACO variance (authors conclude loss comparable to discarding one Big Five scale) Multiple Big Five measures compared with HEXACO-PI-R scales High None Psychometrics
BS-0038 Observational case Psychometric analyses identify Honesty-Humility as a distinct factor, and models that separate it from classic agreeableness-related facets show better prediction of deceit-related variables. Personality / Psychometrics Factor relations; incremental prediction Improved prediction reported for deceit-without-hostility variables after realignment Lexical Big Five + Five-Factor Model facet realignment High None Psychometrics
BS-0039 Observational case A meta-analysis found high rank-order stability of personality traits in adulthood, with higher estimates at older ages and values near 0.74 in later adulthood. Personality / Psychometrics Rank-order consistency (test-retest correlation) r ≈ 0.74 (ages 50-70; ~6.7-year interval at peak) 152 longitudinal studies; 3,217 test-retest coefficients High None Meta-analysis
BS-0040 Observational case Exercise/physical-activity identity correlates with physical activity behavior across studies. Health Behavior / Physical Activity Meta-analytic correlation (r) r = 0.44 Meta-analysis across 62 datasets High None Meta-analysis
BS-0041 Observational case Within an exercise RCT, measured increases in exercise identity were associated with physical activity at follow-up beyond treatment condition; conditions did not significantly differ in identity change. Health Behavior / Maintenance Minutes/week of moderate-to-vigorous physical activity (MVPA) at follow-up b = 16.76 minutes/week per +1 identity point (95% CI: 2.36-31.15) 16-week RCT + follow-up; N = 130 Medium None RCT
BS-0042 Observational case Habit and identity measures are strongly correlated in health behaviors; this association is consistent with, but does not establish, a mutual-reinforcement hypothesis. Health Behavior / Habit Three-level meta-analytic correlation (r) r = 0.55 (95% CI: 0.49-0.60) 19 articles; N = 13,340 High None Meta-analysis
BS-0045 Observational case Conscientiousness is consistently associated with healthier behavior patterns across health behavior domains. Personality / Health Behavior Meta-analytic correlations (r) across behaviors r ≈ 0.05 (physical activity) to r ≈ −0.28 (drug use), directionally consistent across domains 194 studies High None Meta-analysis
BS-0046 Comparative test Personality traits can change via intervention, but average effects are modest and typically require sustained effort. Personality / Change Meta-analytic standardized mean difference (d) in trait change Average trait change d ≈ 0.37 across interventions 207 intervention studies; 20,653 participants High None Meta-analysis
BS-0047 Independent replication or evaluation Personality-life outcome associations replicate at high rates in a preregistered, high-powered replication project (with expected direction and moderate effect-size shrinkage). Personality / Replication Replication rate and relative replication effect size 87% of replications significant in expected direction; replication effect sizes averaged 77% of original effects 78 replications across 4 samples; total N = 6,106 High None Preregistered replication
BS-0048 Observational case Traits can be modeled as density distributions of states: people show substantial within-person variability, yet their average trait levels (distribution means) are highly stable/reliable within experience-sampling windows. Personality / Psychometrics Split-half stability of distribution means (“location stability”) Location stability (Study 1, 2-week ESM): Extraversion 0.90, Agreeableness 0.94, Conscientiousness 0.87, Emotional Stability 0.90, Intellect 0.94 Experience sampling over ~2 weeks; Study 1 N = 127 High None Experience sampling
BS-0049 Observational case Conscientiousness is associated with lower mortality risk across longitudinal cohorts. Personality / Health Meta-analytic correlation (r) with mortality r = −0.11 (95% CI: −0.13 to −0.09) Meta-analysis of 20 independent samples High None Meta-analysis
BS-0050 Observational case General mental ability is a strong predictor of performance in job training programs. Cognitive Ability / Training Meta-analytic validity correlation (r) r ≈ 0.56 for job training performance Meta-analytic synthesis (Schmidt & Hunter, 1998) High None Meta-analysis
BS-0051 Independent replication or evaluation A preregistered, 39-lab replication did not support the classic induced-compliance (high-choice > low-choice) attitude-change effect. Social Psychology / Replication Attitude-change contrast (high choice vs low choice) Primary high-choice vs low-choice prediction not supported; counterattitudinal vs neutral contrast supported 39 labs; 19 countries; N = 4,898 High None Multilab preregistered replication
BS-0052 Observational case A modern archival analysis concludes that key elements of the classic ‘When Prophecy Fails’ account were substantially inaccurate and involved serious ethical problems. History / Research Integrity Archival/historical reassessment Not applicable Archival analysis published 2026 Medium None Historical analysis
BS-0053 Observational case A widely cited behavioral-ethics intervention paper on signature placement and honesty was retracted. Behavioral Ethics / Retraction Retraction notice Not applicable Retraction published 2021 (original 2012) High None Retraction
BS-0054 Comparative test Automatic enrollment dramatically increases 401(k) participation among new hires and anchors many participants to default contribution rates and asset allocations. Finance / Defaults 401(k) participation rate; default contribution and investment choices Participation at 3-15 months tenure: 86% (automatic enrollment) vs 37% (opt-in). Default contribution adherence declines from 92% to 57% over 15 months. 3-15 months of tenure; up to 15 months post-hire High Case Field study, Working paper
BS-0055 Comparative test Automatic enrollment can raise participation, but long-run retirement wealth effects can be modest after accounting for turnover and withdrawals. Finance / Defaults Long-run retirement savings / wealth accumulation +0.6% of income in retirement savings; effects substantially smaller once accounting for job changes and cash-outs Long-run model with turnover and withdrawals Medium Case Working paper
BS-0056 Comparative test Experimental attempts to induce habits in humans often fail; extensive training may not produce outcome-insensitive (habitual) responding. Habit Formation / Neuroscience Outcome devaluation sensitivity after training Five experiments reported failures to induce habitual responding in humans Five experiments High None Journal article
BS-0057 Observational case Acorns attaches automatic round-up contributions to card spending, and company and press sources report about $43 in average monthly contributions; they do not isolate an incremental effect. FinTech / Personal finance Average round-ups contribution per user; early contribution volume Average round-ups: ~$43 per customer per month (reported); >$150 invested in first 4 months from round-ups alone (company-reported) Monthly; first 4 months post-enablement Working Case Company documentation, Company post, Press
BS-0058 Observational case YNAB centers active allocation before spending, while Mint centered passive tracking; the contrast is a hypothesis about durable budgeting behavior, not a comparative test. FinTech / Personal finance Product viability and sustained usage (qualitative); budgeting behavior cadence Mint: 1.5M users at 2009 acquisition (reported) and later shutdown (2024). YNAB: sustained subscription model anchored in repeated allocation behavior. 2007-2024 Working Case Company documentation, Press, Press
BS-0059 Observational case ClassPass offered access to varied fitness venues rather than a single-gym commitment; its reported trajectory is consistent with, but does not test, a variety-seeking fit hypothesis. Fitness / Marketplace Subscriber retention proxies; repeat class attendance Reported partner-side pattern: high share of visits are to new studios early; usage concentrates after initial exploration (company/partner reporting varies). 2012-2018 (model evolution) Working Case Press, Case write-up, Press
BS-0060 Observational case Gym members in transaction data often chose flat-fee contracts that cost more than pay-per-visit alternatives while attending infrequently and delaying cancellation; the authors interpret this pattern as overconfidence. Behavioral Economics / Fitness Attendance vs contract choice; cancellation delay Members overestimated visits (~9.5 forecast vs ~4.2 actual/month); flat-fee members attended ~4.3x/month and paid an effective price >$17/visit (field data). 3 years; N=7,752 members (3 US health clubs) High Case Journal article
BS-0061 Observational case Peloton’s reported growth coincided with an at-home, instructor-led, socially reinforced model; adoption later shifted as pandemic conditions changed. Fitness / Consumer subscription Subscriber growth; engagement and churn (reported) Hardware-subscriber engagement spiked during COVID and normalized afterward (reported). 2018-2023 Working Case Press, Press
BS-0062 Observational case In a longitudinal social-network study, receiving kudos was associated with modestly higher subsequent running frequency; Strava also uses social feedback and competition. Fitness / Social platforms Running frequency and peer influence Receiving kudos is associated with increased running frequency (study of running clubs; magnitude context-specific); Strava reports 180M+ athletes in 185+ countries. Study-specific (running clubs; longitudinal social network analysis) Working Case Journal article, Company
BS-0063 Observational case A single-site observational study enrolled 110 participants in a modified, instructor-supported nine-week Couch-to-5K program. The paper’s abstract reports 27.3% completion, but its Results section reports 64.5% dropout and Figure 1 shows 39 completers (35.5%), so the completion rate is unresolved. Fitness / Program design Completion and dropout in a modified Couch-to-5K program; separate NHS app downloads and completed runs Study-source conflict: 27.3% completion in the abstract versus 39 of 110 completers (35.5%) in Figure 1, consistent with 64.5% dropout in Results. Separately, NHS reported more than 7 million app downloads since 2016 and 8.7 million runs completed in 2024. Modified nine-week program studied from May 2018 to May 2020; separate NHS app reporting through 2024 Working Case Study, Government
BS-0064 Observational case Discord formalized team voice and text coordination for gaming groups and later reported adoption across broader communities; this sequence is consistent with an existing-behavior fit hypothesis. Technology / Communication Adoption and community usage (reported) 200M+ global monthly active users (company-reported, 2025); expansion beyond gaming (reported). 2015-2025 Working Case Company, Reference
BS-0065 Observational case Figma offered real-time multiplayer editing for collaborative design workflows and later achieved broad adoption; source accounts describe less handoff work and same-file co-creation without a causal comparison. Technology / Work tools Workflow friction (qualitative); adoption (reported) Real-time multiplayer editing reduced coordination overhead by eliminating export/sync/email file behaviors (mechanism; quantitative varies). 2016-2023 Working Case Company post, Press
BS-0066 Observational case TikTok says its For You feed ranks videos using user interactions and video information; follower count and prior high-performing videos are not direct ranking factors. Technology / Social media Creator output and consumption behavior (qualitative); adoption (reported) Mechanism-level: recommendation feed makes reach less dependent on follower graphs; video-length flexibility expands viable content behaviors. 2016-2024 Working Case Company post
BS-0067 Observational case By its 2006 acquisition, YouTube allowed general video upload and sharing and reported more than 100 million daily views and 65,000 daily uploads; it added a revenue-sharing Partner Program in 2007. Technology / Creator platforms Upload and viewing volume at acquisition; creator incentives At acquisition (reported): ~100M daily video views; ~65K new videos uploaded daily. 2005-2007 High Case Filing, Press
BS-0068 Observational case Netflix paired subscriptions with no late fees and reported rapid subscriber growth, while Blockbuster relied on store returns and late-fee revenue; the sources do not isolate these model differences as the cause of the outcome. Media / Subscription Business model constraint alignment (qualitative) Netflix reported 857,000 subscribers at year-end 2002, 2.61 million at year-end 2004, and about 20 million at year-end 2010. Blockbuster projected that eliminating late fees would reduce 2005 operating income by $250 million to $300 million. 1999-2010 Working Case Press, Press, Company filing, Company filing, Company filing, Company filing
BS-0069 Observational case Google ended standalone Wave development in 2010; co-creator Lars Rasmussen described the project as ambitious and questioned its rapid, resource-heavy execution. Technology / Collaboration tools Adoption (reported) and usability barriers (qualitative) Creators reported misalignment between their needs and mass-market learning costs; early active usage was low relative to invites (reported). 2009-2012 Working Case Interview, Reference
BS-0070 Observational case Quibi launched as commute-oriented short-form video during the 2020 collapse in commuting, reported low trial conversion, and shut down about six months later; the sources do not isolate a single cause. Media / Failures Trial-to-paid conversion; subscriber counts (reported) Trial-to-paid conversion reported ~8% for early cohorts (Sensor Tower estimate); shutdown announced ~6 months post-launch. Apr-Oct 2020 Working Case Press, Press, Press
BS-0071 Observational case Meditation-app studies report substantial early attrition and low retention; effort, delayed rewards, fit, and scaffolding are hypotheses for this pattern rather than established causes. Health / Digital therapeutics Attrition and retention rates Meta-analytic attrition ~25% (higher in larger studies); median 30-day retention for mindfulness apps reported ~4.7% (app analytics). 30 days and longer; study-dependent Working Case Meta-analysis, Study
BS-0072 Observational case Digital health studies report substantial early dropout and abandonment; a scoping review identified reported reasons including technical and functional issues, privacy concerns, poor user experience, content and features, time and financial costs, and changing user needs. Health / Digital products Dropout and abandonment rates Health app dropout reported ~43%; median abandonment over time can exceed 70% (study-dependent). First weeks to ~100 days (study-dependent) Working Case Study
BS-0073 Comparative test Simplifying HIV treatment behaviors (single-tablet regimens; long-acting injectables) improves adherence and clinical outcomes compared to higher-burden regimens. Healthcare / Adherence Adherence and viral suppression; discontinuation Meta-analysis: discontinuation 36.3% (single-tablet) vs 48.8% (multi-tablet); ~21% higher odds of undetectable viral load (reported). Meta-analysis (study periods vary) High Case Meta-analysis, Meta-analysis, Report
BS-0074 Observational case Blue Apron cohorts showed steep reported retention declines over six months; fit with weekly planning, cooking, and delivery management is a hypothesis for meal-kit churn, not an isolated cause. Consumer subscription / Food Customer retention over time Blue Apron reported ~50% continuation after two weeks and ~10% after six months (press reporting; varies by cohort). First 6 months Working Case Press, Press
BS-0075 Observational case Waze paired low-friction micro-contributions and rewards with large reported growth in crowdsourced traffic data; available case sources do not isolate effects on adoption or navigation outcomes. Technology / Navigation User growth and contribution volume (reported) At acquisition: ~50M users (reported); incident reporting and map edits at large scale (reported). 2009-2013 and beyond Working Case Press, Case write-up, Press
BS-0076 Observational case In two company moves to open-plan offices, face-to-face interaction fell sharply while electronic messaging rose; the before-and-after design does not isolate noise or interruption as the cause. Workplace design Face-to-face interaction time; email/IM volume Face-to-face interaction decreased ~70%; email increased ~56%; IM increased ~67% after open office adoption (field study). Before/after move to open plan (Fortune 500 HQ) High Case Journal article, Review
BS-0077 Comparative test Responsible beverage service (RBS) training reduces over-service and alcohol-related harms by targeting the gatekeeper behavior that can change in the moment. Public health / Alcohol harms Crash and injury outcomes; intoxication measures Oregon server-training mandate associated with a 23% reduction in fatal single-vehicle nighttime crashes; community reviews report reduced intoxicated driving and related harms. 1986-1994 (Oregon); multi-study reviews High Case Systematic review, Study
BS-0078 Observational case In a company-authored account, Proposify reported low completion of its old onboarding and, one month after a guided tutorial launched, a slight increase in trial-to-paid conversion plus associations between tutorial completion, proposal sending, and payment; no denominators, base rates, or causal comparison were reported. SaaS / Sales enablement Onboarding completion, proposal sending, and trial-to-paid conversion (company-reported) Not estimable from the published account; sample counts, base and post rates, uncertainty, and a defined calculation were not reported. Old flow had been in place for over one year, but the analyzed baseline window was not reported; post-launch observation was about one month Working Case Company-reported, Company post
BS-0079 Observational case M-PESA’s public pilot accounts support candidate-use and operating-support hypotheses; they do not establish a representative paid first-use or repeat-opportunity rate. Financial services / Public-source teaching Participant-reported pilot observations and development history No complete opportunity-to-completion rate is recoverable from the reviewed accounts. Pilot began 11 October 2005 and officially ended 1 May 2006; source publication dates differ from event dates Working Case Participant-authored case, Company-published research report, Company retrospective
BS-0080 Observational case Instagram’s early public history supports a qualified reconstruction of product focus and problem selection; it supplies no isolated causal estimate for the pivot. Consumer technology / Product strategy Qualitative product history and founder-reported user count Nearly 4 million users reported in May 2011; no established active-user, completion or retention rate 2010 product decision; founder talk on 11 May 2011; later engineering account in October 2015 Working Case Founder talk, Transcript of the same talk, Contemporaneous product observation, Founder retrospective
BS-0081 Observational case The NAO’s 2019 Verify audit identifies gaps between reported identity-registration success and the user’s complete government-service journey. Public digital services / Identity assurance Reported Verify registration success and the boundaries of its denominator and endpoint 48% reported single-attempt sign-up success in February 2019 versus an earlier 90% projection; not an end-to-end service-completion rate or a treatment effect Performance through February 2019; NAO report published 5 March 2019 Working None Independent government audit, Primary audit report, Summary of the same audit
BS-0082 Comparative test Wikimedia’s Growth experiment found improved first-edit activation; the reported increase in retained editors was inferred through activation rather than directly detected. New accounts on Arabic, Vietnamese, Czech, and Korean Wikipedias First Article or Article Talk edit within 24 hours; constructive first edit unreverted within 48 hours; another edit on a different day in the following two weeks after activation Activation improved; no retention change was directly detected. No feature-specific effect is isolated here. Accounts registered November 21, 2019 to May 14, 2020; November 2020 analysis Working None Original experiment report
BS-0083 Comparative test Six California outreach experiments found no detected increase in tax filing or EITC claiming from the tested letters and text messages. California households sampled for outreach; six partly overlapping trials Administrative household-level filing and credit-claim outcomes; website engagement is an intermediate measure No detected increase in the primary filing or claiming outcomes; no pooled numerical effect asserted here Spring 2018 and spring 2019, addressing tax years 2017 and 2018 Working None Peer-reviewed study, Author working paper with methods and appendix

How to add a row #

  1. Add an entry to /_data/evidence_ledger.yml.
  2. Assign one evidence class from the Editorial Policy.
  3. Add the canonical case path to case_bindings when a public case cites the record.
  4. Use {% include evidence-ref.html id="BS-XXXX" %} next to the claim on the relevant page.
  5. Keep denominators explicit and attach a named calculation or source receipt for each quantitative case.

The YAML ledger is the only public source of evidence rows. Do not copy ledger rows into a second table or page.