Measurement Lab #
Purpose. Practice defining what a result measures, which people or opportunities it includes, and what next decision the evidence supports.
This lab has two public-source exercises and one invented dataset. The public studies did not test DRIVE, BMF, or BFA. The exercises are original teaching analysis. The dataset contains no actual observations, and completing the lab is not a participant study.
Work in workbook sections 6 to 8. Record your answers before opening the separate answer guide. Suggested time is 25 minutes for all three exercises, an instructional estimate. A host can select one exercise for a shorter session.
A. Wikipedia: three different results #
Public source. Wikimedia’s November 2020 report analyzes a randomized comparison of all Growth features with the default experience on four Wikipedias. It includes homepage, suggested tasks, and help features. Accounts that switched the features on or off were excluded from the analysis. The observed period was November 2019 to May 2020. BS-0082
Read the original report’s glossary, findings, and methodology. This source was inspected on September 5, 2026; its visible revision is January 9, 2025. The later page revision is not the experiment date.
Source facts for offline practice. Activation means a first edit to an Article or Article Talk page within 24 hours of registration. Constructive activation adds that the edit is not reverted within 48 hours. Retention involves another edit on a different day in the two weeks after activation. The report detected an activation increase but did not directly detect a retention change; it inferred more retained users through increased activation. These are source definitions and findings, not an answer to the next-test task.
Your task:
- Write separate measures for a first edit, a constructive first edit, and a return edit. Include their windows and the population each rate describes.
- Explain how retained editors per original account differ from returning editors among people who made a first edit.
- State what the comparison can tell you about the full feature bundle. Identify what would require another comparison to isolate suggested tasks.
- Write a next test if the operating goal is useful repeat contributions. Define the return opportunity, quality measure, follow-up, and result that would change your decision.
Deliverable: one measurement table and a bounded next-test proposal. Do not treat an edit that remains unreverted as a complete assessment of its quality, or infer a long-term result from a short observation window.
B. California EITC: contact and completion #
Public source. Linos and colleagues report six randomized outreach experiments conducted in California in 2018 and 2019. Their study used letters and text messages, with filing and credit-claim outcomes linked to administrative records. The samples and control conditions must be read by experiment; the studies partly overlap. BS-0083
Read the November 2022 journal abstract. For methods and engagement details, use the research team’s November 2020 working paper, especially its introduction and methods. Both were inspected on September 5, 2026. The working paper is identified separately from the final publication.
Source facts for offline practice. The tested outreach did not produce a detected increase in filing or claiming. Some recipients visited linked resources. The randomized outcome comparisons used administrative records; visiting a website was a separate engagement measure. The supplied source summary does not establish that everyone received or read a message, obtained usable assistance, or completed each preparation step. No numerical effect calculation is required for this exercise.
Your task:
- Draw a proposed chain from receipt of outreach to the completed claim. Mark which steps the available evidence measures and which remain unknown.
- Explain what visiting a linked website can establish. Name the additional evidence needed to claim that filing or claiming changed because of the outreach.
- Assess this proposed conclusion: “The treatment did not increase claiming, so filing is the wrong behavior to target.” List two other explanations that remain possible and an observation that could help distinguish them.
- Select one next decision from the workbook. Identify the particular treatment or service route affected, and state what remains unresolved about the wider objective.
Deliverable: one behavior chain, one corrected conclusion, and one decision with its evidence limits. An address or offer for help is not evidence that hands-on help was received.
C. Fictional shift-handover data #
Download measurement-records.csv. It contains 24 invented records from 12 fictional teams, with a first and second scheduled handover occasion for each team. The observation unit is a team on an occasion, not an independent person. The records cover a fictional two-round exercise; they are unrelated to Wikipedia or the EITC studies.
Scenario and recording rules #
An outgoing lead saves a form listing the next action and owner for unfinished work. The incoming lead must acknowledge a usable handover before taking over. The useful outcome is that the next shift can act on correct ownership information; the file does not measure that later outcome.
- Eligible: the team belongs to the exercise roster. All rostered team-occasion records are retained.
- Opportunity: unfinished work requires a handover on that occasion. No unfinished work means no applicable opportunity.
- Attempted: the outgoing lead saves the form before the deadline.
- Completed: a handover with the required next-action and named-owner fields is acknowledged by the incoming lead before taking over. These fields define “usable” for this exercise. A saved form alone does not complete it; the later accuracy and use of that information are not measured.
- Normal support: the planned form and ordinary supervisor help.
- Exceptional support: extra coach time outside the planned service. It was not assigned randomly.
- Staff minutes: recorded help for that occasion, including help on failed or unresolved attempts. Zero means a recorded zero; unknown means the help record is missing. It excludes participant time and other operating costs.
Data dictionary #
| Column | Meaning and values |
|---|---|
| dataset_id | Identifies this entirely fictional exercise |
| record_id | Unique observation identifier |
| team_id | Fictional team; appears twice |
| occasion | First or second scheduled handover |
| eligible | Yes: on the exercise roster |
| opportunity | Yes, no, or unknown |
| attempted | Yes, no, unknown, or na for no opportunity |
| completed | Yes, no, unknown, or na for no opportunity |
| support | Normal or exceptional |
| staff_minutes | Nonnegative recorded minutes or unknown |
| observation_note | Invented explanation of the recorded state |
Unknown records remain in the roster. Do not fill them with zeros or delete them without reporting the effect. For this exercise, a verified no-attempt on a known opportunity also means no completion. A missing completion record after a known attempt remains unknown.
Calculate and interpret #
- Count records, unique teams, known opportunities, known absences of opportunity, and unknown opportunities. Keep the units distinct.
- Calculate attempts per known opportunity. Then calculate completion among known outcomes on known opportunities. Report unresolved outcomes beside the rate.
- Bound completion on known opportunities by treating the unresolved completion as a failure and then as a success. Explain why these bounds do not include every source of uncertainty.
- Calculate completion among attempts with known outcomes. Compare its denominator with the opportunity denominator. Explain the question each answers.
- Identify the first-completing teams. Among those with a known second opportunity, calculate repeat completion among known outcomes and its bounds. Name anyone excluded for lacking another opportunity.
- Compare normal and exceptional support descriptively. Sum recorded staff minutes, including time spent on failures and unresolved attempts. Report how many staff-time records remain unknown.
- Evaluate the instructional rules below, then record one of the five next decisions. Include the next observation that would help distinguish a delivery problem from a problem with the selected action.
Illustrative operating rules #
These rules are invented for the exercise, agreed before interpreting its records. They are not universal BMF thresholds or a recommended research sample size.
The fictional team would consider a limited next stage only if at least 70% of applicable handovers complete under normal support, because it assumes its fallback can handle at most the remaining 30%. It also has a total 60-minute staff-help budget for both rounds combined. Missing outcomes, uncertain opportunities, and the precision of a rate must be addressed before expansion. The two-round exercise cannot establish long-term repetition or a population-wide rate.
Apply the rules to the stated support condition, not whichever subgroup has the highest result. A rate above a threshold does not cancel a failed cost or data-quality requirement. Do not claim a causal advantage for exceptional help from these records.
Deliverable: your counts, clearly named fractions, a support-cost note, and a short decision memo. Then open the answer guide to check the calculations and compare your reasoning.
Original exercises and fictional data: Jason Hreha, version 1, 2026-09-05, CC BY 4.0. Linked sources retain their own rights. No real participant data is used.