Measurement Lab Answers #
Purpose. Check the calculations and compare the reasoning after completing the lab. The arithmetic has specified answers; several next decisions can be defensible if their assumptions and evidence limits are explicit.
Return to the student exercise if you have not recorded a response. The public-source sections are instructor analysis. All handover numbers below are invented teaching data, not research findings.
A. Wikipedia: separate first, quality, and return measures #
The activation analysis counts a first edit to an Article or Article Talk page within 24 hours of registration; constructive activation adds no reversion within 48 hours. Retention involves another edit on a different day in the following two weeks after activation. BS-0082
A return rate among activated accounts asks a different question from retained editors per original account. More people can reach the first stage without the conditional return rate improving. Name the population and conditioning rule beside either rate.
The report found improved activation but did not directly detect a retention change. It inferred an increase in retained users through activation. The randomized contrast covers the feature bundle, and exclusion of people who toggled features limits a pure intention-to-treat interpretation. It cannot isolate suggested tasks or establish long-term retention. See the original report.
Defensible next test. Compare a specified task or return-support change while holding the other features constant. Define the intended newcomers, available editing occasions, repeat window, and quality measure before assignment. Retain the original assigned groups for the primary comparison and record actual feature use separately. Longer follow-up and adequate precision would be needed for a longer-term claim. This is a proposed design, not Wikimedia’s historical protocol.
B. California EITC: a treatment result is not a full diagnosis #
The research reports no detected increase in filing or claiming from the tested outreach, despite engagement with linked resources. That narrows the claim to these interventions and settings; it does not establish that no workable route to claiming exists. BS-0083 See the journal record and identified working paper.
Proposed chain: message delivery, attention, eligibility check, access to information and documents, help if needed, preparation, submission, and recorded claim. The chain is an instructor reconstruction. A website visit measures one intermediate step. An outcome comparison needs the assigned populations and administrative results; restricting it to visitors would select a subgroup that chose to visit.
Possible remaining explanations include difficulty preparing the return, inability to reach usable help, and incomplete delivery to some recipients. A null overall result does not identify which explanation dominates. An offer of assistance should not be coded as received assistance.
Defensible decision. Stop extending the tested outreach on an unsupported expectation of increased claims. If further investment is justified, investigate completion barriers and specify a feasible support change before testing it. “Change support” can describe that next service decision; it is not a claim that more intensive help has already been shown effective by this study. A reasoned stop on further investment is also defensible under a stated resource limit.
C. Fictional data: expected counts #
The CSV contains 24 invented team-occasion records from 12 teams. Each team appears twice. Download the machine-readable answer values for checking calculations.
| Count | Expected value |
|---|---|
| Eligible roster records | 24 |
| Unique fictional teams | 12 |
| Known applicable opportunities | 19 |
| Known occasions without an opportunity | 3 |
| Unknown opportunities | 2 |
| Attempts on known opportunities | 16 |
| Known completions on known opportunities | 11 |
| Known noncompletions on known opportunities | 7 |
| Unknown completion on a known opportunity | 1 |
The two unknown opportunities belong to T12. Their attempt and completion states are also unknown. Keep them visible outside the opportunity-specific calculations. T05’s second record has a known opportunity and attempt but an unknown completion.
Completion calculations #
Percentages are rounded to two decimal places. All figures in this section come from the invented file.
| Question | Calculation | Result |
|---|---|---|
| Attempts per known opportunity | 16 / 19 | 84.21% |
| Completion among known outcomes on known opportunities | 11 / 18 | 61.11% |
| Completion on known opportunities, unresolved completion fails | 11 / 19 | 57.89% |
| Completion on known opportunities, unresolved completion succeeds | 12 / 19 | 63.16% |
| Completion among attempts with known outcomes | 11 / 15 | 73.33% |
| Conditional completion, unresolved attempt fails | 11 / 16 | 68.75% |
| Conditional completion, unresolved attempt succeeds | 12 / 16 | 75.00% |
| Observed completions per scheduled roster record | 11 / 24 | 45.83% |
The 73.33% figure omits opportunities with no attempt and the unresolved attempted outcome. It answers a conditional question. It is not the proportion of opportunities successfully completed.
The 45.83% figure uses a legitimate but different denominator: every scheduled roster record. It includes known occasions without a task to do and unknown opportunities. Do not label it completion per applicable opportunity.
The bounds account only for the one unresolved completion among known opportunities. They do not resolve T12’s unknown opportunities, sampling uncertainty, repeated-team dependence, or selection. No inferential confidence interval or causal estimate is supplied by this teaching file.
Repeat completion #
First completers are T01 through T06. Five have a known second opportunity; T06 has no unfinished work on its second occasion and is excluded from this particular denominator. Of those five, T01, T02, and T03 complete again; T04 does not; T05 is unresolved.
| Repeat measure | Calculation | Result |
|---|---|---|
| Repeat among known second outcomes | 3 / 4 | 75.00% |
| Lower bound among known second opportunities | 3 / 5 | 60.00% |
| Upper bound among known second opportunities | 4 / 5 | 80.00% |
Counting T06 as a failed repeat would confuse absence of a task with failure to do it. Counting every team as a first completer would answer a different question. These are two occasions, so no conclusion about long-term repetition follows.
Support and recorded effort #
| Support condition | Known opportunities | Completions | Unresolved completion | Recorded staff minutes |
|---|---|---|---|---|
| Normal | 14 | 7 | 0 | 25 |
| Exceptional | 5 | 4 | 1 | 63 |
Normal-support completion is 7/14, or 50.00%. Exceptional-support completion is bounded by 4/5 and 5/5, or 80.00% to 100.00%. Excluding the unresolved exceptional record produces 4/4, which must be identified as a known-outcome subset.
Support was not randomly assigned. The table cannot show that coaching caused the difference. It also cannot show that exceptional help can be supplied at normal operating capacity.
Recorded staff effort totals 88 minutes, with two unknown staff-time records. Include effort on failures and on T05’s unresolved attempt. The recorded total already exceeds the illustrative 60-minute budget by 28 minutes. Unknown minutes can increase that total. The ratio of recorded staff effort to known completions is 88/11, or 8 minutes per known completion. It is not the average coaching duration of successful teams or the service’s full cost.
Worked next decision #
Change support, with a limited diagnostic step before another trial. The normal-support rate is below the invented 70% operating requirement, and recorded help already exceeds the budget. The file does not support expansion, a claim of BMF, or a causal conclusion about coaching.
Inspect the handoff at acknowledgement and the unresolved task-owner step. Compare the practical requirements of a scheduled handover overlap, a clearer ownership rule, and the current written route. Repair the missing records where possible. Estimate normal staffing and queueing for the proposed change, then agree a new test and decision rule before seeing its results.
This recommendation is instructor analysis. A response that stops the route because no affordable repair fits the stated constraints can also be defensible. Reselection needs evidence favoring an alternative action. “Inconclusive” is appropriate for the precise repeat rate and coaching effect; those uncertainties do not erase the observed normal-support and recorded-budget failures.
Original answer guide and fictional calculations: Jason Hreha, version 1, 2026-09-05, CC BY 4.0. Linked sources retain their own rights. No real participant study is reported.