What receives treatment?
Person, school, branch, machine?
The unit of assignment determines independence and interference assumptions.
Side 77 · Special
A study of creating evidence before observing the answer. Experimental design decides how interventions, comparison groups, randomization, measurement and replication will make causal interpretation possible.
Before choosing sample size or randomization, specify exactly which intervention contrast and outcome matter.
Person, school, branch, machine?
The unit of assignment determines independence and interference assumptions.
Define treatment precisely.
Vague treatment definitions produce ambiguous causal conclusions.
Control condition.
The control defines the counterfactual contrast.
Timing + metric.
Primary outcomes should be chosen before observing results.
Average, subgroup, assignment effect?
Design should match the quantity the study intends to estimate.
Randomization breaks systematic links between treatment assignment and baseline characteristics in expectation.
Easy to implement, but small samples can still be imbalanced by chance.
Blocks protect precision when known variables strongly predict outcomes.
Useful when treatment is delivered collectively or spillover would contaminate individual assignment.
Can improve precision when matching variables are strongly prognostic.
Strong within-unit comparison when carryover and time effects are controlled.
Blinding reduces behavior and measurement changes caused by knowledge of condition.
They estimate main effects and interactions by combining factor levels systematically.
Each factor contains two or more levels.
Collapses across levels of the other factors.
Interactions reveal combinations that simple one-factor studies miss.
Rich information grows expensive as factors multiply.
Efficiency comes by accepting aliasing assumptions about which effects are negligible.
Outcome definitions, instruments and timing determine whether treatment effects can be observed cleanly.
| Design choice | Question | Failure mode |
|---|---|---|
| Primary outcome | Which result matters most? | Outcome switching / multiplicity |
| Instrument | Is the measure reliable and valid? | Noise or construct mismatch |
| Timing | When should effect appear? | Too early or too late measurement |
| Baseline | Should pre-treatment state be measured? | Lost precision or imbalance visibility |
| Protocol | Is measurement identical across groups? | Differential measurement bias |
Detectability depends on effect size, noise, sample size, allocation and significance threshold.
Design around substantively meaningful effects, not any nonzero difference.
More variation requires more information to distinguish treatment from noise.
More units usually increase precision, with diminishing returns.
More stringent thresholds require stronger evidence.
More tests increase the chance of at least one false positive without correction.
Within-cluster similarity reduces effective information.
Strong designs anticipate contamination, attrition, noncompliance and generalization limits.
Did the assigned intervention cause the measured difference?
Did dropout differ by condition in a way that breaks comparability?
Did assigned units actually receive or follow the intervention?
Did one unit’s treatment affect another unit’s outcome?
Would the result travel to other populations, settings or implementations?