Skip to content

Side 77 · Special

Experimental
Design

A study of creating evidence before observing the answer. Experimental design decides how interventions, comparison groups, randomization, measurement and replication will make causal interpretation possible.

question→intervention→comparison→measurement→inference
06design lenses
05allocation tools
05validity threats
77Side

The design begins with the estimand.

Before choosing sample size or randomization, specify exactly which intervention contrast and outcome matter.

01 · Unit

What receives treatment?

Person, school, branch, machine?

The unit of assignment determines independence and interference assumptions.

02 · Intervention

What is manipulated?

Define treatment precisely.

Vague treatment definitions produce ambiguous causal conclusions.

03 · Comparison

Compared with what?

Control condition.

The control defines the counterfactual contrast.

04 · Outcome

What response is measured?

Timing + metric.

Primary outcomes should be chosen before observing results.

05 · Estimand

Which causal quantity?

Average, subgroup, assignment effect?

Design should match the quantity the study intends to estimate.

Allocation manufactures comparability.

Randomization breaks systematic links between treatment assignment and baseline characteristics in expectation.

Simple randomization

Assign independently.

Easy to implement, but small samples can still be imbalanced by chance.

Blocking

Randomize within similar groups.

Blocks protect precision when known variables strongly predict outcomes.

Cluster randomization

Assign groups, not individuals.

Useful when treatment is delivered collectively or spillover would contaminate individual assignment.

Matched pairs

Pair similar units before assignment.

Can improve precision when matching variables are strongly prognostic.

Crossover

Each unit receives multiple conditions.

Strong within-unit comparison when carryover and time effects are controlled.

Blinding

Hide assignment where possible.

Blinding reduces behavior and measurement changes caused by knowledge of condition.

Factorial designs study several interventions at once.

They estimate main effects and interactions by combining factor levels systematically.

Factor

A manipulated variable.

Each factor contains two or more levels.

Main effect

Average effect of one factor.

Collapses across levels of the other factors.

Interaction

Effect of one factor depends on another.

Interactions reveal combinations that simple one-factor studies miss.

Full factorial

Observe every combination.

Rich information grows expensive as factors multiply.

Fractional factorial

Study selected combinations.

Efficiency comes by accepting aliasing assumptions about which effects are negligible.

Good randomization cannot rescue bad measurement.

Outcome definitions, instruments and timing determine whether treatment effects can be observed cleanly.

Design choiceQuestionFailure mode
Primary outcomeWhich result matters most?Outcome switching / multiplicity
InstrumentIs the measure reliable and valid?Noise or construct mismatch
TimingWhen should effect appear?Too early or too late measurement
BaselineShould pre-treatment state be measured?Lost precision or imbalance visibility
ProtocolIs measurement identical across groups?Differential measurement bias

Power is a design property, not a post-hoc excuse.

Detectability depends on effect size, noise, sample size, allocation and significance threshold.

Effect size

How large a difference matters?

Design around substantively meaningful effects, not any nonzero difference.

Variance

How noisy is the outcome?

More variation requires more information to distinguish treatment from noise.

Sample size

How many independent units?

More units usually increase precision, with diminishing returns.

Alpha

How strict is the false-positive threshold?

More stringent thresholds require stronger evidence.

Multiplicity

How many hypotheses are tested?

More tests increase the chance of at least one false positive without correction.

Clustering

Are observations truly independent?

Within-cluster similarity reduces effective information.

Experimental validity can fail before or after assignment.

Strong designs anticipate contamination, attrition, noncompliance and generalization limits.

Internal validity

Did the assigned intervention cause the measured difference?

Attrition

Did dropout differ by condition in a way that breaks comparability?

Compliance

Did assigned units actually receive or follow the intervention?

Interference

Did one unit’s treatment affect another unit’s outcome?

External validity

Would the result travel to other populations, settings or implementations?

Design of ExperimentsFisher · foundations
Experimental and Quasi-Experimental DesignsShadish, Cook & Campbell · validity
Design and Analysis of ExperimentsMontgomery · factorial methods
Field Experimentsrandomization in applied settings