Skip to content

Side 100

Reproducibility &
Replication

A study of whether scientific results survive being reconstructed, repeated and challenged. Reproducibility preserves the path from evidence to result; replication tests whether the phenomenon itself returns under an independent attempt.

claim→record→reproduce→replicate→revise
06reliability lenses
05failure modes
05correction moves
100Side

Reproducibility and replication test different links in the scientific chain.

A result can be perfectly reproducible from the original materials yet fail to replicate independently—and the reverse distinction matters for diagnosis.

01 · Claim

What result is being asserted?

Estimate, effect, model, phenomenon?

The target must be precise enough to determine what successful repetition would mean.

02 · Record

Can the path to the result be reconstructed?

Data, code, protocol, decisions?

Missing procedural detail makes even honest results difficult to audit.

03 · Reproduce

Can the same materials yield the same result?

Same data, same analysis.

Reproduction tests computational and procedural transparency.

04 · Replicate

Does the result return with new evidence?

Independent data or experiment.

Replication tests whether the finding extends beyond one realization.

05 · Revise

What should change when repetition fails?

Method, claim, theory?

Scientific reliability depends on correction, not on never being wrong.

A result is only as inspectable as the record that produced it.

Transparent science preserves enough context for another person to retrace the transformation from raw observation to published claim.

Protocol

What was done?

Procedures, instruments, exclusions and timing should be recorded at operational detail.

Data

What was observed?

Raw and processed data should remain distinguishable where access is ethically and legally possible.

Code

How was analysis executed?

Scripts preserve calculations more faithfully than prose descriptions alone.

Environment

Which software and dependencies mattered?

Version differences can change computational results.

Decisions

Which judgment calls shaped the analysis?

Transformations, exclusions and model choices can alter conclusions materially.

Provenance

Where did every artifact come from?

Traceable lineage makes the evidentiary chain inspectable.

Reproduction asks whether the reported result can be regenerated.

The exercise catches missing steps, software drift, transcription mistakes, hidden defaults and analytical ambiguity.

Same data

Hold the evidence fixed.

Reproduction isolates whether the published analysis is recoverable from the original inputs.

Same method

Follow the documented pipeline.

Ambiguity becomes visible when an independent analyst must interpret incomplete instructions.

Independent run

Do not rely on the original interactive state.

A clean environment tests whether hidden dependencies exist.

Compare

Check outputs at several stages.

Intermediate comparisons locate the point at which results diverge.

Explain

Document differences rather than hiding them.

A discrepancy can reveal software, data or specification dependence.

Replication asks whether the phenomenon survives new evidence.

A useful replication specifies what should remain the same and what may legitimately vary.

Replication typeWhat changes?What it tests
DirectAs little as practicalRepeatability under closely matched conditions
ConceptualOperationalization or methodWhether the theoretical relation survives a different implementation
Multi-siteLocation and sampleGeneralization across contexts
Registered replicationAnalysis fixed in advanceReduces hindsight and selective analysis
Robustness analysisReasonable analytic choicesDependence on specification

Failure to repeat a result has many possible causes.

A non-replication is evidence that requires diagnosis, not an automatic verdict about fraud or falsity.

Sampling variation

Real effects do not appear identically every time.

Small studies can produce unstable estimates even when the underlying effect exists.

Context dependence

The phenomenon may require hidden conditions.

Population, setting or implementation differences can alter outcomes.

Low power

Evidence may be too weak to discriminate.

Both original and replication studies can be underinformative.

Analytic flexibility

Many defensible paths can yield different results.

Undocumented researcher choices can make findings specification-dependent.

Publication selection

Visible studies are not a random sample of conducted studies.

Selective publication can exaggerate apparent reliability.

Error or misconduct

Some failures reveal genuine defects.

Data errors, coding mistakes, fabrication or inappropriate methods remain possibilities that require evidence.

Science becomes reliable through organized correction.

The goal is not to freeze published knowledge but to make claims progressively harder to overturn for ordinary reasons.

Expose

Make methods, data, code and decisions inspectable where possible.

Precommit

Separate planned tests from analyses discovered after seeing results.

Repeat

Reproduce analyses and replicate findings under independent conditions.

Aggregate

Interpret a body of evidence rather than treating one study as final.

Revise

Update claims, methods and theories when accumulated evidence demands it.

Side 100 closes the first hundred with a constraint on every other Side: understanding becomes stronger when another person can retrace it, test it, fail to recover it, and force it to improve.
Reproducible Researchtransparent computational workflow
Replicationindependent testing of empirical findings
Open Sciencerecords, preregistration and accessible evidence
Meta-sciencestudying the reliability of scientific practice