What result is being asserted?
Estimate, effect, model, phenomenon?
The target must be precise enough to determine what successful repetition would mean.
Side 100
A study of whether scientific results survive being reconstructed, repeated and challenged. Reproducibility preserves the path from evidence to result; replication tests whether the phenomenon itself returns under an independent attempt.
A result can be perfectly reproducible from the original materials yet fail to replicate independently—and the reverse distinction matters for diagnosis.
Estimate, effect, model, phenomenon?
The target must be precise enough to determine what successful repetition would mean.
Data, code, protocol, decisions?
Missing procedural detail makes even honest results difficult to audit.
Same data, same analysis.
Reproduction tests computational and procedural transparency.
Independent data or experiment.
Replication tests whether the finding extends beyond one realization.
Method, claim, theory?
Scientific reliability depends on correction, not on never being wrong.
Transparent science preserves enough context for another person to retrace the transformation from raw observation to published claim.
Procedures, instruments, exclusions and timing should be recorded at operational detail.
Raw and processed data should remain distinguishable where access is ethically and legally possible.
Scripts preserve calculations more faithfully than prose descriptions alone.
Version differences can change computational results.
Transformations, exclusions and model choices can alter conclusions materially.
Traceable lineage makes the evidentiary chain inspectable.
The exercise catches missing steps, software drift, transcription mistakes, hidden defaults and analytical ambiguity.
Reproduction isolates whether the published analysis is recoverable from the original inputs.
Ambiguity becomes visible when an independent analyst must interpret incomplete instructions.
A clean environment tests whether hidden dependencies exist.
Intermediate comparisons locate the point at which results diverge.
A discrepancy can reveal software, data or specification dependence.
A useful replication specifies what should remain the same and what may legitimately vary.
| Replication type | What changes? | What it tests |
|---|---|---|
| Direct | As little as practical | Repeatability under closely matched conditions |
| Conceptual | Operationalization or method | Whether the theoretical relation survives a different implementation |
| Multi-site | Location and sample | Generalization across contexts |
| Registered replication | Analysis fixed in advance | Reduces hindsight and selective analysis |
| Robustness analysis | Reasonable analytic choices | Dependence on specification |
A non-replication is evidence that requires diagnosis, not an automatic verdict about fraud or falsity.
Small studies can produce unstable estimates even when the underlying effect exists.
Population, setting or implementation differences can alter outcomes.
Both original and replication studies can be underinformative.
Undocumented researcher choices can make findings specification-dependent.
Selective publication can exaggerate apparent reliability.
Data errors, coding mistakes, fabrication or inappropriate methods remain possibilities that require evidence.
The goal is not to freeze published knowledge but to make claims progressively harder to overturn for ordinary reasons.
Make methods, data, code and decisions inspectable where possible.
Separate planned tests from analyses discovered after seeing results.
Reproduce analyses and replicate findings under independent conditions.
Interpret a body of evidence rather than treating one study as final.
Update claims, methods and theories when accumulated evidence demands it.