What happened?
Define the target phenomenon.
Before measurement, specify what event, property or process the claim is actually about.
Side 16
A study of how observations become reasons for belief. Evidence does not speak alone: it is collected through instruments, interpreted through models, weighed against alternatives and judged under uncertainty.
The same observation can support different claims under different assumptions. To evaluate evidence, trace the entire route from reality to conclusion.
Define the target phenomenon.
Before measurement, specify what event, property or process the claim is actually about.
Direct, recorded or reported?
Observation is selective: conditions, timing and access determine what can appear in the record.
What instrument or proxy?
Measurements transform phenomena into values, categories, images, scores or traces.
Deduction, induction, causal model, comparison?
The evidential force comes from the inference, not from the data alone.
Description, explanation, prediction or generalization?
The broader the claim, the more assumptions must survive scrutiny.
No universal hierarchy works across every domain. The relevant question is what kind of evidence bears on the claim being made.
Useful for describing what occurred under observed conditions, but vulnerable to selection and measurement limits.
Useful for causal questions when treatment assignment and comparison isolate the relevant difference.
Events or rules generate comparison groups without researcher assignment; credibility depends on the identifying assumptions.
Expertise, incentives, access, independence and track record affect how testimony should be weighted.
Documents, residues, logs, fossils and records can support reconstruction when the process creating the trace is understood.
Agreement across different measurement and inference pathways can be stronger than repetition through one shared weakness.
Warrant is the bridge between what was observed and what is being concluded.
Information can be accurate yet irrelevant to the proposition under dispute.
Calibration, error rates, instrument stability and procedure matter.
Strong evidence should be more expected under one explanation than its serious competitors.
A local result may not warrant claims about other populations, settings, periods or mechanisms.
Ten sources repeating one original report do not equal ten independent observations.
Epistemic strength is often comparative and cumulative. The relevant update may be “more plausible than before,” not “proved.”
Replication can test whether a result reappears, whether a method generalizes and whether the original interpretation survives independent scrutiny.
Tests whether the result can recur under similar conditions.
Uses another operationalization or method to examine whether the underlying claim survives.
Tests whether findings generalize to new populations, environments or periods.
Disagreement may disappear once the factual record is shared.
Competing causal assumptions or background theories can produce different conclusions.
Reasonable people can update in the same direction while ending at different confidence levels.
Many failures arise from how observations are generated, selected, transformed or communicated.
Who enters the dataset may already depend on the outcome or exposure of interest.
A convenient metric can become mistaken for the phenomenon it only partially represents.
Unreported comparisons make apparently surprising results less surprising.
Positive, novel or dramatic findings may be more likely to appear than null or routine results.
Apparent corroboration can collapse when reports trace back to the same data or witness.
Association becomes causation, one population becomes universal, or a proxy becomes the underlying construct.
Practice decomposing a statement before deciding whether its evidence is strong.
Possible evidence of association. Ask whether longer-term customers are simply more likely to discover the feature, whether adoption preceded retention, and what comparable non-users looked like before use.
Check independence. Did they inspect separate evidence, or are all three relying on the same original report, model or dataset?
Ask what was replicated: procedure, direction, effect size, population, mechanism or merely a statistically detectable pattern?
Define the prediction horizon, historical sample, false-positive rate, revisions to the indicator, and whether the rule was specified before observing the outcomes.