Skip to content

Side 16

Epistemology
of Evidence

A study of how observations become reasons for belief. Evidence does not speak alone: it is collected through instruments, interpreted through models, weighed against alternatives and judged under uncertainty.

claim←inference←measurement←observation↔background knowledge
06evidence forms
05warrant tests
06failure modes
16Side

Evidence is a relationship, not an object.

The same observation can support different claims under different assumptions. To evaluate evidence, trace the entire route from reality to conclusion.

01 · World

What happened?

Define the target phenomenon.

Before measurement, specify what event, property or process the claim is actually about.

02 · Observation

What was noticed?

Direct, recorded or reported?

Observation is selective: conditions, timing and access determine what can appear in the record.

03 · Measurement

How was it represented?

What instrument or proxy?

Measurements transform phenomena into values, categories, images, scores or traces.

04 · Inference

What bridge connects data to claim?

Deduction, induction, causal model, comparison?

The evidential force comes from the inference, not from the data alone.

05 · Claim

How far does the conclusion travel?

Description, explanation, prediction or generalization?

The broader the claim, the more assumptions must survive scrutiny.

Core habitobservation + measurement + assumptions + inference → warranted claim

Different evidence answers different questions.

No universal hierarchy works across every domain. The relevant question is what kind of evidence bears on the claim being made.

Observation

Direct or instrument-mediated record.

Useful for describing what occurred under observed conditions, but vulnerable to selection and measurement limits.

Experiment

Deliberate intervention.

Useful for causal questions when treatment assignment and comparison isolate the relevant difference.

Natural experiment

World-created contrast.

Events or rules generate comparison groups without researcher assignment; credibility depends on the identifying assumptions.

Testimony

Knowledge carried by people.

Expertise, incentives, access, independence and track record affect how testimony should be weighted.

Trace evidence

Effects left by prior events.

Documents, residues, logs, fossils and records can support reconstruction when the process creating the trace is understood.

Convergence

Independent methods point the same way.

Agreement across different measurement and inference pathways can be stronger than repetition through one shared weakness.

Ask what makes the evidence support the claim.

Warrant is the bridge between what was observed and what is being concluded.

Relevance

Does this evidence bear on this claim?

Information can be accurate yet irrelevant to the proposition under dispute.

Reliability

Would the method tend to get this right?

Calibration, error rates, instrument stability and procedure matter.

Discrimination

Does the evidence distinguish alternatives?

Strong evidence should be more expected under one explanation than its serious competitors.

Scope

How far can the conclusion generalize?

A local result may not warrant claims about other populations, settings, periods or mechanisms.

Independence

Are multiple pieces of evidence genuinely separate?

Ten sources repeating one original report do not equal ten independent observations.

Evidence can raise confidence without settling the question.

Epistemic strength is often comparative and cumulative. The relevant update may be “more plausible than before,” not “proved.”

Repetition tests more than one thing.

Replication can test whether a result reappears, whether a method generalizes and whether the original interpretation survives independent scrutiny.

Replication lenses

Direct

Repeat the procedure closely.

Tests whether the result can recur under similar conditions.

Conceptual

Test the same idea differently.

Uses another operationalization or method to examine whether the underlying claim survives.

External

Move the setting.

Tests whether findings generalize to new populations, environments or periods.

Disagreement lenses

Data

Are parties seeing different evidence?

Disagreement may disappear once the factual record is shared.

Model

Same data, different inference.

Competing causal assumptions or background theories can produce different conclusions.

Prior

Same likelihood, different starting beliefs.

Reasonable people can update in the same direction while ending at different confidence levels.

Evidence can be weakened before analysis begins.

Many failures arise from how observations are generated, selected, transformed or communicated.

Selection

The observed cases are not neutral.

Who enters the dataset may already depend on the outcome or exposure of interest.

Measurement

The proxy drifts from the construct.

A convenient metric can become mistaken for the phenomenon it only partially represents.

Multiplicity

Enough searches eventually find patterns.

Unreported comparisons make apparently surprising results less surprising.

Publication

Visible evidence is filtered.

Positive, novel or dramatic findings may be more likely to appear than null or routine results.

Dependence

Many sources share one origin.

Apparent corroboration can collapse when reports trace back to the same data or witness.

Overclaim

The conclusion outruns the design.

Association becomes causation, one population becomes universal, or a proxy becomes the underlying construct.

Claim clinic.

Practice decomposing a statement before deciding whether its evidence is strong.

“Customers who use feature X stay longer.”

Possible evidence of association. Ask whether longer-term customers are simply more likely to discover the feature, whether adoption preceded retention, and what comparable non-users looked like before use.

“Three experts independently agree.”

Check independence. Did they inspect separate evidence, or are all three relying on the same original report, model or dataset?

“The result replicated.”

Ask what was replicated: procedure, direction, effect size, population, mechanism or merely a statistically detectable pattern?

“This indicator predicts recessions.”

Define the prediction horizon, historical sample, false-positive rate, revisions to the indicator, and whether the rule was specified before observing the outcomes.

The Logic of Scientific DiscoveryKarl Popper · testing and falsifiability
The Structure of Scientific RevolutionsThomas Kuhn · paradigms and scientific change
Evidence and EvolutionElliott Sober · evidential reasoning
Why Trust Science?Naomi Oreskes · expertise and scientific institutions