What happened?
Summarize variation before explaining it.
Center, spread, shape, outliers and relationships turn raw observations into a visible empirical object.
Side 06
A study of uncertainty from both directions: probability starts with a model and asks what data could occur; statistics starts with data and asks what the underlying world may be. The point is not calculation alone. It is learning what evidence can and cannot justify.
The dependency chain matters. Inference is fragile when variation, probability and sampling have not been understood first.
Summarize variation before explaining it.
Center, spread, shape, outliers and relationships turn raw observations into a visible empirical object.
Assign structure to uncertain outcomes.
Events, conditional probability, independence and random variables create the grammar for uncertainty.
Separate population from sample.
Sampling variation explains why repeated studies differ even when nothing fundamental has changed.
Estimate before declaring.
Intervals, likelihoods, tests and posterior distributions quantify uncertainty around claims.
Evidence does not choose the loss function.
Decisions require costs, benefits, thresholds and consequences in addition to statistical evidence.
Probability is the forward problem: assume a model of the world, then reason about possible observations.
A number from 0 to 1 representing uncertainty about an event under a specified model.
The probability of A after restricting attention to cases where B is known to have occurred.
B supplies no information about A under the model. Independence is an assumption to examine, not a default.
The probability-weighted long-run average of a random variable. It need not be a value that can actually occur.
A measure of spread around the expectation. Two processes can share a mean while carrying very different risk.
New evidence updates an existing probability through the likelihood of seeing that evidence under competing possibilities.
A highly sensitive test can still produce mostly false positives when the underlying condition is rare.
Distributions are not decorative curves. Each one encodes assumptions about what can happen and how often.
| Family | Question it models | Structural clue | Typical use |
|---|---|---|---|
| Bernoulli | Did one binary event occur? | Two outcomes; one probability p. | Success/failure, yes/no events. |
| Binomial | How many successes in n trials? | Fixed number of comparable Bernoulli trials. | Counts of successes. |
| Normal | How does continuous variation cluster around a center? | Symmetric bell shape; fully described by mean and variance. | Measurement error and many aggregate phenomena. |
| Poisson | How many events arrive in a fixed interval? | Count data generated by a rate. | Arrivals, incidents, defects. |
| Exponential | How long until the next event? | Waiting time associated with a constant event rate. | Reliability and inter-arrival times. |
| Power law | What if rare large events matter disproportionately? | Heavy tail; scale lacks a single “typical” size. | Some network, wealth and event-size phenomena. |
A named distribution is useful only when its assumptions are plausible enough for the question being asked.
Inference is the reverse problem: data are observed; the generative process is partly unknown.
Parameters are fixed but unknown. Repeated hypothetical samples define the long-run behavior of estimators, confidence procedures and tests.
Report a plausible range and its procedure, not merely a single best estimate.
A p-value describes how surprising the data or more extreme results would be if the null model were true. It is not the probability that the null is true.
Uncertainty about parameters is represented directly with probability distributions. Prior information is updated by the likelihood to obtain a posterior.
Make assumptions visible rather than pretending the analysis began without them.
The result depends jointly on prior information, the model and observed data.
An effect can be precisely estimated and trivial, or large enough to matter while still uncertain.
Causal claims require a model of what would have happened under an alternative exposure or intervention, not merely a correlation.
The observed association can partly or wholly reflect a common cause.
Who enters, remains in or responds to a study can create systematic distortion.
Restricting or controlling for a common effect can manufacture a relationship that was absent.
When noisy measurements are selected for extremity, later measurements tend to be less extreme even without intervention.
Searching many outcomes, subgroups or models raises the chance of apparently notable results.
Error, categorization and construct validity can limit what a measured number actually represents.
Open a claim and interrogate it before deciding whether the result deserves belief, action or another study.
Ask: 23% relative or absolute? How were adopters selected? What was baseline retention? Was adoption randomized? Could more engaged users simply be more likely to adopt? What uncertainty surrounds the estimate?
Ask: What effect size was estimated? What is the interval? How many tests were run? Was the analysis pre-specified? Does the result survive plausible modeling choices? What decision would change if the p-value were 0.06?
Ask: Compared with what baseline? On which population and time period? Is the class distribution imbalanced? Which errors matter most? Was the evaluation set genuinely held out?
Ask: Compared with what counterfactual? Did composition change? Was there a broader economic trend? Which groups moved? Is the mean hiding distributional changes?