Data & target
Define examples, labels and sampling so the training problem represents the intended deployment task.
Subject
Purpose
Predictive and generative models studied through data, objectives, representations, optimization, generalization and deployment under distribution shift.
Structure
Components → constraints → flows → control → failure
Machine learning is a pipeline from data-generating process to objective to trained model to deployment, and every transition can introduce failure even when benchmark accuracy is high.
Define examples, labels and sampling so the training problem represents the intended deployment task.
Choose features or learned architectures that encode useful inductive biases without silently excluding relevant variation.
Train by minimizing a surrogate loss while recognizing that optimization success does not guarantee the right behavioral objective.
Evaluate on unseen data and diagnose overfitting, leakage, calibration and subgroup performance.
Monitor changing inputs, feedback loops, robustness and human use after the model leaves the benchmark environment.
training loss ≠ real-world utility
prediction ≠ causal explanation
benchmark improvement ≠ general capability
What distribution is the model actually expected to generalize to?
Which errors are hidden by aggregate performance metrics?
How should objectives change when model outputs alter the future data it receives?
Use held-out and external evaluation, ablations, calibration and stress tests; claims about mechanisms or causality require evidence beyond predictive performance.