Side 13 · Structured learning experiment
AI
Models
A beginner-first but technically serious study of the machinery beneath modern AI models. The aim is not to memorize terminology. It is to build a causal mental model strong enough to distinguish the model itself from inference, context, memory, tools, product interfaces and the larger system that produces an AI interaction.
One system, eleven variables.
The order follows the causal chain from how information becomes computable, through what training optimizes and stores, to what happens at inference time and how a model becomes part of a larger tool-using system.
Representation
How does language become something a machine can compute over?
02 · DevelopedObjective
What exactly is the machine being trained to accomplish?
03 · NextParameters
Where is what the model has learned actually stored?
04Architecture
What computation transforms input into useful internal representations?
05Learning
How do parameters change so the model becomes better at its objective?
06Inference
What actually happens when a user sends a message?
07Context
How does the model use information currently available to it?
08Reasoning
How can prediction machinery produce behavior that looks like thinking and problem-solving?
09Memory & Retrieval
What differs between parameter knowledge, context, stored memory and retrieved information?
10Tools & Runtime
What happens outside the model when an AI searches, calculates, reads or acts?
11Capability, Emergence & Limits
Why does sophisticated behavior arise, and where does it fail?
Throughout the study, keep five things separate: the model, the inference/runtime system, the product interface, external tools, and the broader system producing the interaction.
Representation
How does something like language become something a machine can compute over?
Developed1 · What it is
A language model does not receive language in the form a human experiences it. Text must first become numerical structure. The useful chain is text → tokens → token IDs → vectors → contextual representations.
A token is not necessarily a word. It can be a word, part of a word, punctuation, a number or another recurring character sequence. Each token receives an integer ID, but that ID is only a label. The mathematically useful representation comes when the ID is mapped to a vector: an embedding.
2 · Why it matters
The model cannot directly manipulate reality. It manipulates representations of information. If something is absent from the input or represented poorly, the model has less to work with. This is the foundation for later distinctions between knowledge encoded in parameters, information supplied in context and information retrieved through tools.
3 · How it works
A token embedding can be represented as a vector such as x = [0.2, -0.7, 0.4, 0.1]. Real model vectors have far more dimensions. Neural networks transform these vectors using learned operations.
where x is the input vector, W a learned matrix, b a learned bias, and y the transformed representation.
These transformations are composed repeatedly, often with nonlinear operations between them. Crucially, humans do not assign meanings such as “dimension 17 = animals.” Training discovers numerical structure that helps the model perform its objective.
Initial token representations are not enough. Consider “The dog chased the cat” versus “The cat chased the dog.” The same words occur, but the relationships differ. Modern models therefore construct contextual representations: a token's internal representation changes according to surrounding tokens. Attention, studied under Architecture, is central to this process.
4 · Concrete mental model
Imagine language projected into a vast learned mathematical landscape. Pieces of language become points and patterns of geometry encode useful relationships. During processing, those representations are repeatedly transformed and contextualized.
human experience
↓
text
↓
tokens
↓
numerical vectors
↓
contextual representations
↓
transformed internal states
Where the analogy breaks: there is no literal 3D map containing objects labelled “cat,” “France” or “democracy.” Representation is distributed across many dimensions, layers, activations and parameters, and it changes with context.
5 · Common misunderstandings
Tokens are not words.
Token boundaries need not match human linguistic units.
Token numbers do not contain meaning.
Numerical proximity between token IDs says nothing about semantic similarity.
There is no neat dictionary entry inside the model.
Useful semantic structure is distributed through learned numerical machinery.
The model does not see a sentence as you do.
By core computation time, the text has undergone several representational transformations.
6 · How it manifests in interaction
Long prompts containing background, constraints, examples, exceptions and objectives do more than “explain the task.” They alter the information structure available to the model. Structuring a problem as objective → constraints → architecture → stages → failure modes gives the model a different computational object than an underspecified request.
7 · What it enables / prevents
Learned representations can support syntax, semantics, entities, relationships, styles, abstractions, code structures and patterns of reasoning without each being manually programmed. But capability depends on usable representation. A model may be capable of answering a question while lacking the relevant information at inference time. Capability is not access.
8 · Connections
The objective determines which representations become useful. Parameters store the learned transformations. Architecture determines how representations interact. Learning adjusts the parameters. Inference executes the transformations. Context changes the representations currently active. Reasoning depends on transformations over representations. Memory asks where information resides. Tools introduce new information and operations.
9 · Personal assessment
Demonstrated: you operationally understand that information structure changes model behavior. You routinely decompose ambiguous problems, specify constraints and restructure inputs when a workflow is failing.
Inference — intuitive: you appear to treat AI as a computational system rather than a database of canned answers, and you recognize a gap between the conversational interface and the machinery underneath it.
Inference — probable misconception: your working model has tended to remain semantic: “I give the AI an idea, it interprets it, reasons, and answers.” Mechanically, that hides the layers of numerical transformation and encourages the false idea that a coherent human concept must correspond to one discrete internal object.
Not yet encountered formally: tokenization, embedding matrices, vector spaces, activations, contextual and distributed representations, representation geometry and the relevant matrix operations.
Relative strength: operational intuition is ahead of mechanistic vocabulary. The gap is not AI fluency versus AI ignorance; it is effective manipulation of a system whose internal representation machinery you have not yet learned.
Knowledge versus access
System A has billions of learned parameters from enormous text training but receives only the current prompt. System B has no learned world knowledge but can perfectly search every document in existence. Asked “What happened in the Canadian economy last Tuesday?”, which system has the knowledge, which has the information, and are knowing something and having access to information the same thing?
Objective
What exactly is the machine being trained to accomplish?
Developed1 · What it is
A machine-learning system needs a criterion by which one output is better than another during training. For a language model, the foundational pretraining task is remarkably simple: given the tokens so far, predict the next token.
The → predict "capital" The capital → predict "of" The capital of → predict "France" The capital of France → predict "is" The capital of France is → predict "Paris"
2 · Why it matters
The objective supplies the learning pressure. Without it, there is no mathematical direction saying one parameter configuration is better than another. This separates machine learning from traditional rule programming: humans design the architecture, data pipeline and optimization process; the training process discovers parameters that perform the objective.
data → model → prediction → compare with target → error → adjust parameters → repeat
3 · How it works
The probability of the next token given all preceding tokens available in context.
The model computes a probability distribution across candidate next tokens. Once a token is selected, it joins the sequence and the model runs again. This is autoregressive generation: previous output becomes part of the input conditioning later output.
L is loss, y is the correct next token, and P(y) is the probability assigned to it. High probability for the observed token means lower loss.
Training then asks how changes to parameters would affect loss. The gradient supplies that direction; gradient descent moves parameters slightly in the direction that reduces error.
θ represents trainable parameters, η the learning rate, and ∇θL the gradient of loss with respect to those parameters.
4 · Why prediction can produce broad capability
To predict language well, a system benefits from learning structure about what language describes. “Water freezes at ___ degrees Celsius” rewards physical knowledge. “Alice is older than Bob; Bob is older than Claire; therefore the oldest is ___” rewards relational structure. Code completion rewards programming structure. The objective is simple; a strong strategy for satisfying it is not.
Saying a model “predicts the next token” does not tell us how complex the computation producing that probability distribution may be. Prediction and reasoning are not automatically mutually exclusive descriptions.
5 · Pretraining is not the whole assistant
A raw model trained only to continue internet text need not behave like an assistant. Modern systems use additional post-training to shape instruction-following, preferences, constraints and useful response behavior.
large-scale pretraining
↓
broad predictive capability
↓
post-training
↓
assistant-like behavior
↓
inference-time system + context + tools
↓
product experience
6 · Concrete mental model
Picture a vast parameter landscape in which height represents prediction error. Training repeatedly estimates a local downhill direction and adjusts the model. The analogy is useful for optimization, but the real space may have billions of dimensions, training uses batches rather than one global view, and optimization need not find a unique or globally best solution.
Training is not inserting answers into storage. It is optimizing a gigantic numerical system against an objective.
7 · Common misunderstandings
“Just autocomplete” is technically suggestive but explanatorily weak.
The output is autoregressive, but the function computing each next-token distribution can be highly complex.
The base model was not fundamentally trained to answer your questions.
Broad prediction comes first; assistant behavior is shaped later.
Training is not database lookup.
A learned relationship can affect output through parameters without the model retrieving a stored sentence from its training corpus.
The foundational objective is not truth itself.
Truthful behavior requires more than observing that next-token prediction rewards statistical fit to training data.
Probability does not mean the whole computation is random.
The network can deterministically compute a probability distribution while token selection uses a separate sampling procedure.
8 · How it manifests in interaction
Your prompt changes the conditioning information, not the model's trained parameters. Correcting a framing inside a conversation can radically change later output without any gradient descent occurring. Because generation is autoregressive, an early mistaken framing can also condition later tokens and create error propagation.
Changes parameters through optimization. Produces learned capability before ordinary interaction.
Changes context at inference time. Uses the already-trained model without rewriting its parameters.
9 · What it enables / prevents
A broad predictive objective allows one learned system to support translation, summarization, coding, classification, explanation, extraction, transformation and some forms of reasoning without a separate hand-coded rule set for every task.
But predictive optimization alone does not guarantee truth, current information, exact arithmetic, logical consistency or recognition of ignorance. A surrounding system can compensate with retrieval, calculators, verification and other tools. At that point, we are analyzing a system containing a model, not the model alone.
10 · Connections
Representation gives the objective something mathematical to operate over. Parameters are what optimization changes. Architecture constrains the computations available. Learning is the mechanism that turns error into parameter updates. Inference uses the trained system without ordinary training updates. Context conditions current behavior. Reasoning asks what computations can occur inside prediction. Memory and tools separate learned structure from external information and capability.
11 · Personal assessment
Demonstrated: you already separate AI use from AI-enabled system design and understand operationally that instructions, information and workflow can change behavior without rebuilding the underlying model.
Inference — intuitive: you appear to reject the crude “database of everything it read” model and to recognize generalization: a model can transform information and attempt tasks that were not individually hard-coded.
Inference — probable misconception: the distinction between a training objective and the apparent goal of an assistant during a conversation was probably not previously sharp. When you ask for “the strongest commercial opportunity,” that becomes an inference-time instruction, not a newly installed training objective.
Not yet mastered: conditional probability notation, cross-entropy, gradients, gradient descent, backpropagation, autoregressive generation, post-training and sampling. You now have their roles; the mathematics and mechanisms remain ahead.
Relative strength: you have observed context sensitivity, error propagation, generalization, decomposition and model-versus-system capability before learning their formal machinery.
Relative weakness: the mathematical mechanism is still new. You cannot yet trace how loss changes a particular parameter or explain where learned knowledge resides. That is the next variable.
The liar
Train Model A on consistently correct geography and Model B on an equally consistent fictional corpus in which Paris is Germany's capital and Berlin is France's. Both achieve equally low prediction loss. Which better achieved the objective? Does Model B “know” something false, or merely learn the statistical structure of its world? If prediction rewards fit rather than truth directly, where can truth enter the larger system?
Parameters
Where is what the model has learned actually stored?
NextWeights, distributed knowledge, parameter count, memorization versus generalization, and why a parameter is not a fact.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Architecture
What computation transforms input into useful internal representations?
AheadNeural layers, transformers, attention, residual streams, position and the computational path from token representations to logits.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Learning
How do parameters change so the model becomes better at its objective?
AheadForward passes, loss, derivatives, backpropagation, optimizers, batches, epochs, data and the practical mechanics of training.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Inference
What actually happens when you send a message?
AheadPrompt assembly, forward computation, logits, softmax, sampling, autoregressive decoding, latency and deterministic versus stochastic generation.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Context
How does the model use information currently available to it?
AheadContext windows, attention over current tokens, instructions, in-context learning, context limits and the difference between context and parameters.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Reasoning
What does it mean for prediction machinery to appear to think, solve, plan or reason?
AheadInternal computation, intermediate representations, test-time computation, decomposition, chain-like processes, verification and the limits of behavioral evidence.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Memory & Retrieval
What is the difference between something encoded in parameters, retrieved from somewhere, and remembered across interactions?
AheadParametric knowledge, working context, external memory, retrieval systems, persistence and apparent versus actual memory.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Tools & Runtime
What happens outside the model when an AI searches, calculates, reads files or takes actions?
AheadTool calling, orchestration, permissions, APIs, code execution, retrieval, product logic and the boundary between model capability and system capability.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
Capability, Emergence & Limits
Why can sophisticated behavior arise, where does it break, and what does intelligence mean here?
AheadScaling, generalization, emergence, hallucination, calibration, brittleness, interpretability and the difference between intelligence-like behavior and claims about inner experience.
When developed, this variable will use the same structure as Variables 1–2: what it is; why it matters; how it works; concrete mental model; common misunderstandings; manifestations in ChatGPT; what it enables and prevents; connections to the other variables; and a personal assessment separating demonstrated understanding from inference.
The distinction to preserve throughout.
“The AI” is not one indivisible thing. Every later variable should make it easier to identify which layer is responsible for an observed behavior.