What outcomes can occur?
Define the alphabet.
A source produces symbols or events according to some probability distribution.
Side 32
A study of how uncertainty becomes information and how messages survive compression and noise. Information theory strips communication down to source, code, channel and limits.
An event carries more information when it was less expected before it occurred.
Define the alphabet.
A source produces symbols or events according to some probability distribution.
Assign probabilities.
Rare outcomes are more surprising and carry more self-information.
−log₂ p(x)
Using base 2 expresses information in bits.
Expected information.
Entropy is highest when outcomes are spread more evenly across possibilities.
Mutual information.
Shared structure can be quantified without requiring a linear relationship.
Good coding exploits probability structure while preserving enough distinction for decoding.
Messages are mapped into sequences from a code alphabet.
This property enables instantaneous unambiguous decoding.
Variable-length coding approaches the entropy limit for known symbol probabilities.
Longer blocks can exploit dependencies that single-symbol codes miss.
Decoding quality depends on code design and channel corruption.
Compression performance is compared against the theoretical source entropy.
If data contain repetition or unequal symbol probabilities, they can often be represented with fewer bits than a naive encoding uses.
No information needed for reconstruction may be discarded.
Repeated patterns and predictable symbols are represented more compactly.
Average code length cannot be driven arbitrarily below source entropy without losing information.
Some detail is intentionally discarded to achieve much larger compression.
Compression quality depends on which differences the application can tolerate.
Rate–distortion theory studies the minimum bitrate needed for a chosen fidelity level.
The central question is not whether noise exists, but how much reliable communication remains possible despite it.
Input probabilities can be chosen to match channel characteristics.
Channel models specify probabilities of output given input.
The receiver sees a noisy version of the transmitted signal.
Channel capacity is the highest rate at which error probability can be made arbitrarily small with suitable coding.
Physical channels impose limits through bandwidth, signal power and noise.
It quantifies how much observing the output tells us about the input.
Compression removes predictable redundancy; error correction deliberately adds structured redundancy so damaged messages can be reconstructed.
Extra bits can detect some transmission errors without identifying all of them.
Hamming distance determines how many bit errors a code can detect or correct.
Structured codes make certain corruption patterns recoverable.
More redundancy generally reduces net information throughput while improving robustness.
Below channel capacity, suitable codes can drive error probability very low; above it, reliability cannot be guaranteed.
Its mathematics connects communication, inference, learning, thermodynamics, biology and computation.
Measures such as KL divergence quantify how one probability distribution differs from another.
Cross-entropy is central to classification and probabilistic learning.
Genetic and signaling systems can be studied through entropy and mutual information.
Neural coding asks how much information activity carries about stimuli or actions.
Statistical mechanics and information theory share mathematical structure while describing different physical and informational contexts.
Natural language contains redundancy, enabling both compression and error recovery.