Local Minimum Pictures presents

How Wrong Am I?

a film by Learning Rate

A machine starts with two numbers and one question. Twenty-five episodes, two and a half centuries, one answer that keeps changing shape.

A drama series in twenty-five episodes. 1 of 25 out now; next: Episode 2, Downhill, coming soon. Every number on screen is real: computed in your browser, or taken from a real model and checked against it. Best on a desktop with a graphics card.

Episodes

Part one · episodes 1–8

Learning

The learner gets a score, a way down and a vote of its neighbours. Then its first crisis: it can fool itself, its score hides which way it is wrong, and it is wrong more for some people than for others.

Every Line Has a Score

“By the square of every miss.”

Linear regression and the loss surface

Sixty points, one line, two knobs. Every line you could draw has a score, and all of those scores together form a bowl whose bottom is the best line.

New ideas modelparameterspredictionresiduallossmean squared errorloss surfaceleast squaresconvex bowl

  • 65,536 lines drawn at once, each inked by its real score, and one line of your own to drag through sixty points.
  • Every miss becomes a square; the score is the area of the average square.
  • Sixty troughs, one per data point, settle onto each other until they make the glass bowl whose bottom is the best line.

7 scenes · where the series begins · after Legendre, “Sur la méthode des moindres quarrés”, 1805

Episode link Not started
coming soon

Downhill

“Less wrong, one step downhill.”

Gradient descent and momentum

A model cannot see the valley it is searching for. It can only feel the slope under its feet, and take a step.

New ideas gradientgradient descentlearning ratelocal minimummomentumadaptive steps (Adam)feature scaling

  • A ball that only feels the slope, stepping by η × gradient
  • Momentum climbing out of a dip that traps plain descent
  • Episode 1's bowl drawn to the walker's scale: a ravine 12 times longer than it is wide that takes 278 zigzag steps, and one step once x is standardised

8 scenes · builds on Every Line Has a Score · after Cauchy, Comptes rendus 25, 1847

Coming soon
coming soon

The Line That Leans

“By how surprised I am.”

Logistic regression and classification

Two crowds of points and a line between them. Distance from the line becomes a probability, and every point that surprises the model pulls the line toward it.

New ideas classificationdot productweight vectorsigmoidpredicted probabilitydecision boundarycross-entropylikelihoodaccuracy

  • A probability sheet rising out of the floor: a score turned into belief by the sigmoid
  • Each point's surprise drawn as a stalk; their average is the loss being trained down
  • The boundary swinging 92° under real gradient descent, its past positions fanned behind it

7 scenes · builds on Every Line Has a Score and Downhill · after Cox, “The regression analysis of binary sequences”, 1958

Coming soon
coming soon

Ask the Neighbours

“By what my neighbours say.”

k-nearest neighbours, distance and scaling

Two hundred and forty houses on a turquoise town plan. Every spot on the map asks its k nearest houses and takes their vote, and turning k wears the town down from jagged islands to smooth hills, then sinks it to sea level.

New ideas distancenearest neighboursk-nearest neighboursmodels without a fixed shape (non-parametric)curse of dimensionality

  • Every one of 3.69 million points on the plan running its own k-nearest vote on the graphics card, the land rising as the vote grows lopsided
  • k = 1 as a map of lots, a Voronoi diagram: 100% right on the town’s own houses, 82.0% on 2,000 new ones
  • North measured in millimetres turning the whole plan into stripes, and 1,000 dimensions where the nearest point is 89% as far as the farthest

7 scenes · builds on The Line That Leans and Every Line Has a Score · after Fix & Hodges (1951) and Cover & Hart (1967)

Coming soon
coming soon

The Trade

“Wrong the same way, or wrong every time.”

The bias–variance trade-off

Two thousand models, each trained on its own thirty noisy points. Too simple and they are wrong the same way. Too flexible and they are wrong every time.

New ideas generalisationtraining errortest erroroverfittingunderfittingbiasvarianceirreducible noisemodel complexitypolynomial features

  • 2,000 polynomial fits drawn at once, degree by degree
  • Bias as a gap, variance as a spread, adding up to test error
  • A complexity dial that moves the sweet spot when you change the data

7 scenes · builds on Every Line Has a Score · after Geman, Bienenstock & Doursat, “Neural networks and the bias/variance dilemma”, 1992

Coming soon
coming soon

The Honest Exam

“On what I haven’t seen.”

Regularisation and cross-validation

A penalty tames a model that bends too much. Folds of held-out data choose how hard to press it. One peek at the answers turns fifty coin flips into a 97% score.

New ideas regularisationridge (L2)lasso (L1)sparsityhyperparametertrain / test splitvalidation setcross-validationdata leakagesplitting by time

  • The Trade's 30 points dealt as index cards into five piles, each sitting out once while its error is chalked in
  • A classifier on pure noise scoring 97% after one peek, and 50% done honestly, over 50 real runs
  • Two years of invented days: shuffled rehearsals promise an error of 25.4, the sealed next ten weeks deliver 53.0

9 scenes · builds on The Trade and Every Line Has a Score · after Stone (1974) and Geisser (1975)

Coming soon
coming soon

The Rare Case

“It depends which way I’m wrong.”

Confusion matrix, precision and recall, ROC, base rates and calibration

A made-up town of 100,000 takes a test that catches 99% of the ill. Everyone who tests positive walks into one queue, and only one in six of them is ill. Then a real thyroid model, its threshold, its curves and its honesty.

New ideas confusion matrixfalse positivefalse negativeprecisionrecalldecision thresholdROC curvearea under the curve (AUC)base rateBayes' ruleclass imbalancecalibration

  • 100,000 people on a night plain: a line of light sweeps across, and 5,940 stay lit
  • Everyone who tested positive in one queue; the ill fill 5 of its 30 lanes: 16.7%
  • A real model's 563 unseen patients on a strip you can cut, its ROC and precision–recall curves, and what a 50/50 training sample does to its probabilities

4 scenes · builds on The Line That Leans and The Honest Exam · after Bayes (read 1763, printed 1764)

Coming soon
coming soon

Wrong for Whom?

“More for some of you than others.”

Fairness: error rates by group and the impossibility

Two invented groups, one risk model, one threshold, and more than twice the false alarms for one group. Give each group its own threshold and try every pair: with different base rates a flag cannot mean the same thing in both groups while the mistakes also fall equally, so fair is a choice someone has to state.

New ideas error rates by groupequal selection rates (demographic parity)equal error rates (equalised odds)the fairness impossibilityproxy features

  • 100,000 invented people on a legal pad, sorted into columns where a false-alarm rate is a height you can see
  • Every pair of thresholds on one inked chart: the curves for equal error rates and equal precision never meet
  • Delete the group column and postcode rebuilds it; delete postcode too and calibration breaks

7 scenes · builds on The Rare Case and The Line That Leans · after Kleinberg, Mullainathan & Raghavan (2016) and Chouldechova (2016)

Coming soon

Part two · episodes 9–18

Bending

Straight lines are not enough. It asks questions, corrects itself and is made to explain itself; then it loses its teacher; then it lifts the data by hand, and learns to bend space itself.

coming soon

Two Hundred Trees

“Wrong alone, right together.”

Decision trees and random forests

One tree remembers every point it studied and stumbles on new ones. Two hundred trees, each a little wrong, agree on something better.

New ideas decision treeimpuritybootstrapensemblemajority votedecorrelated errors

  • One tree that aces its 500 points and misses 39 of 300 new ones
  • Two hundred trees, each grown on its own resample, voting
  • Ask the forest about any spot and read its ballot

7 scenes · builds on The Trade and The Line That Leans · after Breiman, “Random Forests”, 2001

Coming soon
coming soon

The Last Mistake

“By what’s left over.”

Gradient boosting

Boosting grows trees in a line, each one fitted to what the others still get wrong: gradient descent where every step is a tree. Watch Downhill's landscape carved out of flat steps, then see it lose, honestly, to a forest on this terrain.

New ideas boostingweak learnerfitting residualsshrinkageearly stopping

  • Sixteen flat steps rise out of a slab: one tree, crude on its own.
  • What is still wrong, carved flat a quarter-tree at a time, while the model beside it grows into Downhill's canyon.
  • The validation curve turns at round 195, and a sealed test set opened once says the forest wins this terrain.

7 scenes · builds on Downhill and Two Hundred Trees · after Friedman, “Greedy function approximation: a gradient boosting machine”, 2001

Coming soon
coming soon

Why Did It Say That?

“Shared out among my reasons.”

Feature importance, partial dependence and Shapley values

A boosted model prices 20,640 California districts from the 1990 census. Shuffle its inputs, turn one dial for everyone, and share one answer out fairly among its reasons. At dusk the coast lights up with the model's location premium.

New ideas feature importancepermutation importancepartial dependenceShapley valueslocal explanation

  • Shuffle latitude among 4,128 unseen districts and the typical miss quadruples, from $48,600 to $197,400
  • 20,640 threads of what-if answers condense into one partial-dependence line, then fall apart again
  • Santa Monica's $325,200 rebuilt as a stack of Shapley shares, then every district's location share poured into a dusk map until the coast glows

7 scenes · builds on Two Hundred Trees and Every Line Has a Score · after Shapley (RAND, 1951)

Coming soon
coming soon

Five Pins

“By the distance to my pin.”

k-means clustering

Four hundred thousand stars that nobody has named. k-means finds the clumps without being told what a clump is.

New ideas unsupervised learningclusteringcentroidinertiak-meansinitialisationelbow method

  • 400,000 stars sorted by guess, assign, average, repeat
  • Territory lines, and the stars that flip at every round
  • The elbow: how many clumps is the right number?

7 scenes · builds on Ask the Neighbours and Downhill · after Lloyd, “Least squares quantization in PCM”, 1957

Coming soon
coming soon

Soft Edges

“By how unlikely my story is.”

Gaussian mixtures and EM

The same 400,000 stars as Five Pins, but now a star may belong a little to two clumps. Starting from Five Pins' own answer, EM runs live in your browser, ellipses stretch and turn round by round, and the long bar k-means cut in two ends up inside one ink ring.

New ideas Gaussiancovariancemixture modelsoft assignmentexpectation–maximisation

  • Five Pins' hard territories get wet: the dashed cut through the long bar softens into colour that blends where clumps overlap.
  • One star split 51% indigo, 49% viridian, with the formula that splits it and a probe you can drag to any star.
  • 108 real EM rounds on a fixed table: the indigo ellipse swings across the cut until 90% of the bar is under one ring, the likelihood climbing all the way.

7 scenes · builds on Five Pins and The Trade · after Dempster, Laird & Rubin, “Maximum likelihood from incomplete data via the EM algorithm”, 1977

Coming soon
coming soon

Shadows

“By what the shadow loses.”

Principal component analysis

Turn a light around a cloud of points until its shadow is as wide as it can be. That direction keeps the most of the story, and a few dozen such directions stand in for all 256 pixels of a real handwritten digit.

New ideas projectionprincipal componentexplained variancedimensionality reductioneigenvectorreconstruction

  • A light circling 200,000 points while their shadow on a paper screen stretches to its widest, 1.90
  • A rod that finds the cloud's long axis by itself: multiply by the covariance, normalise, repeat
  • 4,000 real handwritten digits (MNIST) sorting themselves by three numbers, then rebuilt from eigen-digits, one term at a time

7 scenes · builds on Soft Edges and Five Pins · after Pearson, “On lines and planes of closest fit”, 1901

Coming soon
coming soon

The Widest Street

“By standing too close.”

Support vector machines and kernels

Hundreds of lines split two crowds perfectly. The support vector machine picks the one with the widest empty street, set by the few points on its kerbs, and a kernel lifts the data until a flat street can curve.

New ideas marginsupport vectorshinge losssoft marginkernelfeature mapkernel trick

  • 400 real, verified perfect lines threading the gap between two crowds, then the one street wider than all of them
  • Drag a point: only the three on the kerb move the street, and the SVM re-solves live
  • Rings lifted onto a bowl, split by a flat street in 3D, then lowered until the street lands as a circle

7 scenes · builds on The Line That Leans and The Honest Exam · after Boser, Guyon & Vapnik, “A training algorithm for optimal margin classifiers”, 1992

Coming soon
coming soon

Folding Space

“Wrong until I bend.”

Hidden layers and nonlinearity

Two spirals no straight line can split. A small network stretches, turns and bends the plane, layer by layer, until a flat cut separates them, and the cut, carried back, is a spiral.

New ideas neuronactivation functionhidden layernonlinearitylearned representationlinear separability

  • Every straight line, swept through 3,600 angles: none gets more than 391 of 600 spirals right
  • A rubber sheet printed with the spirals, folded in 3D by the trained network's own three layers until one flat pane splits coral from ink
  • Unfold it and the flat cut lands on the original plane as a spiral boundary; then train your own, and watch narrow networks stall

7 scenes · builds on The Line That Leans and The Widest Street · after Cybenko, “Approximation by superpositions of a sigmoidal function”, 1989

Coming soon
coming soon

The City That Reads

“Every road shares the blame.”

Neural networks and backpropagation

Each tower is a neuron, each road a weight, each lit window a number it is carrying. The city teaches itself to read digits while you watch.

New ideas backpropagationchain rulesoftmaxmulti-class outputminibatchepoch

  • A digit flowing through 85 towers and 1,384 roads
  • One tower's sum, added up road by road
  • 1,600 training steps to scrub, with the error wave flowing back

7 scenes · builds on Folding Space and The Line That Leans · after Rumelhart, Hinton & Williams, “Learning representations by back-propagating errors”, 1986

Coming soon
coming soon

The Sliding Window

“The moment it moves.”

Convolutional networks

Move a real handwritten digit two pixels and a dense network forgets it. A convolution slides one small filter across the whole image and asks only whether a shape appeared, not where, so the network still recognises a digit after it moves sideways. Up and down it slips, and the film shows why.

New ideas convolutionfilteractivation mapweight sharingpoolingreceptive fieldshift tolerance

  • 1,000 real handwritten test digits (MNIST) on glass plates: move each two pixels and the dense network reads 317 of them, down from 914
  • One 3×3 window sweeping a handwritten 2 while its dot products develop eight learned activation plates, one grain per multiplication
  • The digit slides sideways and every plate slides with it: the CNN keeps 91.0% at 3 px, a flattened CNN 40.8%, the dense network 12.6%

7 scenes · builds on The City That Reads and Folding Space · after LeCun et al., “Backpropagation applied to handwritten zip code recognition”, 1989

Coming soon

Part three · episodes 19–25

Writing

Words become places, words watch each other, and at last it writes. Then it acts, learns what people want, learns to tell cause from coincidence, and makes.

coming soon

A Map of Words

“About the company a word keeps.”

Word embeddings

A word can be a list of numbers learned from the company it keeps. Your browser learns them for the Alice books, then you fly over a map of 10,000 words where closeness is an angle.

New ideas tokenembeddingcosine similaritycontext windowdirections as relations

  • 52,190 words of the Alice books scattering, collapsing into one knot, then bursting apart into company as a model trains live in your browser
  • An island of 10,000 words in contours and ink, its districts lettering themselves from their two commonest words
  • Fourteen words lifting off the map onto a surveyor's plate: king − man + woman lands beside queen, and hot:cold::tall:? fails

7 scenes · builds on The City That Reads and Shadows · after Mikolov et al., word2vec, 2013

Coming soon
coming soon

Every Word Watches

“About who to listen to.”

Self-attention in transformers

Every word watches every other word. Real weights from a real language model, BERT, reading one sentence, drawn as a star chart.

New ideas attentionquery, key, valueattention weightsscaled dot productmultiple headscontextual meaningword order (positional encoding)

  • BERT’s own attention for one sentence, arc by arc
  • Shuffle the words: without positions, dog bites man and man bites dog look the same to it; BERT’s learned positions make the order count
  • Change one word and watch a head swing from animal to street

8 scenes · builds on A Map of Words and The Line That Leans · after Vaswani et al., “Attention is all you need”, 2017

Coming soon
coming soon

The Next Word

“By the surprise of the next word.”

Language models

A real GPT-2 reads "The animal didn't cross the street because it was too" and lights all 50,257 tokens it could say next. Writing is picking one, adding it, and asking again.

New ideas language modelnext-token predictionlogitstemperaturesamplingautoregressioncausal maskperplexity

  • The whole vocabulary as a galaxy of 50,257 lights, burning by probability, warm where "it" would be the animal and rose where it would be the street
  • A seeded number sliding down the stacked lights to land on "big", whose light tears out of the galaxy onto a paper slip
  • Every light re-lighting in a wave from the word just added, then temperature squeezing it into one star or smearing it over thousands

7 scenes · builds on The Line That Leans and A Map of Words · after Shannon, “A mathematical theory of communication”, 1948

Coming soon
coming soon

The Maze That Learns Back

“About the future, a little less each time.”

Reinforcement learning (Q-learning)

No answers, only rewards. An agent wanders a maze, and the value of every square rises backward from the goal until the arrows agree and the path straightens.

New ideas agentstateactionrewardepisodeQ-valueBellman updatediscountexplorationpolicyvalue function

  • A game piece wandering a dawn-lit board for 1,654 moves, the trip counter climbing as 127,025 real moves play back
  • One Bellman update with its real numbers: a guess of −20 nudged to −12.875, wrong about the future by +14.25
  • The value terrain rising backward from the goal until all 108 arrows agree with value iteration, beside the cliff it still falls off

7 scenes · builds on The Next Word and Downhill · after Watkins, “Learning from delayed rewards”, 1989

Coming soon
coming soon

Taught to Help

“By what you would have preferred.”

From next-word model to assistant: fine-tuning and RLHF

The same question goes to two real language models: SmolLM2, trained only to continue text, and its instruction-tuned twin. A reward model learned from people's choices then pulls 64 of the base model's replies on a leash called β, until the light pours into confident wrong answers.

New ideas pre-trainingfine-tuninginstruction tuninghuman preference comparisonsreward modelreinforcement learning from human feedbackKL penalty

  • Two windows on the same 49,152 lights: the base model's reply loops through "1 . 1 . 1", the tuned model's walks across the vocabulary to "Rayleigh scattering"
  • A reward model trained in your browser on 4,000 real human choices, right 66.6% of the time on pairs it never saw, and wrong more often than not when people preferred the shorter reply
  • 64 real replies as paper slips on a desk while 200,000 grains of light pour into the reward model's favourites: "It comes from water." The KL meter counts the cost

8 scenes · builds on The Next Word and The Line That Leans · after Ziegler et al. (2019) and Christiano et al. (2017)

Coming soon
coming soon

Did It Work?

“I can’t tell, until somebody changes something.”

Experiments and causation: randomisation, confounding, Simpson’s paradox

In 1973 Berkeley admitted men at a higher rate across the whole campus, yet women at a higher rate in four of its six largest departments. In an invented shop, keen customers fake a +7.93-point win. A coin flip finds +0.77, and 1,000 resamples say how unsure that is.

New ideas correlation versus causationconfounderrandomisationA/B testSimpson's paradoxconfidence intervalselection bias

  • Berkeley 1973 as 4,526 real applicants: 44.5% of men and 30.4% of women admitted, yet women ahead in four of six departments
  • 100,000 invented shoppers choose for themselves (+7.93 points), then a coin splits them into mirror halves (+0.77, the truth +1.0 inside its interval)
  • 10,000 A/A tests turning red evening by evening: peeking calls a false winner 27.2% of the time, looking once 4.96%

8 scenes · builds on Every Line Has a Score and The Next Word · after Fisher (1925, 1926)

Coming soon
coming soon

From Noise

“About the noise, one step at a time.”

Diffusion models

Dissolve Episode 1's world into static, then teach a small network to undo one sliver of noise at a time. Run it from pure static and 100,000 new points draw that world again, and Episode 1's line fits itself to them.

New ideas generative modelforward noisingnoise predictionreverse processnoise schedule

  • 100,000 grains of silver static on Episode 1's graph paper, lying in a developing tray, step through 100 real denoising passes on your graphics card and develop into a saffron scatter.
  • The network's guess of the noise, drawn as faint arrows that point into the shape, and its error sitting on the exact floor: 0.38, the best any network could do.
  • Episode 1's line refits itself to the drawing: slope 0.894 beside the rule's 0.9, its examples' 0.895 and Episode 1's sixty points' 0.916. Then the print lifts into Episode 1's paper.

7 scenes · builds on The City That Reads and Folding Space · after Ho, Jain & Abbeel, “Denoising diffusion probabilistic models”, 2020

Coming soon