cosmic college · ai 101
ai 101 · thinking in math
Artificial intelligence, taught the way a math class would teach it. Bits, functions, vectors, and learning by small corrections, with simple numbers you can follow by hand. Then a short history of how decades of patient ideas became tools we use every day.
math that talks back
You type a question into a chat window, and a few seconds later a paragraph appears that reads as if someone thought about it. It feels like magic. Underneath, it is arithmetic: multiplication and addition, done an enormous number of times, very quickly, by machines that only understand on and off.
This lesson treats AI the way a math class would. No code and no walls of jargon. We start with bits, build up to functions, vectors, and learning by small corrections, then see how stacking simple pieces gives you something that can write a poem or translate a menu. A short history closes the loop, because every idea here took decades of patient work to arrive.
If you can multiply two decimals and follow a recipe, you can follow every step below.
what computing is
A computer stores everything as bits. A bit is a switch with two states, written 0 and 1. One bit tells you very little. Line bits up and the options multiply: every extra bit doubles the number of patterns you can make.
worked example · counting patterns
patterns = 2number of bits
1 bit gives 2 patterns. 3 bits give 2 × 2 × 2 = 8. 8 bits, one byte, give 256.
In the long-standing ASCII standard, the capital letter A is the number 65, stored as the byte 01000001. Text, photos, and sound are all agreed-upon patterns like this one.
Bits become useful when simple rules combine them. Three rules do most of the work. AND outputs 1 only if both inputs are 1. OR outputs 1 if either input is 1. NOT flips a bit. Wire a few of these together and you can add:
| a | b | sum bit (one or the other) | carry bit (a AND b) |
|---|---|---|---|
| 0 | 0 | 0 | 0 |
| 0 | 1 | 1 | 0 |
| 1 | 0 | 1 | 0 |
| 1 | 1 | 0 | 1 |
Read the last row as 1 + 1 = 10 in binary, which is two. Chain these little adders and you get a circuit that adds any pair of numbers. Chain enough circuits and you get a processor.
That is the central idea of computing: complicated behavior built from very simple rules, repeated and connected. Every program is a function. It takes an input, follows fixed steps, and returns an output.
A computer does a small number of very simple things, extremely fast. The power comes from how many times it does them and how they are arranged.
a model is a function
In math class, a function is a rule that turns an input into an output. The function f(x) = 2x + 1 takes 3 and returns 7. Give it the same input and you always get the same output.
An AI model is a function too. The input might be a photo, a sentence, or a column of sales figures. The output might be a label, the next word, or a forecast. The difference is how the rule gets written. Nobody types it in by hand. The rule contains adjustable numbers, called weights or parameters, and those numbers are set by learning from examples.
same shape · input, rule, output
Both boxes are functions. In the top row a person wrote the rule. In the bottom row the rule is a huge arrangement of multiplications and additions whose numbers were learned from examples.
Take a tiny model that guesses how long a bike ride will take: predicted minutes = w × kilometers. The single weight w is the model's entire knowledge. If w = 4, a 2 km ride is predicted at 8 minutes. Change w and the model changes its mind. Large models work the same way, with billions of weights in place of one.
Model = function + adjustable numbers. Training = choosing the numbers. Using the model = plugging in new inputs.
numbers as vectors
Models only handle numbers, so everything has to become numbers first. The trick is to describe each thing with a list of numbers, called a vector. Each position in the list measures one feature.
Here is a toy version with two made-up features, how alive something is and how much it has an engine, each scored from 0 to 1:
- cat = (0.9, 0.1)
- dog = (0.8, 0.2)
- car = (0.1, 0.9)
Draw each list as an arrow from the origin and similar things point in similar directions. The number that measures this is the dot product: multiply matching positions, then add.
worked example · the dot product
a · b = (a1 × b1) + (a2 × b2)
cat · dog = (0.9 × 0.8) + (0.1 × 0.2) = 0.72 + 0.02 = 0.74
cat · car = (0.9 × 0.1) + (0.1 × 0.9) = 0.09 + 0.09 = 0.18
The bigger score says cat and dog agree more than cat and car do. When arrows have different lengths, dividing the dot product by both lengths gives cosine similarity, a pure measure of direction.
vectors · direction means similarity
Cat and dog sit almost on top of each other, so their dot product is high. Car points toward the engine axis, so its score with either animal is low. Real models use hundreds or thousands of axes, and the geometry works the same way.
Real models learn their own features instead of using ones we name, and each vector has hundreds or thousands of positions instead of two. The geometry still holds: related ideas land near each other. A well-known 2013 result showed word vectors where king − man + woman lands close to queen, a sign that directions in the space can carry meaning.
Keep the dot product in mind. It is the most repeated operation in modern AI, and it comes back twice more in this lesson.
learning by small steps
Back to the bike model, minutes = w × kilometers. Suppose we timed one real ride: 2 km took 6 minutes. We start with a poor guess, w = 1, and let the model learn.
First we need a score for how wrong the model is. A common choice is squared error: (prediction − actual)². Squaring keeps the score positive and punishes big misses more than small ones. That score is called the loss.
Then we ask a simple question: if we nudge w up a little, does the loss go up or down, and how fast? That rate is the slope, also called the gradient. We take a step in the downhill direction, sized by a small number called the learning rate. Then we repeat. That loop is gradient descent.
worked example · gradient descent by hand
loss = (2w − 6)²
slope = 2 × (2w − 6) × 2
new w = w − 0.1 × slope
- start: w = 1. Prediction 2 minutes, error −4, loss 16. Slope = 2 × (−4) × 2 = −16.
- step 1: w = 1 − 0.1 × (−16) = 2.6. Prediction 5.2, error −0.8, loss 0.64.
- step 2: slope = 2 × (−0.8) × 2 = −3.2, so w = 2.6 + 0.32 = 2.92. Prediction 5.84, loss 0.0256.
- step 3: slope = −0.64, so w = 2.984. Loss is about 0.001.
Three steps took the loss from 16 to about 0.001, and w settled toward 3 minutes per kilometer, exactly what the ride showed. Nobody told the model the answer. It followed the slope.
You can check the slope without calculus. At w = 1 the loss is 16. At w = 1.01 the prediction is 2.02, the error is −3.98, and the loss is 15.84. A nudge of 0.01 lowered the loss by about 0.16, which is 16 times the nudge. That is what a slope of −16 means.
gradient descent · rolling downhill
The loss makes a bowl. The dashed line is the slope at the start: steep and pointing downhill to the right. Each step moves w toward the bottom, and the steps shrink as the bowl flattens.
Real training does exactly this with billions of weights at once. The method that computes all those slopes efficiently, layer by layer, is called backpropagation. Every weight gets its own tiny nudge on every step, and the loss drifts down over millions of steps across huge piles of examples.
Learning = measure the error, find the downhill direction, take a small step, repeat. Step size matters: in our example, a learning rate of 0.3 would jump past 3 to w = 5.8, with a loss of 31.36, and bounce farther away each time.
stacked simple functions
One weight can learn one straight-line relationship. The world is rarely a straight line. Neural networks get their flexibility by stacking many small functions, each one easy to understand on its own.
The basic unit is a neuron, a loose borrowing from biology. It does three things: multiply each input by a weight, add the results along with one extra number called a bias, then pass the total through a simple bend. A common bend is ReLU: keep positive numbers as they are and turn negative numbers into 0.
worked example · one neuron
output = max(0, w1x1 + w2x2 + b)
Weights (0.5, −1), bias 4. Inputs (2, 3): 0.5 × 2 + (−1) × 3 + 4 = 1 − 3 + 4 = 2. Positive, so the output is 2.
Inputs (2, 6): 1 − 6 + 4 = −1. Negative, so the bend makes the output 0.
The weighted sum is a dot product, the same one from the cat and dog example, plus the bias.
Why the bend? Without it, stacking adds nothing new. If one layer doubles a number and the next triples it, the pair simply multiplies by 6, one straight line again. The bend lets each layer fold the problem into pieces, and enough pieces together can trace almost any shape. The early single-layer perceptron hit exactly this wall: it could not learn XOR, the rule "one or the other, not both".
a small network · 3 → 4 → 4 → 2
Information flows left to right. The highlighted path is one of many routes a signal can take. Training adjusts all 46 numbers together, using the same downhill steps as the bike model.
That small drawing already has 46 adjustable numbers. Large language models follow the same pattern with many more layers and billions of parameters. Training sets every one of them with the downhill-step method from the last section.
A neural network is simple functions, stacked. Each piece is a dot product and a bend. Depth and width give it flexibility.
why scale mattered
Most of the ideas above existed by the late 1980s. What changed afterward was scale, along three dials that grew together:
- data: the web put enormous amounts of text and images within reach, and projects such as ImageNet (2009) labeled millions of photos for training and testing.
- parameters: bigger networks can hold more patterns and subtler ones.
- compute: graphics chips (GPUs), built to shade millions of pixels in parallel, turned out to be excellent at the same multiply-and-add work that neural networks need.
In 2020 researchers measured something striking. As model size, data, and compute grew together, the loss fell smoothly and predictably, following a simple curve called a power law. That made progress something you could plan. In 2022 a follow-up study found that many large models had been trained on too little data for their size, and that data and parameters should grow roughly in step, at around 20 training tokens per parameter.
worked example · sizing a training run
training compute ≈ 6 × parameters × training tokens
This is a widely used approximation. Assumption: a 1 billion parameter model trained on 20 billion tokens, following the 20-per-parameter guideline.
Compute ≈ 6 × 109 × 2 × 1010 = 1.2 × 1020 arithmetic operations. Assumption: a chip sustaining 1014 useful operations per second. Time = 1.2 × 1020 ÷ 1014 = 1.2 × 106 seconds, about 14 days on one chip.
Make the model 10 times bigger and feed it 10 times more tokens, and the compute grows 100 times.
This is where the track bends toward hardware. Once progress follows compute, chips, power, cooling, and datacenters become part of the story. That is the ground 201 covers.
Three dials: data, parameters, compute. Turn them together and results improve predictably. That predictability is why compute became a planning question.
tokens and next words
Language models read text as tokens: common words, pieces of longer words, punctuation, and spaces. A word like "unbelievable" might become three tokens, such as un · believ · able. Each token has an ID number from a fixed vocabulary, typically tens of thousands of entries, and each ID maps to a learned vector, just like cat and dog earlier.
The training task is almost disarmingly simple: given the tokens so far, predict the next one. The model produces a score for every token in its vocabulary, and those scores are turned into probabilities that add up to 1.
next-token prediction · one step
The model does not store one answer. It assigns a probability to every token, picks one (often a likely one, with a little randomness), appends it, and runs again.
The loss is how surprised the model was by the real next token. If the text actually said "mat" and the model gave mat 0.41, that is a little surprise. If it gave mat 0.01, that is a lot. Gradient descent nudges billions of weights to be a bit less surprised, across trillions of tokens of text.
To write, the model repeats one move: predict a distribution, pick a token, add it to the text, predict again. A paragraph is a few hundred of those loops. To get good at guessing the next word across all kinds of writing, a model has to pick up grammar, facts, and patterns of reasoning along the way. That is why such a plain objective turned out to be so capable.
Transformers, introduced in 2017, add one more idea to the stacked layers: attention. At each layer, every token scores the tokens before it with a dot product and blends in information from the ones that score highest. That is how a word like "it" can find the noun it refers to, many words back.
A language model is a next-token function. It has no lookup table of answers. It generates from patterns stored in its weights, which is why it can be fluent and still wrong. Check the facts that matter.
a brief history
The ideas arrived slowly, with long quiet stretches between the leaps.
timeline · from turing to language models
Grey bands mark the two AI winters, when funding and interest dropped. Dates are the commonly cited years for each milestone.
foundations · 1936 to 1956
In 1936 Alan Turing described an abstract machine that reads and writes symbols on a tape by following simple rules, and showed that such a machine could carry out any step-by-step procedure. It became the blueprint for the idea of a general-purpose computer. In 1943 Warren McCulloch and Walter Pitts described a simplified neuron as a logic unit. In 1950 Turing asked whether machines can think and proposed the imitation game as a way to test it. In 1956 a summer workshop at Dartmouth College gave the field its name: artificial intelligence.
learning, limits, and winters · 1958 to the 1990s
In 1958 Frank Rosenblatt introduced the perceptron, a single layer of weights that learned from examples. In 1969 Marvin Minsky and Seymour Papert laid out what a single layer cannot do, and enthusiasm cooled. Two "AI winters" followed, in the 1970s and again in the late 1980s, when funding and interest dropped after promises ran ahead of results. In between, in 1986, David Rumelhart, Geoffrey Hinton, and Ronald Williams showed how backpropagation could train networks with hidden layers, building on earlier work from the 1970s.
scale arrives · 2012 onward
In 2012 a deep network called AlexNet, trained on two graphics cards, won the ImageNet image recognition challenge by a wide margin. That result convinced much of the field that large networks, large datasets, and GPUs worked together. In 2017 the paper "Attention Is All You Need" introduced the transformer. From 2018 onward, transformer language models grew quickly, from about a hundred million parameters to 175 billion by 2020. In late 2022, chat assistants built on large language models reached a broad public audience.
The big ideas are old. The recent change is scale. Patient research plus abundant data and compute turned math from the 1950s and 1980s into everyday tools.
closing rules of thumb
- Computing is simple rules on bits, repeated and connected.
- A model is a function with adjustable numbers inside. Training picks the numbers.
- Vectors turn things into arrows; dot products measure how much two arrows agree.
- Learning is gradient descent: measure the error, step downhill, repeat.
- Neural networks stack simple pieces, each a dot product and a bend.
- Data, parameters, and compute grew together, and results improved predictably.
- A language model predicts the next token, one loop at a time. It is fluent by design, so check the facts that matter.
- The machinery is math, all the way down, and that is the wonder of it.
Next in this track, 201 · under the hood follows the math onto real hardware and software: chips, memory, power, and what it costs to train and run a model. 301 · datacenters in space asks what changes when that hardware leaves the planet. Both will appear in cosmic college when the drafts are ready.