Nobody Told It Coffee Is Bitter
Ask how a language model knows grammar, or that coffee is bitter, and the answer is that nobody told it. It was only ever asked to guess the next word. Make the guesses, take the nudges, and watch knowledge appear in a handful of numbers, including knowledge that is wrong.
Someone at a talk about LLMs asks: if nobody wrote the rules of English into the model, how does it know that in "Sam was drinking coffee and eating chocolate. It was bitter." the bitter thing is probably the coffee?
"It learned patterns from lots of text" is true, and it explains nothing. Here is what learning means for these models, at a size you can check with a pencil.
01The only job: guess the next word
A language model is never handed rules or facts. Training gives it one job, repeated trillions of times: read some words, give a probability to every word that could come next, and get scored on how much probability it put on the word that really did. Here is a whole training text, six sentences long:
02It doesn't count. It gets nudged.
The model never tallies sentences. It holds a score for each word (the raw numbers are called logits) and turns scores into probabilities with softmax: raise e, about 2.718, to the power of each score, then divide each result by their total so they add up to 100%. Equal scores give equal probabilities, so it starts at 33% each.
After each sentence, every score gets nudged: new score = old score − (probability − target), where the target is 1 for the word that came next and 0 for the others. That one line is what gradient descent works out to for this kind of scoring (with a step size of 1, chosen here to keep the numbers readable).
03Why it isn't a lookup table
So far the model kept a separate row of scores for "coffee is". That's a lookup table, and a table has nothing to say about a phrase it never saw. Real models meet unseen sentences all the time.
So instead of rows, every word gets a short list of numbers, its vector (also called an embedding). Here each list holds just two numbers. To guess, the model adds up the vectors of the words it's reading and compares the total with a vector for each candidate next word (multiply matching numbers and add them up: a dot product). The biggest match gets the highest score. Now the nudges land on the word vectors, and every sentence a word appears in nudges the same vector. The new training text:
04Back to Sam's coffee
The sentence from the talk: "Sam was drinking coffee and eating chocolate. It was bitter." A GPT-style model reads it left to right, one word at a time.
05Which summary is right?
06The same machinery, a wrong answer
A real model does all of this with billions of numbers instead of a few dozen. Inside an LLM runs a tiny trained GPT, about 30,000 numbers, in your browser so you can inspect it stage by stage, and Attention, From the Ground Up shows how the looking-back part works.