Learning LabExplorable explanations
← All artifacts
Machine Learning

Nobody Told It Coffee Is Bitter

Ask how a language model knows grammar, or that coffee is bitter, and the answer is that nobody told it. It was only ever asked to guess the next word. Make the guesses, take the nudges, and watch knowledge appear in a handful of numbers, including knowledge that is wrong.

language-modelstrainingembeddingsgeneralisationlesson
LiveInteractive · drag, toggle, run it

Someone at a talk about LLMs asks: if nobody wrote the rules of English into the model, how does it know that in "Sam was drinking coffee and eating chocolate. It was bitter." the bitter thing is probably the coffee?

"It learned patterns from lots of text" is true, and it explains nothing. Here is what learning means for these models, at a size you can check with a pencil.

01The only job: guess the next word

A language model is never handed rules or facts. Training gives it one job, repeated trillions of times: read some words, give a probability to every word that could come next, and get scored on how much probability it put on the word that really did. Here is a whole training text, six sentences long:

Everything the model ever reads
"coffee is bitter"× 3
"coffee is hot"× 1
"chocolate is sweet"× 2
Your call
If the model guessed as well as possible on this text, what would it say after "coffee is"?
Pick an answer to see what happens.

02It doesn't count. It gets nudged.

The model never tallies sentences. It holds a score for each word (the raw numbers are called logits) and turns scores into probabilities with softmax: raise e, about 2.718, to the power of each score, then divide each result by their total so they add up to 100%. Equal scores give equal probabilities, so it starts at 33% each.

After each sentence, every score gets nudged: new score = old score − (probability − target), where the target is 1 for the word that came next and 0 for the others. That one line is what gradient descent works out to for this kind of scoring (with a step size of 1, chosen here to keep the numbers readable).

Your call
All three words start at 33%. The model reads "coffee is bitter" once and takes one nudge. Bitter is now:
Pick an answer to see what happens.
Your call
Keep reading the four sentences, round after round. Where does bitter end up?
Pick an answer to see what happens.

03Why it isn't a lookup table

So far the model kept a separate row of scores for "coffee is". That's a lookup table, and a table has nothing to say about a phrase it never saw. Real models meet unseen sentences all the time.

So instead of rows, every word gets a short list of numbers, its vector (also called an embedding). Here each list holds just two numbers. To guess, the model adds up the vectors of the words it's reading and compares the total with a vector for each candidate next word (multiply matching numbers and add them up: a dot product). The biggest match gets the highest score. Now the nudges land on the word vectors, and every sentence a word appears in nudges the same vector. The new training text:

Everything the model ever reads
"coffee tastes strong"× 2
"espresso tastes strong"× 2
"coffee is bitter"× 3
"chocolate tastes sweet"× 2
"chocolate is sweet"× 2
"candy tastes sweet"× 2
"salt tastes strong"× 2
"soup tastes salty"× 2
"soup is salty"× 2
Your call
"Espresso is" never appears in this text. After training, what will the model put after "espresso is"?
Pick an answer to see what happens.

04Back to Sam's coffee

The sentence from the talk: "Sam was drinking coffee and eating chocolate. It was bitter." A GPT-style model reads it left to right, one word at a time.

Your call
When the model is at the word "It", can it look at "bitter" to work out what "It" means?
Pick an answer to see what happens.

05Which summary is right?

Your call
Pick the sentence you'd give the person who asked the question.
Pick an answer to see what happens.

06The same machinery, a wrong answer

Your call
"Salt is" never appears in the text either. Salt only appeared in "salt tastes strong", exactly like espresso. What will the model put after "salt is"?
Pick an answer to see what happens.

A real model does all of this with billions of numbers instead of a few dozen. Inside an LLM runs a tiny trained GPT, about 30,000 numbers, in your browser so you can inspect it stage by stage, and Attention, From the Ground Up shows how the looking-back part works.

These are toy models. They use whole words where real models use tokens (word pieces), read at most two words of context, and choose from only a few next words. The vectors start from seeded random numbers. This start was picked because the groups are easy to see, and other starts can settle the less-supported guesses differently. The training step is the gradient of cross-entropy with respect to the scores. The attention diagram mentioned in section 4 is from Google's 2017 Transformer announcement, which shows encoder self-attention. Everything on this page is computed in your browser.