ICS 582Lecture 04Full guide
Word embeddings
The whole lecture on one page, taught concept by concept. Work through the parts in order, mark each concept once you understand it, and open the slide chips when you want the original slides.
- Parts
- 10
- Concepts
- 59
- Slides
- 111
- Reading
- 354 min
Part 01: What words mean: lemmas, senses and synonymy
Why treating words as strings or logical symbols is unsatisfying, what lexical semantics asks of a theory of meaning, how lemmas split into senses, and why perfect synonymy probably does not exist.
6 concepts, slides 1-10
Why this part matters
Every vector model in this lecture, from raw counts to word2vec, is graded against the same exam: does it capture the relations between word meanings that linguists named long before anyone trained an embedding? This part writes that exam. It asks what a word means, how one word splits into several senses, and why two words almost never mean exactly the same thing.
The payoff is practical in three directions. Exams ask you to define sense and polysemy and to explain why water and H₂O are not perfect synonyms. In research, a static embedding gives mouse one vector that blends the rodent and the computer device, which is the core motivation for word sense disambiguation and for contextual embeddings. In real systems, a search engine that expands a query with synonyms can silently shift the sense or the register of what the user asked.
By the end you can
- Explain why vocabulary indices and logical symbols fail as meaning representations, using the one-hot dot product and the DOG example.
- Define lemma, wordform, sense and polysemy using the WordNet entry for mouse.
- State the truth-conditional definition of synonymy and test a candidate pair by substitution.
- Apply the principle of contrast to water and H₂O and similar pairs, naming the dimension (dialect, register or connotation) on which they differ.
- Distinguish synonymy, a relation between senses, from similarity, a graded relation between words, with examples.
Open the vocabulary file of an n-gram model. cat might be entry w_412, dog entry w_977 and spreadsheet entry w_3051. Ask the model which of those three words are alike and it has nothing to say. The numbers are positions in a list, and a list position carries no meaning: 412 is not closer in sense to 977 than to 3051.
Turning each index into a vector does not help. A one-hot vector has a single 1 at the word's index and zeros everywhere else. Two distinct words never share a non-zero position, so their dot product is always 0, whichever pair you pick:
Jurafsky and Martin describe exactly this situation: in the n-gram models of Chapter 3 and in classical NLP applications, the only representation of a word is a string of letters or an index in a vocabulary list. A model trained on the cat sat learns nothing about the dog sat, because the two sentences share no symbol in the position that matters.
The logic class answer, and why it is circular
An introductory logic course offers a different answer: the meaning of dog is the predicate DOG, and the meaning of cat is CAT. Relations between meanings are then stated as axioms written by hand, for instance that every dog is a mammal:
This is more expressive than an index, because the axioms support inference. But the symbol itself still says nothing; it is the word in capital letters. The old semantics joke makes the point. Q: What is the meaning of life? A: LIFE. Jurafsky and Martin attribute it to the semanticist Barbara Partee and call capitalization a pretty unsatisfactory model of meaning. Every relation you want, from similarity to connotation, has to be typed in by a person, and nothing about DOG tells you that it should sit near CAT.
| Representation | What it encodes | What it misses |
|---|---|---|
| String or index w_i | Which word this is: identity, and nothing else | Every relation. cat is exactly as far from dog as from spreadsheet |
| Logical symbol DOG | Whatever axioms someone writes by hand, such as every dog is a mammal | Graded similarity, connotation, and any relation nobody wrote down; the symbol only renames the word |
| Vector (preview of this lecture) | Position in a space learned from how the word is used, so closeness is computable | Sense distinctions, if one vector must serve every sense of the word |
The rest of the lecture fills the third row. Vector semantics represents a word as a point in a space built from the contexts it appears in, and the resulting Embedding makes closeness a number you can compute instead of an axiom someone must write.
Recall
Why are vocabulary indices and logic symbols like DOG unsatisfying meaning representations?
Consider three sentences: Ann bought a car from Bo. Bo sold Ann a car. Ann paid Bo for a car. One event, one car, one transfer of money, described from three positions. Any adequate account of word meaning has to know that buy, sell and pay are tied together this way, and none of the symbol representations from the previous concept does.
Lexical semantics is the linguistic study of word meaning. It is not the study of dictionary definitions one entry at a time; it is the study of how meanings relate to each other. Jurafsky and Martin turn its findings into a list of what a model of word meaning should deliver, and this list is the checklist every later model in the lecture is graded against.
Desiderata from lexical semantics (SLP3 section 5.1)
- Similarity
- cat is similar to dog, and the model should say so without being told
- Antonymy
- hot and cold are opposites on one dimension, temperature, and alike on everything else
- Connotation
- happy carries positive feeling and sad carries negative feeling, beyond what each word refers to
- Perspective
- buy, sell and pay describe one commercial event from the buyer, the seller and the money
- Inference
- from 'Ann sold Bo a car' a question answering system should conclude that Bo bought a car
The first three entries get their own treatment in part 02: similarity as a graded human judgment, Antonymy as opposition on a single feature, and Connotation as affective meaning. Perspective and inference are what make the list more than a thesaurus: a question answering system that is asked who bought the car must connect it to a sentence that only says who sold it.
Keep the list in view as the lecture moves on. Sparse count vectors will turn out to be good at similarity and relatedness and weak at antonymy, because hot and cold occur in the same contexts. Word2vec will improve similarity further and add analogies, yet still give one vector to every sense of a word. Each model earns or loses marks on these rows.
Recall
List the five desiderata for a theory of word meaning, with one example each.
Type mouse into WordNet. The slide quotes two of its meanings: any of numerous small rodents, and a hand-operated device that controls a cursor. The full WordNet 3.0 entry has six: four noun senses and two verb senses. Before naming the parts of this entry, look at how much a single spelling has to carry.
from nltk.corpus import wordnet as wn
for synset in wn.synsets("mouse"):
print(synset.name(), synset.lemma_names())The six synsets NLTK returns for mouse (WordNet 3.0)
- mouse.n.01
- Any of numerous small rodents (the slide's first sense)
- shiner.n.01
- A swollen bruise around the eye; the synset is shiner, black_eye, mouse
- mouse.n.03
- A person who is quiet or timid
- mouse.n.04
- A hand-operated device that controls a cursor (the slide's second sense)
- sneak.v.01
- To go stealthily or furtively
- mouse.v.02
- To manipulate the mouse of a computer
Lemma and wordform
The headword mouse is a Lemma, also called the citation form: the form under which a dictionary lists the word. The plural mice has no entry of its own; it is a wordform of the same lemma. A wordform is any inflected form a lemma takes in running text. The same split holds across verbs and languages, and it is the same lemma idea you met in tokenization and morphology.
| Wordform | Lemma | Inflection |
|---|---|---|
| mice | mouse | Plural noun |
| sang, sung | sing | Past tense and past participle |
| duermes | dormir | Spanish, second person singular present: you sleep |
Sense and polysemy
Each numbered meaning in the entry is a sense: a discrete aspect of the word's meaning. A lemma with several senses is polysemous, and the phenomenon is called Polysemy. WordNet groups senses into synsets, sets of near-synonymous senses that share one gloss and express one concept. That is why the black eye sense appears under the name shiner.n.01: its synset is {shiner, black_eye, mouse}, and WordNet names a synset after its first member.
How do you know two meanings really are separate senses? One practical test is the zeugma. In ?Does Air France serve breakfast and Philadelphia? the two uses of serve (providing food and flying to a city) are forced to share one verb, and the sentence sounds like a pun. That oddness is the evidence for two senses. With one sense, coordination is fine: Air France serves breakfast and lunch.
Why the split matters for systems
Jurafsky and Martin point out that a search for mouse info is ambiguous between a pet owner and a shopper. Deciding which sense a given occurrence uses is the task of word sense disambiguation. And the split sets up a limitation you will meet at the end of this lecture: a Static embedding gives each word type one vector, so the vector for mouse must blend the rodent and the device. A Contextual embedding computes a different vector for each occurrence, which is how modern models separate senses.
Recall
Using mouse and mice, define lemma, wordform, sense and polysemy.
Quick check
Which statement correctly relates a lemma to its senses?
couch and sofa. filbert and hazelnut. car and automobile. vomit and throw up. big and large. Each pair can be swapped in some sentence without anyone noticing a change in what is claimed. That intuition has a precise form.
Two words are synonymous if they can be substituted for each other in any sentence without changing the truth conditions of the sentence, that is, the situations in which the sentence would be true. This is Synonymy in its truth-conditional definition. I sat on the couch and I sat on the sofa are true in exactly the same situations. WordNet even files car and automobile in the same synset, car.n.01.
The slide states the same idea more loosely: synonyms have the same meaning in some or all contexts. All contexts is the strict truth-conditional ideal. Some contexts is what real pairs such as big and large achieve, and that gap is exactly why synonymy is stated between senses rather than words.
The twist: the relation is between senses
Try the substitution with big and large. In Would I be flying on a large or small plane? the swap to big is harmless. In Miss Nelson became a kind of big sister to Benjamin it is not: a large sister is a different claim. The textbook draws the conclusion directly. Synonymy is a relationship between senses rather than words. WordNet makes this visible: big has 17 synsets and large has 11. They share some, such as large.a.01, above average in size, and not others, such as big.s.01, significant. A claim that two words are synonyms is really a claim about one of their senses.
Worked example
Testing a candidate synonym pair
Substitute in several sentences
Take big and large. Swap them in a large plane, a big house, a big decision and my big sister.Check the truth conditions
A large plane and a big plane are true of the same planes. A big decision is an important one, and a large decision is at best odd. My large sister makes a claim about her size, not her age.Find the sense in play
The swaps that succeed all use the size sense (large.a.01). The swaps that fail use senses only big has: significant, and older or grown up.Check register and genre
Even in the size sense, check whether one word belongs to a different style. Here neither is marked, so the pair survives.Result
Of the sentences tested, big and large swap only in the size sense, so they are near-synonyms in that sense, not as words.
Recall
Give the truth-conditional definition of synonymy, and say why it is a relation between senses.
Quick check
Why does 'my large sister' sound wrong while 'a large plane' is fine?
water and H₂O refer to the same substance. Swap them in the glass contains water and the sentence stays true in exactly the same situations. Now imagine a hiking guide that says to carry two litres of H₂O per person. Nothing false was said, and yet the sentence is wrong for its setting. That wrongness is part of what the words mean.
The same thing happens with big and large: even setting aside the older-sibling sense, the pairs that pass the truth test still differ somewhere. Jurafsky and Martin put it carefully: while substitutions between some pairs of words like car and automobile or water and H₂O are truth preserving, the words are still not identical in meaning, and probably no two words are absolutely identical in meaning.
The principle of contrast
The generalization behind this is the Principle of contrast: a difference in linguistic form is always associated with some difference in meaning. The idea has a long history, from Girard in 1718 and Bréal in 1897 to Eve Clark in 1987, who stated it as: every two forms contrast in meaning. Clark used it to explain language acquisition: a child who already knows one word for a thing assumes a new word for the same thing must mean something different. In her words, there are no true synonyms.
Where do apparent synonyms differ, if not in truth? Clark names three dimensions, and the slide's list of politeness, slang, register and genre maps onto them.
- Dialect. autumn and fall, truck and lorry, tap and faucet: the choice tells the listener where the speaker is from.
- Register. die, pass away and pop off; attempt and try. A register is a speech style such as formal, colloquial or technical. Politeness and slang sit here, and so does genre: H₂O is the technical register of a chemistry text.
- Connotation. politician and statesman, skinny and slim: the referent can be the same while the attitude differs. This is Connotation, which part 02 measures.
| Pair | Same truth? | Where they differ | Dimension |
|---|---|---|---|
| water and H₂O | Yes | H₂O belongs to scientific writing; it is odd in a hiking or surfing guide | Genre and register |
| big and large (sister) | No | big has an older or grown-up sense that large lacks | Not a contrast case: a different sense (see concept 4) |
| die and pass away | Yes | pass away is the polite, euphemistic choice; pop off is slang | Register (politeness) |
| politician and statesman | Mostly | statesman praises, politician often does not | Connotation |
| truck and lorry | Yes | American versus British English | Dialect |
The consequence for the rest of the course is terminological: when NLP papers say synonym, they mean approximate synonymy, two senses close enough to substitute in most contexts. It also has an engineering consequence. Query expansion in search, paraphrase generation and data augmentation by synonym replacement all assume substitutability, and the principle of contrast predicts they will shift register, sense or tone some of the time. A paraphraser that rewrites passed away as died has kept the truth and changed the message.
Recall
State the principle of contrast and name the three dimensions along which apparent synonyms usually differ, per Clark.
Recall
Why is water and H₂O not a perfect synonym pair, even though substitution preserves truth?
Quick check
In a hiking guide, what best describes replacing water with H₂O?
car and bicycle are not synonyms. Neither are cow and horse. Yet everyone feels they belong together: both pairs share an element of meaning, a vehicle you ride on a road, a large farm animal. Swap them, though, and the truth changes. I took my car to work describes a different morning from I took my bicycle to work.
This relation is Word similarity. Jurafsky and Martin give the reason it matters: while words don't have many synonyms, most words do have lots of similar words. A model that only knew synonymy would have almost nothing to say about most of the vocabulary; a model of similarity has something to say about every word.
There is a second, quieter shift. Synonymy was a relation between senses, which requires deciding first what the senses of every word are. Similarity is usually stated between words, which avoids committing to a sense inventory at all. That is exactly the quantity vector models compute: one number for a pair of words, later the cosine of the angle between their vectors. Part 02 shows how humans rate it, on datasets such as SimLex-999, and how it differs from relatedness.
| Relation | Holds between | Example | Substitutable? |
|---|---|---|---|
| Synonymy | Senses | couch and sofa | Mostly yes, in the shared sense |
| Similarity | Words | car and bicycle | No, the truth of the sentence changes |
Recall
How does similarity differ from synonymy, and why is it the better target for vector models?
Quick check
Which pair is similar but not synonymous?
Recap
If you remember nothing else
- Treating a word as an index or as a symbol like DOG gives identity, not meaning. Every pair of distinct one-hot words has dot product 0.
- Lexical semantics wants a model that captures similarity, antonymy, connotation, perspective (buy, sell, pay) and inference.
- A lemma (citation form) groups wordforms such as mouse and mice. Its senses are discrete aspects of meaning. Several senses means polysemy.
- WordNet 3.0 lists 4 noun senses and 2 verb senses for mouse. The slide shows only the rodent and the cursor device.
- Synonyms can be substituted without changing truth conditions, and synonymy holds between senses: big and large share size but not older sibling.
- Principle of contrast: every difference in form marks a difference in meaning, so perfect synonyms probably do not exist.
- Apparent synonyms differ in dialect, register or connotation. H₂O belongs to a scientific genre, so it is odd in a hiking guide.
- Similarity (car and bicycle, cow and horse) is a graded relation between words. It is what vector models measure next.
Sources
- Speech and Language Processing, 3rd edition draft, chapter 5: EmbeddingsBookJurafsky and Martin, StanfordDraft of August 19, 2026. Section 5.1 Lexical Semantics: strings and indices, the Partee joke, desiderata, lemma, wordform, senses, synonymy, principle of contrast, similarity(opens in a new tab)
- Speech and Language Processing, 3rd edition draft, Appendix I: Word Senses and WordNetBookJurafsky and Martin, StanfordSense definition, polysemy and homonymy, zeugma, synsets, synonymy between senses with big and large sister(opens in a new tab)
- Vector Semantics and Embeddings (lecture slides)DocsDan Jurafsky, StanfordUpstream of the course deck, including the H20 typo and the 1967 Partee line(opens in a new tab)
- The principle of contrast: A constraint on language acquisitionPaperEve V. Clark, in MacWhinney (ed.), Mechanisms of Language Acquisition, 1987Statement of the principle, dialect, register and connotation, and the rejected Homonymy Assumption (pp. 1 to 5)(opens in a new tab)
- Open English WordNet: mouseDocsOpen English WordNetFour noun senses and two verb senses for mouse(opens in a new tab)
- WordNet Interface (HOWTO)DocsNLTK ProjectThe synsets API used in the mouse example(opens in a new tab)
- Water (CID 962)DocsPubChem, US National Library of MedicineMolecular formula H2O, for the slide errata(opens in a new tab)
Part 02: Similarity, relatedness, antonymy and connotation
The graded relations between word senses (similarity rated by humans, relatedness through semantic fields, antonymy as opposition on one feature) and the affective meaning captured by valence, arousal and dominance.
6 concepts, slides 11-17
Why this part matters
Before we build a single vector, we need to know what a good vector is supposed to capture. This part names the meaning relations that every embedding model in the rest of the lecture is judged against.
Similarity is the target of SimLex-style intrinsic evaluation. Relatedness is what co-occurrence counts actually pick up, whether you wanted it or not. Antonymy is the classic failure case of distributional models. Connotation is the raw material of sentiment and affect lexicons. Exams like to hand you a list of word pairs and ask which relation each one shows, and a research project that reports a score on a word benchmark must know which of these relations that benchmark measures. Part 01 gave us lemmas, senses and synonymy. Here we add the graded, messier relations that sit around them.
By the end you can
- Explain word similarity as a graded, human-rated relation and read SimLex-999 scores.
- Distinguish similarity from relatedness, and identify a semantic field from examples.
- Define antonymy, tell scale or binary opposites from reversives, and explain why antonyms look similar to distributional models.
- Describe connotation and evaluation, and give word sets that differ only in connotation.
- Define valence, arousal and dominance, and interpret NRC VAD scores.
- Label a word pair with the right relation: synonym, similar, related, antonym, or a connotation contrast.
Read these pairs and give each one a number from 0 to 10 for how alike the two meanings are: vanish and disappear, behave and obey, belief and impression, muscle and bone, modest and flexible, hole and agreement. You probably gave the first pair close to 10, the last close to 0, and found the middle harder but not impossible. Hundreds of people did exactly this, and their averages are the numbers below.
| Pair | SimLex similarity (0 to 10) | USF association | POS |
|---|---|---|---|
| vanish / disappear | 9.8 | 2.76 | verb |
| behave / obey | 7.3 | 0.21 | verb |
| belief / impression | 5.95 | 0.10 | noun |
| muscle / bone | 3.65 | 0.13 | noun |
| modest / flexible | 0.98 | 0 | adjective |
| hole / agreement | 0.3 | 0 | noun |
The association column comes from the University of South Florida free-association norms: how often people answer the second word when given the first as a cue. It measures connection, not shared features.
The scores fall away smoothly. There is no point where the pairs stop being similar and start being dissimilar, which is the first thing to notice. Synonymy, from part 01, is close to a yes or no question about two senses. Word similarity is a matter of degree: two words are similar when their meanings share features, and they can share many features, a few, or none. vanish and disappear share nearly all of them. muscle and bone share a few (body tissue, anatomy) and differ on the rest. hole and agreement share essentially nothing.
The second thing to notice is that similarity is a relation between words, not senses. This is the car and bicycle idea from the end of part 01. To say whether two senses are synonyms you need a sense inventory; to ask a person how similar two words feel, you do not. That makes word similarity cheap to collect and directly comparable to anything that produces one number per word pair, which is exactly what a vector model does.
Where the numbers come from
The table is SimLex-999 (Hill, Reichart and Korhonen 2015). It contains 999 pairs: 666 noun pairs, 222 verb pairs and 111 adjective pairs, mixing concrete and abstract words. About 500 Mechanical Turk workers rated them on an integer slider from 0 to 6, and the mean ratings were then linearly rescaled to 0 to 10. The slide shows the rescaled values, which match the released data file exactly. Individual raters disagree a fair amount: the average Spearman correlation between two raters is 0.67, and between one rater and the mean of the others it is 0.78. The average is what is stable.
Why should you care about these particular numbers? Because later in this lecture they become the gold standard for intrinsic evaluation. You compute the cosine between the two vectors of every SimLex pair, rank the pairs by cosine, and report the Spearman rank correlation with the human ranking. A model that agrees with people about which pairs are more alike scores high.
Recall
What scale does SimLex-999 use, and what does 0.3 mean for hole and agreement?
Quick check
Which pair would SimLex-999 annotators rate as most similar?
Take hot and cold. Both are adjectives. Both describe temperature. Both slot into the same frames: hot coffee and cold coffee, hot weather and cold weather, it is too hot today and it is too cold today. Line up everything you know about the two words and they agree on almost every point. They disagree on exactly one: which end of the temperature scale they name.
That is the definition of antonymy. Antonyms are senses that are opposite with respect to only one feature of meaning and otherwise very similar. It sounds paradoxical that opposites are mostly alike, but you cannot be opposite to something unless you are first comparable to it. Hot is not the opposite of Tuesday.
Kinds of opposition
The single opposed feature can be of different types. In the first group the two words name the two values of a binary choice or the two ends of a scale: long and short, fast and slow, big and little. In the second group, the reversives, the two words describe change or movement in opposite directions: rise and fall, up and down. Mohammad, Dorr, Hirst and Turney refine this further (antipodals, complementaries, gradable opposites) and note that many contrasting pairs, such as warm and cold, are not strict opposites at all.
| Kind | Pairs | What is opposed |
|---|---|---|
| Opposite ends of a scale | long / short, fast / slow, hot / cold | A position on one graded dimension (length, speed, temperature) |
| Binary opposition | in / out | Two values with no middle ground |
| Reversive | rise / fall, up / down | The direction of a change or movement |
What humans say, and why models struggle
Because antonyms differ on a feature people care about, human raters call them dissimilar. Because they share everything else, they are among the most strongly associated pairs in the language. SimLex records both, and the gap is striking. Hill and colleagues conclude that antonyms are the most strongly associated word pairs among the finer-grained relations they examined.
| Pair | SimLex similarity | USF association |
|---|---|---|
| night / day | 1.88 | 8.19 |
| old / new | 1.58 | 7.25 |
| short / long | 1.23 | 5.36 |
| bottom / top | 0.70 | 6.96 |
| large / big (synonyms, for contrast) | 9.55 | 0.68 |
Now look ahead. The distributional hypothesis that drives the rest of this lecture says that words in similar contexts have similar meanings. hot and cold occur in almost identical contexts, so a model built on contexts will put them close together. Opposites even co-occur in the same sentence more often than chance would predict (Charles and Miller, cited by Mohammad and colleagues), which pulls them closer still. SLP3 is blunt about the result: automatically distinguishing synonyms from antonyms can be difficult.
Recall
What do antonyms have in common, and why does that matter for distributional models?
Recall
Name the two kinds of antonymy on slide 14, with an example of each.
Quick check
Why do distributional models often place hot close to cold?
A museum shop sells a replica of an ancient vase. A street stall sells a knockoff. Both objects are copies of a real thing, and a description of either would read much the same. But the first word is close to praise and the second is an accusation.
The part of meaning that differs here is connotation: the aspects of a word's meaning tied to a writer's or reader's emotions, sentiment, opinions or evaluations. Some words exist mainly to evaluate: great and love are positive, terrible and hate are negative. Others, like replica and knockoff, describe the same thing while carrying different attitudes toward it. Positive or negative evaluation in language is called sentiment, and connotation is what sentiment analysis, stance detection, and NLP work on political language and consumer reviews all exploit.
Connotation can be measured. Affect lexicons give each word a score, and the one we meet in the next concept, the NRC VAD Lexicon, gives a valence (pleasantness) between 0 and 1. Here is the SLP3 example worked through with its real numbers.
Worked example
Two sets of copies, one difference in feeling
Look up each word's valence
Negative set: fake 0.073, knockoff 0.350, forgery 0.235. Positive set: copy 0.460, replica 0.480, reproduction 0.800.Average each set
Negative mean: (0.073 + 0.350 + 0.235) / 3 ≈ 0.219. Positive mean: (0.460 + 0.480 + 0.800) / 3 = 0.580.Compare, and read the numbers critically
The positive set is about 0.36 higher. But positive is relative here: copy, at 0.460, sits just below the neutral midpoint. And reproduction scores high partly because it is polysemous: its biological sense (having children) is pleasant, and a lexicon with one score per word averages over all senses.Result
Words that refer to nearly the same thing can sit far apart on valence. Other pairs show the same pattern: innocent 0.729 against naive 0.406, great 0.958 against terrible 0.061, love 1.000 against hate 0.031.
Recall
Give two words with nearly the same reference but different connotation, and say how you would measure the difference.
Compare napping and toxic. napping is pleasant and calm, and it puts you in no particular position of control. toxic is unpleasant and agitating. A single positive or negative score would capture the first difference, but not the second, and not the question of who has the power.
Osgood and colleagues (1957) found that people's ratings of words consistently varied along three affective dimensions, now called valence, arousal and dominance:
- Valence is the pleasantness of the stimulus. napping 0.765, toxic 0.008.
- Arousal is the intensity of emotion the stimulus provokes. napping 0.046, toxic 0.885.
- Dominance is the degree of control the stimulus exerts. napping 0.306, toxic 0.492.
Three numbers per word means each word becomes a point in a three-dimensional space. SLP3 calls this the first expression of the idea behind vector semantics: on the 1 to 9 scales of Warriner and colleagues (2013), heartbreak sits at [2.45, 5.65, 3.58]. Part 03 takes that idea and scales it from three hand-chosen dimensions to hundreds learned from text.
How the NRC VAD Lexicon was built
The numbers come from the NRC VAD Lexicon (Mohammad 2018). Version 1 covers about 20,000 English words; the released file has 19,971 entries. Asking people to rate a word on a slider is unreliable, because everyone uses the slider differently. Instead Mohammad used best-worst scaling. An annotator sees four words and picks the one highest on the dimension (say, most pleasant) and the one lowest. Over many such four-word sets, each word is scored by how often it won minus how often it lost.
Comparative judgements are much more consistent than absolute ones. Splitting the annotators into two random halves and correlating the scores each half produces gives split-half reliability of r = 0.95 for valence, 0.90 for arousal and 0.91 for dominance. Version 2, released in March 2025, extends the lexicon to over 55,000 terms (about 10,000 of them multiword phrases) on a -1 to 1 scale, with split-half Spearman of 0.98, 0.97 and 0.96.
| Word | Valence | Arousal | Dominance |
|---|---|---|---|
| love | 1.000 | 0.519 | 0.673 |
| happy | 1.000 | 0.735 | 0.772 |
| toxic | 0.008 | 0.885 | 0.492 |
| nightmare | 0.005 | 0.810 | 0.436 |
| elated | 0.792 | 0.960 | 0.725 |
| frenzy | 0.610 | 0.965 | 0.682 |
| mellow | 0.633 | 0.069 | 0.265 |
| napping | 0.765 | 0.046 | 0.306 |
| calm | 0.875 | 0.100 | 0.282 |
| excited | 0.908 | 0.931 | 0.709 |
| powerful | 0.865 | 0.830 | 0.991 |
| leadership | 0.870 | 0.690 | 0.983 |
| controlling | 0.490 | 0.441 | 0.885 |
| weak | 0.180 | 0.241 | 0.045 |
| empty | 0.188 | 0.183 | 0.081 |
Recall
Define valence, arousal and dominance, and place napping on each.
Recall
How is an NRC VAD score produced, and how reliable is it?
Quick check
In NRC VAD v1, napping scores 0.046 on arousal. What does that tell you?
Step back and look at what we now have. One word can map to many senses: mouse is a rodent or a pointing device. One sense can map to many words: couch and sofa. The mapping between words and concepts is many to many, and on top of it sits a set of relations, some between senses and some between whole words.
| Relation | Level | Graded? | Example | Typical evidence |
|---|---|---|---|---|
| Synonymy | Sense | Rarely exact | couch / sofa | Thesaurus or WordNet synsets |
| Antonymy | Sense | Comes in kinds | hot / cold | WordNet antonym links |
| Similarity | Word | Graded | vanish / disappear | SimLex-999 ratings |
| Relatedness | Word | Graded | coffee / cup | Association norms or WordSim-353 |
| Connotation | Word | Graded | replica / knockoff | NRC VAD Lexicon |
A lemma groups its senses, and polysemy is the fact that it has several. Synonymy and antonymy are relations between senses. Similarity, relatedness and connotation are graded and can be asked of whole words, which is what makes them measurable with ratings and lexicons.
This table is a list of desiderata. Any representation of word meaning we build next should put similar words near each other, keep related words in recognisable neighbourhoods, encode affect in some consistent direction, and ideally tell synonyms from antonyms. Part 03 introduces vectors as that representation. A vector space captures relatedness readily and affect reasonably well; separating true similarity from relatedness is harder, and telling synonyms from antonyms is where it struggles most, as concepts 2 and 3 warned.
Quick check
replica and knockoff differ mainly in which relation?
Recap
If you remember nothing else
- Similarity is graded and word level. SimLex-999 (999 pairs, 0 to 10) runs from vanish/disappear 9.8 down to hole/agreement 0.3.
- Relatedness (association) is broader: coffee/cup are related through a shared event but not similar. Similarity is a special case of relatedness.
- A semantic field is a set of words covering one domain with structured relations: the hospital, restaurant and house fields. Topic models induce fields from text.
- Antonyms are opposite on one feature and alike on the rest. The kinds are binary or scalar opposites (long/short) and reversives (rise/fall).
- Antonyms score low on SimLex similarity but high on association (night/day 1.88 against 8.19), so distributional models tend to put them close together.
- Connotation is affective meaning. Near-synonyms can differ sharply: replica 0.480 against fake 0.073 valence.
- Osgood's three affective dimensions are valence (pleasantness), arousal (intensity) and dominance (control).
- NRC VAD v1 has about 20k words scored 0 to 1 by best-worst scaling. v2 (2025) has over 55k terms on -1 to 1.
- Every relation here is a test that later vector representations must pass.
Sources
- Speech and Language Processing, 3rd edition draft, chapter 5: EmbeddingsBookJurafsky and Martin, Stanford UniversitySection 5.1 lexical semantics: similarity, relatedness, semantic fields, connotation examples, VAD as the first vector semantics(opens in a new tab)
- Speech and Language Processing, 3rd edition draft, appendix I: Word Senses and WordNetBookJurafsky and Martin, Stanford UniversityAntonymy: binary and scalar opposites, reversives, and the difficulty of separating synonyms from antonyms(opens in a new tab)
- Speech and Language Processing, 3rd edition draft, chapter 23: Lexicons for Sentiment, Affect, and ConnotationBookJurafsky and Martin, Stanford UniversityBest-worst scaling procedure, score formula and split-half reliability(opens in a new tab)
- SimLex-999: Evaluating Semantic Models With (Genuine) Similarity EstimationPaperHill, Reichart and Korhonen, Computational Linguistics 41(4), 2015Dataset design, the 0 to 6 slider rescaled to 0 to 10, and why WordSim-353 measures association(opens in a new tab)
- SimLex-999 homepageDocsFelix HillPair counts, inter-annotator agreement; its clothes-closet figure conflicts with the data file, which gives 3.27(opens in a new tab)
- Evaluating WordNet-based Measures of Lexical Semantic RelatednessPaperBudanitsky and Hirst, Computational Linguistics 32(1), 2006Similarity as a special case of relatedness, and Resnik's cars, gasoline and bicycles example(opens in a new tab)
- Computing Lexical ContrastPaperMohammad, Dorr, Hirst and Turney, Computational Linguistics 39(3), 2013Kinds of opposites and contrasting non-opposites(opens in a new tab)
- Integrating Distributional Lexical Contrast into Word Embeddings for Antonym-Synonym DistinctionPaperNguyen, Schulte im Walde and Vu, ACL 2016Distributional models retrieve both synonyms and antonyms as related words(opens in a new tab)
- Obtaining Reliable Human Ratings of Valence, Arousal, and Dominance for 20,000 English WordsPaperSaif M. Mohammad, ACL 2018NRC VAD v1 construction and Table 2 extremes(opens in a new tab)
- NRC Valence, Arousal, and Dominance LexiconDocsNational Research Council CanadaRelease history, v1 and v2 coverage and scales(opens in a new tab)
- NRC VAD Lexicon v2: Norms for Valence, Arousal, and Dominance for over 55k English TermsPaperSaif M. Mohammad, 2025The -1 to 1 scale and split-half reliability of v2(opens in a new tab)
Part 03: Vector semantics and the distributional hypothesis
Defining a word by its contexts (Wittgenstein, Harris, the ongchoi example), combining that with Osgood's meaning as a point in space, and arriving at embeddings, plus why vectors generalize better than word identities and the two kinds (sparse tf-idf, dense word2vec).
6 concepts, slides 18-30
Why this part matters
Every model in this course from here on, word2vec now and BERT and large language models later, begins by turning words into vectors. This part explains why that works at all. Words that keep similar company mean similar things, and once words are points in a space, "similar" becomes something you can measure.
Parts 1 and 2 listed what a model of word meaning should capture: synonymy, similarity, relatedness and connotation. This part introduces the model that meets many of those wishes. It is a core exam topic (state the hypothesis, compare sparse and dense vectors), the basis of retrieval and semantic search in real systems, and the representation behind most NLP research projects you are likely to start.
By the end you can
- State the distributional hypothesis and attribute it to Harris, Firth and Joos, with Wittgenstein's "meaning is use" as its philosophical root.
- Infer the category of an unknown word (ongchoi) from the contexts it shares with known words.
- Represent a word as a point in a space of affective dimensions and compute the distance between two words.
- Define an embedding and read a 2D t-SNE word map without over-reading global distances.
- Explain why vector features generalize to similar unseen words when identity features cannot.
- Compare sparse (tf-idf, PPMI) and dense (word2vec) embeddings on length, sparsity, construction and use.
Take two words, "oculist" and "eye-doctor". Collect every sentence each appears in, and look at the neighbors: eye, examined, prescription, glasses, appointment. The two lists are almost the same. Now compare "oculist" with "lawyer". Some neighbors still overlap (appointment, fee, office), but far fewer. Without a dictionary, the overlap of environments already tells you which pair is closer in meaning.
That observation is the distributional hypothesis. Zellig Harris put it in 1954 using exactly this example: if two words have almost identical environments, meaning neighboring words or the grammatical frames they occur in, we call them synonyms. He then made the claim graded. The difference in meaning between two words corresponds roughly to the amount of difference in their environments. That second sentence matters more than the first, because it turns meaning into a quantity you can estimate by counting.
The idea was in the air in the 1950s. Wittgenstein had argued that, for a large class of cases, the meaning of a word is its use in the language. Joos (1950) described the meaning of a morpheme as the set of conditional probabilities of its occurrence alongside every other morpheme, which is almost a definition of a language model. Firth (1957) gave the line everyone quotes: "You shall know a word by the company it keeps." Firth and Harris are often merged, but they meant different things. Firth cared about situational and cultural context; Harris cared about the formal distribution of words inside text. NLP took Harris's version, because it can be computed from a corpus alone.
| Thinker | Year | Claim | What it contributes |
|---|---|---|---|
| Ludwig Wittgenstein | 1953 | For a large class of cases, the meaning of a word is its use in the language | The philosophical license: stop looking for meaning behind the word and look at how it is used |
| Martin Joos | 1950 | The meaning of a morpheme is the set of conditional probabilities of its occurrence with all other morphemes | A probabilistic statement, decades before anyone could count at scale |
| Zellig Harris | 1954 | Words with almost identical environments are synonyms; the difference in meaning roughly matches the difference in environments | The operational, graded form that NLP actually implements |
| J. R. Firth | 1957 | You shall know a word by the company it keeps | The slogan, from a theory of meaning in situational and cultural context |
This is why the lecture moves to vector semantics. Earlier, lexical semantics gave us a list of relations a good model should respect: synonymy, similarity, relatedness, connotation. Writing those relations down by hand for every pair of words is impossible. The distributional hypothesis says you do not have to: read enough text, record each word's environments, and the relations fall out of the overlaps. Vector semantics is the standard way NLP does this today.
Recall
State the distributional hypothesis, and say who formulated it in the 1950s.
Quick check
Harris (1954) wrote about two words A and B that have almost identical environments. What did he conclude about them?
Suppose you have never seen the word "ongchoi", a recent borrowing into English from Cantonese, and you meet it three times: ongchoi is delicious sauteed with garlic; ongchoi is superb over rice; ongchoi leaves with salty sauces. You do not know what it is yet. But you have read plenty of other sentences: spinach sauteed with garlic over rice, chard stems and leaves are delicious, collard greens and other salty leafy greens.
Put the contexts side by side and the answer is hard to miss. Sauteed, garlic, rice, leaves, delicious and salty all show up around ongchoi and around the leafy greens you already know. Nothing about laptops, contracts or weather. So ongchoi is very likely a leafy green that people cook and eat. It is: the plant is Ipomoea aquatica, a relative of morning glory sometimes called water spinach, and it has other names in Chinese, Malay and Vietnamese.
Contexts that ongchoi shares with known words
- ongchoi
- delicious, sauteed, garlic, superb, rice, leaves, salty, sauces
- spinach
- sauteed, garlic, rice
- chard
- stems, leaves, delicious
- collard greens
- salty, leafy, greens
Worked example
Inferring ongchoi
List the contexts of the unknown word
Around ongchoi: delicious, sauteed, garlic, superb, rice, leaves, salty, sauces.List the contexts of candidate known words
Spinach: sauteed, garlic, rice. Chard: stems, leaves, delicious. Collard greens: salty, leafy, greens. A distractor such as laptop: screen, battery, keyboard.Mark what is shared
Spinach shares 3 context words with ongchoi, chard shares 2, collard greens share 1, and laptop shares 0. Together the greens cover sauteed, garlic, rice, leaves, delicious and salty.Result
Ongchoi sits with the leafy greens and nowhere near laptop. The prediction is "a cooked leafy green", which is right.
This is the computational form of the distributional hypothesis. Define the meaning of a word by its distribution, the neighboring words or grammatical environments it appears in, and then do the obvious thing: count the words in the context of ongchoi and compare those counts with the counts for every other word. A table of such counts, one row per word and one column per context word, is the term-context matrix of part 4, and each row is a word's vector.
Recall
What exactly would a program count to carry out the ongchoi inference?
Here are two words with three numbers each. Heartbreak is [2.45, 5.65, 3.58] and courageous is [8.05, 5.5, 7.38]. The first number is valence (how pleasant), the second is arousal (how intense the emotion), the third is dominance (how much control is exerted). Plot both as points. They sit far apart on valence and dominance, and almost level on arousal: both words are emotionally charged, one pleasantly and with control, the other unpleasantly and with loss of control.
Valence, arousal and dominance ratings on a 1 to 9 scale (Warriner et al. 2013)
- courageous
- [8.05, 5.5, 7.38]
- music
- [7.67, 5.57, 6.5]
- heartbreak
- [2.45, 5.65, 3.58]
- cub
- [6.71, 3.95, 4.24]
The idea of placing a word at a point goes back to Charles Osgood and colleagues in 1957. Osgood asked people to rate words on many bipolar scales, such as good to bad, strong to weak, active to passive, a method he called the semantic differential. Factor analysis showed that most of the variation came from three factors, which he named evaluation, potency and activity. Today the same three are usually called valence, dominance and arousal, the VAD dimensions of a word's connotation from part 2. Osgood noticed that three numbers per word make each word a point in a three-dimensional space, and he proposed that similarity of meaning is nearness in that space. That is the first appearance of vector semantics.
Worked example
How far is courageous from heartbreak?
Subtract coordinate by coordinate
Valence 8.05 − 2.45 = 5.60, arousal 5.5 − 5.65 = −0.15, dominance 7.38 − 3.58 = 3.80.Square and add
5.60² + 0.15² + 3.80² = 31.36 + 0.02 + 14.44 = 45.82.Take the square root
√45.82 ≈ 6.77.Result
For comparison, courageous is 0.96 from music and 3.75 from cub. Heartbreak is the far outlier, and almost all of the gap comes from valence and dominance.
Put the two threads together. Idea 1 is that meaning can be defined by linguistic distribution. Idea 2, from Osgood, is that meaning can be a point in a multidimensional space. Vector semantics joins them: represent each word as a point, but let the word's distribution, not a panel of human raters, decide where the point goes.
Recall
What two 1950s ideas does vector semantics combine?
Recall
Where do the slide 25 numbers really come from, and on what scale?
Quick check
Using the slide 25 ratings and Euclidean distance, which word is farthest from courageous in VAD space?
Look at a map of word vectors squashed onto a page. Good, nice, wonderful, fantastic, amazing and very good huddle in one region. Bad, worst, worse, dislike and not good gather in another. Function words such as to, by, that, is and with sit off by themselves. Nobody placed these words by hand. Their positions came from training on text.
This is vector semantics in its working form. Each word is a vector, a list of numbers, rather than an arbitrary symbol such as the string "good" or the index w45. Similar words end up nearby in what is called semantic space. And the space is built automatically, by seeing which words are nearby in text, so the distributional hypothesis does the placing that Osgood's raters used to do.
Such a vector is called an embedding, because the word is embedded into a space. The term began in the latent semantic analysis community in the late 1990s, where it named the mapping from sparse count space into a smaller dense space, and it later shifted to mean the resulting vector itself. Embeddings are now the standard way to represent word meaning in NLP: practically every modern system, from classifiers to large language models, starts by looking up or computing them. They give a fine-grained model of similarity, a number for every pair of words instead of a yes or no.
Recall
What can and cannot be read from a 2D t-SNE map of word embeddings?
Build a sentiment classifier the traditional way. Feature 5 is "the previous word was terrible", and training learns that it signals a negative review. At test time a review says "awful acting". If awful never appeared in the labeled training data, feature 5 does not fire, no other feature knows about awful, and the classifier has nothing to go on.
Now replace the identity feature with the previous word's embedding. During training the input was terrible's vector, say [35, 22, 17, ...], and the classifier learned weights on those coordinates. At test time awful arrives as [34, 21, 14, ...]. The weights do not care whether the word is the same string; they act on the numbers, and these numbers are nearly the same. The classifier treats awful almost exactly as it treated terrible. It has generalized to a similar but unseen word.
Worked example
How close are terrible and awful?
Dot product
Using only the three coordinates the slide shows: 35×34 + 22×21 + 17×14 = 1190 + 462 + 238 = 1890.Lengths
√(35² + 22² + 17²) = √1998 ≈ 44.70 and √(34² + 21² + 14²) = √1793 ≈ 42.34.Divide
1890 / (44.70 × 42.34) ≈ 0.9986.Result
On these three coordinates the vectors point in almost the same direction, giving a cosine similarity near 1 and an angle of about 3°. Their Euclidean distance is √11 ≈ 3.32, small next to their lengths. Anything the classifier learned about terrible transfers.
| Property | Identity feature | Vector feature |
|---|---|---|
| What is stored | A yes or no: was the previous word exactly terrible | The previous word's vector, for example [35, 22, 17, ...] |
| Test-time match | String equality | Geometry: weights act on every coordinate, so nearby vectors produce nearby scores |
| Unseen similar word (awful) | Feature never fires; the learned weight is wasted | [34, 21, 14] lands almost where terrible did, so the classifier reacts in nearly the same way |
| Number of weights | One per vocabulary word, often 50,000 or more | One per dimension, often 300 |
This matters because of what lecture 2 showed about vocabularies. Zipf's law means most word types are rare, and about half of them appear only once. A classifier trained on a few thousand labeled reviews will meet many words at test time that it never saw with a label. Embeddings, learned from billions of unlabeled words, carry knowledge about those words into the classifier. Dense vectors also need far fewer weights, around 300 per input position instead of one per vocabulary word, and they can place car and automobile close together, which separate identity features never can.
Recall
Why does a word-identity feature fail on an unseen synonym when an embedding feature does not?
Quick check
A sentiment classifier learned that the previous word 'terrible' signals negativity. At test time it meets 'awful', which never appeared in its labeled training data. Why can an embedding-based classifier still handle it?
Picture two vectors for ongchoi. The first has one slot per vocabulary word, tens of thousands of them, and stores how often each word appeared near ongchoi. Garlic, rice and leaves have counts; nearly every other slot is 0. The second has about 300 real numbers, almost none of them zero, some negative, and none of them labeled with a word. Both are embeddings. This lecture covers both families.
Sparse vectors: weighted counts
A sparse vector is built from a simple function of counts. With tf-idf, the classic weighting from information retrieval, the counts come from documents. With PPMI, they come from nearby words, re-weighted so that informative co-occurrences stand out. These vectors are the workhorse of search engines and a strong, cheap baseline in almost any text task. Their dimensions are interpretable, because each one is a specific word or document.
Dense vectors: learned by prediction
A dense vector such as one from word2vec is learned instead of counted. Word2vec trains a simple classifier to predict whether a word is likely to appear near a target word, and keeps the classifier's weights as the embedding. The labels come free from running text, which is self-supervision. Both families give one fixed vector per word type, a static embedding. Later in the course, contextual embeddings such as BERT compute a different vector for each occurrence of a word.
| Property | Sparse (tf-idf, PPMI) | Dense (word2vec) |
|---|---|---|
| Length | |V|, tens of thousands | 50 to 1000 |
| Zeros | Mostly zero | Mostly non-zero, can be negative |
| How built | Weighted co-occurrence counts | A classifier trained to predict whether a word appears near the target |
| What one dimension means | A specific context word or document | Nothing individually |
| Typical use | Information retrieval, a strong baseline | Input features for neural NLP |
| Synonyms such as car and automobile | Separate, unrelated dimensions | Can end up with similar coordinates |
Recall
Contrast sparse and dense embeddings on length, content and construction.
Quick check
Which statement correctly contrasts the two kinds of embeddings on slide 30?
Recap
If you remember nothing else
- Distributional hypothesis: words in similar environments have similar meanings, roughly in proportion to how similar the environments are (Harris 1954, Firth 1957, Joos 1950; Wittgenstein 1953: meaning is use).
- Ongchoi shares sauteed, garlic, rice, leaves, delicious and salty with spinach, chard and collards, so it is a leafy green. It is Ipomoea aquatica, water spinach.
- Osgood (1957) treated a word's connotation as a point in a space of a few rated dimensions, and similarity as distance. The slide 25 numbers are Warriner et al. (2013) ratings on a 1 to 9 scale.
- Vector semantics combines both ideas: a word is a point in a multidimensional space built from its distribution. That vector is an embedding.
- The slide 27 map is a 2D t-SNE projection of 60-dimensional sentiment-trained embeddings (Li et al. 2015), not the embedding itself.
- Vector features generalize: [34, 21, 14] is almost parallel to [35, 22, 17] (cosine ≈ 0.999), while an identity feature needs the exact word.
- Sparse vectors (tf-idf, PPMI) are |V| long, mostly zero, and built from counts. Dense vectors (word2vec) have 50 to 1000 non-zero entries learned by a classifier predicting neighbors. Contextual embeddings come later.
Sources
- Speech and Language Processing, Ch. 5: EmbeddingsBookJurafsky and Martin, StanfordMain textbook: ongchoi, VAD, summary and historical notes(opens in a new tab)
- Speech and Language Processing, 3rd ed. draft (Aug 2024), Ch. 6BookJurafsky and Martin, StanfordFigure 6.1 caption, the source of the slide 27 map(opens in a new tab)
- Ludwig WittgensteinArticleStanford Encyclopedia of PhilosophyFull Philosophical Investigations §43 quote(opens in a new tab)
- Distributional StructurePaperHarris, Word 10 (1954), Taylor & Francis(opens in a new tab)
- What company do words keep? Revisiting the distributional semantics of J.R. Firth & Zellig HarrisPaperBrunila and LaViolette, NAACL 2022(opens in a new tab)
- Quote Origin: You Shall Know a Word by the Company It KeepsArticleQuote InvestigatorProvenance of the Firth 1957 line(opens in a new tab)
- The Measurement of MeaningBookOsgood, Suci and Tannenbaum, University of Illinois Press(opens in a new tab)
- Norms of valence, arousal, and dominance for 13,915 English lemmasPaperWarriner, Kuperman and Brysbaert, Behavior Research Methods (2013)The real source of the slide 25 numbers(opens in a new tab)
- Visualizing and Understanding Neural Models in NLPPaperLi, Chen, Hovy and Jurafsky, NAACL 2016(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov et al., arXiv 2013(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov et al., NeurIPS 2013(opens in a new tab)
- From Frequency to Meaning: Vector Space Models of SemanticsPaperTurney and Pantel, JAIR 37 (2010)(opens in a new tab)
- Embedding ProjectorDocsTensorFlowExplore real embedding neighborhoods(opens in a new tab)
Part 04: Term matrices and cosine similarity
Building term-document and word-word (term-context) co-occurrence matrices, reading rows and columns as vectors, and measuring similarity with the dot product and its length-normalized form, the cosine.
5 concepts, slides 31-44
Why this part matters
The previous part said a word is known by the company it keeps. This part turns that slogan into arithmetic: count tables, row and column vectors, and one number, the cosine, that says how alike two vectors are. Every later topic in the lecture stands on it. tf-idf and PPMI reweight these same matrices, word2vec vectors are compared with this same cosine, and the retrievers behind RAG systems and vector databases rank by cosine or by a dot product on normalized vectors.
The argument runs in five steps. First a table of counts over Shakespeare plays, read by columns as documents. Then the same kind of table read by rows as words, and a second table that counts neighbors instead of documents. Then the dot product, the obvious way to compare two rows, and the flaw that makes it reward frequent words. Then the cosine, which removes length and keeps only direction. Finally the full computation for cherry, digital and information by hand, with the slide's one rounding slip flagged.
By the end you can
- Build and read a term-document matrix and say what its rows and columns represent.
- Explain how a word-word (term-context) matrix is filled from a context window and why it is sparse.
- Compute a dot product and a vector length, and explain why the raw dot product favors frequent words.
- Derive cosine as the normalized dot product and state its range, including why counts give 0 to 1.
- Compute cos(cherry, information) and cos(digital, information) by hand and interpret them as angles.
Take four Shakespeare plays and four words, and count. As You Like It uses battle once, good 114 times, fool 36 times and wit 20 times. Twelfth Night has no battle at all but 58 fools. Julius Caesar and Henry V, the two histories, have almost no fools and plenty of battles. Put the counts in a grid with one row per word and one column per play, and you already have a usable representation of each play.
| Word | As You Like It | Twelfth Night | Julius Caesar | Henry V |
|---|---|---|---|---|
| battle | 1 | 0 | 7 | 13 |
| good | 114 | 80 | 62 | 89 |
| fool | 36 | 58 | 1 | 4 |
| wit | 20 | 15 | 2 | 3 |
This grid is a Term-document matrix. In general it has |V| rows, one per word in the vocabulary, and |D| columns, one per document, and the cell in row w and column d counts how often w occurs in d. Each column is then a list of |V| numbers: a vector. As You Like It becomes [1, 114, 36, 20], a point in a 4-dimensional space whose axes are the words. This is the basic move of Vector semantics, applied first to documents rather than words.
Each play as a column vector over (battle, good, fool, wit)
- As You Like It (comedy)
- [1, 114, 36, 20]
- Twelfth Night (comedy)
- [0, 80, 58, 15]
- Julius Caesar (history)
- [7, 62, 1, 2]
- Henry V (history)
- [13, 89, 4, 3]
Notice what was thrown away. The counts do not record where in the play a word appeared, or what came before it. A document reduced to word frequencies with order discarded is a bag of words, and it is the oldest representation in information retrieval. It sounds crude, but for the question "what is this document about?" it is surprisingly strong, because topic is carried mostly by which words occur and how often.
Seeing the vectors
Four dimensions cannot be drawn, so pick two: fool on the horizontal axis and battle on the vertical one. The comedies become long arrows lying almost flat along fool (As You Like It at [36, 1], Twelfth Night at [58, 0]), while the histories become short arrows pointing up along battle (Julius Caesar at [1, 7], Henry V at [4, 13]). Comedies point one way, histories another. That is the whole promise of the vector view: similar documents have similar columns, so they point in similar directions.
Gerard Salton turned this picture into the vector space model of information retrieval: represent every document and every query as a vector of term weights, measure how close the query vector is to each document vector, and return the closest documents first (Salton, Wong and Yang 1975). A web search for "battle" is, in this model, a vector with a single non-zero entry, and Henry V wins because its column leans hardest in that direction. The measure of closeness is the cosine you will meet in concept 4. Here is a preview of what it says about the plays, computed once on the two plotted dimensions and once on all four.
| Pair | fool and battle only | All four words |
|---|---|---|
| As You Like It and Twelfth Night | 0.9996 | 0.95 |
| Julius Caesar and Henry V | 0.988 | 0.999 |
| As You Like It and Julius Caesar | 0.169 | 0.945 |
| Twelfth Night and Julius Caesar | 0.141 | 0.809 |
On the plane, the pattern is crisp: comedies with comedies near 1, comedy with history near 0.15. On all four dimensions every pair lands between about 0.81 and 0.999, and As You Like It looks 0.945 similar to Julius Caesar. The culprit is good: it is frequent in every play, so it dominates every vector and makes all of them point roughly the same way. A word that occurs everywhere carries no information about which document you are in, and the next part's tf-idf exists precisely to turn its weight down.
Recall
Why does As You Like It look 0.945 similar to Julius Caesar on all four words, but only 0.169 on fool and battle?
Turn the same Shakespeare table sideways and read it by rows. battle becomes [1, 0, 7, 13]: a word that shows up a little in the comedies and a lot in Julius Caesar and Henry V. fool becomes [36, 58, 1, 4], the opposite profile. Without any dictionary, the rows already say that battle is a history word and fool is a comedy word, and two words with similar rows occur in similar documents.
Counting neighbors instead of documents
Documents are a coarse unit of context. A play contains thousands of words, so two words sharing a play says little more than that they share a topic. The finer alternative is to count, for each target word, which words appear right next to it. Slide the target through a large corpus, and every time it occurs, look at the words within a fixed window on each side (SLP uses ±4) and add 1 to the cell for each of them.
The result is a Term-context matrix, also called a word-word or word-context matrix. It is square, |V| x |V|: rows are target words, columns are context words, and the cell (w, c) counts how often c appeared inside the window around w. Here are four rows from Wikipedia counts, restricted to five of the context columns.
| Target | computer | data | result | pie | sugar |
|---|---|---|---|---|---|
| cherry | 2 | 8 | 9 | 442 | 25 |
| strawberry | 0 | 0 | 1 | 60 | 19 |
| digital | 1670 | 1683 | 85 | 5 | 4 |
| information | 3325 | 3982 | 378 | 5 | 13 |
The rows sort themselves into two families. cherry and strawberry have their mass in pie and sugar; digital and information have theirs in computer and data. Plot digital at [1683, 1670] and information at [3982, 3325] on the data and computer axes and the two arrows point almost the same way, about 5° apart, even though information is more than twice as long. That is the Distributional hypothesis made concrete: two words are similar when their context vectors are similar. Each row is an Embedding of its word, a sparse one built by counting, which later parts will reweight and then replace with short dense vectors.
Ahead of concept 4, here is what the cosine, a 0 to 1 score of how closely two rows point the same way, says about these rows.
| Pair | Cosine | Why |
|---|---|---|
| cherry and strawberry | 0.969 | Both live in the pie and sugar columns |
| digital and information | 0.996 | Both live in the computer and data columns |
| cherry and digital | 0.019 | Almost no shared mass |
| strawberry and information | 0.003 | Almost no shared mass |
Two practical facts follow from the shape. First, almost every cell is zero: most of the |V| words never appear within four words of cherry. These are sparse vectors, and real systems store only the non-zero entries (a compressed sparse row format in SciPy, for example). Second, |V| is not the full vocabulary of the corpus. SLP notes that it is usually the 10,000 to 50,000 most frequent words, and that keeping more than about 50,000 rarely helps (SLP 5.3).
| Matrix | Shape | Cell | Similarity it captures |
|---|---|---|---|
| Term-document | |V| x |D| | Count of the word in the document | Topical: which texts a word appears in |
| Word-word (term-context) | |V| x |V| | Count of the context word in a window around the target | Closer, more substitutable similarity |
Recall
What are the rows and columns of a term-document matrix, and of a term-context matrix?
Quick check
A term-document matrix over 37 plays and 20,000 word types has which shape?
Which word is closer to fool: good or wit? Multiply the Shakespeare rows entry by entry and add. For good and fool that gives 9162; for fool and wit only 1604. By this measure good is almost six times closer to fool than wit is, which is clearly wrong: wit and fool are both comedy words, while good is just everywhere.
The measure is the Dot product, the most natural way to compare two vectors:
It is large when both vectors have large values in the same dimensions, so it does behave like a similarity measure. Vectors whose non-zero entries sit in different dimensions get 0: they are orthogonal, which for count vectors means the two words never share a context. The trouble is that a product of two numbers grows with either number. The dot product therefore depends on how big the vectors are, and the size of a vector is its Vector length:
Worked example
Why good beats wit as fool's neighbor
The rows
good [114, 80, 62, 89], fool [36, 58, 1, 4], wit [20, 15, 2, 3] over (As You Like It, Twelfth Night, Julius Caesar, Henry V).Dot products
good · fool = 114×36 + 80×58 + 62×1 + 89×4 = 4104 + 4640 + 62 + 356 = 9162. fool · wit = 36×20 + 58×15 + 1×2 + 4×3 = 720 + 870 + 2 + 12 = 1604.Lengths
|good| = √31161 ≈ 176.5, |fool| = √4677 ≈ 68.4, |wit| = √638 ≈ 25.3. good is seven times longer than wit simply because it is a more frequent word.Divide out the lengths
9162 / (176.5 × 68.4) ≈ 0.759 and 1604 / (68.4 × 25.3) ≈ 0.929.Result
The raw dot product ranks good first, only because good is long. Once length is divided out, wit is the closer neighbor (0.929 against 0.759), which matches intuition. That division is the cosine of the next concept.
The general lesson is SLP's: the dot product favors long vectors, and more frequent words have longer vectors, because they co-occur with many words and do so many times. Words such as of, the and you would therefore come out as the nearest neighbor of almost everything. Information retrieval met the same problem with documents: a long document has larger counts everywhere, and a raw dot product with a query rewards it for being long rather than for being on topic (Manning, Raghavan and Schütze, "Dot products").
Recall
Why does the raw dot product favor frequent words?
Run an experiment on digital. Suppose the corpus were twice as large, so every count in its row doubles, from [5, 1683, 1670] to [10, 3366, 3340] over (pie, data, computer). Its dot product with information doubles too, from 12,254,481 to 24,508,962. Nothing about the meaning of digital changed, yet the similarity score did. The cure is to measure the angle between the vectors rather than their overlap, and the angle does not move at all.
Geometry supplies the formula. For any two vectors, the dot product equals the product of their lengths times the cosine of the angle θ between them. Solve for the cosine and you get Cosine similarity:
The middle form is the useful one to remember. Dividing a vector by its own length gives a unit vector, a vector of length 1 pointing the same way: [3, 4] has length 5, so its unit vector is [0.6, 0.8], and 0.6² + 0.8² = 1. The cosine is simply the dot product of the two unit vectors. Length has been removed before the comparison, so only direction is left.
Three properties follow. The range is -1 to 1: 1 for vectors pointing the same way, 0 for orthogonal vectors, -1 for opposite directions. Scaling is invisible: for any c > 0, cos(cv, w) = cos(v, w), because the c appears once in the numerator and once in |cv| = c|v| and cancels (SLP exercise 5.2). And for raw counts the range shrinks to 0 to 1, because no count is negative.
This is also why real systems often skip the division at query time. If every stored vector is normalized to length 1 once, in advance, then a plain dot product is the cosine. scikit-learn documents exactly this: cosine_similarity is the normalized dot product, and on L2-normalized data it is equivalent to linear_kernel. Vector databases that offer an "inner product" metric rely on the same identity, and it is why the embedding vectors of many retrieval models come out already normalized.
Recall
Compute cos([1, 0], [1, 1]).
Recall
What happens to cos(v, w) if v is multiplied by 3? And for unit vectors, how do the dot product and the cosine relate?
Quick check
Why do we divide the dot product by the two vector lengths?
Quick check
For raw co-occurrence count vectors, what range can the cosine take?
Now do the whole computation once by hand, on three words and three context dimensions (pie, data, computer): cherry [442, 8, 2], digital [5, 1683, 1670] and information [5, 3982, 3325]. The question is which of cherry and digital is closer to information under the cosine.
Worked example
cos(cherry, information) and cos(digital, information)
Lengths
|cherry| ≈ 442.08, |digital| ≈ 2370.95, |information| ≈ 5187.68. Each is the length, the square root of the sum of squared counts.cherry and information
Dot product (numerator) 442×5 + 8×3982 + 2×3325 = 2210 + 31856 + 6650 = 40716. Cosine 40716 / (442.08 × 5187.68) ≈ 0.0178, an angle of about 89.0°.digital and information
Numerator 5×5 + 1683×3982 + 1670×3325 = 25 + 6701706 + 5552750 = 12254481. Cosine 12254481 / (2370.95 × 5187.68) ≈ 0.9963, an angle of about 4.9°.Result
cos(digital, information) ≈ 0.996 and cos(cherry, information) ≈ 0.018. digital is almost exactly aligned with information; cherry is nearly at a right angle to it. For completeness, cos(cherry, digital) ≈ 0.018 as well.
Every intermediate quantity, for checking your own arithmetic
- |cherry|
- √(442² + 8² + 2²) = √195432 ≈ 442.08
- |digital|
- √(5² + 1683² + 1670²) = √5621414 ≈ 2370.95
- |information|
- √(5² + 3982² + 3325²) = √26911974 ≈ 5187.68
- cherry · information
- 442×5 + 8×3982 + 2×3325 = 2210 + 31856 + 6650 = 40716
- digital · information
- 5×5 + 1683×3982 + 1670×3325 = 25 + 6701706 + 5552750 = 12254481
- cos(cherry, digital)
- ≈ 0.018
The picture makes the numbers obvious. Plot the words on two of the dimensions, pie upward and computer to the right. cherry stands nearly upright, because nearly all of its mass is in pie. digital and information both lie almost flat along computer. The angle between cherry and information is close to a right angle, so its cosine is close to 0; the angle between digital and information is a sliver, so its cosine is close to 1.
Recall
Without a calculator, why must cos(cherry, information) be small?
Quick check
Using counts over (pie, data, computer), which word is closest to information?
Recap
If you remember nothing else
- A term-document matrix is |V| x |D|: columns are document vectors, and rows are word vectors over documents.
- A word-word (term-context) matrix is |V| x |V| and counts context words in a window such as ±4. Its rows are sparse word vectors.
- The dot product Σ v_i w_i is high when both vectors are large in the same dimensions, but it grows with vector length, so frequent words win.
- Cosine = v·w / (|v||w|) is the dot product of unit vectors. It ignores length and measures only the angle.
- Cosine runs from -1 to 1 in general and from 0 to 1 for counts. It is unchanged by scaling a vector by a positive number.
- Worked example: cos(digital, information) ≈ 0.996 and cos(cherry, information) ≈ 0.018 (the slide's .017 is a rounding slip).
Sources
- Speech and Language Processing, chapter 5: Embeddings, sections 5.3 and 5.4BookJurafsky and Martin, draft of August 2026Word-context matrix, the ±4 window, |V| of 10,000 to 50,000, the dot product favoring long vectors, the cosine formula, unit vectors, the 0 to 1 range for counts, and the printed .018 and .996(opens in a new tab)
- Speech and Language Processing, chapter 11: Information Retrieval and Retrieval-Augmented Generation, section 11.1.1BookJurafsky and Martin, draft of August 2026Term-document matrix, the Shakespeare plays, figure 11.5 (which also carries the 40 tick) and bag of words(opens in a new tab)
- Introduction to Information Retrieval: Dot productsBookManning, Raghavan and Schütze, Cambridge University PressCosine as length normalization of document vectors and the document length problem(opens in a new tab)
- Introduction to Information Retrieval: Queries as vectorsBookManning, Raghavan and Schütze, Cambridge University PressRanking documents by query-document cosine and top-K retrieval(opens in a new tab)
- A vector space model for automatic indexingPaperCommunications of the ACM 18(11):613 to 620, 1975 (Salton, Wong and Yang)The vector space model of information retrieval(opens in a new tab)
- From Frequency to Meaning: Vector Space Models of SemanticsPaperJournal of Artificial Intelligence Research 37, 2010 (Turney and Pantel)Term-document, word-context and pair-pattern matrices compared(opens in a new tab)
- sklearn.metrics.pairwise.cosine_similarityDocsscikit-learn documentationCosine as the normalized dot product, its equivalence to linear_kernel on L2-normalized data, and sparse input(opens in a new tab)
Part 05: Weighting with tf-idf
Why raw frequency is a poor representation, how log-scaled term frequency and inverse document frequency combine into tf-idf, and a full worked tf-idf table for the Shakespeare example.
5 concepts, slides 45-52
Why this part matters
Part 04 built count vectors and compared them with cosine. Those vectors have a flaw you can see the moment you look at real counts: the words with the biggest numbers are the, it and good, words that sit near everything and therefore say nothing. Any similarity computed from raw counts is dominated by them.
This part fixes the term-document side of that problem with tf-idf, the weighting that has been the baseline of information retrieval since the 1970s and still ships as the default sparse retriever and text feature extractor. Its descendant BM25 is the keyword half of most retrieval-augmented generation pipelines you will meet in research. Exams like this material because it is computable by hand: the difference between document frequency and collection frequency, and a tf-idf cell worked out from a count, are classic questions. The next part does the same job for term-term matrices with PPMI.
By the end you can
- Explain why raw co-occurrence counts over-weight frequent, uninformative words.
- Compute log-scaled term frequency and contrast the log10(count + 1) and 1 + log10(count) variants.
- Distinguish document frequency from collection frequency using Romeo and action.
- Compute idf with N = 37 and explain why a word in every document gets weight zero.
- Reproduce any cell of the slide 52 tf-idf table and explain how the choice of document changes the weights.
Build a Term-context matrix from a large corpus and look at the row for apricot. The context sugar has a healthy count there, and that is genuinely useful: apricots are sweet things you cook with sugar, and a word whose row also has a high sugar count, such as peach, is probably similar. Now look further along the same row. The contexts the, it and they have counts many times larger than sugar. They have similarly huge counts in the row for digital, for information, for every word in the vocabulary.
That is the paradox. Frequency is clearly informative, since co-occurring often is exactly how sugar earns its place. But frequency is not proportional to information. A Dot product sums products dimension by dimension, so a dimension where both vectors carry a count in the thousands swamps a dimension where both carry a count of twenty. Even after normalising by length, Cosine similarity on raw counts mostly measures how much two words share the ubiquitous dimensions, and every word shares those. What we want instead is a weight with two pressures in it: reward a term that is frequent here, and penalise a term that is frequent everywhere.
Two classic reweightings
Which reweighting you pick depends on which matrix you have. For a Term-document matrix, where columns are documents, the standard answer is tf-idf. Jurafsky and Martin describe it as "the product of two terms, the term frequency tf and the inverse document frequency idf", and note that the "-" is a hyphen, not a minus sign. The first factor rewards local frequency; the second penalises spread across the collection.
For a term-term matrix, where both rows and columns are words, the standard answer is Pointwise mutual information. It compares how often two words actually appear together with how often they would if they were independent. Words like good and great earn a high PMI only if they meet more often than their individual frequencies predict, so a context like the that meets everything at the expected rate scores near zero. Part 06 develops PMI in full.
| Scheme | Matrix it suits | Question it answers |
|---|---|---|
| tf-idf | Term-document | Is this term distinctive for this document? |
| PMI | Term-term (word-context) | Do these two words co-occur more than chance would predict? |
The idea behind the idf half is older than vector semantics. Karen Sparck Jones argued in 1972 that "matches on less frequent, more specific, terms are of greater value than matches on frequent terms", and proposed weighting each term by how few documents it appears in. Every term-weighting scheme in this lecture, PPMI included, is a variation on her insight.
Recall
Why does a cosine between two raw count vectors tend to be high for almost any pair of words?
Suppose a word appears 0, 1, 9, 99 or 999 times in a document. Under the slides' formula those counts become a term frequency of 0, 0.301, 1, 2 and 3. A hundredfold increase in occurrences, from 9 to 999, adds only 2 to the weight.
Raw count to log-scaled term frequency, tf = log10(count + 1)
- count 0
- tf = log10(1) = 0
- count 1
- tf = log10(2) = 0.301
- count 9
- tf = log10(10) = 1
- count 99
- tf = log10(100) = 2
- count 999
- tf = log10(1000) = 3
The simplest term frequency would be the raw count itself, tf = count(t, d). The trouble is that significance does not grow linearly with repetition. Manning, Raghavan and Schütze put it plainly: "It seems unlikely that twenty occurrences of a term in a document truly carry twenty times the significance of a single occurrence." The second mention of battle in a play tells you a lot (this play has a battle in it); the hundredth mention adds far less. So we squash the count with a logarithm.
The +1 is there because log 0 is undefined. Adding one before taking the log sends a count of zero to log10(1) = 0, which is exactly the weight an absent term should have, while barely changing large counts.
The other formula you will meet
The current SLP3 draft (now Chapter 11, on retrieval and RAG) and the IR book both use a slightly different squash. SLP3 notes in a footnote that log10(count + 1) is "this alternative formulation" used in its earlier editions, which is where the slides come from. Neither is an error; they are two members of the same sublinear family.
The two agree on the shape but not the numbers. Under the variant, a count of 1 gives 1 + log10 1 = 1 rather than 0.301, a count of 10 gives 2 rather than 1.041, and a count of 7 gives 1.845 rather than 0.903. scikit-learn offers a third version through sublinear_tf=True, which uses 1 + ln(tf) with the natural log.
Recall
Compute tf = log10(count + 1) for counts 0, 9 and 99.
Across the complete works of Shakespeare, the words Romeo and action each occur exactly 113 times. By raw volume they are identical. Yet every occurrence of Romeo sits in one play, Romeo and Juliet, while action is spread across 31 different plays. If you are handed a document and told it contains Romeo, you know which play it is. If you are told it contains action, you have learned almost nothing.
The two numbers in that story have names. The total number of times a term occurs across the whole collection is its Collection frequency, cf. The number of documents that contain the term at least once is its Document frequency, df_t. The IR book defines them exactly that way and concludes that for discriminating between documents it is "better to use a document-level statistic ... than to use a collection-wide statistic". SLP3 gives the reason: "Terms that occur in only a few documents are useful for discriminating those documents from the rest of the collection."
| Word | cf | df | idf |
|---|---|---|---|
| Romeo | 113 | 1 | log10(37 / 1) = 1.57 |
| action | 113 | 31 | log10(37 / 31) = 0.077 |
The Shakespeare pair is not a curiosity. The IR book finds the same pattern in the Reuters newswire collection, where try and insurance have almost the same collection frequency but very different document frequencies. Insurance clusters in the articles that are actually about insurance; try is sprinkled through everything.
| Word | cf | df |
|---|---|---|
| try | 10,422 | 8,760 |
| insurance | 10,440 | 3,997 |
Recall
Romeo and action both have collection frequency 113. Give their df values and say which gets the larger idf with N = 37.
Quick check
Romeo and action both occur 113 times across Shakespeare. Why does Romeo receive a much higher idf?
Treat each of Shakespeare's 37 plays as a document and walk down a ladder of words ordered by how many plays they appear in. Romeo is in one play, salad in two, Falstaff in four, forest in twelve, battle in twenty-one, wit in thirty-four, fool in thirty-six, and good and sweet in all thirty-seven. We want a weight that is large at the top of the ladder and vanishes at the bottom.
The ratio N / df_t of collection size to Document frequency does that: it is 37 for Romeo and 1 for good. Raw ratios grow too fast, though (a word in one document out of a million would get a ratio of a million), so we take the logarithm, exactly as we did for tf. The result is the inverse document frequency.
| Word | df | idf |
|---|---|---|
| Romeo | 1 | 1.57 |
| salad | 2 | 1.27 |
| Falstaff | 4 | 0.966 (slide prints 0.967) |
| forest | 12 | 0.489 |
| battle | 21 | 0.246 |
| wit | 34 | 0.037 |
| fool | 36 | 0.012 |
| good, sweet | 37 | 0 |
Two boundary values are worth memorising. The largest possible idf is log10 N, reached when a term occurs in a single document: log10 37 = 1.57 here. The smallest is 0, reached when a term occurs in every document, since log10(N / N) = log10 1 = 0. SLP3 says it directly: the lowest weight, 0, is "assigned to terms that occur in every document", words "like good or sweet". Notice that idf is not linear in df. Going from df = 1 to df = 2 costs 0.30, the same as going from df = 12 to df = 24: halving the spread always buys the same amount of weight.
Worked example
Checking the two ends of the ladder
Romeo, df = 1
idf = log10(37 / 1) = log10 37 = 1.568, which rounds to the slide's 1.57.fool, df = 36
37 / 36 = 1.0278, so idf = log10 1.0278 = 0.0119, which rounds to 0.012.good, df = 37
idf = log10(37 / 37) = log10 1 = 0.Result
A word in one play is worth 1.57; a word missing from just one play is worth about 0.012, over a hundred times less; a word in every play is worth nothing.
What counts as a document?
Nothing in the formula says a document must be a play. SLP3 defines a document as "whatever unit of text the system indexes and retrieves (web pages, scientific papers, news articles, or even shorter passages like paragraphs)". It could be a Wikipedia article, a tweet, a paragraph, or a 300-token chunk in a RAG index. The choice is a modelling decision, and it changes everything downstream: N changes, every df changes, and so every idf changes.
Switch Shakespeare from plays to paragraphs and N grows into the thousands. good is no longer in every document, because most paragraphs do not contain it, so it gets a positive idf. A word like battle that clusters in a few scenes now looks rarer relative to the collection, and its weight rises. The right unit is the one that matches what you retrieve or compare: if your system returns passages, compute df over passages.
Recall
What happens to the idf values if you treat each paragraph as a document instead of each play?
Take battle in Julius Caesar. The play uses the word 7 times, so its term frequency is log10(7 + 1) = log10 8 = 0.903. battle appears in 21 of the 37 plays, so its idf is 0.246. Multiply: 0.903 × 0.246 = 0.222, which the slide prints as 0.22. That is one cell of the tf-idf matrix, and every other cell is computed the same way.
The rule is just the definition applied cell by cell: squash the count into Term frequency, look up the word's Inverse document frequency, multiply. The table below combines the raw counts on slide 52 with the idf column of slide 50 and shows the intermediate tf values the slide skips. All sixteen weights agree with the slide.
| Word | Raw counts | tf = log10(count + 1) | idf | tf-idf |
|---|---|---|---|---|
| battle | 1, 0, 7, 13 | 0.301, 0, 0.903, 1.146 | 0.246 | 0.074, 0, 0.22, 0.28 |
| good | 114, 80, 62, 89 | 2.061, 1.908, 1.799, 1.954 | 0 | 0, 0, 0, 0 |
| fool | 36, 58, 1, 4 | 1.568, 1.771, 0.301, 0.699 | 0.0119 | 0.019, 0.021, 0.0036, 0.0083 |
| wit | 20, 15, 2, 3 | 1.322, 1.204, 0.477, 0.602 | 0.0367 | 0.049, 0.044, 0.018, 0.022 |
Worked example
Two cells from the slide 52 table
battle in Julius Caesar
tf = log10 8 = 0.903, idf = log10(37 / 21) = 0.246, so w = 0.903 × 0.246 = 0.222 ≈ 0.22.fool in Henry V
The count is 4, so tf = log10 5 = 0.699. fool is in 36 plays, so idf = 0.0119, and w = 0.699 × 0.0119 = 0.0083.Result
battle in Julius Caesar outweighs fool in Henry V by a factor of about 27, although the raw counts differ only by 7 against 4. The difference comes almost entirely from idf.
What the table teaches
Look at the good row. It has the largest counts in the table, from 62 to 114, and its tf values are all around 2. Yet every tf-idf cell is 0, because good occurs in all 37 plays and its idf is log10 1 = 0. The word that dominated the raw vectors has been removed from the comparison without anyone writing a stop list.
Now look at battle. In raw counts it was a minor dimension, dwarfed by good in every play and by fool in the comedies. After weighting it is the largest value in every column where it appears, and it cleanly separates Julius Caesar (0.22) and Henry V (0.28), plays full of war, from the comedies As You Like It (0.074) and Twelfth Night (0). The IR book summarises the pattern: a weight is highest when a term occurs many times in a small number of documents, lower when it occurs fewer times or in many documents, and lowest when it occurs in virtually all documents.
The resulting document vectors are still sparse vectors of vocabulary length, mostly zeros, and you compare them with Cosine similarity exactly as in part 04. The difference is that the dimensions now carry information in proportion to how distinctive they are.
tf-idf in real systems
The formula is the same everywhere, but the details are not, and the differences matter when you reproduce a paper or debug a pipeline. scikit-learn's TfidfVectorizer by default uses raw counts for tf, a smoothed natural-log idf, ln((1 + n) / (1 + df)) + 1, and then L2-normalises each document vector. Because of the +1, a term in every document gets idf 1, not 0, so good would survive. The documentation's own example turns the vector [3, 0, 1.8473] into [0.8515, 0, 0.5243] after normalisation.
| Source | tf | idf | idf of a term in every document |
|---|---|---|---|
| Slides (SLP3 earlier editions) | log10(count + 1) | log10(N / df) | 0 |
| SLP3 current draft, IIR | 1 + log10(count), or 0 | log10(N / df) | 0 |
| scikit-learn default | raw count (sublinear_tf=False) | ln((1 + n) / (1 + df)) + 1 | 1 |
BM25 is the modern member of the family and the standard keyword retriever in search engines and hybrid RAG systems. It keeps idf but replaces log tf with a saturating function controlled by a parameter k, usually between 1.2 and 2, and normalises for document length with b = 0.75. Salton and Buckley's 1988 comparison of term-weighting schemes is the classic study of why these choices matter.
Recall
Compute the tf-idf of battle in Henry V (count 13, df 21, N 37).
Recall
Why does every cell in the good row of slide 52 become 0?
Quick check
With N = 37 and tf = log10(count + 1), what is the tf-idf of battle in Julius Caesar (count 7, df 21)?
Quick check
Why is every entry in the good row of the slide 52 tf-idf table equal to zero?
Quick check
In scikit-learn's default TfidfVectorizer, what idf does a term that appears in every document receive?
Recap
If you remember nothing else
- Raw frequency is informative but dominated by ubiquitous words. tf-idf (term-document) and PMI (term-term) reweight counts.
- The slides use tf = log10(count + 1), so 0 → 0, 9 → 1, 99 → 2. The current SLP3 draft and IIR use 1 + log10(count) for count > 0.
- df counts documents, cf counts tokens. Romeo and action share cf 113 but have df 1 versus 31.
- idf = log10(N / df). With N = 37 plays it ranges from 1.57 (Romeo) to 0 (good, sweet). The slides never state N.
- w = tf × idf. Battle in Julius Caesar is log10 8 × 0.246 = 0.22, and the whole good row becomes 0.
- A document is any unit you choose. Changing it changes N, df and every weight.
- Libraries differ: scikit-learn uses ln((1 + n) / (1 + df)) + 1 with L2 normalization, so a term in every document keeps idf 1 instead of 0. BM25 adds tf saturation and length normalization.
Sources
- Speech and Language Processing, 3rd ed. draft, Chapter 11: Information Retrieval and Retrieval-Augmented GenerationBookJurafsky and Martin, draft of August 2026Section 11.1.2 on tf-idf: the log tf formula and footnote on the slides' variant, the Shakespeare idf table, the definition of a document, stop words and BM25 parameters(opens in a new tab)
- Introduction to Information Retrieval, section 6.2.1: Inverse document frequencyBookManning, Raghavan and Schütze, Cambridge University PressCollection frequency versus document frequency, and the Reuters try and insurance example(opens in a new tab)
- Introduction to Information Retrieval, section 6.2.2: Tf-idf weightingBookManning, Raghavan and Schütze, Cambridge University PressWhen a tf-idf weight is high, lower and lowest(opens in a new tab)
- Introduction to Information Retrieval, section 6.4.1: Sublinear tf scalingBookManning, Raghavan and Schütze, Cambridge University PressWhy twenty occurrences are not twenty times as significant, and the 1 + log tf variant(opens in a new tab)
- A Statistical Interpretation of Term Specificity and Its Application in RetrievalPaperSparck Jones, Journal of Documentation 28(1):11-21, 1972The original proposal of inverse document frequency(opens in a new tab)
- Term-weighting approaches in automatic text retrievalPaperSalton and Buckley, Information Processing and Management 24(5):513-523, 1988The classic comparison of tf, idf and normalisation choices(opens in a new tab)
- Understanding inverse document frequency: on theoretical arguments for IDFPaperRobertson, Journal of Documentation 60(5), 2004Why the probabilistic retrieval model, not information theory, grounds idf(opens in a new tab)
- Feature extraction: Tf-idf term weightingDocsscikit-learn documentationsmooth_idf formula with the natural log, sublinear_tf, and default L2 normalisation(opens in a new tab)
Part 06: Pointwise mutual information and PPMI
Measuring whether two words co-occur more than chance with PMI, clipping negatives to get PPMI, computing it step by step on a term-context matrix, and correcting PMI's bias toward rare words with alpha-weighted context probabilities.
6 concepts, slides 53-60
Why this part matters
Part 05 fixed raw counts on the term-document side with tf-idf. Term-context matrices have the same disease: the biggest numbers belong to frequent words such as information and data, which sit next to almost everything. A raw count cannot tell you whether two words genuinely belong together or merely happen to be common.
Positive pointwise mutual information is the standard cure. It rescales every cell of a term-context matrix by what chance alone would predict, keeping only above-chance association, and it turns that matrix into meaningful sparse word vectors. It is also the bridge to the rest of this lecture: skip-gram with negative sampling turns out to factorize a shifted PMI matrix, and the magic number 0.75 you will meet in word2vec first appears here. Outside embeddings, PMI drives collocation extraction, lexicography and feature selection, so it will show up in your research. Exams almost always ask for one PPMI cell by hand.
By the end you can
- Explain PMI as observed versus chance co-occurrence, measured in bits.
- Justify clipping negative PMI to zero to obtain PPMI.
- Compute joint and marginal probabilities from a term-context count matrix.
- Compute any single PPMI cell by hand from counts.
- Explain PMI's bias toward rare contexts and two standard fixes.
- Apply alpha = 0.75 context smoothing and predict its effect on each cell.
Start with numbers. In the small Term-context matrix used throughout this part, the word information accounts for 7703 of the N = 11716 counted word-context pairs, and the context data for 5673 of them. If the two words had nothing to do with each other, how often would you expect to see them together? The full table, with every row and column sum, is in the counts-to-probabilities section below.
Independence answers that. If knowing one word tells you nothing about the other, the probability of the pair is just the product of the separate probabilities: P(information) × P(data) = .6575 × .4842 = .3184. So about 31.8% of all pairs should be (information, data) by chance alone. The observed share is 3982 / 11716 = .3399. The ratio of observed to expected is 1.0676: the pair occurs only 6.8% more often than chance, and log2 1.0676 = .0944 bits. Compare cherry and pie: there the ratio is 20.8, and the log is 4.38 bits. That number is pointwise mutual information.
Read the fraction one piece at a time. The numerator is what actually happened: how often the pair appeared together. The denominator is what independence predicts, because under independence P(x, y) = P(x) P(y). The ratio therefore says how many times more often the pair occurs than chance, and the base-2 logarithm turns that ratio into bits. A ratio of 1 gives 0 bits (independent), a ratio above 1 gives a positive score (attraction), and a ratio below 1 gives a negative score (avoidance). Each extra bit means the pair is twice as over-represented.
Where the measure comes from
Church and Hanks (1989, journal version 1990) brought the measure to lexicography. They estimated P(x, y) by counting how often x is followed by y within a window of w = 5 words, and they read the score exactly as above: well above zero for genuine association, near zero for no relation, well below zero for words in complementary distribution. Their rough rule of thumb was that pairs with a score above 3 tend to be interesting, and they ignored pairs seen 5 times or fewer because the ratio was unstable, an early sign of the rare-event problem later in this part. One subtlety a PhD reader should notice: because their count encoded word order (x before y), their association ratio was not symmetric. The term-context version on these slides counts a symmetric window, so PMI(w, c) depends only on which pair you pick.
Recall
In one sentence, what does PMI(w, c) = 0 mean?
Suppose two words each have probability 10^-6, which is typical of most of the vocabulary. Independence predicts that they appear together with probability 10^-12. To say with confidence that they appear together less than that, you must be able to estimate probabilities well below 10^-12, which takes a corpus on the order of trillions of pairs. With a corpus of a few million tokens you will simply never see them together, and you cannot tell "these words avoid each other" apart from "these words never happened to meet".
That is the first reason negative PMI is unreliable: the evidence needed to estimate below-chance co-occurrence grows with the rarity of the words, and for most word pairs it is never available. The second reason is about evaluation: it is not clear that people can even judge "unrelatedness" reliably, so there is no good gold standard to check negative scores against. The practical answer, used since Church and Hanks and later Dagan and colleagues (1993) and Niwa and Nitta (1994), is to keep only the positive side.
Positive PMI keeps the attraction signal and floors everything else at zero. It has a welcome side effect. A pair that never co-occurs has P(w, c) = 0, so its Pointwise mutual information is log2 0 = -∞, a value that would wreck any later arithmetic. Clipping maps it cleanly to 0. In the slide table this is exactly what happens to strawberry/computer and strawberry/data, both with count 0. The result is a sparse vector per word: long, mostly zero, with a few positive entries for the contexts that really characterise it.
| PMI | PPMI | |
|---|---|---|
| Range | -∞ to +∞ (asymptotically) | 0 to +∞ |
| Unseen pair (count 0) | log2 0 = -∞ | 0 |
| What a value says | Attraction (positive) or avoidance (negative) | Strength of attraction only; 0 means no evidence of it |
| Reliability | Negative side needs enormous corpora | Keeps only the side small corpora can estimate |
| Matrix shape | Dense, with many large negative entries | Sparse: most entries are exactly 0 |
Empirically the choice pays off. Levy, Goldberg and Dagan (2015) report that Bullinaria and Levy (2007) found PPMI outperforms plain PMI on semantic similarity tasks, and PPMI became the default count-based baseline against which neural embeddings were later compared.
Recall
Give two reasons PPMI discards negative PMI.
Quick check
Why does PPMI replace negative PMI values with zero?
Everything so far needs three probabilities per cell: the joint P(w, c) and the two marginals P(w) and P(c). All three come from one count matrix F. Here is the matrix the slides use, taken from SLP3, whose counts come from Wikipedia: four target words as rows, five context words as columns, with the row and column sums added.
| computer | data | result | pie | sugar | row sum | |
|---|---|---|---|---|---|---|
| cherry | 2 | 8 | 9 | 442 | 25 | 486 |
| strawberry | 0 | 0 | 1 | 60 | 19 | 80 |
| digital | 1670 | 1683 | 85 | 5 | 4 | 3447 |
| information | 3325 | 3982 | 378 | 5 | 13 | 7703 |
| column sum | 4997 | 5673 | 473 | 512 | 61 | N = 11716 |
| p(c) | .4265 | .4842 | .0404 | .0437 | .0052 | 1 |
Add every cell and you get the grand total N = 11716. Each probability is then a share of that one total. The joint probability of a cell is the cell divided by N. The marginal probability of a target word is its row sum divided by N, and the marginal probability of a context is its column sum divided by N.
Worked example
Three probabilities for (information, data)
Find the total
Sum all twenty cells, or equivalently the four row sums: 486 + 80 + 3447 + 7703 = 11716.Joint
The cell count is 3982, so p(information, data) = 3982 / 11716 = .3399.Row marginal
The information row sums to 7703, so p(information) = 7703 / 11716 = .6575.Column marginal
The data column sums to 5673, so p(data) = 5673 / 11716 = .4842.Result
- Joint p(information, data)
- 3982 / 11716 = .3399
- Marginal p(information)
- 7703 / 11716 = .6575
- Marginal p(data)
- 5673 / 11716 = .4842
- Chance prediction p(information) p(data)
- .6575 × .4842 = .3184
Because the joint and both marginals come from the same matrix and the same N, they are consistent: each row of joint probabilities sums to its row marginal, each column to its column marginal, and everything sums to 1. SLP3 is candid that this "pretends" the five listed contexts are the only ones in the world. In a real system the matrix has tens of thousands of columns and the same recipe applies unchanged.
Recall
From the count table, compute p(strawberry) and p(sugar).
With the three probabilities in hand, one cell of the Pointwise mutual information matrix is a division and a log. Work it once by hand for information/data, then notice a shortcut that skips the rounding entirely.
Worked example
PPMI(information, data) by hand
Ratio of observed to expected
.3399 / (.6575 × .4842) = .3399 / .3184 = 1.0675.Same ratio straight from counts
The Ns cancel into one: f × N / (row × col) = 3982 × 11716 / (7703 × 5673) = 1.0676. This route avoids rounding the probabilities first.Take log base 2
Most calculators lack log2, so use log2 x = ln x / ln 2: ln 1.0676 / 0.6931 = 0.0654 / 0.6931 = .0944.Clip
.0944 > 0, so the clip changes nothing.Result
PPMI(information, data) = .09, matching the slide.
Repeat for every cell and you get the full PMI matrix below. Every Positive PMI value on slide 58 is this matrix after clipping; SLP3 quotes PMI(cherry, computer) = -6.7 in its figure caption as an example of a large negative value that PPMI discards.
| computer | data | result | pie | sugar | |
|---|---|---|---|---|---|
| cherry | -6.70 | -4.88 | -1.12 | 4.38 | 3.30 |
| strawberry | -∞ | -∞ | -1.69 | 4.10 | 5.51 |
| digital | 0.18 | 0.01 | -0.71 | -4.91 | -2.17 |
| information | 0.02 | 0.09 | 0.28 | -6.07 | -1.63 |
| computer | data | result | pie | sugar | |
|---|---|---|---|---|---|
| cherry | 0 | 0 | 0 | 4.38 | 3.30 |
| strawberry | 0 | 0 | 0 | 4.10 | 5.51 |
| digital | 0.18 | 0.01 | 0 | 0 | 0 |
| information | 0.02 | 0.09 | 0.28 | 0 | 0 |
Now read the PPMI matrix as a set of word vectors. The fruit rows light up only on pie and sugar; the technology rows only on computer, data and result. The two groups share no non-zero dimension, so the cosine similarity between cherry and digital is exactly 0, while cherry and strawberry have a cosine of about 0.96. The raw counts were noisier: cherry co-occurs with computer and data a few times, and in a full-size matrix such incidental counts, scaled by frequent contexts, blur every row. PPMI vectors are still long and sparse; the second half of this lecture learns short dense vectors instead.
Recall
Compute PPMI(cherry, pie) from count 442, row sum 486, column sum 512 and N = 11716.
Quick check
Using the slide counts, what is PMI(information, data)?
Add one more row and one more column to the slide table: a word w and a context c, each seen exactly once, and that once together. The shortcut gives PMI = log2(1 × 11716 / (1 × 1)) = log2 11716 = 13.5 bits. That is more than double strawberry/sugar (5.51), from a single observation that may be a typo or a coincidence. One accident beats thousands of genuine co-occurrences.
The mechanism is the denominator. When a context is rare, P(c) is tiny, so even one co-occurrence produces a large ratio. Levy, Goldberg and Dagan (2015) call this Pointwise mutual information's Achilles' heel, citing Turney and Pantel (2010): a word's highest-scoring dimensions become obscure contexts it met once or twice, and since similar words rarely share those accidents, their vectors look less alike under cosine than they should. Clipping does not help, because Positive PMI only touches negative values and this bias inflates positive ones.
Two fixes
- Give rare contexts a little more probability. If P(c) for a rare context is nudged upward, its PMI falls. This is the alpha weighting of the next concept.
- Add-k smoothing. Add a small constant k to every count before computing probabilities, exactly as in the n-gram language models of Lecture 03. Slide 59 names the simplest case, add-one smoothing, which is k = 1 and lowers strawberry/sugar from 5.51 to 5.41. SLP3 gives k = 0.1 to 3 as common choices, and notes that the larger the k, the more the non-zero counts are discounted. A count of 1 becomes 3 with k = 2, tripling, while a count of 442 barely moves. On the slide table, add-2 lowers strawberry/sugar from 5.51 to 5.31 and leaves the overall pattern intact.
The oldest fix is the bluntest. Church and Hanks simply discarded pairs seen 5 times or fewer. Frequency cut-offs of that kind are still common in collocation tools.
| Add-k smoothing | Context alpha | |
|---|---|---|
| What changes | Every count gets + k before probabilities are computed | Only the context distribution P(c) is reshaped |
| Which counts move most | Small counts, proportionally (1 becomes 3 with k = 2) | Rare contexts gain probability, frequent ones lose a little |
| Typical value | k from 0.1 to 3 | alpha = 0.75 |
| strawberry/sugar | 5.51 → 5.31 (k = 2) | 5.51 → 4.01 |
| Origin | Laplace smoothing, as in n-gram language models | word2vec negative sampling, carried over by Levy et al. (2015) |
Recall
Why does clipping to PPMI not remove PMI's bias toward rare contexts?
Take two contexts with P(a) = .99 and P(b) = .01. Raising both to 0.75 and renormalising roughly triples the rare one, so every Pointwise mutual information involving b falls by about 1.6 bits. The worked example below shows each step.
The exponent Alpha-weighted context probability flattens the context distribution. Because 0 < α < 1 shrinks large counts proportionally more than small ones, after renormalising P_alpha(c) > P(c) for rare contexts and P_alpha(c) < P(c) for frequent ones. A larger denominator means a lower PMI, so rare contexts lose exactly the inflated advantage the previous concept described, and Positive PMI computed with P_alpha(c) keeps fewer spurious rare dimensions. Only P(c) changes; P(w, c) and P(w) stay as before.
On the slide table the effect is easy to predict. Sugar, with only 61 counts, goes from P = .0052 to .0148, so strawberry/sugar drops from 5.51 to 4.01 and cherry/sugar from 3.30 to 1.80. The frequent context computer goes from .4265 to .4019, so digital/computer actually rises from 0.18 to 0.27. Result goes from .0404 to .0686, which pushes information/result from 0.28 to a PMI of -0.48, so its PPMI becomes 0.
| computer | pie | sugar | |
|---|---|---|---|
| P(c) | .4265 → .4019 | .0437 → .0728 | .0052 → .0148 |
| cherry | -6.70 → -6.61 | 4.38 → 3.64 | 3.30 → 1.80 |
| strawberry | -∞ → -∞ | 4.10 → 3.37 | 5.51 → 4.01 |
| digital | 0.18 → 0.27 | -4.91 → -5.65 | -2.17 → -3.67 |
| information | 0.02 → 0.10 | -6.07 → -6.81 | -1.63 → -3.13 |
Worked example
Slide 60's two-context example
Raise to alpha
.99^.75 = .9925 and .01^.75 = .0316.Find the shared normaliser
.9925 + .0316 = 1.0241. Both contexts are divided by this same sum.Renormalise
P_alpha(a) = .9925 / 1.0241 = .97 and P_alpha(b) = .0316 / 1.0241 = .03.Result
The rare context's probability rises from .01 to .03; the frequent one falls from .99 to .97.
Why 0.75, and the road to word2vec
The number is borrowed. Mikolov and colleagues (2013) found that drawing negative samples for word2vec from the unigram distribution raised to the 3/4 power, U(w)^(3/4) / Z, worked significantly better than either the plain unigram or the uniform distribution. Levy, Goldberg and Dagan (2015) carried the trick over to count-based PMI as "context distribution smoothing", tested α ∈ {1, 0.75}, and found that it alleviates PMI's bias toward rare words and consistently improves performance across tasks, methods and configurations.
The connection runs deeper than a shared constant. Levy and Goldberg (2014) showed that Skip-gram with negative sampling with k negative samples implicitly factorizes a word-context matrix whose cells are PMI(w, c) - log k. The sparse analogue is shifted PPMI, max(PMI(w, c) - log k, 0), which they found competitive on word similarity. So PPMI and word2vec are close relatives: Negative sampling is, in a precise sense, a smoothed and compressed way of estimating the same quantity you just computed by hand. Keep this in mind for the skip-gram parts that follow.
Recall
Why does raising context counts to 0.75 lower the PMI of rare contexts?
Recall
Why is "raise the probabilities" the same as "raise the counts" to alpha?
Quick check
With P(a) = .99, P(b) = .01 and alpha = 0.75, what is P_alpha(b)?
Quick check
Applying alpha = 0.75 to the slide's context counts does what to PPMI(strawberry, sugar)?
Recap
If you remember nothing else
- PMI(w, c) = log2 P(w, c) / (P(w)P(c)) measures how far a pair sits above or below chance, in bits. Zero means independent.
- Negative PMI is unreliable without enormous corpora, so PPMI = max(PMI, 0). Clipping also maps unseen pairs from minus infinity to 0.
- All probabilities come from one matrix total N: the joint is the cell over N, the marginals are the row and column sums over N.
- Shortcut: PMI = log2(f_ij × N / (row_i × col_j)). In the slide example, information/data = 0.09 and strawberry/sugar = 5.51.
- PMI favours rare events: one co-occurrence with a rare context can outscore thousands of genuine ones. PPMI does not fix this.
- Fixes: add-k smoothing (add-one on slide 59 is k = 1; SLP3 uses k from 0.1 to 3), or P_alpha(c) with alpha = 0.75, which raises rare-context probabilities and lowers their PMI.
- alpha = 0.75 comes from word2vec negative sampling, and SGNS itself approximates a PMI matrix shifted by log k.
Sources
- Speech and Language Processing, 3rd ed. draft, Section 6.6: Pointwise Mutual Information (PMI)BookJurafsky and Martin, draft of August 20, 2024pp. 114-116: PMI and PPMI definitions, the Fano terminology footnote, the 10^-6 argument, footnote on log 0, the count matrix with N = 11716, the PPMI matrix, add-k values and alpha = 0.75(opens in a new tab)
- Speech and Language Processing, 3rd ed. draft, Chapter 5: EmbeddingsBookJurafsky and Martin, draft of August 2026Keeps alpha = 0.75 for negative sampling; the PPMI section was dropped in this draft(opens in a new tab)
- Word Association Norms, Mutual Information, and LexicographyPaperChurch and Hanks, Computational Linguistics 16(1), 1990The association ratio with log2, window w = 5, the I > 3 rule of thumb, ignoring f(x, y) ≤ 5, and asymmetry from word order(opens in a new tab)
- Word Association Norms, Mutual Information, and LexicographyPaperChurch and Hanks, ACL 1989, pp. 76-83The conference version cited on slide 54(opens in a new tab)
- Co-Occurrence Vectors From Corpora vs. Distance Vectors From DictionariesPaperNiwa and Nitta, COLING 1994Early use of positive PMI co-occurrence vectors(opens in a new tab)
- Improving Distributional Similarity with Lessons Learned from Word EmbeddingsPaperLevy, Goldberg and Dagan, TACL 3, 2015PMI's bias toward rare contexts, context distribution smoothing with alpha in {1, 0.75}, shifted PPMI, and the report of Bullinaria and Levy (2007) that PPMI beats PMI(opens in a new tab)
- Neural Word Embedding as Implicit Matrix FactorizationPaperLevy and Goldberg, NeurIPS 27, 2014SGNS implicitly factorizes a PMI matrix shifted by log k(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov, Sutskever, Chen, Corrado and Dean, NeurIPS 26, 2013The unigram distribution raised to the 3/4 power as the negative-sampling noise distribution(opens in a new tab)
- From Frequency to Meaning: Vector Space Models of SemanticsPaperTurney and Pantel, Journal of Artificial Intelligence Research 37, 2010Survey of count-based models, cited by Levy et al. for PMI's rare-event bias(opens in a new tab)
Part 07: Dense vectors and the skip-gram classifier
Why short dense vectors beat long sparse ones, the main ways to get them, and how word2vec turns embedding learning into a self-supervised binary task scored by the sigmoid of a dot product.
6 concepts, slides 61-74
Why this part matters
This is where the lecture switches from counting to learning. Until now every vector was a row of a count table, reweighted by tf-idf or PPMI. From here on the vectors are parameters that a small classifier learns, and the classifier is trained on a task that running text answers for free. Skip-gram with negative sampling is an early, influential template for the pretext tasks that followed in NLP, including the masked-word objective of BERT, and it is a staple exam item: explain self-supervision, compute σ(c·w).
It also matters for research and real systems. The downloadable static vectors you will reach for (GoogleNews 300-d, GloVe 6B) and the gensim switches you will set (sg=1, negative=5, window=5) all come from the ideas below. The argument runs in six steps: why short dense vectors beat long sparse ones, where dense vectors come from, the idea of predicting instead of counting, how a window turns text into pairs, how a dot product becomes a probability, and how one window is scored as a whole. How the vectors are actually learned is the next part.
By the end you can
- Contrast sparse and dense vectors by length, zeros and value range, and give three reasons dense vectors help, including the car and automobile argument.
- Name the main sources of dense vectors (word2vec, GloVe, SVD/LSA, contextual models) and distinguish static from contextual embeddings.
- Explain self-supervision and the four-step skip-gram recipe, and say why the classifier is discarded while its weights are kept.
- Generate the positive (target, context) pairs for a ±m window, and relate m to the slides' L = 2m context words per target.
- Compute P(+|w,c) = σ(c·w) and P(-|w,c) = σ(-c·w) for given vectors.
- Score a whole window with the product of sigmoids and its log sum, and state the independence assumption behind it.
One sentence says "she drove the car to work"; another says "he parked the automobile outside". In a count space, drove collects weight on the car dimension and parked collects weight on the automobile dimension. Those are two different coordinates. If those are their only informative neighbors, the dot product of the two verbs is exactly 0, so their cosine is 0 too, even though any reader sees that the two sentences describe the same kind of event.
Worked example
Two verbs, two synonyms, zero similarity
Keep only the two relevant dimensions
Use the axes (car, automobile). drove = [1, 0] and parked = [0, 1].Dot product
1×0 + 0×1 = 0.Cosine
0 / (1 × 1) = 0. The vectors are orthogonal: as far as the space can tell, the two verbs share nothing.Result
Sparse vectors treat car and automobile as unrelated axes, so words that differ only in which synonym they co-occur with look completely dissimilar.
Part 03 previewed this contrast on slide 30; here it gets its full argument. The vectors of the last three parts were sparse vectors. Built from a tf-idf or Pointwise mutual information weighting, they have one dimension per vocabulary word, so their length is |V|, typically 20,000 to 50,000, and nearly every entry is zero. A Dense vector is the opposite on every count. It has a small fixed number of dimensions d, usually between 50 and 1000; almost all of its entries are non-zero; and they are real numbers that can be negative. The price is interpretability: the d dimensions do not correspond to context words or to anything else nameable (SLP 5.5). Such a vector is still an Embedding, just a learned one rather than a counted one.
| Property | Sparse (tf-idf, PPMI) | Dense (word2vec) |
|---|---|---|
| Length | |V| ≈ 20,000 to 50,000 | d ≈ 50 to 1000 |
| Zeros | Almost every entry | Almost none |
| Values | Counts or non-negative weights | Real numbers, positive or negative |
| Dimension meaning | Each axis is one context word | No clear interpretation |
| Weights a one-word classifier needs | 50,000 | 300 |
| Synonyms | Separate, orthogonal axes | Can share the same directions |
Three reasons dense vectors win
- Fewer weights to learn. A classifier that uses one word as a feature needs one weight per dimension: 300 weights for a dense vector instead of 50,000 for a sparse one. Fewer parameters may generalize better and overfit less.
- Better Synonymy. In a dense space, car and automobile can occupy nearby directions, so a word whose neighbors include car is automatically close to a word whose neighbors include automobile. The Cosine similarity of drove and parked is no longer forced to zero.
- Empirically better. SLP states it bluntly: dense vectors work better in every NLP task than sparse vectors. The slide says the same more cautiously: in practice, they work better.
Recall
Why do sparse vectors fail on car and automobile, and give two other reasons dense vectors help.
Quick check
One word occurs only near car, another only near automobile. What cosine do their sparse count vectors give?
Suppose you need word vectors for a project tomorrow. You do not have to train anything. You can download the word2vec GoogleNews vectors, or one of the GloVe sets from Stanford, and have a dense vector for millions of words in a few minutes. Knowing where each set comes from tells you what it can and cannot do.
Pretrained static vectors you can download today
- word2vec GoogleNews
- 300-d vectors for 3,000,000 words and phrases, trained on part of Google News (about 100 billion words); 1,662 MB as word2vec-google-news-300 in gensim-data
- GloVe 6B
- Wikipedia 2014 plus Gigaword 5, 400K vocabulary, in 50, 100, 200 and 300 dimensions
- GloVe, larger sets
- Common Crawl (42B and 840B tokens), Twitter (27B), and newer 2024 Dolma and Wikipedia plus Gigaword releases
The lecture sorts dense vectors into three families. The first is inspired by neural language models: learn vectors by predicting words from their neighbors. Word2vec is the flagship, and it is a toolkit with two algorithms, Skip-gram (predict the neighbors from the word) and Continuous bag of words (predict the word from its neighbors). The slide also places GloVe here, which is a simplification noted below. The second family is factorization: take a count matrix and compress it with singular value decomposition. Latent semantic analysis (LSA) is SVD applied to a term-document matrix with the first 300 or so dimensions kept (Deerwester et al. 1990, in SLP's history section). The slide then presents contextual models such as ELMo and BERT as an alternative to these static embeddings. Both the SVD bullet and the contextual bullet are greyed out, because this lecture concentrates on word2vec.
| Family | Example | Learns from | One vector per |
|---|---|---|---|
| Prediction (neural-LM inspired) | word2vec SGNS, CBOW | Local windows, by predicting neighbors | Word type |
| Global co-occurrence regression | GloVe | Global co-occurrence counts | Word type |
| Matrix factorization | SVD, LSA | Term-document (or PPMI) matrix | Word type |
| Contextual | ELMo, BERT | The whole sentence, at run time | Token in context |
The last column is the line that matters most. Word2vec, GloVe and LSA all produce a Static embedding: one fixed vector per word type, looked up from a table. The word bank gets the same vector in "river bank" and "bank loan", so Polysemy is averaged into a single point. A Contextual embedding is computed by running the whole sentence through a network, so each token occurrence gets its own vector (Peters et al. 2018; Devlin et al. 2019). Those models come later in the course; this lecture is about static vectors.
In practice you train your own vectors with gensim. Its defaults are worth memorizing because they bite: vector_size=100, window=5, negative=5 (the docs say usually between 5 and 20), ns_exponent=0.75, min_count=5, and sg=0, which means CBOW.
from gensim.models import Word2Vec
import gensim.downloader as api
model = Word2Vec(
sentences=corpus,
vector_size=300,
window=5,
sg=1,
negative=5,
min_count=5,
)
google_news = api.load("word2vec-google-news-300")
google_news.most_similar("automobile", topn=3)Recall
Static versus contextual embeddings: give one example of each.
Read the phrase "...a tablespoon of apricot jam...". Without anyone annotating anything, it hands you a labeled fact: (apricot, jam) is a yes, these two words occur near each other. Now pick a random word from the lexicon, say aardvark. (apricot, aardvark) is almost certainly a no. Every sentence of every book and web page produces such facts by the dozen.
That is the whole idea behind Skip-gram with negative sampling. Instead of counting how often a word c occurs near apricot, train a binary classifier on the question "is c likely to show up near apricot?" Nobody actually cares about that question. What we keep are the weights the classifier learns in order to answer it, and those weights are the embeddings. Because the gold answers come from running text rather than from people, this is Self-supervision.
The four-step recipe
- Treat the target word and each neighboring context word as a positive example.
- Randomly sample other words from the lexicon to make negative examples.
- Train logistic regression to tell the two kinds of pair apart.
- Throw the classifier away and use its learned weights as the embeddings.
The idea has a lineage. Bengio et al. (2003) and Collobert et al. (2011) had already shown that the next word in running text is a free supervision signal for learning word vectors, and Collobert et al. stressed the "vast amounts of mostly unlabeled training data" this unlocks. Word2vec simplifies their neural language models in two ways (SLP 5.5). It replaces the hard task of predicting the next word over the whole vocabulary with a yes-or-no question about a single pair, and it replaces a multi-layer network with plain logistic regression. The payoff was speed: Mikolov et al. (2013a) learned high quality vectors from a 1.6 billion word data set in less than a day.
The binary framing is also what keeps training cheap. The basic skip-gram formulation defines its probability with a softmax over the whole vocabulary, whose cost grows with W, "often large (10^5 to 10^7 terms)" (Mikolov et al. 2013b). Negative sampling replaces that sum with a handful of k sampled noise words, 5 to 20 for small data sets and 2 to 5 for large ones. How those negatives are drawn and how the weights are updated is the subject of the next part.
Recall
What is self-supervision in word2vec, and why does it need no human labels?
Recall
List the four steps of the skip-gram intuition.
Quick check
In skip-gram training, where do the gold labels come from?
Take the running text "...lemon, a tablespoon of apricot jam, a pinch..." and put a window of two words on each side of apricot. The target is w = apricot, and the four context words are c1 = tablespoon, c2 = of, c3 = jam and c4 = a. Each one forms a positive training pair with the target.
Worked example
Positive pairs from a ±2 window
Target apricot
(apricot, tablespoon), (apricot, of), (apricot, jam), (apricot, a).Slide one word to the right: target jam
(jam, of), (jam, apricot), (jam, a), (jam, pinch).Result
Every token takes a turn as the target. A window of ±2 gives up to 4 positive pairs per target, and in general a half-width of m gives up to 2m. Fewer appear only at the edges of a text.
A note on letters: the slides and SLP write L for the number of context words being scored, so a ±2 window has L = 4. Keep that in mind when you meet c_1:L in the last concept of this part. The half-width is written m throughout this lecture, so L = 2m. Try the window yourself: click any word to make it the target and change the half-width.
- (apricot, tablespoon)
- (apricot, of)
- (apricot, jam)
- (apricot, a)
Once the pairs exist, the classifier's job is simple to state. Given a candidate pair (w, c), return the probability that c is a real context word of w. For (apricot, jam) that probability should be high; for (apricot, aardvark) it should be low. Since there are only two outcomes, the probability of the negative answer is whatever is left over:
Recall
How many positive pairs does a ±3 window give for a target in mid-sentence, and what is the slides' L for it?
Quick check
With a ±1 window over '...tablespoon of apricot jam, a pinch...', which positive pairs does target jam produce?
Give apricot a toy three-dimensional vector w = (1, 0.5, -1). Give jam the context vector c = (2, 1, -0.5) and aardvark c = (-1, 0.5, 1). Multiply apricot by jam elementwise and add: 2 + 0.5 + 0.5 = 3. Do the same with aardvark: -1 + 0.25 - 1 = -1.75. The real neighbor scores high, the random word scores low. The question is how to turn those two scores into probabilities.
The intuition the classifier uses is the one from the cosine parts: two vectors are similar when their Dot product is high, and the Cosine similarity is just a dot product divided by the two vector lengths. So skip-gram takes the dot product c·w as its similarity score. But neither a dot product nor a cosine is a probability. Because embedding entries can be negative, the dot product can be anything from -∞ to ∞ (SLP 5.5.1).
Logistic regression already has the fix. The sigmoid function maps any real number into the open interval (0, 1), so the classifier passes the dot product through it:
The identity 1 - σ(x) = σ(-x) is worth checking once, because it is used constantly: the two cases always sum to 1, and flipping the sign of the score swaps them. A few anchor values make hand computation quick: σ(0) = 0.5, σ(2) = 0.8808, σ(-2) = 0.1192, σ(5) = 0.9933, σ(-5) = 0.0067.
Worked example
P(+) for jam and for aardvark, with w = (1, 0.5, -1)
Multiply elementwise
jam: (1×2, 0.5×1, -1×-0.5) = (2, 0.5, 0.5). aardvark: (1×-1, 0.5×0.5, -1×1) = (-1, 0.25, -1).Sum
c·w = 3 for jam and c·w = -1.75 for aardvark.Exponentiate the negated score
e^-3 = 0.0498 and e^1.75 = 5.7546.Invert one plus that
jam: 1 / 1.0498 = 0.9526. aardvark: 1 / 6.7546 = 0.1480.Result
P(+ | apricot, jam) = 0.9526 and P(+ | apricot, aardvark) = 0.1480, so P(- | apricot, aardvark) = 0.8520.
Recall
Compute P(+|w,c) for w = (1, 0.5, -1) and c = (2, 1, -0.5).
Quick check
With w = (1, 0.5, -1) and c = (2, 1, -0.5), what is P(+|w,c)?
One pair at a time is not enough: apricot has four neighbors in its ±2 window, and the classifier should score the whole window. Keep w = (1, 0.5, -1) and give each of the four context words a toy vector.
| Context | c | c·w | σ(c·w) | log σ |
|---|---|---|---|---|
| tablespoon | (0.5, 0, -0.5) | 1.0 | 0.7311 | -0.3133 |
| of | (0.2, -0.4, 0.1) | -0.1 | 0.4750 | -0.7444 |
| jam | (2, 1, -0.5) | 3.0 | 0.9526 | -0.0486 |
| a | (0, 0.2, 0) | 0.1 | 0.5250 | -0.6444 |
Skip-gram makes one simplifying move: it assumes the context words are independent of each other given the target. Under that assumption the probability that all of them are real neighbors is the product of the individual sigmoid probabilities, and its logarithm is a sum.
Worked example
Product and log sum for the apricot window
Multiply the four sigmoids
0.7311 × 0.4750 × 0.9526 × 0.5250 = 0.1737.Add the four logs
-0.3133 - 0.7444 - 0.0486 - 0.6444 ≈ -1.7507.Check they agree
e^-1.7507 ≈ 0.1737. The log sum is the log of the product.Result
Four true neighbors give a joint probability of only 0.1737, dragged down by the weakly scored function words of and a. With hundreds of factors the product would underflow, which is why training works in log space.
That last sentence hides a detail that sets up the next part. SLP's figure 5.6 shows that skip-gram stores two embeddings per word: one for when the word is a target and one for when it is a context. They live in a target matrix W and a context matrix C, the Target and context matrices, so the parameters are 2|V| vectors of dimension d. Learning those two matrices from positive and negative pairs is what part 8 is about.
Recall
Write P(+|w,c_1:L), say which assumption it relies on, and why we take logs.
Quick check
Why does skip-gram multiply the sigmoid scores of the context words?
Recap
If you remember nothing else
- Sparse tf-idf and PPMI vectors have length |V| (20,000 to 50,000) and are mostly zeros. Dense embeddings have 50 to 1000 dimensions with real values that can be negative.
- Dense vectors need far fewer classifier weights (300 vs 50,000), capture synonymy that separate dimensions miss (car and automobile), and work better in practice.
- The dense families are word2vec (skip-gram, CBOW), GloVe (global co-occurrence), SVD/LSA, and contextual models (ELMo, BERT). The first three are static: one vector per word type.
- Word2vec predicts rather than counts. A binary classifier asks "is c near w?", and its learned weights become the embeddings.
- Self-supervision: neighbors in running text are the gold positives and random lexicon words are the negatives, so no human labels are needed.
- A ±2 window around apricot gives four positives: tablespoon, of, jam, a.
- P(+|w,c) = σ(c·w) = 1/(1 + e^(-c·w)) and P(-|w,c) = σ(-c·w). σ(0) = 0.5.
- Assuming independent context words, P(+|w,c_1:L) = ∏σ(c_i·w) and log P = Σ log σ(c_i·w).
- SGNS stores two vectors per word: a target matrix W and a context matrix C.
Sources
- Speech and Language Processing, chapter 5: Embeddings, section 5.5 Word2vecBookJurafsky and Martin, draft of August 2026Dense versus sparse, the 300 vs 50,000 weights, car and automobile, self-supervision and the four steps, eqs 5.11 to 5.18, the -∞ to ∞ remark, and figure 5.6 with W and C(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov, Chen, Corrado and Dean, arXiv 2013CBOW and skip-gram; high quality vectors from 1.6 billion words in less than a day(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov, Sutskever, Chen, Corrado and Dean, NIPS 2013Negative sampling, k of 5 to 20 or 2 to 5, softmax cost of 10^5 to 10^7 terms, subsampling(opens in a new tab)
- GloVe: Global Vectors for Word RepresentationPaperPennington, Socher and Manning, EMNLP 2014(opens in a new tab)
- GloVe project pageDocsStanford NLPPretrained sets: 6B, 42B, 840B, Twitter, and the 2024 releases(opens in a new tab)
- word2vec project archiveDocsGoogle Code ArchiveGoogleNews vectors: 300-d, 3 million words and phrases, about 100 billion words(opens in a new tab)
- gensim Word2Vec documentationDocsgensimDefaults vector_size=100, window=5, sg=0 (CBOW), negative=5, ns_exponent=0.75, min_count=5(opens in a new tab)
- gensim-data model catalogueDocsGitHub (piskvorky)word2vec-google-news-300: 3,000,000 vectors, 1,662 MB(opens in a new tab)
- Don't count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectorsPaperBaroni, Dinu and Kruszewski, ACL 2014(opens in a new tab)
- Improving Distributional Similarity with Lessons Learned from Word EmbeddingsPaperLevy, Goldberg and Dagan, TACL 2015(opens in a new tab)
- Neural Word Embedding as Implicit Matrix FactorizationPaperLevy and Goldberg, NIPS 2014(opens in a new tab)
- A Neural Probabilistic Language ModelPaperBengio, Ducharme, Vincent and Jauvin, JMLR 2003(opens in a new tab)
- Natural Language Processing (Almost) from ScratchPaperCollobert et al., JMLR 2011(opens in a new tab)
- Deep Contextualized Word Representations (ELMo)PaperPeters et al., NAACL 2018(opens in a new tab)
- BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingPaperDevlin et al., NAACL 2019(opens in a new tab)
Part 08: Learning skip-gram embeddings with negative sampling
The two embedding matrices W and C, how positive and k negative training pairs are built, the cross-entropy loss for SGNS, its gradients, and the SGD updates that pull true neighbors together and push sampled noise apart.
6 concepts, slides 75-88
Why this part matters
Part 07 built a classifier, P(+ | w, c) = σ(c · w), that scores whether c is a real neighbor of w. It assumed the vectors already existed. This part is where they come from: the Skip-gram with negative sampling training loop that starts from random numbers and, one window at a time, turns them into embeddings in which apricot sits near jam.
Three reasons to know this cold. The Stochastic gradient descent update for skip-gram is a standard exam derivation (it is SLP3 Exercise 5.3). The same pattern of positive and noise pairs, dot-product scores and a cross-entropy loss is the backbone of contrastive learning in modern retrieval and sentence embedding models, so it shows up in research papers well beyond word2vec. And when you train or debug real embeddings, the defaults matter: k = 5 negatives, a 0.75 exponent on the noise distribution, and a learning rate that starts at 0.025 and decays.
By the end you can
- Describe θ as W stacked over C, with 2|V| rows of dimension d, and say what each matrix is for.
- Build positive pairs from a ±2 window and draw k noise words from the α = 0.75 weighted unigram.
- Derive the SGNS cross-entropy loss from the independence assumption and σ(−x) = 1 − σ(x).
- Derive the three gradients and apply one SGD update by hand.
- Explain why W and C are trained separately and then summed or reduced to W alone.
Take a vocabulary of 10,000 words and embeddings of dimension d = 300. How many numbers does Skip-gram learn? The natural guess is one vector per word, so 3,000,000. The real answer is twice that, 6,000,000, because every word owns two rows.
The parameters θ are two matrices stacked on top of each other. The top block W has one row per vocabulary word, indexed 1 to |V|; these are the target embeddings, used when a word sits at the center of a window. The bottom block C has another row per word, indexed |V| + 1 to 2|V|; these are the context embeddings, used when a word appears as a neighbor or is drawn as a noise word. So apricot has a row w_apricot in W and a separate row c_apricot in C, and together these are the Target and context matrices. SLP3 (Fig. 5.6) calls the W rows the input embeddings and the C rows the output embeddings, which is also the vocabulary of Mikolov et al. and of Rong's derivation notes.
What each block of θ holds
- W, rows 1 to |V|
- Target embeddings (also called input embeddings). Row w_i is used when word i sits at the center of a window.
- C, rows |V| + 1 to 2|V|
- Context embeddings (also called output embeddings). Row c_i is used when word i is a neighbor or a sampled noise word.
- Shape of θ
- 2|V| × d
- Size for |V| = 10,000, d = 300
- 2 × 10,000 × 300 = 6,000,000 parameters
- What the classifier reads
- Only dot products c · w, with c taken from C and w taken from W. A row of W is never dotted with another row of W.
The reason for two tables becomes clear once you look at what the classifier from part 07 actually computes. It only ever takes a Dot product between a context row and a target row, c · w. A target never meets another target, and a context never meets another context. W and C are therefore two different roles, and the model is free to learn different coordinates for each role. Every row is a Dense vector of length d; nothing here is sparse or counted.
Recall
How many vectors does θ hold, and how many parameters is that for |V| = 10,000 and d = 300?
Take the sentence fragment "lemon, a tablespoon of apricot jam, a pinch" with apricot as the target and a ±2 window. The four words inside the window give four positive pairs: (apricot, tablespoon), (apricot, of), (apricot, jam) and (apricot, a). No human labeled them; the corpus did, which is the Self-supervision idea from part 07.
A classifier trained only on positives would learn to say "yes" to everything. So each positive pair is matched with k negative pairs, built by keeping the target and replacing the context with a noise word drawn at random from the lexicon. With k = 2 the four positives above get eight negatives, for example aardvark, my, where, coaxial, seven, forever, dear and if. The slide lists these eight as one pool, so the way the table below assigns two to each positive is only illustrative. This is Negative sampling: the model learns to tell real neighbors from random words, rather than to predict the exact neighbor out of all |V| words.
| Positive context (label 1) | Two noise words drawn for it (label 0, illustrative) |
|---|---|
| tablespoon | aardvark, my |
| of | where, coaxial |
| jam | seven, forever |
| a | dear, if |
Which random words? The flattened unigram
Noise words are not drawn uniformly, and not quite in proportion to frequency either. They are drawn from the unigram distribution raised to the power α = 0.75 and renormalized, the same Alpha-weighted context probability trick part 06 used to stop Pointwise mutual information from overrating rare contexts. The only constraint SLP3 imposes is that the noise word is not the target itself.
Raising probabilities to a power below one compresses their range. Frequent words lose a little mass, rare words gain proportionally much more, so the model sees rare words as negatives often enough to learn something about them, while frequent function words still dominate. The two-word example from part 06 (slide 60, SLP3 eq. 5.20) applies unchanged: with P(a) = 0.99 and P(b) = 0.01, the rare word roughly triples its chance of being drawn, to about 0.03, while a drops only to 0.97. The arithmetic is identical; what changes is the role. In part 06 the flattened distribution sat in a PMI denominator, here it decides which words are sampled as noise.
Mikolov et al. (2013) report that U(w)^{3/4} / Z "outperformed significantly" both the raw unigram and the uniform distribution. They recommend k between 5 and 20 for small datasets and 2 to 5 for large ones. The reference C code and gensim both default to negative = 5 and an exponent of 0.75 (gensim calls it ns_exponent). Levy, Goldberg and Dagan (2015) named the same move context distribution smoothing and showed it consistently improves count-based PMI vectors too, not just word2vec.
Recall
How are noise words chosen, and what does α = 0.75 do?
Quick check
Why does SGNS draw noise words from the unigram raised to 0.75?
Fix one training instance: target w = apricot, positive context c_pos = jam, and k = 2 noise words matrix and Tolstoy. The loss for this instance is
The first term is small when the classifier is confident that jam is a neighbor. Each of the other two is small when the classifier is confident that the noise word is not. Where does this shape come from? We want the classifier to give high probability to the correct label of every pair in the instance: label + for the positive and label − for each noise word. Treating the k + 1 decisions as independent, and minimizing the negative log of their joint probability, gives the cross-entropy loss in four short steps.
- Independence turns the joint probability of the k + 1 labels into a product, and the minus log turns maximizing probability into minimizing a loss.
- The log of a product is a sum of logs.
- A pair is either a neighbor or not, so P(− | w, c) = 1 − P(+ | w, c).
- Plug in P(+ | w, c) = σ(c · w) from part 07 and use the Sigmoid identity 1 − σ(x) = σ(−x).
The negative terms are not decoration. Goldberg and Levy (2014) point out that with positives alone the objective has a trivial solution: make every vector the same, with a large enough norm that every dot product is huge, and every positive gets probability 1 (they note this happens once the dot product reaches about 40). Such vectors are useless, since every word looks like every other. The noise terms make that collapse expensive, because identical vectors would also give every noise pair probability 1.
Recall
Write the SGNS loss for one target w with one positive and k negatives.
Take the instance from the previous concept, "...apricot jam..." with k = 2 noise words matrix and Tolstoy. One learning step does three things at once. It moves apricot's target vector w and jam's context vector c_jam toward each other, so c_pos · w rises. It moves w and c_matrix apart, and w and c_Tolstoy apart, so both c_neg · w values fall. Every other row of θ, including aardvark and zebra in both W and C, is left exactly as it was.
The rule behind the picture is gradient descent. The gradient ∇_θ L is the vector of partial derivatives of the loss with respect to every parameter; it points in the direction in which the loss grows fastest. To reduce the loss, step the opposite way, by an amount scaled by the Learning rate η.
"Stochastic" means the gradient is computed from one training instance (or a small batch) at a time, not from the whole corpus. The training loop walks through the corpus, takes each target and each of its window positives, samples k noise words, computes the loss for that small instance, and steps. Because the loss of one instance involves only w, c_pos and the k noise contexts, the gradient is zero for every other row, and the update touches just 2 + k rows out of 2|V|. That sparsity is what makes word2vec fast enough to train on billions of tokens.
Recall
Which rows of θ change after one (w, c_pos) step with k = 2?
Quick check
For one positive pair with k = 2, which vectors does a single SGD step change?
The fastest way to trust the update equations is to run one by hand. Use tiny vectors so every number is visible, then read the general rule off the arithmetic.
Worked example
One SGD step in two dimensions
Setup
d = 2, η = 0.5, w = (1, 0), c_pos = (0, 1) for jam, c_neg1 = (1, 1) for matrix, c_neg2 = (0.5, −1) for Tolstoy.Dot products and sigmoids
c_pos · w = 0, c_neg1 · w = 1, c_neg2 · w = 0.5. So σ(0) = 0.5, σ(1) = 0.731, σ(0.5) = 0.622. The classifier is unsure about jam and wrongly leans toward calling both noise words neighbors.Loss before the step
L = −[log 0.5 + log(1 − 0.731) + log(1 − 0.622)] = 0.693 + 1.313 + 0.974 ≈ 2.980.Gradients (all at time t)
Each gradient is (σ(c · w) − y) times the partner vector, with y = 1 for jam and y = 0 for noise words; the derivation follows below. ∂L/∂c_pos = (0.5 − 1) w = (−0.5, 0); ∂L/∂c_neg1 = 0.731 w = (0.731, 0); ∂L/∂c_neg2 = 0.622 w = (0.622, 0); ∂L/∂w = −0.5 c_pos + 0.731 c_neg1 + 0.622 c_neg2 = (1.042, −0.391).Apply θ − η ∇L
c_pos = (0.25, 1.0), c_neg1 = (0.634, 1.0), c_neg2 = (0.189, −1.0), w = (0.479, 0.196).Check the effect
New dot products: c_pos · w = 0.315 (up from 0), c_neg1 · w = 0.500 (down from 1), c_neg2 · w = −0.105 (down from 0.5).Result
The loss falls from 2.980 to 2.163 in one step. The positive was pulled up and both negatives were pushed down, exactly the picture in the previous concept.
Where the gradients come from
Two derivative facts do all the work, both about the Sigmoid derivative. From dσ/dz = σ(z)(1 − σ(z)) it follows that d/dz log σ(z) = 1 − σ(z) and d/dz log σ(−z) = −σ(z). Apply them with the chain rule, remembering that z = c · w has derivative w with respect to c and c with respect to w, and the minus sign in front of the loss flips the signs.
All three have the same shape: (prediction minus label) times the other vector. Write y = 1 for a positive pair and y = 0 for a noise pair; each gradient is (σ(c · w) − y) times the partner vector. The reference word2vec.c computes exactly this as g = (label − sigmoid) * alpha, folding the learning rate and the sign into one number. Plugging the gradients into θ^{t+1} = θ^t − η ∇L gives the three update rules.
| Parameter | Label y | Gradient | Direction of the update |
|---|---|---|---|
| c_pos | 1 | [σ(c_pos·w) − 1] w | Toward w (the factor is negative) |
| c_neg_i | 0 | σ(c_neg_i·w) w | Away from w |
| w | both | [σ(c_pos·w) − 1] c_pos + Σ σ(c_neg_i·w) c_neg_i | Toward c_pos, away from each c_neg_i |
Recall
Derive ∂L/∂c_pos.
Quick check
The update for a negative context is η·σ(c_neg·w)·w. When is it nearly zero?
When training ends, apricot still has two rows, w_apricot and c_apricot. SLP3 says the common choice is to add them and represent apricot by w_apricot + c_apricot; the alternative is to keep only w_apricot and throw C away. Either way the result is a Static embedding: one fixed vector per word type, compared with Cosine similarity.
Why keep the tables apart during training at all? Goldberg and Levy (2014) give the argument. Suppose dog had a single vector v used in both roles. Then the score for the pair (dog, dog) would be v · v = |v|², which is large for any vector with a large norm, so the model would believe dog is its own most likely neighbor. Real text rarely says "dog dog", so the model would have to keep word norms small to avoid that, fighting its own objective. Two tables remove the conflict.
Why does adding them afterwards help? Levy, Goldberg and Dagan (2015) show that the cosine of two summed vectors includes terms like w_x · c_y, which measure whether x and y appear in each other's contexts. Adding context vectors therefore adds first-order similarity (co-occurrence) to the second-order similarity (shared neighbors) that W alone captures. The trick comes from GloVe, whose authors summed the two sets of vectors as a cheap way of combining two models, much like an ensemble.
| Option | Vector for word i | What it captures | In practice |
|---|---|---|---|
| Keep W only | w_i | Second-order similarity: two words are close when they have similar neighbors | Smallest vectors; the default in gensim's wv |
| Sum | w_i + c_i | Adds first-order similarity terms: words that co-occur also get closer | SLP3's common choice; behaves like an ensemble |
| Concatenate | [w_i ; c_i] | Keeps both views separately, doubling the dimension | Rarely used; costs 2d per word |
The whole recipe
- Initialize W and C randomly: 2|V| vectors of dimension d. (word2vec.c draws W uniformly in ±0.5 / d and starts C at zeros.)
- Slide a window over the corpus. Each (target, neighbor) pair is a positive; for each, draw k noise words from P_α as negatives.
- Train the logistic classifier σ(c · w) to separate positives from negatives, minimizing L_CE with SGD.
- Throw the classifier away and keep the learned rows as the embeddings, summed or W only.
Step 4 is the heart of Self-supervision: the yes-or-no prediction task was never the goal. It was a pretext that forced the model to place words with similar contexts close together, and the Embedding is the by-product we wanted.
Recall
Why keep W and C separate during training, then sum them?
Quick check
Why does SGNS train separate W and C tables instead of one vector per word?
Recap
If you remember nothing else
- θ = [W; C]: 2|V| vectors of size d. W holds targets and C holds contexts and noise words.
- Each window positive gets k noise words from P_α(w) ∝ count(w)^0.75, never the target itself.
- L_CE = −[log σ(c_pos·w) + Σ log σ(−c_neg_i·w)], minimized with SGD.
- The gradients are (σ − y) times the other vector: [σ(c_pos·w) − 1]w, σ(c_neg·w)w, and the combined sum for w.
- Updates take the form θ^{t+1} = θ^t − η∇L. Use time-t values throughout. Only 2 + k rows change per pair.
- After training, discard the classifier and keep w_i + c_i, or w_i alone.
Sources
- Speech and Language Processing (3rd ed. draft), Chapter 5: EmbeddingsBookJurafsky and Martin, StanfordSection 5.5.2, eqs. 5.19 to 5.27, Figs. 5.6 and 5.7, Exercise 5.3(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov et al., NeurIPS 2013Negative sampling objective (eq. 4), k of 5 to 20 or 2 to 5, U(w)^(3/4) noise distribution(opens in a new tab)
- word2vec Explained: Deriving Mikolov et al.'s Negative-Sampling Word-Embedding MethodPaperGoldberg and Levy, arXiv 2014Derivation, why separate target and context vectors, the trivial solution without negatives(opens in a new tab)
- Neural Word Embedding as Implicit Matrix FactorizationPaperLevy and Goldberg, NeurIPS 2014SGNS factorizes a shifted PMI matrix(opens in a new tab)
- Improving Distributional Similarity with Lessons Learned from Word EmbeddingsPaperLevy, Goldberg and Dagan, TACL 2015Adding context vectors (w + c), context distribution smoothing 0.75, PMI − log k(opens in a new tab)
- word2vec Parameter Learning ExplainedPaperRong, arXiv 2014Input and output vectors, update equations with learning rate η and prediction error(opens in a new tab)
- word2vec reference C implementationDocsMikolov, GitHubnegative = 5, power = 0.75, alpha = 0.025 decaying linearly toward zero, initialization, g = (label − sigmoid) * alpha(opens in a new tab)
- gensim Word2Vec APIDocsgensimnegative = 5, ns_exponent = 0.75, alpha 0.025 decaying linearly to gensim's own min_alpha 0.0001(opens in a new tab)
Part 09: Hyperparameters, word2vec variants and window size
Standard SGNS settings, the CBOW, GloVe, FastText and skip-thought alternatives, and how the context window size decides whether neighbors are syntactic or topical.
6 concepts, slides 89-97
Why this part matters
Picking k, d and the window size is the first thing you do when you train embeddings on a thesis corpus or an Arabic system, and the defaults you inherit from a library are not always the ones the papers recommend. CBOW, GloVe, FastText and skip-thought are the standard comparisons on an exam, and the window size quietly decides whether your vectors find synonyms or topics.
The part opens with the settings that make skip-gram work in practice and what each one trades. It then walks through four relatives of skip-gram, each changing one design decision: CBOW flips the prediction direction, GloVe replaces the sliding window with a global count matrix, FastText breaks words into character pieces, and skip-thought moves the whole idea from words to sentences. It closes with the one hyperparameter that changes not how good the vectors are but what kind of similarity they encode.
By the end you can
- State the standard SGNS hyperparameters and explain what raising k, d or the window costs and buys.
- Contrast CBOW with skip-gram by input, output, speed and rare-word quality.
- Write the GloVe objective and explain each part of its weighting function f.
- Decompose a word into FastText n-grams and explain how unseen words get vectors.
- Explain skip-thought vectors as skip-gram lifted to sentences.
- Predict whether a small or large window yields functional or topical neighbors.
Two students train skip-gram on the same afternoon. One has a clinical corpus of 5 million tokens; the other has 6 billion tokens of web text. Should they use the same number of negative samples? No. The small corpus gives each word only a few true context pairs, so each of those precious pairs has to do more work: contrasting it against many noise words, k = 15 to 20 (the paper allows 5 to 20), squeezes more signal out of it. On the huge corpus every word already sees thousands of real contexts, and k = 2 to 5 is enough while costing a fraction of the compute.
That is the pattern for every knob in skip-gram with negative sampling: each one trades quality against compute or memory, and the right value depends on how much data you have. The lecture gives a "somewhat standard" setting: the model is SGNS, 15 to 20 negative samples for smaller datasets and 2 to 5 for the huge datasets that are usually used, a dense vector of 300 dimensions (100 or 50 also work), and a sliding context window of 5 to 10.
The SGNS knobs, their usual values, and what raising each one buys and costs
- Model
- Skip-gram with negative sampling (SGNS). The alternatives in this part (CBOW, GloVe, FastText) are judged against it.
- Negatives k
- 15 to 20 on small corpora per the slide (5 to 20 per Mikolov et al.), 2 to 5 on very large corpora. Raising k sharpens the contrast for each true pair and costs one more dot product per positive.
- Dimension d
- 300 is the common choice; 100 or 50 also work. Raising d gives room for more distinctions but grows memory and every dot product linearly, with small gains past 300.
- Window half-width m
- 5 to 10 on the slide, read as words per side, matching word2vec's and Gensim's window parameter (SLP3 says 1 to 10 per side). A ±m window gives 2m context words, which is part 07's L. Raising it adds more pairs per position and shifts neighbors from functional to topical (last concept of this part).
- Noise distribution
- Unigram counts raised to 3/4, the α of weighted PPMI, so rare words are drawn as negatives a little more often than their raw frequency.
- Subsampling threshold t
- Around 10^-5 in Mikolov et al.: very frequent words such as the are randomly dropped, which speeds training and helps rare words.
What each knob costs
The cost of negatives is easy to count. For every true (target, context) pair, the skip-gram classifier computes one dot product for the positive and one for each of the k noise words, and each dot product sends a gradient into one row of the context matrix:
Going from k = 5 to k = 20 therefore makes training about 3.5 times slower (21 / 6). The dimension sets the memory. The model holds two matrices, the target and context matrices W and C, each with one row of length d per vocabulary word:
Two smaller settings travel with these. Negatives are not drawn by raw frequency but from the unigram distribution raised to 3/4, the same α = 0.75 trick that weighted PPMI uses, which gives rare words a slightly better chance of being picked as noise. And very frequent words are randomly discarded with a threshold around 10^-5, so that the model does not spend most of its updates on pairs like (cat, the) (Mikolov et al. 2013b).
The defaults you meet in practice differ from the slide, which matters when you compare your results to a paper. Gensim, the usual Python library, defaults to CBOW (the variant explained in the next concept), not skip-gram, with 100 dimensions.
| Setting | d | Window | k | Model |
|---|---|---|---|---|
| Mikolov et al. 2013b experiments | 300 | 5 | 5 to 20 (small data) | skip-gram |
| Gensim Word2Vec | 100 | 5 | 5 | CBOW (sg=0) |
| fastText | 100 | 5 | 5 | skip-gram |
| This slide | 300 (or 100, 50) | 5 to 10 | 15 to 20 (small), 2 to 5 (huge) | SGNS |
Quick check
According to the lecture, which range of k suits a small training corpus?
Recall
State the standard SGNS hyperparameters from the lecture.
Take the sentence "I saw a cute grey cat playing in the garden" with a ±2 window around cat. Skip-gram, the model of the last two parts, turns this position into four separate training pairs: (cat, cute), (cat, grey), (cat, playing), (cat, in). The continuous bag of words model, the other half of word2vec, runs the arrow the other way. It looks up the four context vectors for cute, grey, playing and in, combines them into a single hidden vector h, scores every vocabulary word against h, and is trained so that the highest score goes to cat. One position, one prediction.
The combination step is what gives CBOW its name. The context vectors are pooled into one, which throws away their order (a bag), and they are dense real-valued vectors rather than counts (continuous). In Mikolov et al.'s original description the projection layer is shared so that "all words get projected into the same position (their vectors are averaged)", and "the order of words in the history does not influence the projection":
Worked example
One sentence, two models
Pick the window
Target position: cat. With m = 2, the context words are cute, grey (left) and playing, in (right).
CBOW builds one training instance
h = (w_cute + w_grey + w_playing + w_in) / 4. The model scores h against the output vectors and the loss pushes the score of cat up. Four vectors go in, one gradient signal comes out, and it is shared equally by the four context words.
Skip-gram builds four
The pairs (cat, cute), (cat, grey), (cat, playing) and (cat, in) are each scored and updated separately, each with its own k negatives. The center vector of cat receives four updates.
Result
CBOW does one prediction per position; skip-gram does 2m = 4. That factor is why CBOW trains faster and why skip-gram gives each word, including rare ones, more direct updates.
The consequences follow from that count. With negative sampling, CBOW does 1 + k dot products per position and skip-gram 2m × (1 + k), so skip-gram is about 2m times slower (Mikolov et al. 2013a report the same gap with hierarchical softmax). CBOW's averaging also smooths: a rare word that appears as context is blended with its frequent neighbors before any prediction is made, so its own vector gets a diluted signal. Skip-gram gives every occurrence of a rare word its own pairs, which is why it handles rare words better and in Mikolov et al.'s experiments did better on semantic analogies. The fastText tutorial adds a practical note: skip-gram "works better with subword information than cbow".
| CBOW | Skip-gram | |
|---|---|---|
| Input | The bag of 2m context vectors, averaged into one h | One center vector |
| Output | A score for the center word | A score for each context word, one pair at a time |
| Predictions per position | 1 | 2m |
| Dot products per position (SGNS) | 1 + k | 2m × (1 + k) |
| Word order inside the window | Ignored | Ignored (each pair is independent) |
| Rare words | Smoothed away by averaging with frequent neighbors | Each occurrence gives its own updates; better |
| Strength | Speed on large corpora, frequent words | Rare words, semantic analogies, subword extensions |
Quick check
In CBOW, what is the input and what is the prediction target?
Recall
In one sentence each, what do CBOW and skip-gram predict, and which is better for rare words?
Consider three cells of a word-word co-occurrence matrix. The pair (ice, solid) was seen 10 times, (the, of) 1000 times, and (ice, fashion) never. GloVe wants a dot product for each pair that matches the log of its count: about log 10 ≈ 2.30 for (ice, solid) and log 1000 ≈ 6.91 for (the, of). It does not trust every cell equally. With the published settings, the ice and solid pair gets weight 0.178, the pair the and of gets the maximum weight 1 and no more, and the ice and fashion pair gets weight 0, so its undefined log 0 never enters the loss.
That is the whole method. GloVe, Global Vectors (Pennington, Socher and Manning 2014), first counts the term-context matrix N(w, c) over the corpus, the same matrix that PPMI reweights. It then fits a Dot product of a word vector and a context vector, plus two learned biases, to the log count in every cell, as a weighted least-squares regression:
The weighting function f is where the design lives. Pennington et al. ask three things of it. It must vanish at zero, because log 0 is undefined and because zero cells are 75 to 95 percent of the matrix. It must not decrease, "so that rare co-occurrences are not overweighted", since a pair seen once is mostly noise. And it must stay "relatively small for large values of x, so that frequent co-occurrences are not overweighted". Their choice is a power curve that flattens into a cap:
| Count x | Weight f(x) | Example |
|---|---|---|
| 1 | 0.0316 | A single co-occurrence: barely trusted |
| 5 | 0.1057 | |
| 10 | 0.1778 | (ice, solid) |
| 50 | 0.5946 | |
| 100 | 1 | x = x_max: full weight reached |
| 1000 | 1 | (the, of): capped, cannot dominate |
Worked example
Three pairs through the loss
(ice, solid), 10 co-occurrences
Target log 10 ≈ 2.30. Weight (10/100)^0.75 ≈ 0.178. If the current prediction u·v + b + b̄ is 1.30, the term contributes 0.178 × 1.0² = 0.178.
(the, of), 1000 co-occurrences
Target log 1000 ≈ 6.91. The count is above x_max, so the weight is capped at 1. Without the cap, the same curve would give (1000/100)^0.75 ≈ 5.6, and this one function-word pair would count as much as about 32 pairs like (ice, solid).
(ice, fashion), 0 co-occurrences
Weight f(0) = 0. The term vanishes, so the undefined log 0 is never evaluated and the optimizer only visits the nonzero cells.
Result
Rare pairs count a little, mid-frequency pairs count a lot, and very frequent pairs are capped. The slide's sum over all w, c ∈ V silently relies on f(0) = 0 to skip the empty cells.
GloVe sits between the two families in this lecture. Like PPMI it is built on global matrix statistics collected once, and like word2vec it learns dense vectors whose dot products carry the meaning. Jurafsky and Martin describe it as "based on ratios of probabilities from the word-word co-occurrence matrix", capturing global corpus statistics as count methods do while learning dense vectors as word2vec does. Like SGNS it ends with two vectors per word, and the authors use the sum W + W̃ as the final embedding, the same trick as adding w and c in skip-gram. The slide's b̄_w is just the second bias set, written b̃_j in the paper. Pennington et al. also chose 300 dimensions for their main results, and they report that α = 3/4 gave "a modest improvement over a linear version with α = 1".
| SGNS | GloVe | |
|---|---|---|
| Data it reads | A stream of (target, context) pairs from sliding windows | The global co-occurrence matrix, counted once |
| Objective | Logistic loss: true pairs up, sampled noise pairs down | Weighted squared error between u·v + biases and log N(w,c) |
| Negatives | k sampled per positive pair | None: zero cells get weight f(0) = 0 and drop out |
| Frequency control | Noise drawn from U(w)^(3/4), subsampling of frequent words | Weight f(x) = (x/x_max)^(3/4), capped at 1 |
| Final vectors | Target matrix W, or W + C | W + W̃ (target plus context) |
Recall
Why does GloVe's weighting function need f(0) = 0, and what does capping f at 1 above x_max do?
Take the word where. FastText first wraps it in boundary symbols, <where>, so that prefixes and suffixes can be told apart from the middle of a word. With n = 3 it then slides a three-character window across: <wh, whe, her, ere, re>. It also keeps the whole word <where> as one more unit. Each of those six units has its own vector, and the vector of where is their sum.
Bojanowski et al. 2017 built this on the skip-gram model. Write 𝒢_w for the set of n-grams of word w, including the word itself, and z_g for the vector of n-gram g. The word's vector and its skip-gram score with a context word c become:
Worked example
All the n-grams of where at the real settings
Pad
where becomes <where>, 7 characters.
n = 3 (5 n-grams)
<wh whe her ere re>
n = 4, 5 and 6 (4 + 3 + 2 n-grams)
<whe wher here ere>, then <wher where here>, then <where where>.
Add the special whole-word sequence
<where> is added as its own unit, distinct from the n-gram where found inside it.
Result
The paper extracts "all the n-grams for n greater or equal to 3 and smaller or equal to 6", which gives 14 n-grams here, plus the whole word: 15 vectors summed into one.
The boundary symbols matter more than they look. The paper's own example: "the sequence <her>, corresponding to the word her, is different from the tri-gram her from the word where". So the pronoun and the piece of where get separate vectors, while the prefix <wh is shared by where, when, what and which.
What the pieces buy
Plain word2vec gives each word type its own row, so a word missing from training has no vector at all. FastText builds a vector for an unseen word such as wherever by summing the vectors of its n-grams, most of which (<wh, whe, her, ere) were learned from other words. The same sharing helps rare inflections in morphologically rich languages such as Arabic, Turkish and German: a rarely seen form borrows statistics from frequent forms that share its stem and affixes. Pretrained FastText vectors exist for 157 languages.
The price is computation, which the slide calls "a lot of additional computation": every update now touches about fifteen vectors for the target instead of one. Memory is kept bounded by hashing all n-grams into a fixed table of 2,000,000 buckets (the bucket default), so collisions are allowed rather than storing every possible substring.
Quick check
FastText meets the unseen word 'wherever' at test time. Which vector does it return?
Recall
List the FastText 3-gram units for 'where' and explain how FastText builds a vector for an unseen word.
Take three consecutive sentences from a novel: "I got back home. I could see the cat on the steps. This was strange." A skip-thought model reads the middle sentence and compresses it into one vector. Two decoders then have to regenerate the neighbors from that vector alone, word by word: one writes "I got back home", the other writes "This was strange".
Generates the previous sentence, "I got back home <eos>", conditioned on the vector.
Reads "I could see the cat on the steps" and outputs one sentence vector.
Generates the next sentence, "This was strange <eos>", conditioned on the vector.
The idea is skip-gram moved up one level. Skip-gram uses a word to predict its neighboring words; Kiros et al. 2015 use a sentence to predict its neighboring sentences. The distributional hypothesis carries over: sentences that appear in similar discourse surroundings, with similar sentences before and after them, are pushed toward similar vectors. In the paper's words, "the sentence s_i is encoded and tries to reconstruct the previous sentences_i−1 and next sentence s_i+1".
Three things change on the way up. A sentence is not one vocabulary item that can be looked up, so the encoder is a recurrent network with GRU units that reads the words in order, which makes the vector order-sensitive. The targets are whole sentences, so the decoders generate text instead of classifying pairs, and there is no negative sampling. And the vectors are big: the model was trained on the BookCorpus (11,038 books, 74,004,228 sentences), the unidirectional encoder gives 2400 dimensions and the combined model 4800. The authors froze these vectors and trained only linear classifiers on top of them for 8 evaluation tasks.
| Skip-gram | Skip-thought | |
|---|---|---|
| Unit | A word | A sentence |
| Context | Words within ±m | The previous and the next sentence |
| Encoder | A lookup in W (one row per word) | A GRU recurrent network over the words, order-sensitive |
| Objective | Classify (target, context) pairs as real or noise | Generate each neighbor sentence word by word |
| Negatives | k sampled per positive | None: decoders use a softmax over the vocabulary |
| Output size | d, typically 100 to 300 | 2400 (uni-skip), 4800 (combine-skip) |
Recall
What does a skip-thought model encode and what does it predict?
Properties of embeddings start with the window. Levy and Goldberg 2014 trained skip-gram on English Wikipedia twice, changing only the window, and looked up the nearest neighbors of Hogwarts. With a ±2 window they were evernight, sunnydale, garderobe, blandings and collinwood: other fictional schools and houses, words that fill the same slot in a sentence. With a ±5 window they were dumbledore, hallows, half-blood, malfoy and snape: the world of Harry Potter. Same corpus, same model, a different idea of what "similar" means.
The reason is what each window can see. Two words on each side of a noun are mostly its syntactic frame: the determiner before it, the preposition, the verb that takes it as an object ("students at Hogwarts", "returned to Sunnydale"). Words that share those frames are words of the same type and function, so a short window rewards similarity. Linguists call this a paradigmatic, or second-order, association: the two words rarely appear together but appear in the same surroundings. Five words on each side reach past the frame into the topic of the passage, so a long window rewards relatedness, words from the same semantic field. That is a syntagmatic, or first-order, association: the words appear near each other.
The slides show the same split with Voita's examples. With larger windows, dog groups with bark and leash, and walking with walked and run: words about the same activity. With smaller windows, Poodle groups with Pitbull and Rottweiler, and walking with running and approaching: words that could replace one another. Jurafsky and Martin summarize it as shorter windows giving representations that are "a bit more syntactic", with neighbors that are "semantically similar words with the same parts of speech", while longer windows give words that are "topically related but not similar".
| Small window | Large window | |
|---|---|---|
| Window | ±2 (small) | ±5 or more (large) |
| What the window sees | The syntactic frame: determiners, prepositions, the verb that takes the word | The whole topic of the passage |
| Neighbor kind | Functional similarity: same slot, same part of speech | Topical relatedness: same semantic field |
| Hogwarts (Levy and Goldberg) | sunnydale, evernight, garderobe, blandings, collinwood | dumbledore, hallows, half-blood, malfoy, snape |
| Dog example (Voita) | Poodle, Pitbull, Rottweiler | dog, bark, leash |
| Verb example (Voita) | walking, running, approaching | walking, walked, run |
| Good for | Synonym finding, POS-like features, slot filling | Topic modelling, retrieval, query expansion |
The same logic explains the window in the term-context matrix of the count methods: it is one hyperparameter shared by PPMI, SGNS and GloVe, and it changes what all of them learn. Levy and Goldberg push it one step further by replacing the window with syntactic dependency contexts. Hogwarts' neighbors then become sunnydale, collinwood, calarts, greendale and millfield: even more purely functional, all schools and fictional places.
Quick check
Skip-gram is trained with a context window of plus or minus 2. Which neighbors does Hogwarts get?
Recall
With a ±2 versus a ±5 window, what are Hogwarts' nearest neighbors, and why?
Recap
If you remember nothing else
- Standard SGNS: k = 15 to 20 on small data (5 to 20 in Mikolov), 2 to 5 on huge data, d = 300 (or 100, 50), window 5 to 10.
- Each extra negative adds one dot product per positive pair; d sets the 2|V|d parameter count.
- CBOW predicts the center from the averaged context bag: faster, smoother. Skip-gram predicts the context from the center: better for rare words.
- GloVe fits u_c·v_w + b_c + b_w to log N(w,c) by weighted least squares; f(x) = (x/100)^0.75 below 100, then 1, and f(0) = 0 skips zero cells.
- FastText adds < and > boundaries, n-grams of 3 to 6 characters and the whole word; a word is the sum of its n-gram vectors, so unseen words still get vectors.
- Skip-thought encodes a sentence with a GRU and decodes the previous and next sentences, giving 2400-dimensional sentence vectors.
- A ±2 window yields functional look-alikes (Hogwarts near Sunnydale); a ±5 window yields topic-mates (Hogwarts near Dumbledore).
Sources
- Speech and Language Processing, chapter 5: EmbeddingsBookJurafsky and Martin, draft of August 2026Window of 1 to 10 per side, short versus long windows, the Hogwarts example, first- and second-order co-occurrence, GloVe and FastText summaries(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov, Chen, Corrado and Dean, 2013CBOW averages context vectors and ignores order; training complexity of CBOW and skip-gram; sampling distant words less(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov, Sutskever, Chen, Corrado and Dean, NeurIPS 2013k of 5 to 20 for small data and 2 to 5 for large, U(w)^(3/4) noise, subsampling threshold near 10^-5, d = 300(opens in a new tab)
- word2vec Parameter Learning ExplainedPaperXin Rong, 2014The CBOW network figure on the slide; CBOW takes the average of the context vectors(opens in a new tab)
- GloVe: Global Vectors for Word RepresentationPaperPennington, Socher and Manning, EMNLP 2014Equations 8 and 9, x_max = 100, α = 3/4, zero entries are 75 to 95 percent of X, final vectors W + W̃(opens in a new tab)
- GloVe project pageDocsStanford NLP GroupPretrained GloVe vectors(opens in a new tab)
- Enriching Word Vectors with Subword InformationPaperBojanowski, Grave, Joulin and Mikolov, TACL 2017Boundary symbols, n from 3 to 6, the whole word included, <her> versus her, sum of n-gram vectors, OOV words(opens in a new tab)
- Word representations tutorialDocsfastTextDefault dimension 100, subwords of 3 to 6 characters, skip-gram works better with subwords than CBOW(opens in a new tab)
- fastText source: getWordVectorDocsFacebook ResearchgetWordVector divides the summed n-gram vectors by their count; Model::computeHidden in src/model.cc averages the input rows during training(opens in a new tab)
- Skip-Thought VectorsPaperKiros, Zhu, Salakhutdinov, Zemel, Torralba, Urtasun and Fidler, NeurIPS 2015GRU encoder, two decoders, BookCorpus, 2400 and 4800 dimensional vectors, 8 evaluation tasks(opens in a new tab)
- Dependency-Based Word EmbeddingsPaperLevy and Goldberg, ACL 2014Table 1 Hogwarts neighbors for BoW2, BoW5 and dependency contexts; uniform window sampling(opens in a new tab)
- models.word2vec: Word2vec embeddingsDocsGensimDefaults vector_size 100, window 5, negative 5, sg 0 (CBOW), cbow_mean 1, ns_exponent 0.75(opens in a new tab)
- NLP Course For You: Word EmbeddingsArticleLena VoitaSource of the slide 89 wording, the CBOW sum figure and the window size examples on slide 97(opens in a new tab)
Part 10: Analogies, bias and evaluating embeddings
The parallelogram method for analogies and its limits, embeddings as a lens on historical meaning change and cultural bias, visualizing and evaluating embeddings intrinsically and extrinsically, and when to use pretrained embeddings.
7 concepts, slides 98-111
Why this part matters
You have trained an embedding. Four questions follow immediately, and this part answers each. What has the space learned? Relations such as male to female or country to capital show up as directions you can probe with one subtraction and one addition. Can you trust it? Analogy scores flatter it, and it carries the biases of its corpus, which matters for any hiring, search or Arabic NLP system built on top of it. How do you measure it? Intrinsic tests against human judgments, or extrinsic tests inside a real task. And should you train it yourself at all, or take vectors someone else trained?
For exams, the parallelogram formula, the intrinsic versus extrinsic split and allocational versus representational harm are standard items. For research, diachronic embeddings and bias measurement are live tools: the same machinery that tracks how awful changed meaning also measures a century of gender stereotypes.
By the end you can
- Compute an analogy answer with the parallelogram method, using argmin distance or argmax cosine, and exclude the input words.
- Explain why analogy accuracy overstates relational knowledge, citing the exclusion effect and relation-dependence.
- Describe how aligned decade embeddings reveal semantic change, and state the laws of conformity and innovation.
- Explain how embeddings encode and amplify cultural bias, distinguish allocational from representational harm, and describe Garg's relative norm measure.
- Classify evaluations as intrinsic or extrinsic, and compute a Spearman correlation against human ratings.
- Choose between frozen, fine-tuned and jointly trained embeddings given data size and task difficulty.
Apple is to tree as grape is to what? Picture the arrow that starts at apple and ends at tree. It means something like "fruit to the plant it grows on". Now pick that same arrow up, keep its length and direction, and set its tail down on grape. Its head lands near vine. That is the whole method.
The same move works on the famous examples. Take king, subtract man, add woman, and the point you reach is close to queen. Take Paris, subtract France, add Italy, and you land close to Rome. In each case the subtraction isolates a relation (royalty without the maleness, the capital-of relation without the particular country) and the addition applies it to a new word. Because the four points form a parallelogram when it works, this is called the Parallelogram method.
The rule, in two equivalent forms
The slides write an analogy as a : a* :: b : b*, read "a is to a* as b is to b*". The relation is the offset from a to a*, so the point to search around is t = a* − a + b. No vocabulary word sits exactly at t, so the answer is the word nearest to it, and the three question words themselves are taken out of the candidate pool (the next concepts show why that matters so much). With Euclidean distance, nearest means smallest distance, so the operator is an argmin:
Mikolov, Yih and Zweig, who made the method famous for dense vectors, wrote it the other way round. They normalize every vector to unit length, compute the same target, and return the word with the largest Cosine similarity to it. Nearest by distance and most similar by cosine are the same idea, one written as a minimization and the other as a maximization. When the candidate vectors have unit length, the two rankings are identical, because ‖x − t‖² = 1 + ‖t‖² − 2 x·t.
Worked example
man : woman :: king : ? in two dimensions
Place the words
man (1, 1), woman (1, 3), king (4, 1), queen (4.2, 3.1), princess (3, 3.5).Build the target
a = man, a* = woman, b = king, so t = (1, 3) − (1, 1) + (4, 1) = (4, 3). The offset (0, 2) is the "male to female" arrow.Measure every candidate
Distances to t: queen 0.224, princess 1.118, king 2.0, woman 3.0, man 3.606.Result
The nearest allowed word is queen. Note that it is near t, not on it: the method always ends in a nearest-neighbour search.
- 1. queen0.224
- 2. princess1.118
- 3. prince1.803
- 4. crown2.28
- 5. throne2.786
Try it above. Pick any three words, watch the a to a* arrow get copied onto b, and read the ranking. Switch the space to the small offset and the candidate pool to "allow input words" and keep the playground in mind for the third concept of this part.
Where the idea came from
The parallelogram is older than embeddings. Rumelhart and Abrahamson proposed it in 1973 as a model of how people solve analogies, working in a space of mammal names built from human similarity judgments (apple : tree :: grape : vine is the illustration SLP3 uses for it). Turney and Littman showed in 2005 that sparse count vectors could solve SAT-style analogies, and Mikolov and colleagues brought it to dense neural embeddings in 2013. Their NAACL paper, titled Linguistic Regularities in Continuous Space Word Representations, built a syntactic test set of 8,000 questions and found its recurrent network vectors answered almost 40% correctly. The slide labels it "Mikolov et al. 2013b" and SLP3's bibliography labels the same paper 2013c, so cite it by title.
Recall
Write the parallelogram method for a : a* :: b : b*, and state the two things you must change on slide 99's version.
Quick check
For man : woman :: king : ?, which point does the parallelogram method search around?
Plot man, woman, uncle, aunt, king and queen from Mikolov, Yih and Zweig's recurrent-network language model (the precursor of Word2vec) in two dimensions and draw an arrow from each male word to its female partner. The three arrows come out roughly parallel and roughly the same length. In a second projection of the same space, king to kings and queen to queens are parallel to each other too, and that plural direction cuts across the gender direction.
This is what it means for a relation to be linear in an Embedding space: the offset vector between the two words of a pair is nearly the same for every pair that stands in that relation. Mikolov and colleagues found such offsets for gender (man to woman), verb tense (walking to walked) and country to capital (Spain to Madrid), and their larger 2013 test set organized analogy questions by exactly these semantic and syntactic families. The same picture holds for GloVe: the GloVe project page shows man to woman, sir to madam, heir to heiress, king to queen, uncle to aunt, nephew to niece, brother to sister, earl to countess, duke to duchess and emperor to empress as ten segments that all tilt the same way, and Pennington and colleagues report that offsets also capture comparative and superlative forms.
| Relation | Example pair | What the offset means | Slides |
|---|---|---|---|
| Gender | man → woman | Male form to female form | 100 to 102 |
| Number | king → kings | Singular to plural | 100 |
| Tense | walking → walked | Progressive to past | 101 |
| Capital | Spain → Madrid | Country to its capital city | 101 |
| Title | earl → countess | Male noble title to female counterpart | 102 |
One word can take part in many relations at once. King is the male member of a gender pair, the singular of a number pair, and a royal term next to throne and crown. A 300-dimensional space has room for all of these as different directions, and that is the sentence Mikolov and colleagues put under their figure: in high-dimensional space, multiple relations can be embedded for a single word. Any 2D picture is one projection chosen to show one of those directions, which is why the slides need two panels to show gender and number for the same words.
Recall
How can king lie on a gender direction and a number direction at the same time?
Change the toy space from the first concept so the gender offset is small: man (1, 1), woman (1.4, 1.3), king (4, 1), queen (4.6, 1.9). Now ask man : woman :: king : ?and let every word compete.
Worked example
The exclusion trap
Build the target
t = (1.4, 1.3) − (1, 1) + (4, 1) = (4.4, 1.3).Rank with input words allowed
Distance to king is 0.5, distance to queen is 0.632. The method answers king, the word you gave it.Exclude a, a* and b
Remove man, woman and king from the pool. The nearest remaining word is queen at 0.632.Result
The right answer appears only after exclusion. With a small offset the target barely moves away from b, so b itself is the nearest point.
This is not a toy artefact. Linzen (2016) ran the standard Word2vec analogy benchmark without excluding the inputs: the nearest neighbour of a* − a + b was b in 93% of cases, a* in 5%, and never a. SLP3 makes the same point with cherry : red :: potato : x, which returns potato or potatoes instead of brown unless those are forbidden. Every published analogy accuracy therefore depends on the exclusion rule, and part of the credit belongs to b*'s simply being b's nearest neighbour. Linzen's baselines make this concrete: a method that ignores a, or even both a and a*, and just returns the neighbour of b, scores very high on plurals.
Where the method works and where it does not
- It works for frequent words, for pairs where b* already sits close to b (SLP3's "small distances", which is also why b wins unless it is excluded), and for certain relations (SLP3): country to capital, and inflections such as plural and tense.
- It does poorly on many lexicographic and derivational relations. The BATS set of Gladkova and colleagues has 99,200 questions in 40 categories, against only 15 relations in the Google set, and accuracy varies widely across them.
- Reversing an analogy uses the same offset with the sign flipped, yet Linzen found accuracy dropped in most categories (mean −0.11): US cities fell from .69 to .17 and common capitals from .9 to .53.
- As a model of human analogy making, the parallelogram is too simple: Peterson, Chen and Griffiths (2020) show it cannot account for how people form even simple analogies.
3CosAdd versus 3CosMul
Levy and Goldberg (2014) rewrote the cosine objective with unit vectors and saw it as a balance: two attractors (b* should resemble b and a*) and one repeller (b* should not resemble a). Added together, one large similarity can swamp the others. Their multiplicative version keeps each term in check and generally does better.
Recall
What does the offset method return most often if a, a* and b are allowed as answers?
Quick check
If the input words stay in the candidate pool, what does the offset method usually return?
Follow three words through two centuries of books. In the 1900s gay sits near daft, flaunting, sweet and cheerful; by the 1990s its neighbours are homosexual and lesbian. In the 1850s broadcast sits near sow and seed, a farmer scattering grain; by the 1990s it sits near newspapers, radio and bbc. In the 1850s awful sits near majestic, awe and solemn, full of awe; by the 1900s it sits near terrible and appalling, and by the 1990s near weird and wonderful. That slide from praise to blame is called pejoration.
Hamilton, Leskovec and Jurafsky (2016) produced these pictures with diachronic embeddings. The recipe has three steps. First, train a separate embedding for each decade of text. They compared Positive PMI, SVD and Skip-gram with negative sampling on six historical corpora in four languages, with a window of 4 and 300 dimensions; the English Google Books corpus alone has 8.5 × 1011 tokens covering 1800 to 1999, and COHA has 4.1 × 108 tokens covering 1810 to 2009. Second, align the decades so their axes mean the same thing. Third, measure how far each word moved between aligned decades, and read its old and new neighbours.
Why alignment is needed
SVD and SGNS only care about dot products between vectors, and any rotation of the whole space preserves every dot product. So the 1900 run and the 1990 run can come out rotated relative to each other for no linguistic reason, and comparing the raw coordinates of gay in the two runs measures that arbitrary rotation. Hamilton and colleagues fix this with orthogonal Procrustes: find the orthogonal matrix that best maps one decade's matrix onto the next. Because the matrix is orthogonal it is a rotation (possibly with a reflection), so cosines within each decade are unchanged. PPMI vectors need no alignment, since their dimensions are context words that mean the same thing in every decade.
Two statistical laws
Measuring displacement for thousands of words let Hamilton and colleagues state two laws. The law of conformity: the rate of semantic change scales with an inverse power of word frequency, so frequent words change slowly. The law of innovation: holding frequency fixed, words with more senses (higher Polysemy) change faster.
Recall
Why must decade-specific embeddings be aligned before you measure semantic change, and how?
Recall
State the law of conformity and the law of innovation.
Bolukbasi and colleagues ran the Parallelogram method on Word2vec trained on Google News (3 million words and phrases, 300 dimensions). Asked Paris : France :: Tokyo : x, it answers Japan. Asked father : doctor :: mother : x, it answers nurse. Asked man : computer programmer :: woman : x, it answers homemaker. The same machinery that captured capitals captured stereotypes.
The reason is unsurprising once said aloud. An Embedding is a compressed summary of co-occurrence statistics, so if the training text talks about women and men in different contexts, that difference becomes geometry. Bolukbasi and colleagues found that gender bias is largely captured by a single direction, roughly the she minus he offset, onto which occupation words project unevenly.
Two kinds of harm
| Harm | Definition | Example |
|---|---|---|
| Allocational | A system distributes a resource or opportunity (jobs, loans, search exposure) unfairly across groups. | A resume search that ranks documents by embedding similarity to programmer pushes women's resumes down. |
| Representational | A system demeans, stereotypes or erases a group, whether or not any resource is at stake. | Caliskan et al. find African American names closer to unpleasant words than European American names. |
Embeddings do not just mirror the bias in their text; they can amplify it, exaggerating an association beyond its strength in the corpus or in the world (Zhao et al. 2017, Ethayarajh et al. 2019, Jia et al. 2020, as summarized in SLP3). Caliskan, Bryson and Narayanan (2017) built the Word Embedding Association Test (WEAT) and reproduced classic Implicit Association Test results with GloVe, including the finding that African American names sit closer to unpleasant words. Debiasing methods such as Bolukbasi's neutralize and equalize steps remove the component of gender-neutral words along the gender direction. They reduce measured bias, but Gonen and Goldberg (2019) show that the stereotyped words still cluster together afterwards: the bias is hidden, not removed.
Measuring a century of stereotypes
Garg, Schiebinger, Jurafsky and Zou (2018) turned diachronic embeddings into a tool for social history. Using the decade embeddings from Hamilton and colleagues, they built a group vector for women (the average of words like she, her, woman) and one for men, and scored each neutral word (an adjective or an occupation) by its relative norm difference: its average distance to the men vector minus its average distance to the women vector. A negative score means the word sits closer to men.
Worked example
Relative norm difference in 2D
Place the vectors
Women group vector (0, 2), men group vector (2, 0), adjective smart (1.6, 0.6).Measure both distances
‖smart − men‖ = √(0.4² + 0.6²) = 0.721 and ‖smart − women‖ = √(1.6² + 1.4²) = 2.126.Result
0.721 − 2.126 = −1.405. A negative value means smart is closer to the men vector.
Run over each decade, the measure tells a story. Competence and intelligence adjectives (smart, wise, thoughtful, logical) were biased toward men, and that bias has been decreasing since the 1960s. Words used to describe outsiders (barbaric, monstrous, hateful, bizarre) were most associated with Asian last names before 1950 and declined steadily afterwards. The embeddings reproduce a 1933 survey of ethnic stereotypes. Occupation bias in the Google News embeddings tracks the 2015 share of women in each occupation (r² = .46), and the decade-by-decade trend in the historical embeddings follows US Census data. Embeddings, in other words, are a usable instrument for measuring culture.
| Query | Constrained answer | Unconstrained answer |
|---|---|---|
| man : doctor :: woman : x | gynecologist | doctor |
| man : computer programmer :: woman : x | homemaker | computer programmer |
Recall
Give one allocational harm and one representational harm caused by embeddings.
Recall
Compute Garg's relative norm difference for w = (1, 1), women = (0, 2), men = (2, 0), and interpret it.
Before testing an embedding, look at it. Slide 107 shows a hierarchical clustering of noun embeddings (from Rohde et al., reproduced in SLP3). Read it bottom-up. Body parts merge early: wrist with ankle, then shoulder, arm and leg. Animals form their own branch: dog with cat, then puppy and kitten. Places split off near the top. And look closely at the places: Tokyo sits with Chicago and the other US cities, while Moscow and Hawaii sit with countries and continents.
The height of each join is the dissimilarity at the moment the two clusters merged, so low joins mean close vectors. The odd placements are a lesson in what an embedding encodes: the tree reflects the contexts words appear in, not a geography ontology. Other ways to look include the nearest-neighbour lists you saw for awful, and 2D projections such as t-SNE (van der Maaten and Hinton 2008), which keep local neighbourhoods but distort global distances.
Intrinsic evaluation: test the vectors directly
Intrinsic evaluation scores the vectors on a small task designed to probe them, without building a full system. The most common form compares model similarity with human judgments.
- WordSim-353 (Finkelstein et al. 2002) asks people to rate 353 word pairs from 0 (totally unrelated) to 10 (very much related or identical); plane and car get 5.77. Because the instruction is about relatedness, cup and coffee can score as high as cup and mug. That is Word relatedness, not Word similarity.
- SimLex-999 (Hill, Reichart and Korhonen 2015) was built to fix that: 999 adjective, noun and verb pairs rated for genuine similarity, so cup and mug score high and cup and coffee score low.
- The TOEFL synonym test has 80 questions with 4 choices each: which word is closest to levied? The model picks the choice with the highest cosine. Latent semantic analysis scored 64.4% (Landauer and Dumais 1997), almost exactly the 64.5% average of non-native college applicants in the US.
- Analogy sets, with all the caveats of the earlier concept.
For a similarity dataset the score is the Spearman rank correlation between the model's cosines and the human ratings. Spearman compares orders, not values, so it does not matter that cosines live in [−1, 1] and ratings in [0, 10].
Worked example
Spearman correlation on four WordSim pairs
Rank both columns
Human ratings are the real WordSim-353 values; the model cosines are illustrative.Pair Human rating Model cosine Human rank Model rank d drink, ear 1.31 0.08 1 1 0 plane, car 5.77 0.42 2 2 0 drink, eat 6.87 0.61 3 4 1 planet, star 8.45 0.55 4 3 1 Sum the squared rank differences
Σd² = 0 + 0 + 1 + 1 = 2, with n = 4.Result
ρ = 1 − (6 · 2) / (4 · 15) = 1 − 0.2 = 0.8. The model orders the pairs almost like people do; it only swaps drink-eat and planet-star.
Worked example
Answering a TOEFL item
The question
Which word is closest in meaning to levied: imposed, believed, requested or correlated?Score each choice
Compute cos(levied, choice) for all four choices.Result
Return the argmax. A good space puts imposed first, because both words appear around taxes, fines and duties.
Extrinsic evaluation: test inside a real task
Extrinsic evaluation plugs each candidate embedding into an actual system (named entity recognition, machine translation, coreference), trains it, and compares the task metric. SLP3 calls this the most important evaluation for vector models, because the task is what you care about. It is also slow and noisy, which is why intrinsic tests remain popular for quick comparisons.
| Evaluation | Type | Data | Score |
|---|---|---|---|
| WordSim-353 | Intrinsic | 353 pairs, 0 to 10, relatedness | Spearman ρ between cosine and human ratings |
| SimLex-999 | Intrinsic | 999 pairs, similarity only | Spearman ρ; penalizes cup and coffee scoring like cup and mug |
| TOEFL synonyms | Intrinsic | 80 items, 4 choices each | Accuracy of argmax cosine over the choices |
| Analogy sets (Google, BATS) | Intrinsic | a : a* :: b : ? questions | Accuracy of the parallelogram method |
| NER, MT, coreference | Extrinsic | A full task with its own labelled data | Task metric (F1, BLEU) with each embedding plugged in |
Recall
Intrinsic or extrinsic: (a) Spearman ρ with SimLex-999, (b) NER F1 with GloVe versus word2vec inputs, (c) TOEFL synonym accuracy.
Quick check
A team reports Spearman correlation with SimLex-999 ratings for new embeddings. What evaluation is this?
Feed "I saw a cat." into a neural network. Each token passes through an Embedding layer, a lookup table from word to vector, and the network sits on top. Where should that table come from? There are three answers: copy it from Word2vec or GloVe and freeze it; copy it and keep updating it on your task; or start it from random numbers and learn it together with the network.
The slides give the rule. When there is not enough labelled data, or the task is simple, use embeddings pretrained on another task. Pretrained vectors bring knowledge distilled by Self-supervision from billions of unlabelled tokens, which your few thousand labelled examples could never teach. When there is enough data and the task is hard, such as language modelling or machine translation, train the embeddings with the model, because the task itself supplies the signal and task-specific vectors fit it better. Fine-tuning pretrained vectors is the middle path.
| Option | Where vectors come from | Updated during training? | When to use | Example |
|---|---|---|---|---|
| Frozen pretrained | word2vec or GloVe trained on another corpus | No | Little labelled data, simple task | Kim's CNN-static |
| Fine-tuned pretrained | Copied from word2vec or GloVe | Yes, starting from the pretrained values | Moderate data, want task-specific nuance | Kim's CNN-non-static |
| Joint from random | Learned from scratch with the network | Yes, from random values | Large data and a hard task (LM, MT) | Kim's CNN-rand; large LMs and MT systems |
Two studies give the evidence. Kim (2014) trained a CNN sentence classifier on small benchmarks three ways. CNN-rand, with random embeddings learned jointly, did poorly. CNN-static, with frozen word2vec vectors, "performs remarkably well", and CNN-non-static, which fine-tunes them, improves further; Kim concludes that pretrained vectors are good, universal feature extractors. Qi and colleagues (2018) asked the same question for neural machine translation and found gains of up to 20 BLEU in the most favourable setting. The gains have a sweet spot: they are largest when training data is scarce, but not so scarce that the system cannot be trained at all (baseline BLEU around 3 to 4), and they shrink as parallel data grows.
Everything in this lecture has been a Static embedding: one fixed vector per word type, whatever the sentence. Later lectures replace it with a Contextual embedding, where the vector for bank depends on its sentence. In those language models the embedding layer is trained jointly with the network, exactly the right-hand side of slide 110.
Recall
You have 500k parallel sentences for a hard MT task. Pretrained or joint, and why?
Quick check
You build a dialect sentiment classifier from 2,000 labeled tweets. How should its embeddings start?
Recap
If you remember nothing else
- Parallelogram method: b̂* = argmin over x of distance(x, a* − a + b), or argmax of cosine, with a, a* and b excluded. Slide 99's argmax of distance is a typo.
- Relations such as gender, number, tense and country to capital appear as roughly constant offsets; one word can sit on several such directions at once.
- Without exclusion the method returns b 93% of the time. It works best for frequent words, answers that already sit near b, and inflectional or capital relations; 3CosMul improves on 3CosAdd.
- Decade-specific embeddings, aligned by orthogonal Procrustes, show gay, broadcast and awful shifting. Frequent words change slowly; polysemous words change fast.
- Embeddings reproduce and amplify stereotypes (father : doctor :: mother : nurse), causing allocational and representational harms. Debiasing hides more than it removes.
- Garg et al. track bias over a century with relative norm differences: competence adjectives lean male (weakening since the 1960s), outsider words were tied to Asian names before 1950, and occupation bias tracks census data.
- Visualize with dendrograms, nearest-neighbour lists or t-SNE; all of them show usage, not an ontology.
- Intrinsic evaluation: WordSim-353, SimLex-999, TOEFL, analogies. Extrinsic evaluation is a downstream task and the more important one.
- Use pretrained vectors with little data or a simple task; train embeddings jointly with a large dataset and a hard task; fine-tuning sits between.
Sources
- Speech and Language Processing (3rd ed. draft), ch. 5 EmbeddingsBookJurafsky and Martin, StanfordSections 5.6 to 5.9: visualizing, the parallelogram, bias and evaluation(opens in a new tab)
- Linguistic Regularities in Continuous Space Word RepresentationsPaperMikolov, Yih and Zweig, NAACL 2013(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov et al., 2013(opens in a new tab)
- GloVe: Global Vectors for Word RepresentationPaperPennington, Socher and Manning, EMNLP 2014(opens in a new tab)
- GloVe project page (linear substructures)DocsStanford NLP Group(opens in a new tab)
- A model for analogical reasoningPaperRumelhart and Abrahamson, Cognitive Psychology 5(1), 1973(opens in a new tab)
- Linguistic Regularities in Sparse and Explicit Word RepresentationsPaperLevy and Goldberg, CoNLL 2014(opens in a new tab)
- Issues in evaluating semantic spaces using word analogiesPaperLinzen, RepEval 2016(opens in a new tab)
- Analogy-based detection of morphological and semantic relations with word embeddings (BATS)PaperGladkova, Drozd and Matsuoka, NAACL SRW 2016(opens in a new tab)
- Towards Understanding Linear Word AnalogiesPaperEthayarajh, Duvenaud and Hirst, ACL 2019(opens in a new tab)
- Parallelograms revisited: Exploring the limitations of vector space models for simple analogiesPaperPeterson, Chen and Griffiths, Cognition 205, 2020(opens in a new tab)
- Diachronic Word Embeddings Reveal Statistical Laws of Semantic ChangePaperHamilton, Leskovec and Jurafsky, ACL 2016(opens in a new tab)
- Syntactic Annotations for the Google Books Ngram CorpusPaperLin et al., ACL 2012 demo(opens in a new tab)
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word EmbeddingsPaperBolukbasi et al., NIPS 2016(opens in a new tab)
- Word embeddings quantify 100 years of gender and ethnic stereotypesPaperGarg, Schiebinger, Jurafsky and Zou, PNAS 2018(opens in a new tab)
- Semantics derived automatically from language corpora contain human-like biasesPaperCaliskan, Bryson and Narayanan, Science 2017(opens in a new tab)
- Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove ThemPaperGonen and Goldberg, NAACL 2019(opens in a new tab)
- Fair Is Better than Sensational: Man Is to Doctor as Woman Is to DoctorPaperNissim, van Noord and van der Goot, Computational Linguistics 46(2), 2020(opens in a new tab)
- The WordSimilarity-353 Test CollectionDocsFinkelstein et al., 2002(opens in a new tab)
- SimLex-999: Evaluating Semantic Models With (Genuine) Similarity EstimationPaperHill, Reichart and Korhonen, Computational Linguistics 41(4), 2015(opens in a new tab)
- Problems With Evaluation of Word Embeddings Using Word Similarity TasksPaperFaruqui et al., RepEval 2016(opens in a new tab)
- TOEFL Synonym Questions (State of the art)DocsACL Wiki(opens in a new tab)
- Visualizing Data using t-SNEPapervan der Maaten and Hinton, JMLR 9, 2008(opens in a new tab)
- Convolutional Neural Networks for Sentence ClassificationPaperKim, EMNLP 2014(opens in a new tab)
- When and Why Are Pre-Trained Word Embeddings Useful for Neural Machine Translation?PaperQi, Sachan, Felix, Padmanabhan and Neubig, NAACL 2018(opens in a new tab)