ICS 582Lecture 04Part 03
Vector semantics and the distributional hypothesis
Defining a word by its contexts (Wittgenstein, Harris, the ongchoi example), combining that with Osgood's meaning as a point in space, and arriving at embeddings, plus why vectors generalize better than word identities and the two kinds (sparse tf-idf, dense word2vec).
- Concepts
- 6
- Slides
- 18-30
- Reading
- 36 min
Why this part matters
Every model in this course from here on, word2vec now and BERT and large language models later, begins by turning words into vectors. This part explains why that works at all. Words that keep similar company mean similar things, and once words are points in a space, "similar" becomes something you can measure.
Parts 1 and 2 listed what a model of word meaning should capture: synonymy, similarity, relatedness and connotation. This part introduces the model that meets many of those wishes. It is a core exam topic (state the hypothesis, compare sparse and dense vectors), the basis of retrieval and semantic search in real systems, and the representation behind most NLP research projects you are likely to start.
By the end you can
- State the distributional hypothesis and attribute it to Harris, Firth and Joos, with Wittgenstein's "meaning is use" as its philosophical root.
- Infer the category of an unknown word (ongchoi) from the contexts it shares with known words.
- Represent a word as a point in a space of affective dimensions and compute the distance between two words.
- Define an embedding and read a 2D t-SNE word map without over-reading global distances.
- Explain why vector features generalize to similar unseen words when identity features cannot.
- Compare sparse (tf-idf, PPMI) and dense (word2vec) embeddings on length, sparsity, construction and use.
Take two words, "oculist" and "eye-doctor". Collect every sentence each appears in, and look at the neighbors: eye, examined, prescription, glasses, appointment. The two lists are almost the same. Now compare "oculist" with "lawyer". Some neighbors still overlap (appointment, fee, office), but far fewer. Without a dictionary, the overlap of environments already tells you which pair is closer in meaning.
That observation is the distributional hypothesis. Zellig Harris put it in 1954 using exactly this example: if two words have almost identical environments, meaning neighboring words or the grammatical frames they occur in, we call them synonyms. He then made the claim graded. The difference in meaning between two words corresponds roughly to the amount of difference in their environments. That second sentence matters more than the first, because it turns meaning into a quantity you can estimate by counting.
The idea was in the air in the 1950s. Wittgenstein had argued that, for a large class of cases, the meaning of a word is its use in the language. Joos (1950) described the meaning of a morpheme as the set of conditional probabilities of its occurrence alongside every other morpheme, which is almost a definition of a language model. Firth (1957) gave the line everyone quotes: "You shall know a word by the company it keeps." Firth and Harris are often merged, but they meant different things. Firth cared about situational and cultural context; Harris cared about the formal distribution of words inside text. NLP took Harris's version, because it can be computed from a corpus alone.
| Thinker | Year | Claim | What it contributes |
|---|---|---|---|
| Ludwig Wittgenstein | 1953 | For a large class of cases, the meaning of a word is its use in the language | The philosophical license: stop looking for meaning behind the word and look at how it is used |
| Martin Joos | 1950 | The meaning of a morpheme is the set of conditional probabilities of its occurrence with all other morphemes | A probabilistic statement, decades before anyone could count at scale |
| Zellig Harris | 1954 | Words with almost identical environments are synonyms; the difference in meaning roughly matches the difference in environments | The operational, graded form that NLP actually implements |
| J. R. Firth | 1957 | You shall know a word by the company it keeps | The slogan, from a theory of meaning in situational and cultural context |
This is why the lecture moves to vector semantics. Earlier, lexical semantics gave us a list of relations a good model should respect: synonymy, similarity, relatedness, connotation. Writing those relations down by hand for every pair of words is impossible. The distributional hypothesis says you do not have to: read enough text, record each word's environments, and the relations fall out of the overlaps. Vector semantics is the standard way NLP does this today.
Recall
State the distributional hypothesis, and say who formulated it in the 1950s.
Quick check
Harris (1954) wrote about two words A and B that have almost identical environments. What did he conclude about them?
Suppose you have never seen the word "ongchoi", a recent borrowing into English from Cantonese, and you meet it three times: ongchoi is delicious sauteed with garlic; ongchoi is superb over rice; ongchoi leaves with salty sauces. You do not know what it is yet. But you have read plenty of other sentences: spinach sauteed with garlic over rice, chard stems and leaves are delicious, collard greens and other salty leafy greens.
Put the contexts side by side and the answer is hard to miss. Sauteed, garlic, rice, leaves, delicious and salty all show up around ongchoi and around the leafy greens you already know. Nothing about laptops, contracts or weather. So ongchoi is very likely a leafy green that people cook and eat. It is: the plant is Ipomoea aquatica, a relative of morning glory sometimes called water spinach, and it has other names in Chinese, Malay and Vietnamese.
Contexts that ongchoi shares with known words
- ongchoi
- delicious, sauteed, garlic, superb, rice, leaves, salty, sauces
- spinach
- sauteed, garlic, rice
- chard
- stems, leaves, delicious
- collard greens
- salty, leafy, greens
Worked example
Inferring ongchoi
List the contexts of the unknown word
Around ongchoi: delicious, sauteed, garlic, superb, rice, leaves, salty, sauces.List the contexts of candidate known words
Spinach: sauteed, garlic, rice. Chard: stems, leaves, delicious. Collard greens: salty, leafy, greens. A distractor such as laptop: screen, battery, keyboard.Mark what is shared
Spinach shares 3 context words with ongchoi, chard shares 2, collard greens share 1, and laptop shares 0. Together the greens cover sauteed, garlic, rice, leaves, delicious and salty.Result
Ongchoi sits with the leafy greens and nowhere near laptop. The prediction is "a cooked leafy green", which is right.
This is the computational form of the distributional hypothesis. Define the meaning of a word by its distribution, the neighboring words or grammatical environments it appears in, and then do the obvious thing: count the words in the context of ongchoi and compare those counts with the counts for every other word. A table of such counts, one row per word and one column per context word, is the term-context matrix of part 4, and each row is a word's vector.
Recall
What exactly would a program count to carry out the ongchoi inference?
Here are two words with three numbers each. Heartbreak is [2.45, 5.65, 3.58] and courageous is [8.05, 5.5, 7.38]. The first number is valence (how pleasant), the second is arousal (how intense the emotion), the third is dominance (how much control is exerted). Plot both as points. They sit far apart on valence and dominance, and almost level on arousal: both words are emotionally charged, one pleasantly and with control, the other unpleasantly and with loss of control.
Valence, arousal and dominance ratings on a 1 to 9 scale (Warriner et al. 2013)
- courageous
- [8.05, 5.5, 7.38]
- music
- [7.67, 5.57, 6.5]
- heartbreak
- [2.45, 5.65, 3.58]
- cub
- [6.71, 3.95, 4.24]
The idea of placing a word at a point goes back to Charles Osgood and colleagues in 1957. Osgood asked people to rate words on many bipolar scales, such as good to bad, strong to weak, active to passive, a method he called the semantic differential. Factor analysis showed that most of the variation came from three factors, which he named evaluation, potency and activity. Today the same three are usually called valence, dominance and arousal, the VAD dimensions of a word's connotation from part 2. Osgood noticed that three numbers per word make each word a point in a three-dimensional space, and he proposed that similarity of meaning is nearness in that space. That is the first appearance of vector semantics.
Worked example
How far is courageous from heartbreak?
Subtract coordinate by coordinate
Valence 8.05 − 2.45 = 5.60, arousal 5.5 − 5.65 = −0.15, dominance 7.38 − 3.58 = 3.80.Square and add
5.60² + 0.15² + 3.80² = 31.36 + 0.02 + 14.44 = 45.82.Take the square root
√45.82 ≈ 6.77.Result
For comparison, courageous is 0.96 from music and 3.75 from cub. Heartbreak is the far outlier, and almost all of the gap comes from valence and dominance.
Put the two threads together. Idea 1 is that meaning can be defined by linguistic distribution. Idea 2, from Osgood, is that meaning can be a point in a multidimensional space. Vector semantics joins them: represent each word as a point, but let the word's distribution, not a panel of human raters, decide where the point goes.
Recall
What two 1950s ideas does vector semantics combine?
Recall
Where do the slide 25 numbers really come from, and on what scale?
Quick check
Using the slide 25 ratings and Euclidean distance, which word is farthest from courageous in VAD space?
Look at a map of word vectors squashed onto a page. Good, nice, wonderful, fantastic, amazing and very good huddle in one region. Bad, worst, worse, dislike and not good gather in another. Function words such as to, by, that, is and with sit off by themselves. Nobody placed these words by hand. Their positions came from training on text.
This is vector semantics in its working form. Each word is a vector, a list of numbers, rather than an arbitrary symbol such as the string "good" or the index w45. Similar words end up nearby in what is called semantic space. And the space is built automatically, by seeing which words are nearby in text, so the distributional hypothesis does the placing that Osgood's raters used to do.
Such a vector is called an embedding, because the word is embedded into a space. The term began in the latent semantic analysis community in the late 1990s, where it named the mapping from sparse count space into a smaller dense space, and it later shifted to mean the resulting vector itself. Embeddings are now the standard way to represent word meaning in NLP: practically every modern system, from classifiers to large language models, starts by looking up or computing them. They give a fine-grained model of similarity, a number for every pair of words instead of a yes or no.
Recall
What can and cannot be read from a 2D t-SNE map of word embeddings?
Build a sentiment classifier the traditional way. Feature 5 is "the previous word was terrible", and training learns that it signals a negative review. At test time a review says "awful acting". If awful never appeared in the labeled training data, feature 5 does not fire, no other feature knows about awful, and the classifier has nothing to go on.
Now replace the identity feature with the previous word's embedding. During training the input was terrible's vector, say [35, 22, 17, ...], and the classifier learned weights on those coordinates. At test time awful arrives as [34, 21, 14, ...]. The weights do not care whether the word is the same string; they act on the numbers, and these numbers are nearly the same. The classifier treats awful almost exactly as it treated terrible. It has generalized to a similar but unseen word.
Worked example
How close are terrible and awful?
Dot product
Using only the three coordinates the slide shows: 35×34 + 22×21 + 17×14 = 1190 + 462 + 238 = 1890.Lengths
√(35² + 22² + 17²) = √1998 ≈ 44.70 and √(34² + 21² + 14²) = √1793 ≈ 42.34.Divide
1890 / (44.70 × 42.34) ≈ 0.9986.Result
On these three coordinates the vectors point in almost the same direction, giving a cosine similarity near 1 and an angle of about 3°. Their Euclidean distance is √11 ≈ 3.32, small next to their lengths. Anything the classifier learned about terrible transfers.
| Property | Identity feature | Vector feature |
|---|---|---|
| What is stored | A yes or no: was the previous word exactly terrible | The previous word's vector, for example [35, 22, 17, ...] |
| Test-time match | String equality | Geometry: weights act on every coordinate, so nearby vectors produce nearby scores |
| Unseen similar word (awful) | Feature never fires; the learned weight is wasted | [34, 21, 14] lands almost where terrible did, so the classifier reacts in nearly the same way |
| Number of weights | One per vocabulary word, often 50,000 or more | One per dimension, often 300 |
This matters because of what lecture 2 showed about vocabularies. Zipf's law means most word types are rare, and about half of them appear only once. A classifier trained on a few thousand labeled reviews will meet many words at test time that it never saw with a label. Embeddings, learned from billions of unlabeled words, carry knowledge about those words into the classifier. Dense vectors also need far fewer weights, around 300 per input position instead of one per vocabulary word, and they can place car and automobile close together, which separate identity features never can.
Recall
Why does a word-identity feature fail on an unseen synonym when an embedding feature does not?
Quick check
A sentiment classifier learned that the previous word 'terrible' signals negativity. At test time it meets 'awful', which never appeared in its labeled training data. Why can an embedding-based classifier still handle it?
Picture two vectors for ongchoi. The first has one slot per vocabulary word, tens of thousands of them, and stores how often each word appeared near ongchoi. Garlic, rice and leaves have counts; nearly every other slot is 0. The second has about 300 real numbers, almost none of them zero, some negative, and none of them labeled with a word. Both are embeddings. This lecture covers both families.
Sparse vectors: weighted counts
A sparse vector is built from a simple function of counts. With tf-idf, the classic weighting from information retrieval, the counts come from documents. With PPMI, they come from nearby words, re-weighted so that informative co-occurrences stand out. These vectors are the workhorse of search engines and a strong, cheap baseline in almost any text task. Their dimensions are interpretable, because each one is a specific word or document.
Dense vectors: learned by prediction
A dense vector such as one from word2vec is learned instead of counted. Word2vec trains a simple classifier to predict whether a word is likely to appear near a target word, and keeps the classifier's weights as the embedding. The labels come free from running text, which is self-supervision. Both families give one fixed vector per word type, a static embedding. Later in the course, contextual embeddings such as BERT compute a different vector for each occurrence of a word.
| Property | Sparse (tf-idf, PPMI) | Dense (word2vec) |
|---|---|---|
| Length | |V|, tens of thousands | 50 to 1000 |
| Zeros | Mostly zero | Mostly non-zero, can be negative |
| How built | Weighted co-occurrence counts | A classifier trained to predict whether a word appears near the target |
| What one dimension means | A specific context word or document | Nothing individually |
| Typical use | Information retrieval, a strong baseline | Input features for neural NLP |
| Synonyms such as car and automobile | Separate, unrelated dimensions | Can end up with similar coordinates |
Recall
Contrast sparse and dense embeddings on length, content and construction.
Quick check
Which statement correctly contrasts the two kinds of embeddings on slide 30?
Recap
If you remember nothing else
- Distributional hypothesis: words in similar environments have similar meanings, roughly in proportion to how similar the environments are (Harris 1954, Firth 1957, Joos 1950; Wittgenstein 1953: meaning is use).
- Ongchoi shares sauteed, garlic, rice, leaves, delicious and salty with spinach, chard and collards, so it is a leafy green. It is Ipomoea aquatica, water spinach.
- Osgood (1957) treated a word's connotation as a point in a space of a few rated dimensions, and similarity as distance. The slide 25 numbers are Warriner et al. (2013) ratings on a 1 to 9 scale.
- Vector semantics combines both ideas: a word is a point in a multidimensional space built from its distribution. That vector is an embedding.
- The slide 27 map is a 2D t-SNE projection of 60-dimensional sentiment-trained embeddings (Li et al. 2015), not the embedding itself.
- Vector features generalize: [34, 21, 14] is almost parallel to [35, 22, 17] (cosine ≈ 0.999), while an identity feature needs the exact word.
- Sparse vectors (tf-idf, PPMI) are |V| long, mostly zero, and built from counts. Dense vectors (word2vec) have 50 to 1000 non-zero entries learned by a classifier predicting neighbors. Contextual embeddings come later.
Sources
- Speech and Language Processing, Ch. 5: EmbeddingsBookJurafsky and Martin, StanfordMain textbook: ongchoi, VAD, summary and historical notes(opens in a new tab)
- Speech and Language Processing, 3rd ed. draft (Aug 2024), Ch. 6BookJurafsky and Martin, StanfordFigure 6.1 caption, the source of the slide 27 map(opens in a new tab)
- Ludwig WittgensteinArticleStanford Encyclopedia of PhilosophyFull Philosophical Investigations §43 quote(opens in a new tab)
- Distributional StructurePaperHarris, Word 10 (1954), Taylor & Francis(opens in a new tab)
- What company do words keep? Revisiting the distributional semantics of J.R. Firth & Zellig HarrisPaperBrunila and LaViolette, NAACL 2022(opens in a new tab)
- Quote Origin: You Shall Know a Word by the Company It KeepsArticleQuote InvestigatorProvenance of the Firth 1957 line(opens in a new tab)
- The Measurement of MeaningBookOsgood, Suci and Tannenbaum, University of Illinois Press(opens in a new tab)
- Norms of valence, arousal, and dominance for 13,915 English lemmasPaperWarriner, Kuperman and Brysbaert, Behavior Research Methods (2013)The real source of the slide 25 numbers(opens in a new tab)
- Visualizing and Understanding Neural Models in NLPPaperLi, Chen, Hovy and Jurafsky, NAACL 2016(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov et al., arXiv 2013(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov et al., NeurIPS 2013(opens in a new tab)
- From Frequency to Meaning: Vector Space Models of SemanticsPaperTurney and Pantel, JAIR 37 (2010)(opens in a new tab)
- Embedding ProjectorDocsTensorFlowExplore real embedding neighborhoods(opens in a new tab)