Majid Al-RaimiVector semantics and the distributional hypothesis

ICS 582Lecture 04Part 03

Vector semantics and the distributional hypothesis

Defining a word by its contexts (Wittgenstein, Harris, the ongchoi example), combining that with Osgood's meaning as a point in space, and arriving at embeddings, plus why vectors generalize better than word identities and the two kinds (sparse tf-idf, dense word2vec).

Concepts
6
Slides
18-30
Reading
36 min
Understood
0/6 concepts

Why this part matters

Every model in this course from here on, word2vec now and BERT and large language models later, begins by turning words into vectors. This part explains why that works at all. Words that keep similar company mean similar things, and once words are points in a space, "similar" becomes something you can measure.

Parts 1 and 2 listed what a model of word meaning should capture: synonymy, similarity, relatedness and connotation. This part introduces the model that meets many of those wishes. It is a core exam topic (state the hypothesis, compare sparse and dense vectors), the basis of retrieval and semantic search in real systems, and the representation behind most NLP research projects you are likely to start.

By the end you can

  1. State the distributional hypothesis and attribute it to Harris, Firth and Joos, with Wittgenstein's "meaning is use" as its philosophical root.
  2. Infer the category of an unknown word (ongchoi) from the contexts it shares with known words.
  3. Represent a word as a point in a space of affective dimensions and compute the distance between two words.
  4. Define an embedding and read a 2D t-SNE word map without over-reading global distances.
  5. Explain why vector features generalize to similar unseen words when identity features cannot.
  6. Compare sparse (tf-idf, PPMI) and dense (word2vec) embeddings on length, sparsity, construction and use.

Meaning from the company a word keeps

Take two words, "oculist" and "eye-doctor". Collect every sentence each appears in, and look at the neighbors: eye, examined, prescription, glasses, appointment. The two lists are almost the same. Now compare "oculist" with "lawyer". Some neighbors still overlap (appointment, fee, office), but far fewer. Without a dictionary, the overlap of environments already tells you which pair is closer in meaning.

That observation is the distributional hypothesis. Zellig Harris put it in 1954 using exactly this example: if two words have almost identical environments, meaning neighboring words or the grammatical frames they occur in, we call them synonyms. He then made the claim graded. The difference in meaning between two words corresponds roughly to the amount of difference in their environments. That second sentence matters more than the first, because it turns meaning into a quantity you can estimate by counting.

The idea was in the air in the 1950s. Wittgenstein had argued that, for a large class of cases, the meaning of a word is its use in the language. Joos (1950) described the meaning of a morpheme as the set of conditional probabilities of its occurrence alongside every other morpheme, which is almost a definition of a language model. Firth (1957) gave the line everyone quotes: "You shall know a word by the company it keeps." Firth and Harris are often merged, but they meant different things. Firth cared about situational and cultural context; Harris cared about the formal distribution of words inside text. NLP took Harris's version, because it can be computed from a corpus alone.

ThinkerYearClaimWhat it contributes
Ludwig Wittgenstein1953For a large class of cases, the meaning of a word is its use in the languageThe philosophical license: stop looking for meaning behind the word and look at how it is used
Martin Joos1950The meaning of a morpheme is the set of conditional probabilities of its occurrence with all other morphemesA probabilistic statement, decades before anyone could count at scale
Zellig Harris1954Words with almost identical environments are synonyms; the difference in meaning roughly matches the difference in environmentsThe operational, graded form that NLP actually implements
J. R. Firth1957You shall know a word by the company it keepsThe slogan, from a theory of meaning in situational and cultural context
Four roots of one idea

This is why the lecture moves to vector semantics. Earlier, lexical semantics gave us a list of relations a good model should respect: synonymy, similarity, relatedness, connotation. Writing those relations down by hand for every pair of words is impossible. The distributional hypothesis says you do not have to: read enough text, record each word's environments, and the relations fall out of the overlaps. Vector semantics is the standard way NLP does this today.

Recall

State the distributional hypothesis, and say who formulated it in the 1950s.

Words that occur in similar environments (neighboring words or grammatical contexts) have similar meanings, and the difference in meaning roughly matches the difference in environments. Harris (1954), Joos (1950) and Firth (1957, "You shall know a word by the company it keeps"), with philosophical roots in Wittgenstein's "meaning is use" (1953).

Quick check

Harris (1954) wrote about two words A and B that have almost identical environments. What did he conclude about them?

Guessing ongchoi from the words around it

Suppose you have never seen the word "ongchoi", a recent borrowing into English from Cantonese, and you meet it three times: ongchoi is delicious sauteed with garlic; ongchoi is superb over rice; ongchoi leaves with salty sauces. You do not know what it is yet. But you have read plenty of other sentences: spinach sauteed with garlic over rice, chard stems and leaves are delicious, collard greens and other salty leafy greens.

Put the contexts side by side and the answer is hard to miss. Sauteed, garlic, rice, leaves, delicious and salty all show up around ongchoi and around the leafy greens you already know. Nothing about laptops, contracts or weather. So ongchoi is very likely a leafy green that people cook and eat. It is: the plant is Ipomoea aquatica, a relative of morning glory sometimes called water spinach, and it has other names in Chinese, Malay and Vietnamese.

Contexts that ongchoi shares with known words

ongchoi
delicious, sauteed, garlic, superb, rice, leaves, salty, sauces
spinach
sauteed, garlic, rice
chard
stems, leaves, delicious
collard greens
salty, leafy, greens
Two rows of context words, ongchoi above and the known greens spinach, chard and collards below. The six shared words (sauteed, garlic, rice, leaves, delicious, salty) light up and link across; superb and stems stay dim because only one row has them.

Worked example

Inferring ongchoi

  1. List the contexts of the unknown word

    Around ongchoi: delicious, sauteed, garlic, superb, rice, leaves, salty, sauces.
  2. List the contexts of candidate known words

    Spinach: sauteed, garlic, rice. Chard: stems, leaves, delicious. Collard greens: salty, leafy, greens. A distractor such as laptop: screen, battery, keyboard.
  3. Mark what is shared

    Spinach shares 3 context words with ongchoi, chard shares 2, collard greens share 1, and laptop shares 0. Together the greens cover sauteed, garlic, rice, leaves, delicious and salty.
  4. Result

    Ongchoi sits with the leafy greens and nowhere near laptop. The prediction is "a cooked leafy green", which is right.

This is the computational form of the distributional hypothesis. Define the meaning of a word by its distribution, the neighboring words or grammatical environments it appears in, and then do the obvious thing: count the words in the context of ongchoi and compare those counts with the counts for every other word. A table of such counts, one row per word and one column per context word, is the term-context matrix of part 4, and each row is a word's vector.

Recall

What exactly would a program count to carry out the ongchoi inference?

For each word, how often every other word appears within a small window around it. Comparing ongchoi's counts with those of spinach, chard and collards shows heavy overlap on sauteed, garlic, rice, leaves, delicious and salty. That count table is the term-context matrix, and each row is the word's vector.

Here are two words with three numbers each. Heartbreak is [2.45, 5.65, 3.58] and courageous is [8.05, 5.5, 7.38]. The first number is valence (how pleasant), the second is arousal (how intense the emotion), the third is dominance (how much control is exerted). Plot both as points. They sit far apart on valence and dominance, and almost level on arousal: both words are emotionally charged, one pleasantly and with control, the other unpleasantly and with loss of control.

Valence, arousal and dominance ratings on a 1 to 9 scale (Warriner et al. 2013)

courageous
[8.05, 5.5, 7.38]
music
[7.67, 5.57, 6.5]
heartbreak
[2.45, 5.65, 3.58]
cub
[6.71, 3.95, 4.24]

The idea of placing a word at a point goes back to Charles Osgood and colleagues in 1957. Osgood asked people to rate words on many bipolar scales, such as good to bad, strong to weak, active to passive, a method he called the semantic differential. Factor analysis showed that most of the variation came from three factors, which he named evaluation, potency and activity. Today the same three are usually called valence, dominance and arousal, the VAD dimensions of a word's connotation from part 2. Osgood noticed that three numbers per word make each word a point in a three-dimensional space, and he proposed that similarity of meaning is nearness in that space. That is the first appearance of vector semantics.

d(u,v)=∑i(ui−vi)2d(\mathbf{u},\mathbf{v})=\sqrt{\sum_i (u_i-v_i)^2}
Euclidean distance between two points

Worked example

How far is courageous from heartbreak?

  1. Subtract coordinate by coordinate

    Valence 8.05 − 2.45 = 5.60, arousal 5.5 − 5.65 = −0.15, dominance 7.38 − 3.58 = 3.80.
  2. Square and add

    5.60² + 0.15² + 3.80² = 31.36 + 0.02 + 14.44 = 45.82.
  3. Take the square root

    √45.82 ≈ 6.77.
  4. Result

    For comparison, courageous is 0.96 from music and 3.75 from cub. Heartbreak is the far outlier, and almost all of the gap comes from valence and dominance.
Four words placed by their valence (V), arousal (A) and dominance (D) ratings. Drop lines show each word's height on the arousal axis; courageous and heartbreak light up and the long segment between them is their distance, about 6.77.

Put the two threads together. Idea 1 is that meaning can be defined by linguistic distribution. Idea 2, from Osgood, is that meaning can be a point in a multidimensional space. Vector semantics joins them: represent each word as a point, but let the word's distribution, not a panel of human raters, decide where the point goes.

Recall

What two 1950s ideas does vector semantics combine?

(1) Meaning defined by linguistic distribution (Harris, Firth, Joos). (2) Meaning as a point in a multidimensional space (Osgood et al. 1957). An embedding is a point in space whose position is derived from the word's distribution.

Recall

Where do the slide 25 numbers really come from, and on what scale?

From Warriner, Kuperman and Brysbaert (2013), human ratings of about 14,000 lemmas on a 1 to 9 scale. The VAD dimensions trace back to Osgood's semantic differential (evaluation, potency, activity).

Quick check

Using the slide 25 ratings and Euclidean distance, which word is farthest from courageous in VAD space?

Embeddings: points placed by distribution

Look at a map of word vectors squashed onto a page. Good, nice, wonderful, fantastic, amazing and very good huddle in one region. Bad, worst, worse, dislike and not good gather in another. Function words such as to, by, that, is and with sit off by themselves. Nobody placed these words by hand. Their positions came from training on text.

Twenty-four words as dots. They start loosely spread, then tighten into three groups: positive words, negative words and function words, each outlined once it forms.

This is vector semantics in its working form. Each word is a vector, a list of numbers, rather than an arbitrary symbol such as the string "good" or the index w45. Similar words end up nearby in what is called semantic space. And the space is built automatically, by seeing which words are nearby in text, so the distributional hypothesis does the placing that Osgood's raters used to do.

Such a vector is called an embedding, because the word is embedded into a space. The term began in the latent semantic analysis community in the late 1990s, where it named the mapping from sparse count space into a smaller dense space, and it later shifted to mean the resulting vector itself. Embeddings are now the standard way to represent word meaning in NLP: practically every modern system, from classifiers to large language models, starts by looking up or computing them. They give a fine-grained model of similarity, a number for every pair of words instead of a yes or no.

Recall

What can and cannot be read from a 2D t-SNE map of word embeddings?

Local neighborhoods (which words are near each other) are roughly preserved. Distances between clusters and absolute positions are distorted by the projection from 60 dimensions, so do not read them as real distances.

Why vectors beat word identities

Build a sentiment classifier the traditional way. Feature 5 is "the previous word was terrible", and training learns that it signals a negative review. At test time a review says "awful acting". If awful never appeared in the labeled training data, feature 5 does not fire, no other feature knows about awful, and the classifier has nothing to go on.

Now replace the identity feature with the previous word's embedding. During training the input was terrible's vector, say [35, 22, 17, ...], and the classifier learned weights on those coordinates. At test time awful arrives as [34, 21, 14, ...]. The weights do not care whether the word is the same string; they act on the numbers, and these numbers are nearly the same. The classifier treats awful almost exactly as it treated terrible. It has generalized to a similar but unseen word.

Left: an identity feature for terrible rejects the key awful, because only an exact match fires. Right: the vector for awful swings in beside the vector for terrible, separated by an angle of about 3 degrees (cosine about 0.999).
cos⁡(u,v)=u⋅v∣u∣ ∣v∣\cos(\mathbf{u},\mathbf{v})=\frac{\mathbf{u}\cdot\mathbf{v}}{|\mathbf{u}|\,|\mathbf{v}|}
Cosine similarity, developed fully in part 4

Worked example

How close are terrible and awful?

  1. Dot product

    Using only the three coordinates the slide shows: 35×34 + 22×21 + 17×14 = 1190 + 462 + 238 = 1890.
  2. Lengths

    √(35² + 22² + 17²) = √1998 ≈ 44.70 and √(34² + 21² + 14²) = √1793 ≈ 42.34.
  3. Divide

    1890 / (44.70 × 42.34) ≈ 0.9986.
  4. Result

    On these three coordinates the vectors point in almost the same direction, giving a cosine similarity near 1 and an angle of about 3°. Their Euclidean distance is √11 ≈ 3.32, small next to their lengths. Anything the classifier learned about terrible transfers.
PropertyIdentity featureVector feature
What is storedA yes or no: was the previous word exactly terribleThe previous word's vector, for example [35, 22, 17, ...]
Test-time matchString equalityGeometry: weights act on every coordinate, so nearby vectors produce nearby scores
Unseen similar word (awful)Feature never fires; the learned weight is wasted[34, 21, 14] lands almost where terrible did, so the classifier reacts in nearly the same way
Number of weightsOne per vocabulary word, often 50,000 or moreOne per dimension, often 300
Identity features against vector features

This matters because of what lecture 2 showed about vocabularies. Zipf's law means most word types are rare, and about half of them appear only once. A classifier trained on a few thousand labeled reviews will meet many words at test time that it never saw with a label. Embeddings, learned from billions of unlabeled words, carry knowledge about those words into the classifier. Dense vectors also need far fewer weights, around 300 per input position instead of one per vocabulary word, and they can place car and automobile close together, which separate identity features never can.

Recall

Why does a word-identity feature fail on an unseen synonym when an embedding feature does not?

The identity feature fires only if the exact same word appeared in training. An embedding feature compares vectors, so a test word whose vector is close to a trained word's vector (for example [34, 21, 14] against [35, 22, 17], cosine ≈ 0.999) gets similar treatment.

Quick check

A sentiment classifier learned that the previous word 'terrible' signals negativity. At test time it meets 'awful', which never appeared in its labeled training data. Why can an embedding-based classifier still handle it?

Picture two vectors for ongchoi. The first has one slot per vocabulary word, tens of thousands of them, and stores how often each word appeared near ongchoi. Garlic, rice and leaves have counts; nearly every other slot is 0. The second has about 300 real numbers, almost none of them zero, some negative, and none of them labeled with a word. Both are embeddings. This lecture covers both families.

Sparse vectors: weighted counts

A sparse vector is built from a simple function of counts. With tf-idf, the classic weighting from information retrieval, the counts come from documents. With PPMI, they come from nearby words, re-weighted so that informative co-occurrences stand out. These vectors are the workhorse of search engines and a strong, cheap baseline in almost any text task. Their dimensions are interpretable, because each one is a specific word or document.

Dense vectors: learned by prediction

A dense vector such as one from word2vec is learned instead of counted. Word2vec trains a simple classifier to predict whether a word is likely to appear near a target word, and keeps the classifier's weights as the embedding. The labels come free from running text, which is self-supervision. Both families give one fixed vector per word type, a static embedding. Later in the course, contextual embeddings such as BERT compute a different vector for each occurrence of a word.

PropertySparse (tf-idf, PPMI)Dense (word2vec)
Length|V|, tens of thousands50 to 1000
ZerosMostly zeroMostly non-zero, can be negative
How builtWeighted co-occurrence countsA classifier trained to predict whether a word appears near the target
What one dimension meansA specific context word or documentNothing individually
Typical useInformation retrieval, a strong baselineInput features for neural NLP
Synonyms such as car and automobileSeparate, unrelated dimensionsCan end up with similar coordinates
The two families of embeddings in this lecture

Recall

Contrast sparse and dense embeddings on length, content and construction.

Sparse vectors (tf-idf, PPMI) are |V|-long, mostly zero, and built from weighted co-occurrence counts. Dense vectors (word2vec) have 50 to 1000 real-valued, mostly non-zero entries, and are learned by training a classifier to predict whether a word appears near the target.

Quick check

Which statement correctly contrasts the two kinds of embeddings on slide 30?

Recap

If you remember nothing else

  • Distributional hypothesis: words in similar environments have similar meanings, roughly in proportion to how similar the environments are (Harris 1954, Firth 1957, Joos 1950; Wittgenstein 1953: meaning is use).
  • Ongchoi shares sauteed, garlic, rice, leaves, delicious and salty with spinach, chard and collards, so it is a leafy green. It is Ipomoea aquatica, water spinach.
  • Osgood (1957) treated a word's connotation as a point in a space of a few rated dimensions, and similarity as distance. The slide 25 numbers are Warriner et al. (2013) ratings on a 1 to 9 scale.
  • Vector semantics combines both ideas: a word is a point in a multidimensional space built from its distribution. That vector is an embedding.
  • The slide 27 map is a 2D t-SNE projection of 60-dimensional sentiment-trained embeddings (Li et al. 2015), not the embedding itself.
  • Vector features generalize: [34, 21, 14] is almost parallel to [35, 22, 17] (cosine ≈ 0.999), while an identity feature needs the exact word.
  • Sparse vectors (tf-idf, PPMI) are |V| long, mostly zero, and built from counts. Dense vectors (word2vec) have 50 to 1000 non-zero entries learned by a classifier predicting neighbors. Contextual embeddings come later.

Sources