Majid Al-RaimiWord embeddings

ICS 582Lecture 04

Word embeddings

How to represent word meaning as vectors: lexical semantics desiderata, the distributional hypothesis, sparse count vectors weighted by tf-idf and PPMI with cosine similarity, dense word2vec skip-gram embeddings trained with negative sampling, their variants (CBOW, GloVe, FastText), and the properties, biases and evaluation of embeddings.

Parts
10
Concepts
59
Slides
111
Reading
354 min
Understood
0/59 concepts
Read the full guideEvery part on one long page: 10 parts, 59 concepts, about 354 min.

AOverview

This lecture answers the question every later model in the course depends on: how can a machine represent what a word means? It starts by showing that a string or an index carries no meaning at all, lists what a theory of word meaning must deliver (senses, synonymy, similarity, relatedness, antonymy and connotation), and then builds that theory out of one idea: a word is known by the company it keeps. From there it turns co-occurrence counts into vectors, reweights them with tf-idf and PPMI, learns short dense vectors with word2vec skip-gram and negative sampling, and finally probes what those vectors know, and what prejudices they absorb, through analogies, bias tests and evaluation.

As a PhD student you need it in three places. Exam questions ask you to compute, not to recall: a cosine between two count vectors, a tf-idf cell for a word in a play, a PMI cell from a small co-occurrence table, a sigmoid probability for a target and context pair, and one stochastic gradient step that pulls a true neighbor closer and pushes noise words away. Your research will use embeddings as the input layer of almost every model you train or compare, so you must be able to defend a choice of window size, weighting, dimensionality and evaluation, and to recognize when a static vector is too blunt for a word with many senses. And any system you build on retrieval, classification or Arabic text inherits both the strengths and the biases of the vectors it starts from, so you should be able to inspect, test and explain them.

From what a word means to what its vector knows: the lecture in five stops

The path has five stops. You start with meaning as linguists describe it: lemmas and senses, and the relations between words that any representation must respect. You then count, turning documents and context windows into term-document and word-word matrices and comparing rows with the cosine. You weight those counts so that frequent but uninformative words stop dominating, first with tf-idf and then with PPMI. You learn dense vectors by training a binary classifier that tells real neighbors from sampled noise, and you see how CBOW, GloVe and FastText vary that recipe. Finally you probe the result: analogies as parallel directions, meaning change over time, cultural bias, and the intrinsic and extrinsic tests that decide whether a set of embeddings is any good.

Success looks like

  • Distinguish lemma, sense, synonymy, similarity, relatedness, antonymy and connotation on a new word pair, and explain why the principle of contrast makes perfect synonyms unlikely.
  • Build a term-document or word-word matrix from a short corpus and compute a dot product, a vector length and a cosine by hand, explaining why the cosine corrects for frequency.
  • Compute a tf-idf weight with log term frequency and idf = log10(N / df), and say why document frequency, not collection frequency, decides how informative a word is.
  • Compute PMI and PPMI cells from a co-occurrence table, and explain both why negative values are clipped and why raising context counts to alpha = 0.75 tames the bias toward rare words.
  • Generate positive and negative skip-gram pairs from a window, write the SGNS loss, and apply one gradient update to a target and its context vectors.
  • Compare skip-gram, CBOW, GloVe and FastText, and predict how window size changes whether neighbors are look-alikes or topic-mates.
  • Solve an analogy with the parallelogram method, state why its accuracy overstates what embeddings know, and describe how to detect bias and evaluate embeddings intrinsically and extrinsically.

How to study this lecture

  1. Choose your route. Read the full guide in one sitting for the big picture, or work through one part at a time when you want depth. The count-based parts and the word2vec parts each make a natural single session.
  2. Answer every recall prompt in your head, or on paper, before you reveal it. Cosines, tf-idf cells, PMI cells and gradient steps only stick if you have computed them yourself at least once.
  3. Take each quiz and read the explanation even when you are right. The explanations carry the distinctions an examiner probes, such as similarity against relatedness or document against collection frequency.
  4. Use the simulators: move the sigmoid inputs, slide the context window, and try analogies in the playground. Seeing a vector move is faster than reading that it moves.
  5. Mark a concept as understood only when you could explain it to a classmate without looking. Unmarked concepts show you where to return.
  6. Keep the glossary and the reference sheet open as you revise. Every formula in the lecture is collected there with its symbols, so you can check notation before an exam.
  7. Come back after a few days and retry the recall prompts and quizzes cold. Spaced practice builds the long-term memory an exam needs.

Sources

BThe 10 parts

  1. 01What words mean: lemmas, senses and synonymyWhy treating words as strings or logical symbols is unsatisfying, what lexical semantics asks of a theory of meaning, how lemmas split into senses, and why perfect synonymy probably does not exist.6 conceptsSlides 1-1036 min
    1. 1.1Strings, indices and capitalized symbols are not meaning
    2. 1.2What a theory of word meaning must deliver
    3. 1.3One lemma, many senses: the mouse entry
    4. 1.4Synonymy is a relation between senses, not words
    5. 1.5Why perfect synonyms probably do not exist: the principle of contrast
    6. 1.6From synonymy to similarity: car and bicycle
  2. 02Similarity, relatedness, antonymy and connotationThe graded relations between word senses (similarity rated by humans, relatedness through semantic fields, antonymy as opposition on one feature) and the affective meaning captured by valence, arousal and dominance.6 conceptsSlides 11-1736 min
    1. 2.1Similarity is graded, and humans can put a number on it
    2. 2.2Relatedness: words that belong to the same scene
    3. 2.3Antonymy: opposite on one feature, alike on all the rest
    4. 2.4Connotation: the feeling a word carries
    5. 2.5Valence, arousal, dominance: affect as three numbers
    6. 2.6The map so far: senses, words and five relations
  3. 03Vector semantics and the distributional hypothesisDefining a word by its contexts (Wittgenstein, Harris, the ongchoi example), combining that with Osgood's meaning as a point in space, and arriving at embeddings, plus why vectors generalize better than word identities and the two kinds (sparse tf-idf, dense word2vec).6 conceptsSlides 18-3036 min
    1. 3.1Meaning from the company a word keeps
    2. 3.2Guessing ongchoi from the words around it
    3. 3.3Osgood's idea: a word's connotation as a point in space
    4. 3.4Embeddings: points placed by distribution
    5. 3.5Why vectors beat word identities
    6. 3.6Two families: sparse counts and dense predictions
  4. 04Term matrices and cosine similarityBuilding term-document and word-word (term-context) co-occurrence matrices, reading rows and columns as vectors, and measuring similarity with the dot product and its length-normalized form, the cosine.5 conceptsSlides 31-4430 min
    1. 4.1Documents as columns of counts
    2. 4.2Words as rows: term-document and word-word matrices
    3. 4.3The dot product and why it rewards frequent words
    4. 4.4Cosine: the length-normalized dot product
    5. 4.5Computing and seeing cosines: cherry, digital, information
  5. 05Weighting with tf-idfWhy raw frequency is a poor representation, how log-scaled term frequency and inverse document frequency combine into tf-idf, and a full worked tf-idf table for the Shakespeare example.5 conceptsSlides 45-5230 min
    1. 5.1Why raw counts mislead, and the two ways to reweight them
    2. 5.2Squashing term frequency with a logarithm
    3. 5.3Document frequency, not collection frequency: Romeo versus action
    4. 5.4Inverse document frequency, and what counts as a document
    5. 5.5The full tf-idf table for four Shakespeare plays
  6. 06Pointwise mutual information and PPMIMeasuring whether two words co-occur more than chance with PMI, clipping negatives to get PPMI, computing it step by step on a term-context matrix, and correcting PMI's bias toward rare words with alpha-weighted context probabilities.6 conceptsSlides 53-6036 min
    1. 6.1PMI: do two words meet more often than chance predicts?
    2. 6.2Why negative PMI is unreliable, and clipping it to zero gives PPMI
    3. 6.3From a count matrix to joint and marginal probabilities
    4. 6.4Computing one PMI cell and reading the PPMI matrix
    5. 6.5PMI's bias toward rare events
    6. 6.6Raising context counts to alpha = 0.75
  7. 07Dense vectors and the skip-gram classifierWhy short dense vectors beat long sparse ones, the main ways to get them, and how word2vec turns embedding learning into a self-supervised binary task scored by the sigmoid of a dot product.6 conceptsSlides 61-7436 min
    1. 7.1Short dense vectors versus long sparse ones
    2. 7.2Where dense vectors come from: predict, factorize, or contextualize
    3. 7.3Predict rather than count: self-supervision turns raw text into labels
    4. 7.4From a context window to (target, context) pairs
    5. 7.5Similarity as a dot product, squashed into a probability by the sigmoid
    6. 7.6Scoring a whole window: independence, products and log sums
  8. 08Learning skip-gram embeddings with negative samplingThe two embedding matrices W and C, how positive and k negative training pairs are built, the cross-entropy loss for SGNS, its gradients, and the SGD updates that pull true neighbors together and push sampled noise apart.6 conceptsSlides 75-8836 min
    1. 8.1Two tables for every word: target matrix W and context matrix C
    2. 8.2Positive pairs from the window, k noise pairs from a flattened unigram
    3. 8.3The SGNS cross-entropy loss: one neighbor up, k noise words down
    4. 8.4One step of stochastic gradient descent, seen as pulling and pushing
    5. 8.5Deriving the gradients and applying one SGD update by hand
    6. 8.6After training: summing w_i + c_i and the whole recipe
  9. 09Hyperparameters, word2vec variants and window sizeStandard SGNS settings, the CBOW, GloVe, FastText and skip-thought alternatives, and how the context window size decides whether neighbors are syntactic or topical.6 conceptsSlides 89-9736 min
    1. 9.1The standard SGNS settings and what each number trades
    2. 9.2CBOW: flipping skip-gram so the context predicts the center
    3. 9.3GloVe: fitting dot products to log co-occurrence counts
    4. 9.4FastText: a word is the sum of its character n-grams
    5. 9.5Skip-thought: the skip-gram idea one level up, for sentences
    6. 9.6Window size decides whether neighbors are look-alikes or topic-mates
  10. 10Analogies, bias and evaluating embeddingsThe parallelogram method for analogies and its limits, embeddings as a lens on historical meaning change and cultural bias, visualizing and evaluating embeddings intrinsically and extrinsically, and when to use pretrained embeddings.7 conceptsSlides 98-11142 min
    1. 10.1The parallelogram method: solve an analogy with one subtraction and one addition
    2. 10.2Relations as consistent directions: gender, tense and capitals in word2vec and GloVe
    3. 10.3Why analogy accuracy overstates what embeddings know
    4. 10.4Diachronic embeddings: watching gay, broadcast and awful change meaning
    5. 10.5Embeddings inherit and amplify cultural bias
    6. 10.6Looking at and testing embeddings: dendrograms, intrinsic and extrinsic evaluation
    7. 10.7Pretrained or trained with the model: choosing where embeddings come from

CGlossary and reference