ICS 582Lecture 04
Word embeddings
How to represent word meaning as vectors: lexical semantics desiderata, the distributional hypothesis, sparse count vectors weighted by tf-idf and PPMI with cosine similarity, dense word2vec skip-gram embeddings trained with negative sampling, their variants (CBOW, GloVe, FastText), and the properties, biases and evaluation of embeddings.
- Parts
- 10
- Concepts
- 59
- Slides
- 111
- Reading
- 354 min
AOverview
This lecture answers the question every later model in the course depends on: how can a machine represent what a word means? It starts by showing that a string or an index carries no meaning at all, lists what a theory of word meaning must deliver (senses, synonymy, similarity, relatedness, antonymy and connotation), and then builds that theory out of one idea: a word is known by the company it keeps. From there it turns co-occurrence counts into vectors, reweights them with tf-idf and PPMI, learns short dense vectors with word2vec skip-gram and negative sampling, and finally probes what those vectors know, and what prejudices they absorb, through analogies, bias tests and evaluation.
As a PhD student you need it in three places. Exam questions ask you to compute, not to recall: a cosine between two count vectors, a tf-idf cell for a word in a play, a PMI cell from a small co-occurrence table, a sigmoid probability for a target and context pair, and one stochastic gradient step that pulls a true neighbor closer and pushes noise words away. Your research will use embeddings as the input layer of almost every model you train or compare, so you must be able to defend a choice of window size, weighting, dimensionality and evaluation, and to recognize when a static vector is too blunt for a word with many senses. And any system you build on retrieval, classification or Arabic text inherits both the strengths and the biases of the vectors it starts from, so you should be able to inspect, test and explain them.
The path has five stops. You start with meaning as linguists describe it: lemmas and senses, and the relations between words that any representation must respect. You then count, turning documents and context windows into term-document and word-word matrices and comparing rows with the cosine. You weight those counts so that frequent but uninformative words stop dominating, first with tf-idf and then with PPMI. You learn dense vectors by training a binary classifier that tells real neighbors from sampled noise, and you see how CBOW, GloVe and FastText vary that recipe. Finally you probe the result: analogies as parallel directions, meaning change over time, cultural bias, and the intrinsic and extrinsic tests that decide whether a set of embeddings is any good.
Success looks like
- Distinguish lemma, sense, synonymy, similarity, relatedness, antonymy and connotation on a new word pair, and explain why the principle of contrast makes perfect synonyms unlikely.
- Build a term-document or word-word matrix from a short corpus and compute a dot product, a vector length and a cosine by hand, explaining why the cosine corrects for frequency.
- Compute a tf-idf weight with log term frequency and idf = log10(N / df), and say why document frequency, not collection frequency, decides how informative a word is.
- Compute PMI and PPMI cells from a co-occurrence table, and explain both why negative values are clipped and why raising context counts to alpha = 0.75 tames the bias toward rare words.
- Generate positive and negative skip-gram pairs from a window, write the SGNS loss, and apply one gradient update to a target and its context vectors.
- Compare skip-gram, CBOW, GloVe and FastText, and predict how window size changes whether neighbors are look-alikes or topic-mates.
- Solve an analogy with the parallelogram method, state why its accuracy overstates what embeddings know, and describe how to detect bias and evaluate embeddings intrinsically and extrinsically.
How to study this lecture
- Choose your route. Read the full guide in one sitting for the big picture, or work through one part at a time when you want depth. The count-based parts and the word2vec parts each make a natural single session.
- Answer every recall prompt in your head, or on paper, before you reveal it. Cosines, tf-idf cells, PMI cells and gradient steps only stick if you have computed them yourself at least once.
- Take each quiz and read the explanation even when you are right. The explanations carry the distinctions an examiner probes, such as similarity against relatedness or document against collection frequency.
- Use the simulators: move the sigmoid inputs, slide the context window, and try analogies in the playground. Seeing a vector move is faster than reading that it moves.
- Mark a concept as understood only when you could explain it to a classmate without looking. Unmarked concepts show you where to return.
- Keep the glossary and the reference sheet open as you revise. Every formula in the lecture is collected there with its symbols, so you can check notation before an exam.
- Come back after a few days and retry the recall prompts and quizzes cold. Spaced practice builds the long-term memory an exam needs.
Sources
- Speech and Language Processing, 3rd edition draft, chapter 5: EmbeddingsBookStanford University (Jurafsky and Martin), release of 19 August 2026Free and current. The chapter behind this lecture: lexical semantics, vector semantics, cosine, tf-idf, PPMI, word2vec skip-gram with negative sampling, visualizing embeddings, the parallelogram method, bias and evaluation(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperAdvances in NeurIPS 2013 (Mikolov, Sutskever, Chen, Corrado and Dean)The paper that introduced skip-gram with negative sampling, the unigram distribution raised to the 3/4 power for drawing noise words, and subsampling of frequent words(opens in a new tab)
BThe 10 parts
- 01What words mean: lemmas, senses and synonymyWhy treating words as strings or logical symbols is unsatisfying, what lexical semantics asks of a theory of meaning, how lemmas split into senses, and why perfect synonymy probably does not exist.6 conceptsSlides 1-1036 min
- 1.1Strings, indices and capitalized symbols are not meaning
- 1.2What a theory of word meaning must deliver
- 1.3One lemma, many senses: the mouse entry
- 1.4Synonymy is a relation between senses, not words
- 1.5Why perfect synonyms probably do not exist: the principle of contrast
- 1.6From synonymy to similarity: car and bicycle
- 02Similarity, relatedness, antonymy and connotationThe graded relations between word senses (similarity rated by humans, relatedness through semantic fields, antonymy as opposition on one feature) and the affective meaning captured by valence, arousal and dominance.6 conceptsSlides 11-1736 min
- 2.1Similarity is graded, and humans can put a number on it
- 2.2Relatedness: words that belong to the same scene
- 2.3Antonymy: opposite on one feature, alike on all the rest
- 2.4Connotation: the feeling a word carries
- 2.5Valence, arousal, dominance: affect as three numbers
- 2.6The map so far: senses, words and five relations
- 03Vector semantics and the distributional hypothesisDefining a word by its contexts (Wittgenstein, Harris, the ongchoi example), combining that with Osgood's meaning as a point in space, and arriving at embeddings, plus why vectors generalize better than word identities and the two kinds (sparse tf-idf, dense word2vec).6 conceptsSlides 18-3036 min
- 04Term matrices and cosine similarityBuilding term-document and word-word (term-context) co-occurrence matrices, reading rows and columns as vectors, and measuring similarity with the dot product and its length-normalized form, the cosine.5 conceptsSlides 31-4430 min
- 05Weighting with tf-idfWhy raw frequency is a poor representation, how log-scaled term frequency and inverse document frequency combine into tf-idf, and a full worked tf-idf table for the Shakespeare example.5 conceptsSlides 45-5230 min
- 06Pointwise mutual information and PPMIMeasuring whether two words co-occur more than chance with PMI, clipping negatives to get PPMI, computing it step by step on a term-context matrix, and correcting PMI's bias toward rare words with alpha-weighted context probabilities.6 conceptsSlides 53-6036 min
- 6.1PMI: do two words meet more often than chance predicts?
- 6.2Why negative PMI is unreliable, and clipping it to zero gives PPMI
- 6.3From a count matrix to joint and marginal probabilities
- 6.4Computing one PMI cell and reading the PPMI matrix
- 6.5PMI's bias toward rare events
- 6.6Raising context counts to alpha = 0.75
- 07Dense vectors and the skip-gram classifierWhy short dense vectors beat long sparse ones, the main ways to get them, and how word2vec turns embedding learning into a self-supervised binary task scored by the sigmoid of a dot product.6 conceptsSlides 61-7436 min
- 7.1Short dense vectors versus long sparse ones
- 7.2Where dense vectors come from: predict, factorize, or contextualize
- 7.3Predict rather than count: self-supervision turns raw text into labels
- 7.4From a context window to (target, context) pairs
- 7.5Similarity as a dot product, squashed into a probability by the sigmoid
- 7.6Scoring a whole window: independence, products and log sums
- 08Learning skip-gram embeddings with negative samplingThe two embedding matrices W and C, how positive and k negative training pairs are built, the cross-entropy loss for SGNS, its gradients, and the SGD updates that pull true neighbors together and push sampled noise apart.6 conceptsSlides 75-8836 min
- 8.1Two tables for every word: target matrix W and context matrix C
- 8.2Positive pairs from the window, k noise pairs from a flattened unigram
- 8.3The SGNS cross-entropy loss: one neighbor up, k noise words down
- 8.4One step of stochastic gradient descent, seen as pulling and pushing
- 8.5Deriving the gradients and applying one SGD update by hand
- 8.6After training: summing w_i + c_i and the whole recipe
- 09Hyperparameters, word2vec variants and window sizeStandard SGNS settings, the CBOW, GloVe, FastText and skip-thought alternatives, and how the context window size decides whether neighbors are syntactic or topical.6 conceptsSlides 89-9736 min
- 9.1The standard SGNS settings and what each number trades
- 9.2CBOW: flipping skip-gram so the context predicts the center
- 9.3GloVe: fitting dot products to log co-occurrence counts
- 9.4FastText: a word is the sum of its character n-grams
- 9.5Skip-thought: the skip-gram idea one level up, for sentences
- 9.6Window size decides whether neighbors are look-alikes or topic-mates
- 10Analogies, bias and evaluating embeddingsThe parallelogram method for analogies and its limits, embeddings as a lens on historical meaning change and cultural bias, visualizing and evaluating embeddings intrinsically and extrinsically, and when to use pretrained embeddings.7 conceptsSlides 98-11142 min
- 10.1The parallelogram method: solve an analogy with one subtraction and one addition
- 10.2Relations as consistent directions: gender, tense and capitals in word2vec and GloVe
- 10.3Why analogy accuracy overstates what embeddings know
- 10.4Diachronic embeddings: watching gay, broadcast and awful change meaning
- 10.5Embeddings inherit and amplify cultural bias
- 10.6Looking at and testing embeddings: dendrograms, intrinsic and extrinsic evaluation
- 10.7Pretrained or trained with the model: choosing where embeddings come from