Word meaning and synonymy Indices and logical symbols give identity, not meaning. A usable representation makes relations computable from the representation itself. Part 01: Word meaning and synonymy
e cat ⋅ e dog = 0 = e cat ⋅ e spreadsheet \mathbf{e}_{\text{cat}} \cdot \mathbf{e}_{\text{dog}} = 0 = \mathbf{e}_{\text{cat}} \cdot \mathbf{e}_{\text{spreadsheet}} e cat ⋅ e dog = 0 = e cat ⋅ e spreadsheet One-hot vectors: every pair of distinct words is equally unrelated Definitions worth memorizing
Lemma The citation form that groups wordforms: mouse for mouse and mice, sing for sang and sung.
Wordform An inflected surface form found in text, such as mice or duermes.
Sense One discrete aspect of a lemma's meaning, such as the rodent or the cursor device.
Polysemy A lemma with several senses; WordNet 3.0 gives mouse 4 noun and 2 verb senses.
Synonymy Substitutable without changing truth conditions; a relation between senses, not words.
Principle of contrast Every difference in form marks a difference in meaning, so perfect synonyms probably do not exist.
One-hot failure Distinct one-hot vectors have dot product 0 : cat is as far from dog as from spreadsheet. Pair Difference Why water / H₂O Genre and register H₂O is odd in a hiking guide die / pass away Register (politeness) pass away is euphemistic, pop off is slang politician / statesman Connotation statesman praises, politician often does not truck / lorry Dialect American against British English big / large Different sense big sister means older; large sister does not
Apparent synonyms and the dimension they differ on Exam tip
The five desiderata, one example each
Similarity (cat and dog), antonymy (hot and cold), connotation (happy and sad), perspective (buy, sell, pay), inference (Ann sold Bo a car, so Bo bought a car).
Similarity, relatedness and connotation Similarity is graded and word level; relatedness is broader and includes words that share an event or a semantic field. Part 02: Similarity, relatedness, connotation
Relation Holds between Nature Example Resource Synonymy Sense Rarely exact couch / sofa WordNet synsets Antonymy Sense Comes in kinds hot / cold WordNet antonym links Similarity Word Graded vanish / disappear SimLex-999 Relatedness Word Graded coffee / cup WordSim-353, association norms Connotation Word Graded replica / knockoff NRC VAD Lexicon
The relation map Kind Examples What is opposed Scale opposites long / short, hot / cold Two ends of one graded dimension Binary opposition in / out Two values, no middle ground Reversives rise / fall, up / down Opposite direction of change
Kinds of antonym Pair Similarity Note vanish / disappear 9.8 Top similarity night / day 1.88 Antonyms: association 8.19 short / long 1.23 Antonyms: association 5.36 large / big 9.55 Synonyms: association 0.68 hole / agreement 0.3 Neither similar nor related
SimLex-999 similarity, 0 to 10 Key idea
Why antonyms fool distributional models
Antonyms score low on similarity but high on association: they share almost every context, so count and prediction models place them close together.
Osgood's affective dimensions and the NRC VAD Lexicon
Valence Pleasantness. love 1.000 , nightmare 0.005 .
Arousal Intensity. frenzy 0.965 , napping 0.046 .
Dominance Control. powerful 0.991 , weak 0.045 .
NRC VAD v1 About 20k words, scored 0 to 1 by best-worst scaling.
NRC VAD v2 Over 55k terms (2025 ), scored -1 to 1 .
Near-synonym valence replica 0.480 against fake 0.073 . score ( w ) = # best ( w ) # seen ( w ) − # worst ( w ) # seen ( w ) \text{score}(w) = \frac{\#\text{best}(w)}{\#\text{seen}(w)} - \frac{\#\text{worst}(w)}{\#\text{seen}(w)} score ( w ) = # seen ( w ) # best ( w ) − # seen ( w ) # worst ( w ) Best-worst scaling, before rescaling to 0 to 1 Vector semantics Words in similar environments have similar meanings, roughly in proportion to how similar the environments are. A word becomes a point in a space built from its distribution, and that vector is an embedding. Part 03: Vector semantics
Thinker Year Statement Joos 1950 Meaning is the set of conditional probabilities of co-occurrence Wittgenstein 1953 The meaning of a word is its use in the language Harris 1954 Difference in meaning roughly matches difference in environments Firth 1957 You shall know a word by the company it keeps
The distributional hypothesis and its authors d ( u , v ) = ∑ i ( u i − v i ) 2 d(\mathbf{u},\mathbf{v})=\sqrt{\sum_i (u_i-v_i)^2} d ( u , v ) = i ∑ ( u i − v i ) 2 Euclidean distance between two points in an affective space Axis Sparse (tf-idf, PPMI) Dense (word2vec) Length |V| ≈ 20,000 to 50,000 d ≈ 50 to 1000 Zeros Almost every entry Almost none, values can be negative How built Weighted counts (tf-idf, PPMI) Classifier predicting neighbors (word2vec) One dimension means A specific context word or document Nothing individually Classifier weights per feature word 50,000 300 car and automobile Separate, orthogonal axes Can share directions Typical use IR, strong baseline Input features for neural NLP
Sparse against dense embeddings Key idea
Why vectors generalize
An identity feature fires only for the exact word. A vector feature lets awful [34, 21, 14] land almost where terrible [35, 22, 17] did (cosine ≈ 0.999 ), so a classifier reacts the same way.
Exam tip
Ongchoi in one line
Ongchoi shares sauteed, garlic, rice, leaves, delicious and salty with spinach, chard and collards, so it is a leafy green (water spinach). The 2D word map on slide 27 is a t-SNE projection, not the embedding itself.
Term matrices and cosine One count matrix gives two readings: columns are document vectors and rows are word vectors. Part 04: Term matrices and cosine
Matrix Shape Cell Similarity it captures Term-document |V| x |D| Count of word in document Topical: which texts a word appears in Word-word (term-context) |V| x |V| Count of context word within a window such as ±4 Closer, more substitutable similarity
The two count matrices Word As You Like It Twelfth Night Julius Caesar Henry V battle 1 0 7 13 good 114 80 62 89 fool 36 58 1 4 wit 20 15 2 3
Shakespeare term-document counts v ⋅ w = ∑ i = 1 N v i w i ∣ v ∣ = ∑ i = 1 N v i 2 \begin{gathered} \mathbf{v}\cdot\mathbf{w} = \sum_{i=1}^{N} v_i w_i \\ |\mathbf{v}| = \sqrt{\sum_{i=1}^{N} v_i^2} \end{gathered} v ⋅ w = i = 1 ∑ N v i w i ∣ v ∣ = i = 1 ∑ N v i 2 Dot product and vector length cos ( v , w ) = v ⋅ w ∣ v ∣ ∣ w ∣ = ∑ i v i w i ∑ i v i 2 ∑ i w i 2 \cos(\mathbf{v},\mathbf{w}) = \frac{\mathbf{v}\cdot\mathbf{w}}{|\mathbf{v}|\,|\mathbf{w}|} = \frac{\sum_{i} v_i w_i}{\sqrt{\sum_{i} v_i^2}\,\sqrt{\sum_{i} w_i^2}} cos ( v , w ) = ∣ v ∣ ∣ w ∣ v ⋅ w = ∑ i v i 2 ∑ i w i 2 ∑ i v i w i Cosine: the dot product of unit vectors, range -1 to 1, and 0 to 1 for counts The worked cosine, every intermediate quantity
Vectors (pie, data, computer) cherry [442, 8, 2] , digital [5, 1683, 1670] , information [5, 3982, 3325]
|cherry| √195432 ≈ 442.08
|digital| √5621414 ≈ 2370.95
|information| √26911974 ≈ 5187.68
cherry · information 2210 + 31856 + 6650 = 40716
digital · information 25 + 6701706 + 5552750 = 12254481
cos(cherry, information) 40716 / 2293351.5 ≈ 0.018 (slide prints .017 )
cos(digital, information) ≈ 0.996 Exam tip
Why cosine and not the dot product
The raw dot product grows with vector length, so frequent words win. Cosine divides out both lengths and keeps only the angle. On counts every product is non-negative, and Cauchy-Schwarz gives |v·w| ≤ |v||w| , so the value lies in [0, 1] .
tf-idf Reward words frequent in this document, penalize words frequent everywhere. Part 05: tf-idf
t f t , d = log 10 ( c o u n t ( t , d ) + 1 ) \mathrm{tf}_{t,d} = \log_{10}\big(\mathrm{count}(t,d) + 1\big) tf t , d = log 10 ( count ( t , d ) + 1 ) Slide form of term frequency t f t , d = { 1 + log 10 c o u n t ( t , d ) if c o u n t ( t , d ) > 0 0 otherwise \mathrm{tf}_{t,d} = \begin{cases} 1 + \log_{10} \mathrm{count}(t,d) & \text{if } \mathrm{count}(t,d) > 0 \\ 0 & \text{otherwise} \end{cases} tf t , d = { 1 + log 10 count ( t , d ) 0 if count ( t , d ) > 0 otherwise Current SLP3 and IIR form i d f t = log 10 ( N d f t ) w t , d = t f t , d × i d f t \mathrm{idf}_t = \log_{10}\!\left(\frac{N}{\mathrm{df}_t}\right) \qquad w_{t,d} = \mathrm{tf}_{t,d} \times \mathrm{idf}_t idf t = log 10 ( df t N ) w t , d = tf t , d × idf t idf and the tf-idf weight; a term in every document gets weight 0 tf ladder and the cf against df test
count 0 log10(1) = 0
count 1 log10(2) = 0.301
count 9 log10(10) = 1
count 99 log10(100) = 2
cf against df Romeo and action both have cf 113 ; df 1 against 31 , so idf 1.57 against 0.077 . Word df idf Romeo 1 1.57 salad 2 1.27 Falstaff 4 0.966 forest 12 0.489 battle 21 0.246 wit 34 0.037 fool 36 0.012 good, sweet 37 0
idf ladder, N = 37 plays Word Counts idf tf-idf battle 1, 0, 7, 13 0.246 0.074, 0, 0.22, 0.28 good 114, 80, 62, 89 0 0, 0, 0, 0 fool 36, 58, 1, 4 0.012 0.019, 0.021, 0.0036, 0.0083 wit 20, 15, 2, 3 0.037 0.049, 0.044, 0.018, 0.022
Counts to tf-idf weights (AYLI, TN, JC, H5) Source tf idf idf of a term in every document Slides (earlier SLP3) log10(count + 1) log10(N / df) 0 SLP3 current draft, IIR 1 + log10(count), or 0 log10(N / df) 0 scikit-learn default raw count ln((1 + n) / (1 + df)) + 1 1
Three formulations you will meet Exam tip
One tf-idf cell in three lines
tf = log10(count + 1) , idf = log10(N / df) , w = tf × idf , and state N . Battle in Julius Caesar: log10 8 × 0.246 = 0.903 × 0.246 ≈ 0.22 . Keep four significant figures until the last step.
PMI and PPMI PMI compares observed co-occurrence with what independence predicts, in bits. Part 06: PPMI
PMI ( w , c ) = log 2 P ( w , c ) P ( w ) P ( c ) \operatorname{PMI}(w, c) = \log_2 \frac{P(w, c)}{P(w)\,P(c)} PMI ( w , c ) = log 2 P ( w ) P ( c ) P ( w , c ) Zero means chance, positive attraction, negative avoidance PPMI ( w , c ) = max ( log 2 P ( w , c ) P ( w ) P ( c ) , 0 ) \operatorname{PPMI}(w, c) = \max\!\left(\log_2 \frac{P(w, c)}{P(w)\,P(c)},\ 0\right) PPMI ( w , c ) = max ( log 2 P ( w ) P ( c ) P ( w , c ) , 0 ) Clip the unreliable negative side p i j = f i j N , p i ∗ = ∑ j f i j N , p ∗ j = ∑ i f i j N , PMI i j = log 2 f i j N f i ∗ f ∗ j \begin{gathered} p_{ij} = \frac{f_{ij}}{N}, \qquad p_{i*} = \frac{\sum_j f_{ij}}{N}, \\ p_{*j} = \frac{\sum_i f_{ij}}{N}, \\ \operatorname{PMI}_{ij} = \log_2 \frac{f_{ij}\, N}{f_{i*}\, f_{*j}} \end{gathered} p ij = N f ij , p i ∗ = N ∑ j f ij , p ∗ j = N ∑ i f ij , PMI ij = log 2 f i ∗ f ∗ j f ij N Probabilities from one matrix total, and the count shortcut PMI PPMI Range -∞ to +∞ 0 to +∞ Unseen pair log2 0 = -∞ 0 Negative side Needs enormous corpora to estimate Discarded Matrix shape Dense, many large negatives Sparse, mostly exact zeros
PMI against PPMI Word computer data result pie sugar row sum cherry 2 8 9 442 25 486 strawberry 0 0 1 60 19 80 digital 1670 1683 85 5 4 3447 information 3325 3982 378 5 13 7703 column sum 4997 5673 473 512 61 N = 11716
Term-context counts Word computer data result pie sugar cherry 0 0 0 4.38 3.30 strawberry 0 0 0 4.10 5.51 digital 0.18 0.01 0 0 0 information 0.02 0.09 0.28 0 0
PPMI matrix, log base 2 Exam tip
One PPMI cell by hand
information and data: p(w, c) = 3982 / 11716 = .3399 , p(w) p(c) = .6575 × .4842 = .3184 , so PMI = log2(.3399 / .3184) ≈ 0.09 . strawberry and sugar gives 5.51 . Always state the clip, and say "0 because the count is zero" for empty cells.
P α ( c ) = count ( c ) α ∑ c ′ count ( c ′ ) α , α = 0.75 P_\alpha(c) = \frac{\operatorname{count}(c)^\alpha}{\sum_{c'} \operatorname{count}(c')^\alpha}, \qquad \alpha = 0.75 P α ( c ) = ∑ c ′ count ( c ′ ) α count ( c ) α , α = 0.75 Context smoothing: rare contexts gain probability, so their PMI falls Add-k smoothing Context alpha What changes Every count gets + k Only P(c) is reshaped Typical value k from 0.1 to 3 alpha = 0.75 strawberry / sugar 5.51 → 5.31 (k = 2) 5.51 → 4.01 Origin Laplace smoothing from n-gram models word2vec negative sampling
Two fixes for PMI's rare-event bias The skip-gram classifier Word2vec predicts rather than counts: a binary classifier asks whether c is near w, and its weights become the embeddings. Neighbors in running text are gold positives, so no human labels are needed. Part 07: The skip-gram classifier
Family Models Reads Vector per Prediction word2vec SGNS, CBOW Local windows Word type (static) Global regression GloVe Global co-occurrence counts Word type (static) Factorization SVD, LSA Term-document or PPMI matrix Word type (static) Contextual ELMo, BERT The whole sentence at run time Token in context
Dense vector families σ ( x ) = 1 1 + exp ( − x ) σ ( 0 ) = 0.5 1 − σ ( x ) = σ ( − x ) \begin{gathered} \sigma(x)=\frac{1}{1+\exp(-x)} \\ \sigma(0) = 0.5 \qquad 1 - \sigma(x) = \sigma(-x) \end{gathered} σ ( x ) = 1 + exp ( − x ) 1 σ ( 0 ) = 0.5 1 − σ ( x ) = σ ( − x ) The logistic sigmoid and its complement rule P ( + ∣ w , c ) = σ ( c ⋅ w ) P ( − ∣ w , c ) = σ ( − c ⋅ w ) \begin{gathered} P(+\mid w,c)=\sigma(\mathbf{c}\cdot\mathbf{w}) \\ P(-\mid w,c)=\sigma(-\mathbf{c}\cdot\mathbf{w}) \end{gathered} P ( + ∣ w , c ) = σ ( c ⋅ w ) P ( − ∣ w , c ) = σ ( − c ⋅ w ) Similarity as a dot product, squashed into a probability P ( + ∣ w , c 1 : L ) = ∏ i = 1 L σ ( c i ⋅ w ) log P ( + ∣ w , c 1 : L ) = ∑ i = 1 L log σ ( c i ⋅ w ) \begin{gathered} P(+\mid w,c_{1:L})=\prod_{i=1}^{L}\sigma(\mathbf{c}_i\cdot\mathbf{w}) \\ \log P(+\mid w,c_{1:L})=\sum_{i=1}^{L}\log\sigma(\mathbf{c}_i\cdot\mathbf{w}) \end{gathered} P ( + ∣ w , c 1 : L ) = i = 1 ∏ L σ ( c i ⋅ w ) log P ( + ∣ w , c 1 : L ) = i = 1 ∑ L log σ ( c i ⋅ w ) A whole window of L context words (L = 2m for a ±m window), assuming independent context words Context c · w σ(c · w) log σ tablespoon 1.0 0.7311 -0.3133 of -0.1 0.4750 -0.7444 jam 3.0 0.9526 -0.0486 a 0.1 0.5250 -0.6444 Window total product ≈ 0.174 sum ≈ -1.75
Scoring the apricot window, ±2 (natural log) Learning SGNS embeddings One positive pulled up, k noise words pushed down, by SGD on two matrices. Part 08: Learning skip-gram embeddings
θ = [ W C ] ∈ R 2 ∣ V ∣ × d \theta=\begin{bmatrix}W\\C\end{bmatrix}\in\mathbb{R}^{2|V|\times d} θ = [ W C ] ∈ R 2∣ V ∣ × d Target and context matrices stacked Parameters and training pairs
W, rows 1 to |V| Target (input) embeddings, used when the word is at the center.
C, rows |V| + 1 to 2|V| Context (output) embeddings, used for neighbors and noise words.
Size 2|V| × d ; |V| = 10,000 at d = 300 gives 6,000,000 parameters.
Rows touched per pair 2 + k : one target, one positive context, k noise contexts.
Noise words Drawn from P_α with α = 0.75 , never the target, never checked for co-occurrence. L C E = − [ log σ ( c p o s ⋅ w ) + ∑ i = 1 k log σ ( − c n e g i ⋅ w ) ] \begin{aligned} L_{CE}=-\Big[&\log\sigma(\mathbf{c}_{pos}\cdot\mathbf{w}) \\ &+\sum_{i=1}^{k}\log\sigma(-\mathbf{c}_{neg_i}\cdot\mathbf{w})\Big] \end{aligned} L C E = − [ log σ ( c p os ⋅ w ) + i = 1 ∑ k log σ ( − c n e g i ⋅ w ) ] Independence, log of a product, P(-) = 1 - P(+), then 1 - σ(x) = σ(-x) Parameter Label y Gradient Update moves it c_pos 1 [σ(c_pos·w) − 1] w Toward w c_neg_i 0 σ(c_neg_i·w) w Away from w w both [σ(c_pos·w) − 1] c_pos + Σ σ(c_neg_i·w) c_neg_i Toward c_pos, away from each c_neg_i
The gradients: (σ - y) times the partner vector c p o s t + 1 = c p o s t − η [ σ ( c p o s t ⋅ w t ) − 1 ] w t c n e g i t + 1 = c n e g i t − η [ σ ( c n e g i t ⋅ w t ) ] w t w t + 1 = w t − η [ [ σ ( c p o s t ⋅ w t ) − 1 ] c p o s t + ∑ i = 1 k [ σ ( c n e g i t ⋅ w t ) ] c n e g i t ] \begin{aligned}
\mathbf{c}_{pos}^{t+1}&=\mathbf{c}_{pos}^{t}-\eta\big[\sigma(\mathbf{c}_{pos}^{t}\cdot\mathbf{w}^{t})-1\big]\mathbf{w}^{t}\\
\mathbf{c}_{neg_i}^{t+1}&=\mathbf{c}_{neg_i}^{t}-\eta\big[\sigma(\mathbf{c}_{neg_i}^{t}\cdot\mathbf{w}^{t})\big]\mathbf{w}^{t}\\
\mathbf{w}^{t+1}&=\mathbf{w}^{t}-\eta\Big[\big[\sigma(\mathbf{c}_{pos}^{t}\cdot\mathbf{w}^{t})-1\big]\mathbf{c}_{pos}^{t}\\&\qquad+\sum_{i=1}^{k}\big[\sigma(\mathbf{c}_{neg_i}^{t}\cdot\mathbf{w}^{t})\big]\mathbf{c}_{neg_i}^{t}\Big]
\end{aligned} c p os t + 1 c n e g i t + 1 w t + 1 = c p os t − η [ σ ( c p os t ⋅ w t ) − 1 ] w t = c n e g i t − η [ σ ( c n e g i t ⋅ w t ) ] w t = w t − η [ [ σ ( c p os t ⋅ w t ) − 1 ] c p os t + i = 1 ∑ k [ σ ( c n e g i t ⋅ w t ) ] c n e g i t ] SGD updates, every right-hand side at time t Choice Vector Effect Keep W only w_i Second-order similarity; gensim wv default Sum w_i + c_i Adds first-order co-occurrence; SLP3's common choice Concatenate [w_i ; c_i] Both views, dimension 2d ; rarely used
What to ship after training Exam tip
Justify every line of the derivation
Independence for the product, log turns it into a sum, the complement rule for negatives, then 1 − σ(x) = σ(−x) . For gradients use dσ/dz = σ(z)(1 − σ(z)) .
Hyperparameters and variants The standard settings, and the three variants judged against SGNS: CBOW, GloVe and FastText. Part 09: Hyperparameters and variants
The SGNS knobs
Model Skip-gram with negative sampling (SGNS).
Negatives k 15 to 20 on small data (5 to 20 in Mikolov), 2 to 5 on huge data.
Dimension d 300 , or 100 or 50 ; little gain past 300 .
Window half-width m 5 to 10 words per side, so 2m context words.
Noise distribution Unigram counts raised to 3/4 .
Subsampling t About 10^-5 : very frequent words are randomly dropped. dot products per positive pair = 1 + k parameters = 2 ∣ V ∣ d \begin{gathered} \text{dot products per positive pair} = 1 + k \\ \text{parameters} = 2\,|V|\,d \end{gathered} dot products per positive pair = 1 + k parameters = 2 ∣ V ∣ d What k and d cost CBOW Skip-gram Input The 2m context vectors, averaged into h One center vector Output A score for the center word A score for each context word Predictions per position 1 2m Dot products per position 1 + k 2m × (1 + k) Rare words Smoothed away by averaging Better: each occurrence updates them Strength Speed, frequent words Rare words, analogies, subword extensions
CBOW against skip-gram h = 1 2 m ∑ − m ≤ j ≤ m , j ≠ 0 w t + j \mathbf{h}=\tfrac{1}{2m}\sum_{-m\le j\le m,\,j\ne 0}\mathbf{w}_{t+j} h = 2 m 1 − m ≤ j ≤ m , j = 0 ∑ w t + j CBOW averages the 2m context vectors of a ±m window into one vector J ( θ ) = ∑ w , c f ( N ( w , c ) ) ( u c ⊤ v w + b c + b ˉ w − log N ( w , c ) ) 2 \begin{aligned} J(\theta)=\sum_{w,c} f\big(N(w,c)\big)\,\big(&\mathbf{u}_c^\top\mathbf{v}_w \\ &+b_c+\bar b_w \\ &-\log N(w,c)\big)^2 \end{aligned} J ( θ ) = w , c ∑ f ( N ( w , c ) ) ( u c ⊤ v w + b c + b ˉ w − log N ( w , c ) ) 2 GloVe: weighted least squares on log co-occurrence f ( x ) = { ( x / x max ) α x < x max 1 otherwise x max = 100 , α = 3 4 \begin{gathered} f(x)=\begin{cases}(x/x_{\max})^{\alpha} & x<x_{\max}\\ 1 & \text{otherwise}\end{cases} \\ x_{\max}=100,\ \alpha=\tfrac{3}{4} \end{gathered} f ( x ) = { ( x / x m a x ) α 1 x < x m a x otherwise x m a x = 100 , α = 4 3 GloVe weighting: f(0) = 0 skips empty cells, the cap stops frequent pairs dominating Count x Weight f(x) 1 0.0316 10 0.1778 50 0.5946 100 1 1000 1 (capped)
f(x) at the published settings SGNS GloVe Data Stream of (target, context) pairs Global co-occurrence matrix, counted once Objective Logistic loss, true up, noise down Weighted squared error to log N(w,c) Negatives k sampled per positiveNone: f(0) = 0 drops zero cells Frequency control Noise from U(w)^(3/4) , subsampling Weight (x/x_max)^(3/4) , capped at 1 Final vectors W, or W + C W + W̃
SGNS against GloVe u w = ∑ g ∈ G w z g \mathbf{u}_w=\sum_{g\in\mathcal{G}_w}\mathbf{z}_g u w = g ∈ G w ∑ z g FastText: a word is the sum of its character n-gram vectors Exam tip
FastText and unseen words
Pad with < and > , take character n-grams of length 3 to 6 plus the whole word, and sum their vectors. An unseen word still has n-grams learned from other words.
Skip-gram Skip-thought Unit A word A sentence Context ±m words Previous and next sentence Encoder Lookup in W GRU over the words Output size 100 to 300 2400 (uni-skip), 4800 (combine-skip)
Skip-gram against skip-thought Small window Large window Window ±2 ±5 or more Sees The syntactic frame The topic of the passage Neighbor kind Functional: same slot, same POS Topical: same semantic field Hogwarts Sunnydale, Evernight, Blandings Dumbledore, half-blood, Malfoy Good for Synonyms, POS-like features, slot filling Topic modelling, retrieval, query expansion
Window size decides the neighbor type Analogy, bias and evaluation A relation is a roughly constant offset; the same geometry also encodes and amplifies social bias. Part 10: Analogy, bias and evaluation
b ^ ∗ = argmin x ∈ V ∖ { a , a ∗ , b } ∥ x − ( a ∗ − a + b ) ∥ \hat{b}^{*}=\operatorname*{argmin}_{x\in V\setminus\{a,a^{*},b\}} \lVert \mathbf{x}-(\mathbf{a}^{*}-\mathbf{a}+\mathbf{b})\rVert b ^ ∗ = x ∈ V ∖ { a , a ∗ , b } argmin ∥ x − ( a ∗ − a + b )∥ Parallelogram method by distance; with cosine, take the argmax 3CosMul: argmax b ∗ ∈ V cos ( b ∗ , b ) cos ( b ∗ , a ∗ ) cos ( b ∗ , a ) + ε , ε = 0.001 \begin{gathered} \text{3CosMul: }\operatorname*{argmax}_{b^{*}\in V}\frac{\cos(b^{*},b)\,\cos(b^{*},a^{*})}{\cos(b^{*},a)+\varepsilon}, \\ \varepsilon=0.001 \end{gathered} 3CosMul: b ∗ ∈ V argmax cos ( b ∗ , a ) + ε cos ( b ∗ , b ) cos ( b ∗ , a ∗ ) , ε = 0.001 3CosMul, which improves on the additive 3CosAdd Relation Offset Gender man → woman Number king → kings Tense walking → walked Capital Spain → Madrid Title earl → countess
Relations as offsets Without excluding the inputs, the method returns b about 93% of the time. It works best for frequent words, inflectional relations and country to capital. Accuracy depends on the relation type and even its direction. R ( t ) = argmin Q ⊤ Q = I ∥ W ( t ) Q − W ( t + 1 ) ∥ F R^{(t)}=\operatorname*{argmin}_{Q^{\top}Q=I}\lVert W^{(t)}Q-W^{(t+1)}\rVert_F R ( t ) = Q ⊤ Q = I argmin ∥ W ( t ) Q − W ( t + 1 ) ∥ F Orthogonal Procrustes aligns decade embeddings before measuring change Key idea
Two laws of semantic change
Conformity: frequent words change slowly. Innovation: polysemous words change fast. Examples of shift: gay, broadcast, awful.
Harm Definition Example Allocational Resources or opportunities distributed unfairly Resume search ranking by similarity to programmer pushes women down Representational A group demeaned, stereotyped or erased African American names closer to unpleasant words (Caliskan et al.)
Harm types bias ( w ) = ∥ w − v men ∥ − ∥ w − v women ∥ \text{bias}(w)=\lVert \mathbf{w}-\mathbf{v}_{\text{men}}\rVert-\lVert \mathbf{w}-\mathbf{v}_{\text{women}}\rVert bias ( w ) = ∥ w − v men ∥ − ∥ w − v women ∥ Garg et al. relative norm difference, tracked per decade Benchmark Kind Data Score WordSim-353 Intrinsic 353 pairs, 0 to 10, relatedness Spearman ρ SimLex-999 Intrinsic 999 pairs, similarity only Spearman ρ TOEFL synonyms Intrinsic 80 items, 4 choices Accuracy of argmax cosine Analogy sets (Google, BATS) Intrinsic a : a* :: b : ? Parallelogram accuracy NER, MT, coreference Extrinsic A full task with labelled data Task metric (F1, BLEU)
Intrinsic and extrinsic evaluation ρ = 1 − 6 ∑ i d i 2 n ( n 2 − 1 ) \rho = 1-\frac{6\sum_i d_i^{2}}{n(n^{2}-1)} ρ = 1 − n ( n 2 − 1 ) 6 ∑ i d i 2 Spearman's rank correlation without ties Pair Human Model Human rank Model rank d drink, ear 1.31 0.08 1 1 0 plane, car 5.77 0.42 2 2 0 drink, eat 6.87 0.61 3 4 1 planet, star 8.45 0.55 4 3 1
Spearman worked: Σd² = 2, n = 4, ρ = 1 - 12/60 = 0.8 Option Updated in training When Evidence Frozen pretrained No Little labelled data, simple task Kim's CNN-static Fine-tuned pretrained Yes, from pretrained values Moderate data, task nuance Kim's CNN-non-static Joint from random Yes, from random values Large data, hard task (LM, MT) Kim's CNN-rand
Where the embeddings come from Exam tip
Three analogy caveats and the evaluation split
Exclude the inputs, compare with an only-b baseline, and report per relation. Extrinsic evaluation on a real task matters more than intrinsic scores; SimLex-999 was built after WordSim-353 to separate similarity from relatedness.
Slide errata Mistakes on the original slides, with the correction to use in the exam.
Slide by slide
Slides 6 and 8 H₂0 is written with the digit zero; the formula is H₂O . Part 01 Slide 13 The restaurant field lists menu twice; a harmless duplicate. Part 02 Slide 16 controlling is shown at 0.983 and influenced at 0.081 ; in NRC VAD v1 controlling is 0.885 and influenced has no entry (0.983 is leadership, 0.081 is empty). Part 02 Slide 25 Captioned Osgood et al. (1957 ), but the numbers are Warriner et al. (2013 ) ratings on a 1 to 9 scale. Part 03 Slide 33 The top battle axis tick reads 40 ; it should be 20 (inherited from SLP). Part 04 Slide 43 cos(cherry, information) = .017 is a truncation; it rounds to .018 . Part 04 Slide 47 PMI is written with a bare log; slides 54 to 60 use log2 . Part 05 Slide 50 N = 37 is never stated; Falstaff's idf is 0.966 , not 0.967 . Part 05 Slide 57 The joint is divided by 111716 ; the total is 11716 , which gives the printed .3399 . Part 06 Slide 60 The denominator of P_alpha(b) reads .01^.75 + .01^.75 ; it must be .99^.75 + .01^.75 . The answer .03 is right. Part 06 Slide 75 The last C index reads 2V ; it is 2|V| . Part 08 Slide 84 The learning rate renders as h ; it is η . Part 08 Slide 86 The w update drops the superscript t on c_pos and c_neg_i ; the old values are meant. Part 08 Slide 88 "V random vectors" should be 2|V| , and negatives are random draws from P_α , not pairs checked not to co-occur. Part 08 Slide 93 Says n-gram vectors are averaged, then summed. The paper and SLP3 define the sum (the library averages); write the sum. Part 09 Slide 99 argmax of distance should be argmin of distance, or argmax of cosine, with a, a*, b excluded. Part 10