Majid Al-RaimiReference sheet

ICS 582Lecture 04Reference

Reference sheet

Word embeddings compressed onto one page: the definitions, formulas and numbers to have in your head before a quiz or exam.

Word meaning and synonymy

Indices and logical symbols give identity, not meaning. A usable representation makes relations computable from the representation itself. Part 01: Word meaning and synonymy

ecat⋅edog=0=ecat⋅espreadsheet\mathbf{e}_{\text{cat}} \cdot \mathbf{e}_{\text{dog}} = 0 = \mathbf{e}_{\text{cat}} \cdot \mathbf{e}_{\text{spreadsheet}}
One-hot vectors: every pair of distinct words is equally unrelated

Definitions worth memorizing

Lemma
The citation form that groups wordforms: mouse for mouse and mice, sing for sang and sung.
Wordform
An inflected surface form found in text, such as mice or duermes.
Sense
One discrete aspect of a lemma's meaning, such as the rodent or the cursor device.
Polysemy
A lemma with several senses; WordNet 3.0 gives mouse 4 noun and 2 verb senses.
Synonymy
Substitutable without changing truth conditions; a relation between senses, not words.
Principle of contrast
Every difference in form marks a difference in meaning, so perfect synonyms probably do not exist.
One-hot failure
Distinct one-hot vectors have dot product 0: cat is as far from dog as from spreadsheet.
PairDifferenceWhy
water / H₂OGenre and registerH₂O is odd in a hiking guide
die / pass awayRegister (politeness)pass away is euphemistic, pop off is slang
politician / statesmanConnotationstatesman praises, politician often does not
truck / lorryDialectAmerican against British English
big / largeDifferent sensebig sister means older; large sister does not
Apparent synonyms and the dimension they differ on

Similarity, relatedness and connotation

Similarity is graded and word level; relatedness is broader and includes words that share an event or a semantic field. Part 02: Similarity, relatedness, connotation

RelationHolds betweenNatureExampleResource
SynonymySenseRarely exactcouch / sofaWordNet synsets
AntonymySenseComes in kindshot / coldWordNet antonym links
SimilarityWordGradedvanish / disappearSimLex-999
RelatednessWordGradedcoffee / cupWordSim-353, association norms
ConnotationWordGradedreplica / knockoffNRC VAD Lexicon
The relation map
KindExamplesWhat is opposed
Scale oppositeslong / short, hot / coldTwo ends of one graded dimension
Binary oppositionin / outTwo values, no middle ground
Reversivesrise / fall, up / downOpposite direction of change
Kinds of antonym
PairSimilarityNote
vanish / disappear9.8Top similarity
night / day1.88Antonyms: association 8.19
short / long1.23Antonyms: association 5.36
large / big9.55Synonyms: association 0.68
hole / agreement0.3Neither similar nor related
SimLex-999 similarity, 0 to 10

Osgood's affective dimensions and the NRC VAD Lexicon

Valence
Pleasantness. love 1.000, nightmare 0.005.
Arousal
Intensity. frenzy 0.965, napping 0.046.
Dominance
Control. powerful 0.991, weak 0.045.
NRC VAD v1
About 20k words, scored 0 to 1 by best-worst scaling.
NRC VAD v2
Over 55k terms (2025), scored -1 to 1.
Near-synonym valence
replica 0.480 against fake 0.073.
score(w)=#best(w)#seen(w)−#worst(w)#seen(w)\text{score}(w) = \frac{\#\text{best}(w)}{\#\text{seen}(w)} - \frac{\#\text{worst}(w)}{\#\text{seen}(w)}
Best-worst scaling, before rescaling to 0 to 1

Vector semantics

Words in similar environments have similar meanings, roughly in proportion to how similar the environments are. A word becomes a point in a space built from its distribution, and that vector is an embedding. Part 03: Vector semantics

ThinkerYearStatement
Joos1950Meaning is the set of conditional probabilities of co-occurrence
Wittgenstein1953The meaning of a word is its use in the language
Harris1954Difference in meaning roughly matches difference in environments
Firth1957You shall know a word by the company it keeps
The distributional hypothesis and its authors
d(u,v)=∑i(ui−vi)2d(\mathbf{u},\mathbf{v})=\sqrt{\sum_i (u_i-v_i)^2}
Euclidean distance between two points in an affective space
AxisSparse (tf-idf, PPMI)Dense (word2vec)
Length|V| ≈ 20,000 to 50,000d ≈ 50 to 1000
ZerosAlmost every entryAlmost none, values can be negative
How builtWeighted counts (tf-idf, PPMI)Classifier predicting neighbors (word2vec)
One dimension meansA specific context word or documentNothing individually
Classifier weights per feature word50,000300
car and automobileSeparate, orthogonal axesCan share directions
Typical useIR, strong baselineInput features for neural NLP
Sparse against dense embeddings

Term matrices and cosine

One count matrix gives two readings: columns are document vectors and rows are word vectors. Part 04: Term matrices and cosine

MatrixShapeCellSimilarity it captures
Term-document|V| x |D|Count of word in documentTopical: which texts a word appears in
Word-word (term-context)|V| x |V|Count of context word within a window such as ±4Closer, more substitutable similarity
The two count matrices
WordAs You Like ItTwelfth NightJulius CaesarHenry V
battle10713
good114806289
fool365814
wit201523
Shakespeare term-document counts
v⋅w=∑i=1Nviwi∣v∣=∑i=1Nvi2\begin{gathered} \mathbf{v}\cdot\mathbf{w} = \sum_{i=1}^{N} v_i w_i \\ |\mathbf{v}| = \sqrt{\sum_{i=1}^{N} v_i^2} \end{gathered}
Dot product and vector length
cos⁡(v,w)=v⋅w∣v∣ ∣w∣=∑iviwi∑ivi2 ∑iwi2\cos(\mathbf{v},\mathbf{w}) = \frac{\mathbf{v}\cdot\mathbf{w}}{|\mathbf{v}|\,|\mathbf{w}|} = \frac{\sum_{i} v_i w_i}{\sqrt{\sum_{i} v_i^2}\,\sqrt{\sum_{i} w_i^2}}
Cosine: the dot product of unit vectors, range -1 to 1, and 0 to 1 for counts

The worked cosine, every intermediate quantity

Vectors (pie, data, computer)
cherry [442, 8, 2], digital [5, 1683, 1670], information [5, 3982, 3325]
|cherry|
√195432 ≈ 442.08
|digital|
√5621414 ≈ 2370.95
|information|
√26911974 ≈ 5187.68
cherry · information
2210 + 31856 + 6650 = 40716
digital · information
25 + 6701706 + 5552750 = 12254481
cos(cherry, information)
40716 / 2293351.5 ≈ 0.018 (slide prints .017)
cos(digital, information)
≈ 0.996

tf-idf

Reward words frequent in this document, penalize words frequent everywhere. Part 05: tf-idf

tft,d=log⁡10(count(t,d)+1)\mathrm{tf}_{t,d} = \log_{10}\big(\mathrm{count}(t,d) + 1\big)
Slide form of term frequency
tft,d={1+log⁡10count(t,d)if count(t,d)>00otherwise\mathrm{tf}_{t,d} = \begin{cases} 1 + \log_{10} \mathrm{count}(t,d) & \text{if } \mathrm{count}(t,d) > 0 \\ 0 & \text{otherwise} \end{cases}
Current SLP3 and IIR form
idft=log⁡10 ⁣(Ndft)wt,d=tft,d×idft\mathrm{idf}_t = \log_{10}\!\left(\frac{N}{\mathrm{df}_t}\right) \qquad w_{t,d} = \mathrm{tf}_{t,d} \times \mathrm{idf}_t
idf and the tf-idf weight; a term in every document gets weight 0

tf ladder and the cf against df test

count 0
log10(1) = 0
count 1
log10(2) = 0.301
count 9
log10(10) = 1
count 99
log10(100) = 2
cf against df
Romeo and action both have cf 113; df 1 against 31, so idf 1.57 against 0.077.
Worddfidf
Romeo11.57
salad21.27
Falstaff40.966
forest120.489
battle210.246
wit340.037
fool360.012
good, sweet370
idf ladder, N = 37 plays
WordCountsidftf-idf
battle1, 0, 7, 130.2460.074, 0, 0.22, 0.28
good114, 80, 62, 8900, 0, 0, 0
fool36, 58, 1, 40.0120.019, 0.021, 0.0036, 0.0083
wit20, 15, 2, 30.0370.049, 0.044, 0.018, 0.022
Counts to tf-idf weights (AYLI, TN, JC, H5)
Sourcetfidfidf of a term in every document
Slides (earlier SLP3)log10(count + 1)log10(N / df)0
SLP3 current draft, IIR1 + log10(count), or 0log10(N / df)0
scikit-learn defaultraw countln((1 + n) / (1 + df)) + 11
Three formulations you will meet

PMI and PPMI

PMI compares observed co-occurrence with what independence predicts, in bits. Part 06: PPMI

PMI⁡(w,c)=log⁡2P(w,c)P(w) P(c)\operatorname{PMI}(w, c) = \log_2 \frac{P(w, c)}{P(w)\,P(c)}
Zero means chance, positive attraction, negative avoidance
PPMI⁡(w,c)=max⁡ ⁣(log⁡2P(w,c)P(w) P(c), 0)\operatorname{PPMI}(w, c) = \max\!\left(\log_2 \frac{P(w, c)}{P(w)\,P(c)},\ 0\right)
Clip the unreliable negative side
pij=fijN,pi∗=∑jfijN,p∗j=∑ifijN,PMI⁡ij=log⁡2fij Nfi∗ f∗j\begin{gathered} p_{ij} = \frac{f_{ij}}{N}, \qquad p_{i*} = \frac{\sum_j f_{ij}}{N}, \\ p_{*j} = \frac{\sum_i f_{ij}}{N}, \\ \operatorname{PMI}_{ij} = \log_2 \frac{f_{ij}\, N}{f_{i*}\, f_{*j}} \end{gathered}
Probabilities from one matrix total, and the count shortcut
PMIPPMI
Range-∞ to +∞0 to +∞
Unseen pairlog2 0 = -∞0
Negative sideNeeds enormous corpora to estimateDiscarded
Matrix shapeDense, many large negativesSparse, mostly exact zeros
PMI against PPMI
Wordcomputerdataresultpiesugarrow sum
cherry28944225486
strawberry001601980
digital1670168385543447
information332539823785137703
column sum4997567347351261N = 11716
Term-context counts
Wordcomputerdataresultpiesugar
cherry0004.383.30
strawberry0004.105.51
digital0.180.01000
information0.020.090.2800
PPMI matrix, log base 2
Pα(c)=count⁡(c)α∑c′count⁡(c′)α,α=0.75P_\alpha(c) = \frac{\operatorname{count}(c)^\alpha}{\sum_{c'} \operatorname{count}(c')^\alpha}, \qquad \alpha = 0.75
Context smoothing: rare contexts gain probability, so their PMI falls
Add-k smoothingContext alpha
What changesEvery count gets + kOnly P(c) is reshaped
Typical valuek from 0.1 to 3alpha = 0.75
strawberry / sugar5.51 → 5.31 (k = 2)5.51 → 4.01
OriginLaplace smoothing from n-gram modelsword2vec negative sampling
Two fixes for PMI's rare-event bias

The skip-gram classifier

Word2vec predicts rather than counts: a binary classifier asks whether c is near w, and its weights become the embeddings. Neighbors in running text are gold positives, so no human labels are needed. Part 07: The skip-gram classifier

FamilyModelsReadsVector per
Predictionword2vec SGNS, CBOWLocal windowsWord type (static)
Global regressionGloVeGlobal co-occurrence countsWord type (static)
FactorizationSVD, LSATerm-document or PPMI matrixWord type (static)
ContextualELMo, BERTThe whole sentence at run timeToken in context
Dense vector families
σ(x)=11+exp⁡(−x)σ(0)=0.51−σ(x)=σ(−x)\begin{gathered} \sigma(x)=\frac{1}{1+\exp(-x)} \\ \sigma(0) = 0.5 \qquad 1 - \sigma(x) = \sigma(-x) \end{gathered}
The logistic sigmoid and its complement rule
P(+∣w,c)=σ(c⋅w)P(−∣w,c)=σ(−c⋅w)\begin{gathered} P(+\mid w,c)=\sigma(\mathbf{c}\cdot\mathbf{w}) \\ P(-\mid w,c)=\sigma(-\mathbf{c}\cdot\mathbf{w}) \end{gathered}
Similarity as a dot product, squashed into a probability
P(+∣w,c1:L)=∏i=1Lσ(ci⋅w)log⁡P(+∣w,c1:L)=∑i=1Llog⁡σ(ci⋅w)\begin{gathered} P(+\mid w,c_{1:L})=\prod_{i=1}^{L}\sigma(\mathbf{c}_i\cdot\mathbf{w}) \\ \log P(+\mid w,c_{1:L})=\sum_{i=1}^{L}\log\sigma(\mathbf{c}_i\cdot\mathbf{w}) \end{gathered}
A whole window of L context words (L = 2m for a ±m window), assuming independent context words
Contextc · wσ(c · w)log σ
tablespoon1.00.7311-0.3133
of-0.10.4750-0.7444
jam3.00.9526-0.0486
a0.10.5250-0.6444
Window totalproduct ≈ 0.174sum ≈ -1.75
Scoring the apricot window, ±2 (natural log)

Learning SGNS embeddings

One positive pulled up, k noise words pushed down, by SGD on two matrices. Part 08: Learning skip-gram embeddings

θ=[WC]∈R2∣V∣×d\theta=\begin{bmatrix}W\\C\end{bmatrix}\in\mathbb{R}^{2|V|\times d}
Target and context matrices stacked

Parameters and training pairs

W, rows 1 to |V|
Target (input) embeddings, used when the word is at the center.
C, rows |V| + 1 to 2|V|
Context (output) embeddings, used for neighbors and noise words.
Size
2|V| × d; |V| = 10,000 at d = 300 gives 6,000,000 parameters.
Rows touched per pair
2 + k: one target, one positive context, k noise contexts.
Noise words
Drawn from P_α with α = 0.75, never the target, never checked for co-occurrence.
LCE=−[log⁡σ(cpos⋅w)+∑i=1klog⁡σ(−cnegi⋅w)]\begin{aligned} L_{CE}=-\Big[&\log\sigma(\mathbf{c}_{pos}\cdot\mathbf{w}) \\ &+\sum_{i=1}^{k}\log\sigma(-\mathbf{c}_{neg_i}\cdot\mathbf{w})\Big] \end{aligned}
Independence, log of a product, P(-) = 1 - P(+), then 1 - σ(x) = σ(-x)
ParameterLabel yGradientUpdate moves it
c_pos1[σ(c_pos·w) − 1] wToward w
c_neg_i0σ(c_neg_i·w) wAway from w
wboth[σ(c_pos·w) − 1] c_pos + Σ σ(c_neg_i·w) c_neg_iToward c_pos, away from each c_neg_i
The gradients: (σ - y) times the partner vector
cpost+1=cpost−η[σ(cpost⋅wt)−1]wtcnegit+1=cnegit−η[σ(cnegit⋅wt)]wtwt+1=wt−η[[σ(cpost⋅wt)−1]cpost+∑i=1k[σ(cnegit⋅wt)]cnegit]\begin{aligned} \mathbf{c}_{pos}^{t+1}&=\mathbf{c}_{pos}^{t}-\eta\big[\sigma(\mathbf{c}_{pos}^{t}\cdot\mathbf{w}^{t})-1\big]\mathbf{w}^{t}\\ \mathbf{c}_{neg_i}^{t+1}&=\mathbf{c}_{neg_i}^{t}-\eta\big[\sigma(\mathbf{c}_{neg_i}^{t}\cdot\mathbf{w}^{t})\big]\mathbf{w}^{t}\\ \mathbf{w}^{t+1}&=\mathbf{w}^{t}-\eta\Big[\big[\sigma(\mathbf{c}_{pos}^{t}\cdot\mathbf{w}^{t})-1\big]\mathbf{c}_{pos}^{t}\\&\qquad+\sum_{i=1}^{k}\big[\sigma(\mathbf{c}_{neg_i}^{t}\cdot\mathbf{w}^{t})\big]\mathbf{c}_{neg_i}^{t}\Big] \end{aligned}
SGD updates, every right-hand side at time t
ChoiceVectorEffect
Keep W onlyw_iSecond-order similarity; gensim wv default
Sumw_i + c_iAdds first-order co-occurrence; SLP3's common choice
Concatenate[w_i ; c_i]Both views, dimension 2d; rarely used
What to ship after training

Hyperparameters and variants

The standard settings, and the three variants judged against SGNS: CBOW, GloVe and FastText. Part 09: Hyperparameters and variants

The SGNS knobs

Model
Skip-gram with negative sampling (SGNS).
Negatives k
15 to 20 on small data (5 to 20 in Mikolov), 2 to 5 on huge data.
Dimension d
300, or 100 or 50; little gain past 300.
Window half-width m
5 to 10 words per side, so 2m context words.
Noise distribution
Unigram counts raised to 3/4.
Subsampling t
About 10^-5: very frequent words are randomly dropped.
dot products per positive pair=1+kparameters=2 ∣V∣ d\begin{gathered} \text{dot products per positive pair} = 1 + k \\ \text{parameters} = 2\,|V|\,d \end{gathered}
What k and d cost
CBOWSkip-gram
InputThe 2m context vectors, averaged into hOne center vector
OutputA score for the center wordA score for each context word
Predictions per position12m
Dot products per position1 + k2m × (1 + k)
Rare wordsSmoothed away by averagingBetter: each occurrence updates them
StrengthSpeed, frequent wordsRare words, analogies, subword extensions
CBOW against skip-gram
h=12m∑−m≤j≤m, j≠0wt+j\mathbf{h}=\tfrac{1}{2m}\sum_{-m\le j\le m,\,j\ne 0}\mathbf{w}_{t+j}
CBOW averages the 2m context vectors of a ±m window into one vector
J(θ)=∑w,cf(N(w,c)) (uc⊤vw+bc+bˉw−log⁡N(w,c))2\begin{aligned} J(\theta)=\sum_{w,c} f\big(N(w,c)\big)\,\big(&\mathbf{u}_c^\top\mathbf{v}_w \\ &+b_c+\bar b_w \\ &-\log N(w,c)\big)^2 \end{aligned}
GloVe: weighted least squares on log co-occurrence
f(x)={(x/xmax⁡)αx<xmax⁡1otherwisexmax⁡=100, α=34\begin{gathered} f(x)=\begin{cases}(x/x_{\max})^{\alpha} & x<x_{\max}\\ 1 & \text{otherwise}\end{cases} \\ x_{\max}=100,\ \alpha=\tfrac{3}{4} \end{gathered}
GloVe weighting: f(0) = 0 skips empty cells, the cap stops frequent pairs dominating
Count xWeight f(x)
10.0316
100.1778
500.5946
1001
10001 (capped)
f(x) at the published settings
SGNSGloVe
DataStream of (target, context) pairsGlobal co-occurrence matrix, counted once
ObjectiveLogistic loss, true up, noise downWeighted squared error to log N(w,c)
Negativesk sampled per positiveNone: f(0) = 0 drops zero cells
Frequency controlNoise from U(w)^(3/4), subsamplingWeight (x/x_max)^(3/4), capped at 1
Final vectorsW, or W + CW + W̃
SGNS against GloVe
uw=∑g∈Gwzg\mathbf{u}_w=\sum_{g\in\mathcal{G}_w}\mathbf{z}_g
FastText: a word is the sum of its character n-gram vectors
Skip-gramSkip-thought
UnitA wordA sentence
Context±m wordsPrevious and next sentence
EncoderLookup in WGRU over the words
Output size100 to 3002400 (uni-skip), 4800 (combine-skip)
Skip-gram against skip-thought
Small windowLarge window
Window±2±5 or more
SeesThe syntactic frameThe topic of the passage
Neighbor kindFunctional: same slot, same POSTopical: same semantic field
HogwartsSunnydale, Evernight, BlandingsDumbledore, half-blood, Malfoy
Good forSynonyms, POS-like features, slot fillingTopic modelling, retrieval, query expansion
Window size decides the neighbor type

Analogy, bias and evaluation

A relation is a roughly constant offset; the same geometry also encodes and amplifies social bias. Part 10: Analogy, bias and evaluation

b^∗=argmin⁡x∈V∖{a,a∗,b}∥x−(a∗−a+b)∥\hat{b}^{*}=\operatorname*{argmin}_{x\in V\setminus\{a,a^{*},b\}} \lVert \mathbf{x}-(\mathbf{a}^{*}-\mathbf{a}+\mathbf{b})\rVert
Parallelogram method by distance; with cosine, take the argmax
3CosMul: argmax⁡b∗∈Vcos⁡(b∗,b) cos⁡(b∗,a∗)cos⁡(b∗,a)+ε,ε=0.001\begin{gathered} \text{3CosMul: }\operatorname*{argmax}_{b^{*}\in V}\frac{\cos(b^{*},b)\,\cos(b^{*},a^{*})}{\cos(b^{*},a)+\varepsilon}, \\ \varepsilon=0.001 \end{gathered}
3CosMul, which improves on the additive 3CosAdd
RelationOffset
Genderman → woman
Numberking → kings
Tensewalking → walked
CapitalSpain → Madrid
Titleearl → countess
Relations as offsets
  • Without excluding the inputs, the method returns b about 93% of the time.
  • It works best for frequent words, inflectional relations and country to capital.
  • Accuracy depends on the relation type and even its direction.
R(t)=argmin⁡Q⊤Q=I∥W(t)Q−W(t+1)∥FR^{(t)}=\operatorname*{argmin}_{Q^{\top}Q=I}\lVert W^{(t)}Q-W^{(t+1)}\rVert_F
Orthogonal Procrustes aligns decade embeddings before measuring change
HarmDefinitionExample
AllocationalResources or opportunities distributed unfairlyResume search ranking by similarity to programmer pushes women down
RepresentationalA group demeaned, stereotyped or erasedAfrican American names closer to unpleasant words (Caliskan et al.)
Harm types
bias(w)=∥w−vmen∥−∥w−vwomen∥\text{bias}(w)=\lVert \mathbf{w}-\mathbf{v}_{\text{men}}\rVert-\lVert \mathbf{w}-\mathbf{v}_{\text{women}}\rVert
Garg et al. relative norm difference, tracked per decade
BenchmarkKindDataScore
WordSim-353Intrinsic353 pairs, 0 to 10, relatednessSpearman ρ
SimLex-999Intrinsic999 pairs, similarity onlySpearman ρ
TOEFL synonymsIntrinsic80 items, 4 choicesAccuracy of argmax cosine
Analogy sets (Google, BATS)Intrinsica : a* :: b : ?Parallelogram accuracy
NER, MT, coreferenceExtrinsicA full task with labelled dataTask metric (F1, BLEU)
Intrinsic and extrinsic evaluation
ρ=1−6∑idi2n(n2−1)\rho = 1-\frac{6\sum_i d_i^{2}}{n(n^{2}-1)}
Spearman's rank correlation without ties
PairHumanModelHuman rankModel rankd
drink, ear1.310.08110
plane, car5.770.42220
drink, eat6.870.61341
planet, star8.450.55431
Spearman worked: Σd² = 2, n = 4, ρ = 1 - 12/60 = 0.8
OptionUpdated in trainingWhenEvidence
Frozen pretrainedNoLittle labelled data, simple taskKim's CNN-static
Fine-tuned pretrainedYes, from pretrained valuesModerate data, task nuanceKim's CNN-non-static
Joint from randomYes, from random valuesLarge data, hard task (LM, MT)Kim's CNN-rand
Where the embeddings come from

Slide errata

Mistakes on the original slides, with the correction to use in the exam.

Slide by slide

Slides 6 and 8
H₂0 is written with the digit zero; the formula is H₂O. Part 01
Slide 13
The restaurant field lists menu twice; a harmless duplicate. Part 02
Slide 16
controlling is shown at 0.983 and influenced at 0.081; in NRC VAD v1 controlling is 0.885 and influenced has no entry (0.983 is leadership, 0.081 is empty). Part 02
Slide 25
Captioned Osgood et al. (1957), but the numbers are Warriner et al. (2013) ratings on a 1 to 9 scale. Part 03
Slide 33
The top battle axis tick reads 40; it should be 20 (inherited from SLP). Part 04
Slide 43
cos(cherry, information) = .017 is a truncation; it rounds to .018. Part 04
Slide 47
PMI is written with a bare log; slides 54 to 60 use log2. Part 05
Slide 50
N = 37 is never stated; Falstaff's idf is 0.966, not 0.967. Part 05
Slide 57
The joint is divided by 111716; the total is 11716, which gives the printed .3399. Part 06
Slide 60
The denominator of P_alpha(b) reads .01^.75 + .01^.75; it must be .99^.75 + .01^.75. The answer .03 is right. Part 06
Slide 75
The last C index reads 2V; it is 2|V|. Part 08
Slide 84
The learning rate renders as h; it is η. Part 08
Slide 86
The w update drops the superscript t on c_pos and c_neg_i; the old values are meant. Part 08
Slide 88
"V random vectors" should be 2|V|, and negatives are random draws from P_α, not pairs checked not to co-occur. Part 08
Slide 93
Says n-gram vectors are averaged, then summed. The paper and SLP3 define the sum (the library averages); write the sum. Part 09
Slide 99
argmax of distance should be argmin of distance, or argmax of cosine, with a, a*, b excluded. Part 10