Majid Al-RaimiAnalogies, bias and evaluating embeddings

ICS 582Lecture 04Part 10

Analogies, bias and evaluating embeddings

The parallelogram method for analogies and its limits, embeddings as a lens on historical meaning change and cultural bias, visualizing and evaluating embeddings intrinsically and extrinsically, and when to use pretrained embeddings.

Concepts
7
Slides
98-111
Reading
42 min
Understood
0/7 concepts

Why this part matters

You have trained an embedding. Four questions follow immediately, and this part answers each. What has the space learned? Relations such as male to female or country to capital show up as directions you can probe with one subtraction and one addition. Can you trust it? Analogy scores flatter it, and it carries the biases of its corpus, which matters for any hiring, search or Arabic NLP system built on top of it. How do you measure it? Intrinsic tests against human judgments, or extrinsic tests inside a real task. And should you train it yourself at all, or take vectors someone else trained?

For exams, the parallelogram formula, the intrinsic versus extrinsic split and allocational versus representational harm are standard items. For research, diachronic embeddings and bias measurement are live tools: the same machinery that tracks how awful changed meaning also measures a century of gender stereotypes.

By the end you can

  1. Compute an analogy answer with the parallelogram method, using argmin distance or argmax cosine, and exclude the input words.
  2. Explain why analogy accuracy overstates relational knowledge, citing the exclusion effect and relation-dependence.
  3. Describe how aligned decade embeddings reveal semantic change, and state the laws of conformity and innovation.
  4. Explain how embeddings encode and amplify cultural bias, distinguish allocational from representational harm, and describe Garg's relative norm measure.
  5. Classify evaluations as intrinsic or extrinsic, and compute a Spearman correlation against human ratings.
  6. Choose between frozen, fine-tuned and jointly trained embeddings given data size and task difficulty.

Apple is to tree as grape is to what? Picture the arrow that starts at apple and ends at tree. It means something like "fruit to the plant it grows on". Now pick that same arrow up, keep its length and direction, and set its tail down on grape. Its head lands near vine. That is the whole method.

The apple to tree arrow is copied to start at grape; its tip lands beside vine, and the dashed edges close the parallelogram.

The same move works on the famous examples. Take king, subtract man, add woman, and the point you reach is close to queen. Take Paris, subtract France, add Italy, and you land close to Rome. In each case the subtraction isolates a relation (royalty without the maleness, the capital-of relation without the particular country) and the addition applies it to a new word. Because the four points form a parallelogram when it works, this is called the Parallelogram method.

The rule, in two equivalent forms

The slides write an analogy as a : a* :: b : b*, read "a is to a* as b is to b*". The relation is the offset from a to a*, so the point to search around is t = a* − a + b. No vocabulary word sits exactly at t, so the answer is the word nearest to it, and the three question words themselves are taken out of the candidate pool (the next concepts show why that matters so much). With Euclidean distance, nearest means smallest distance, so the operator is an argmin:

b^∗=argmin⁡x∈V∖{a,a∗,b}∥x−(a∗−a+b)∥\hat{b}^{*}=\operatorname*{argmin}_{x\in V\setminus\{a,a^{*},b\}} \lVert \mathbf{x}-(\mathbf{a}^{*}-\mathbf{a}+\mathbf{b})\rVert
Parallelogram method with Euclidean distance (SLP3 eq. 5.28, in the slide's a : a* :: b : b* order)

Mikolov, Yih and Zweig, who made the method famous for dense vectors, wrote it the other way round. They normalize every vector to unit length, compute the same target, and return the word with the largest Cosine similarity to it. Nearest by distance and most similar by cosine are the same idea, one written as a minimization and the other as a maximization. When the candidate vectors have unit length, the two rankings are identical, because ‖x − t‖² = 1 + ‖t‖² − 2 x·t.

b^∗=argmax⁡x∈V∖{a,a∗,b}cos⁡(x, a∗−a+b)\hat{b}^{*}=\operatorname*{argmax}_{x\in V\setminus\{a,a^{*},b\}} \cos(\mathbf{x},\ \mathbf{a}^{*}-\mathbf{a}+\mathbf{b})
The cosine form used by Mikolov, Yih and Zweig (2013), with all vectors normalized to unit length

Worked example

man : woman :: king : ? in two dimensions

  1. Place the words

    man (1, 1), woman (1, 3), king (4, 1), queen (4.2, 3.1), princess (3, 3.5).
  2. Build the target

    a = man, a* = woman, b = king, so t = (1, 3) − (1, 1) + (4, 1) = (4, 3). The offset (0, 2) is the "male to female" arrow.
  3. Measure every candidate

    Distances to t: queen 0.224, princess 1.118, king 2.0, woman 3.0, man 3.606.
  4. Result

    The nearest allowed word is queen. Note that it is near t, not on it: the method always ends in a nearest-neighbour search.
SimulatorParallelogram playground: a is to a* as b is to ?
Target t(4.00, 3.00)t = a* − a + b, the point the method searches around.
AnswerqueenNearest remaining word to t.
  1. 1. queen0.224
  2. 2. princess1.118
  3. 3. prince1.803
  4. 4. crown2.28
  5. 5. throne2.786

Try it above. Pick any three words, watch the a to a* arrow get copied onto b, and read the ranking. Switch the space to the small offset and the candidate pool to "allow input words" and keep the playground in mind for the third concept of this part.

Where the idea came from

The parallelogram is older than embeddings. Rumelhart and Abrahamson proposed it in 1973 as a model of how people solve analogies, working in a space of mammal names built from human similarity judgments (apple : tree :: grape : vine is the illustration SLP3 uses for it). Turney and Littman showed in 2005 that sparse count vectors could solve SAT-style analogies, and Mikolov and colleagues brought it to dense neural embeddings in 2013. Their NAACL paper, titled Linguistic Regularities in Continuous Space Word Representations, built a syntactic test set of 8,000 questions and found its recurrent network vectors answered almost 40% correctly. The slide labels it "Mikolov et al. 2013b" and SLP3's bibliography labels the same paper 2013c, so cite it by title.

Recall

Write the parallelogram method for a : a* :: b : b*, and state the two things you must change on slide 99's version.

b̂* = argmin over x ∉ {a, a*, b} of distance(x, a* − a + b). Change argmax to argmin (or switch to cosine and keep argmax), and exclude the three input words.

Quick check

For man : woman :: king : ?, which point does the parallelogram method search around?

Plot man, woman, uncle, aunt, king and queen from Mikolov, Yih and Zweig's recurrent-network language model (the precursor of Word2vec) in two dimensions and draw an arrow from each male word to its female partner. The three arrows come out roughly parallel and roughly the same length. In a second projection of the same space, king to kings and queen to queens are parallel to each other too, and that plural direction cuts across the gender direction.

This is what it means for a relation to be linear in an Embedding space: the offset vector between the two words of a pair is nearly the same for every pair that stands in that relation. Mikolov and colleagues found such offsets for gender (man to woman), verb tense (walking to walked) and country to capital (Spain to Madrid), and their larger 2013 test set organized analogy questions by exactly these semantic and syntactic families. The same picture holds for GloVe: the GloVe project page shows man to woman, sir to madam, heir to heiress, king to queen, uncle to aunt, nephew to niece, brother to sister, earl to countess, duke to duchess and emperor to empress as ten segments that all tilt the same way, and Pennington and colleagues report that offsets also capture comparative and superlative forms.

RelationExample pairWhat the offset meansSlides
Genderman → womanMale form to female form100 to 102
Numberking → kingsSingular to plural100
Tensewalking → walkedProgressive to past101
CapitalSpain → MadridCountry to its capital city101
Titleearl → countessMale noble title to female counterpart102
Relation families that show up as consistent offsets

One word can take part in many relations at once. King is the male member of a gender pair, the singular of a number pair, and a royal term next to throne and crown. A 300-dimensional space has room for all of these as different directions, and that is the sentence Mikolov and colleagues put under their figure: in high-dimensional space, multiple relations can be embedded for a single word. Any 2D picture is one projection chosen to show one of those directions, which is why the slides need two panels to show gender and number for the same words.

Recall

How can king lie on a gender direction and a number direction at the same time?

They are different directions in a high-dimensional space; any 2D plot is one projection chosen to show one of them.

Change the toy space from the first concept so the gender offset is small: man (1, 1), woman (1.4, 1.3), king (4, 1), queen (4.6, 1.9). Now ask man : woman :: king : ?and let every word compete.

Worked example

The exclusion trap

  1. Build the target

    t = (1.4, 1.3) − (1, 1) + (4, 1) = (4.4, 1.3).
  2. Rank with input words allowed

    Distance to king is 0.5, distance to queen is 0.632. The method answers king, the word you gave it.
  3. Exclude a, a* and b

    Remove man, woman and king from the pool. The nearest remaining word is queen at 0.632.
  4. Result

    The right answer appears only after exclusion. With a small offset the target barely moves away from b, so b itself is the nearest point.

This is not a toy artefact. Linzen (2016) ran the standard Word2vec analogy benchmark without excluding the inputs: the nearest neighbour of a* − a + b was b in 93% of cases, a* in 5%, and never a. SLP3 makes the same point with cherry : red :: potato : x, which returns potato or potatoes instead of brown unless those are forbidden. Every published analogy accuracy therefore depends on the exclusion rule, and part of the credit belongs to b*'s simply being b's nearest neighbour. Linzen's baselines make this concrete: a method that ignores a, or even both a and a*, and just returns the neighbour of b, scores very high on plurals.

Where the method works and where it does not

  • It works for frequent words, for pairs where b* already sits close to b (SLP3's "small distances", which is also why b wins unless it is excluded), and for certain relations (SLP3): country to capital, and inflections such as plural and tense.
  • It does poorly on many lexicographic and derivational relations. The BATS set of Gladkova and colleagues has 99,200 questions in 40 categories, against only 15 relations in the Google set, and accuracy varies widely across them.
  • Reversing an analogy uses the same offset with the sign flipped, yet Linzen found accuracy dropped in most categories (mean −0.11): US cities fell from .69 to .17 and common capitals from .9 to .53.
  • As a model of human analogy making, the parallelogram is too simple: Peterson, Chen and Griffiths (2020) show it cannot account for how people form even simple analogies.

3CosAdd versus 3CosMul

Levy and Goldberg (2014) rewrote the cosine objective with unit vectors and saw it as a balance: two attractors (b* should resemble b and a*) and one repeller (b* should not resemble a). Added together, one large similarity can swamp the others. Their multiplicative version keeps each term in check and generally does better.

3CosAdd: argmax⁡b∗∈V(cos⁡(b∗,b)−cos⁡(b∗,a)+cos⁡(b∗,a∗))\begin{aligned} \text{3CosAdd: }\operatorname*{argmax}_{b^{*}\in V}\big(&\cos(b^{*},b) \\ &-\cos(b^{*},a) \\ &+\cos(b^{*},a^{*})\big) \end{aligned}
The additive objective, equivalent to the cosine parallelogram method for unit vectors
3CosMul: argmax⁡b∗∈Vcos⁡(b∗,b) cos⁡(b∗,a∗)cos⁡(b∗,a)+ε,ε=0.001\begin{gathered} \text{3CosMul: }\operatorname*{argmax}_{b^{*}\in V}\frac{\cos(b^{*},b)\,\cos(b^{*},a^{*})}{\cos(b^{*},a)+\varepsilon}, \\ \varepsilon=0.001 \end{gathered}
Levy and Goldberg's multiplicative objective, in the slide's a : a* :: b : b* order; for dense embeddings each cosine is first mapped to [0, 1] by (cos + 1)/2

Recall

What does the offset method return most often if a, a* and b are allowed as answers?

b itself, 93% of the time in Linzen (2016); a* 5%, never a. A small offset leaves the target closest to b.

Quick check

If the input words stay in the candidate pool, what does the offset method usually return?

Follow three words through two centuries of books. In the 1900s gay sits near daft, flaunting, sweet and cheerful; by the 1990s its neighbours are homosexual and lesbian. In the 1850s broadcast sits near sow and seed, a farmer scattering grain; by the 1990s it sits near newspapers, radio and bbc. In the 1850s awful sits near majestic, awe and solemn, full of awe; by the 1900s it sits near terrible and appalling, and by the 1990s near weird and wonderful. That slide from praise to blame is called pejoration.

awful drifts from majestic and awe (1850s) through terrible and appalling (1900s) to weird and wonderful (1990s).

Hamilton, Leskovec and Jurafsky (2016) produced these pictures with diachronic embeddings. The recipe has three steps. First, train a separate embedding for each decade of text. They compared Positive PMI, SVD and Skip-gram with negative sampling on six historical corpora in four languages, with a window of 4 and 300 dimensions; the English Google Books corpus alone has 8.5 × 1011 tokens covering 1800 to 1999, and COHA has 4.1 × 108 tokens covering 1810 to 2009. Second, align the decades so their axes mean the same thing. Third, measure how far each word moved between aligned decades, and read its old and new neighbours.

Why alignment is needed

SVD and SGNS only care about dot products between vectors, and any rotation of the whole space preserves every dot product. So the 1900 run and the 1990 run can come out rotated relative to each other for no linguistic reason, and comparing the raw coordinates of gay in the two runs measures that arbitrary rotation. Hamilton and colleagues fix this with orthogonal Procrustes: find the orthogonal matrix that best maps one decade's matrix onto the next. Because the matrix is orthogonal it is a rotation (possibly with a reflection), so cosines within each decade are unchanged. PPMI vectors need no alignment, since their dimensions are context words that mean the same thing in every decade.

R(t)=argmin⁡Q⊤Q=I∥W(t)Q−W(t+1)∥FR^{(t)}=\operatorname*{argmin}_{Q^{\top}Q=I}\lVert W^{(t)}Q-W^{(t+1)}\rVert_F
Orthogonal Procrustes alignment between decade t and decade t + 1 (Hamilton et al. 2016, eq. 4)

Two statistical laws

Measuring displacement for thousands of words let Hamilton and colleagues state two laws. The law of conformity: the rate of semantic change scales with an inverse power of word frequency, so frequent words change slowly. The law of innovation: holding frequency fixed, words with more senses (higher Polysemy) change faster.

Recall

Why must decade-specific embeddings be aligned before you measure semantic change, and how?

SVD and SGNS spaces can be arbitrarily rotated, so coordinates are not comparable across decades. Orthogonal Procrustes finds the rotation that best maps one decade onto the next while keeping within-decade cosines unchanged.

Recall

State the law of conformity and the law of innovation.

Frequent words change meaning more slowly (the rate scales with an inverse power of frequency). Controlling for frequency, more polysemous words change faster.

Embeddings inherit and amplify cultural bias

Bolukbasi and colleagues ran the Parallelogram method on Word2vec trained on Google News (3 million words and phrases, 300 dimensions). Asked Paris : France :: Tokyo : x, it answers Japan. Asked father : doctor :: mother : x, it answers nurse. Asked man : computer programmer :: woman : x, it answers homemaker. The same machinery that captured capitals captured stereotypes.

The reason is unsurprising once said aloud. An Embedding is a compressed summary of co-occurrence statistics, so if the training text talks about women and men in different contexts, that difference becomes geometry. Bolukbasi and colleagues found that gender bias is largely captured by a single direction, roughly the she minus he offset, onto which occupation words project unevenly.

Occupations drop onto a she to he axis and land at uneven offsets; neutralizing pulls them to the midpoint, yet a small residue stays split.

Two kinds of harm

HarmDefinitionExample
AllocationalA system distributes a resource or opportunity (jobs, loans, search exposure) unfairly across groups.A resume search that ranks documents by embedding similarity to programmer pushes women's resumes down.
RepresentationalA system demeans, stereotypes or erases a group, whether or not any resource is at stake.Caliskan et al. find African American names closer to unpleasant words than European American names.
Harm types (SLP3 section 5.8)

Embeddings do not just mirror the bias in their text; they can amplify it, exaggerating an association beyond its strength in the corpus or in the world (Zhao et al. 2017, Ethayarajh et al. 2019, Jia et al. 2020, as summarized in SLP3). Caliskan, Bryson and Narayanan (2017) built the Word Embedding Association Test (WEAT) and reproduced classic Implicit Association Test results with GloVe, including the finding that African American names sit closer to unpleasant words. Debiasing methods such as Bolukbasi's neutralize and equalize steps remove the component of gender-neutral words along the gender direction. They reduce measured bias, but Gonen and Goldberg (2019) show that the stereotyped words still cluster together afterwards: the bias is hidden, not removed.

Measuring a century of stereotypes

Garg, Schiebinger, Jurafsky and Zou (2018) turned diachronic embeddings into a tool for social history. Using the decade embeddings from Hamilton and colleagues, they built a group vector for women (the average of words like she, her, woman) and one for men, and scored each neutral word (an adjective or an occupation) by its relative norm difference: its average distance to the men vector minus its average distance to the women vector. A negative score means the word sits closer to men.

bias(w)=∥w−vmen∥−∥w−vwomen∥\text{bias}(w)=\lVert \mathbf{w}-\mathbf{v}_{\text{men}}\rVert-\lVert \mathbf{w}-\mathbf{v}_{\text{women}}\rVert
Relative norm difference for one neutral word, averaged over a word list in Garg et al. (positive means closer to women, negative closer to men)

Worked example

Relative norm difference in 2D

  1. Place the vectors

    Women group vector (0, 2), men group vector (2, 0), adjective smart (1.6, 0.6).
  2. Measure both distances

    ‖smart − men‖ = √(0.4² + 0.6²) = 0.721 and ‖smart − women‖ = √(1.6² + 1.4²) = 2.126.
  3. Result

    0.721 − 2.126 = −1.405. A negative value means smart is closer to the men vector.

Run over each decade, the measure tells a story. Competence and intelligence adjectives (smart, wise, thoughtful, logical) were biased toward men, and that bias has been decreasing since the 1960s. Words used to describe outsiders (barbaric, monstrous, hateful, bizarre) were most associated with Asian last names before 1950 and declined steadily afterwards. The embeddings reproduce a 1933 survey of ethnic stereotypes. Occupation bias in the Google News embeddings tracks the 2015 share of women in each occupation (r² = .46), and the decade-by-decade trend in the historical embeddings follows US Census data. Embeddings, in other words, are a usable instrument for measuring culture.

QueryConstrained answerUnconstrained answer
man : doctor :: woman : xgynecologistdoctor
man : computer programmer :: woman : xhomemakercomputer programmer
3CosAdd with and without the input-exclusion constraint (Nissim et al. 2020, Table 1)

Recall

Give one allocational harm and one representational harm caused by embeddings.

Allocational: a hiring search built on embeddings ranks documents with women's names lower. Representational: African American names sit closer to unpleasant words (Caliskan et al. 2017).

Recall

Compute Garg's relative norm difference for w = (1, 1), women = (0, 2), men = (2, 0), and interpret it.

√2 − √2 = 0: equally close to both groups, no gender lean.

Before testing an embedding, look at it. Slide 107 shows a hierarchical clustering of noun embeddings (from Rohde et al., reproduced in SLP3). Read it bottom-up. Body parts merge early: wrist with ankle, then shoulder, arm and leg. Animals form their own branch: dog with cat, then puppy and kitten. Places split off near the top. And look closely at the places: Tokyo sits with Chicago and the other US cities, while Moscow and Hawaii sit with countries and continents.

Leaves join bottom-up in order of merge height: wrist and ankle, dog and cat, Chicago and Tokyo, China and Russia, then the branches.

The height of each join is the dissimilarity at the moment the two clusters merged, so low joins mean close vectors. The odd placements are a lesson in what an embedding encodes: the tree reflects the contexts words appear in, not a geography ontology. Other ways to look include the nearest-neighbour lists you saw for awful, and 2D projections such as t-SNE (van der Maaten and Hinton 2008), which keep local neighbourhoods but distort global distances.

Intrinsic evaluation: test the vectors directly

Intrinsic evaluation scores the vectors on a small task designed to probe them, without building a full system. The most common form compares model similarity with human judgments.

  • WordSim-353 (Finkelstein et al. 2002) asks people to rate 353 word pairs from 0 (totally unrelated) to 10 (very much related or identical); plane and car get 5.77. Because the instruction is about relatedness, cup and coffee can score as high as cup and mug. That is Word relatedness, not Word similarity.
  • SimLex-999 (Hill, Reichart and Korhonen 2015) was built to fix that: 999 adjective, noun and verb pairs rated for genuine similarity, so cup and mug score high and cup and coffee score low.
  • The TOEFL synonym test has 80 questions with 4 choices each: which word is closest to levied? The model picks the choice with the highest cosine. Latent semantic analysis scored 64.4% (Landauer and Dumais 1997), almost exactly the 64.5% average of non-native college applicants in the US.
  • Analogy sets, with all the caveats of the earlier concept.

For a similarity dataset the score is the Spearman rank correlation between the model's cosines and the human ratings. Spearman compares orders, not values, so it does not matter that cosines live in [−1, 1] and ratings in [0, 10].

ρ=1−6∑idi2n(n2−1)\rho = 1-\frac{6\sum_i d_i^{2}}{n(n^{2}-1)}
Spearman's rank correlation without ties, where d_i is the difference between the two ranks of pair i

Worked example

Spearman correlation on four WordSim pairs

  1. Rank both columns

    Human ratings are the real WordSim-353 values; the model cosines are illustrative.
    PairHuman ratingModel cosineHuman rankModel rankd
    drink, ear1.310.08110
    plane, car5.770.42220
    drink, eat6.870.61341
    planet, star8.450.55431
  2. Sum the squared rank differences

    Σd² = 0 + 0 + 1 + 1 = 2, with n = 4.
  3. Result

    ρ = 1 − (6 · 2) / (4 · 15) = 1 − 0.2 = 0.8. The model orders the pairs almost like people do; it only swaps drink-eat and planet-star.

Worked example

Answering a TOEFL item

  1. The question

    Which word is closest in meaning to levied: imposed, believed, requested or correlated?
  2. Score each choice

    Compute cos(levied, choice) for all four choices.
  3. Result

    Return the argmax. A good space puts imposed first, because both words appear around taxes, fines and duties.

Extrinsic evaluation: test inside a real task

Extrinsic evaluation plugs each candidate embedding into an actual system (named entity recognition, machine translation, coreference), trains it, and compares the task metric. SLP3 calls this the most important evaluation for vector models, because the task is what you care about. It is also slow and noisy, which is why intrinsic tests remain popular for quick comparisons.

EvaluationTypeDataScore
WordSim-353Intrinsic353 pairs, 0 to 10, relatednessSpearman ρ between cosine and human ratings
SimLex-999Intrinsic999 pairs, similarity onlySpearman ρ; penalizes cup and coffee scoring like cup and mug
TOEFL synonymsIntrinsic80 items, 4 choices eachAccuracy of argmax cosine over the choices
Analogy sets (Google, BATS)Intrinsica : a* :: b : ? questionsAccuracy of the parallelogram method
NER, MT, coreferenceExtrinsicA full task with its own labelled dataTask metric (F1, BLEU) with each embedding plugged in
Common evaluations and what they measure

Recall

Intrinsic or extrinsic: (a) Spearman ρ with SimLex-999, (b) NER F1 with GloVe versus word2vec inputs, (c) TOEFL synonym accuracy.

(a) intrinsic, (b) extrinsic, (c) intrinsic.

Quick check

A team reports Spearman correlation with SimLex-999 ratings for new embeddings. What evaluation is this?

Feed "I saw a cat." into a neural network. Each token passes through an Embedding layer, a lookup table from word to vector, and the network sits on top. Where should that table come from? There are three answers: copy it from Word2vec or GloVe and freeze it; copy it and keep updating it on your task; or start it from random numbers and learn it together with the network.

The slides give the rule. When there is not enough labelled data, or the task is simple, use embeddings pretrained on another task. Pretrained vectors bring knowledge distilled by Self-supervision from billions of unlabelled tokens, which your few thousand labelled examples could never teach. When there is enough data and the task is hard, such as language modelling or machine translation, train the embeddings with the model, because the task itself supplies the signal and task-specific vectors fit it better. Fine-tuning pretrained vectors is the middle path.

OptionWhere vectors come fromUpdated during training?When to useExample
Frozen pretrainedword2vec or GloVe trained on another corpusNoLittle labelled data, simple taskKim's CNN-static
Fine-tuned pretrainedCopied from word2vec or GloVeYes, starting from the pretrained valuesModerate data, want task-specific nuanceKim's CNN-non-static
Joint from randomLearned from scratch with the networkYes, from random valuesLarge data and a hard task (LM, MT)Kim's CNN-rand; large LMs and MT systems
Three ways to obtain the embedding layer

Two studies give the evidence. Kim (2014) trained a CNN sentence classifier on small benchmarks three ways. CNN-rand, with random embeddings learned jointly, did poorly. CNN-static, with frozen word2vec vectors, "performs remarkably well", and CNN-non-static, which fine-tunes them, improves further; Kim concludes that pretrained vectors are good, universal feature extractors. Qi and colleagues (2018) asked the same question for neural machine translation and found gains of up to 20 BLEU in the most favourable setting. The gains have a sweet spot: they are largest when training data is scarce, but not so scarce that the system cannot be trained at all (baseline BLEU around 3 to 4), and they shrink as parallel data grows.

Everything in this lecture has been a Static embedding: one fixed vector per word type, whatever the sentence. Later lectures replace it with a Contextual embedding, where the vector for bank depends on its sentence. In those language models the embedding layer is trained jointly with the network, exactly the right-hand side of slide 110.

Recall

You have 500k parallel sentences for a hard MT task. Pretrained or joint, and why?

Joint training, possibly initialized from pretrained vectors. Enough data and a hard task let the model learn task-specific embeddings, and Qi et al. found pretrained gains largest in low-resource settings.

Quick check

You build a dialect sentiment classifier from 2,000 labeled tweets. How should its embeddings start?

Recap

If you remember nothing else

  • Parallelogram method: b̂* = argmin over x of distance(x, a* − a + b), or argmax of cosine, with a, a* and b excluded. Slide 99's argmax of distance is a typo.
  • Relations such as gender, number, tense and country to capital appear as roughly constant offsets; one word can sit on several such directions at once.
  • Without exclusion the method returns b 93% of the time. It works best for frequent words, answers that already sit near b, and inflectional or capital relations; 3CosMul improves on 3CosAdd.
  • Decade-specific embeddings, aligned by orthogonal Procrustes, show gay, broadcast and awful shifting. Frequent words change slowly; polysemous words change fast.
  • Embeddings reproduce and amplify stereotypes (father : doctor :: mother : nurse), causing allocational and representational harms. Debiasing hides more than it removes.
  • Garg et al. track bias over a century with relative norm differences: competence adjectives lean male (weakening since the 1960s), outsider words were tied to Asian names before 1950, and occupation bias tracks census data.
  • Visualize with dendrograms, nearest-neighbour lists or t-SNE; all of them show usage, not an ontology.
  • Intrinsic evaluation: WordSim-353, SimLex-999, TOEFL, analogies. Extrinsic evaluation is a downstream task and the more important one.
  • Use pretrained vectors with little data or a simple task; train embeddings jointly with a large dataset and a hard task; fine-tuning sits between.

Sources