ICS 582Lecture 04Part 10
Analogies, bias and evaluating embeddings
The parallelogram method for analogies and its limits, embeddings as a lens on historical meaning change and cultural bias, visualizing and evaluating embeddings intrinsically and extrinsically, and when to use pretrained embeddings.
- Concepts
- 7
- Slides
- 98-111
- Reading
- 42 min
Why this part matters
You have trained an embedding. Four questions follow immediately, and this part answers each. What has the space learned? Relations such as male to female or country to capital show up as directions you can probe with one subtraction and one addition. Can you trust it? Analogy scores flatter it, and it carries the biases of its corpus, which matters for any hiring, search or Arabic NLP system built on top of it. How do you measure it? Intrinsic tests against human judgments, or extrinsic tests inside a real task. And should you train it yourself at all, or take vectors someone else trained?
For exams, the parallelogram formula, the intrinsic versus extrinsic split and allocational versus representational harm are standard items. For research, diachronic embeddings and bias measurement are live tools: the same machinery that tracks how awful changed meaning also measures a century of gender stereotypes.
By the end you can
- Compute an analogy answer with the parallelogram method, using argmin distance or argmax cosine, and exclude the input words.
- Explain why analogy accuracy overstates relational knowledge, citing the exclusion effect and relation-dependence.
- Describe how aligned decade embeddings reveal semantic change, and state the laws of conformity and innovation.
- Explain how embeddings encode and amplify cultural bias, distinguish allocational from representational harm, and describe Garg's relative norm measure.
- Classify evaluations as intrinsic or extrinsic, and compute a Spearman correlation against human ratings.
- Choose between frozen, fine-tuned and jointly trained embeddings given data size and task difficulty.
Apple is to tree as grape is to what? Picture the arrow that starts at apple and ends at tree. It means something like "fruit to the plant it grows on". Now pick that same arrow up, keep its length and direction, and set its tail down on grape. Its head lands near vine. That is the whole method.
The same move works on the famous examples. Take king, subtract man, add woman, and the point you reach is close to queen. Take Paris, subtract France, add Italy, and you land close to Rome. In each case the subtraction isolates a relation (royalty without the maleness, the capital-of relation without the particular country) and the addition applies it to a new word. Because the four points form a parallelogram when it works, this is called the Parallelogram method.
The rule, in two equivalent forms
The slides write an analogy as a : a* :: b : b*, read "a is to a* as b is to b*". The relation is the offset from a to a*, so the point to search around is t = a* − a + b. No vocabulary word sits exactly at t, so the answer is the word nearest to it, and the three question words themselves are taken out of the candidate pool (the next concepts show why that matters so much). With Euclidean distance, nearest means smallest distance, so the operator is an argmin:
Mikolov, Yih and Zweig, who made the method famous for dense vectors, wrote it the other way round. They normalize every vector to unit length, compute the same target, and return the word with the largest Cosine similarity to it. Nearest by distance and most similar by cosine are the same idea, one written as a minimization and the other as a maximization. When the candidate vectors have unit length, the two rankings are identical, because ‖x − t‖² = 1 + ‖t‖² − 2 x·t.
Worked example
man : woman :: king : ? in two dimensions
Place the words
man (1, 1), woman (1, 3), king (4, 1), queen (4.2, 3.1), princess (3, 3.5).Build the target
a = man, a* = woman, b = king, so t = (1, 3) − (1, 1) + (4, 1) = (4, 3). The offset (0, 2) is the "male to female" arrow.Measure every candidate
Distances to t: queen 0.224, princess 1.118, king 2.0, woman 3.0, man 3.606.Result
The nearest allowed word is queen. Note that it is near t, not on it: the method always ends in a nearest-neighbour search.
- 1. queen0.224
- 2. princess1.118
- 3. prince1.803
- 4. crown2.28
- 5. throne2.786
Try it above. Pick any three words, watch the a to a* arrow get copied onto b, and read the ranking. Switch the space to the small offset and the candidate pool to "allow input words" and keep the playground in mind for the third concept of this part.
Where the idea came from
The parallelogram is older than embeddings. Rumelhart and Abrahamson proposed it in 1973 as a model of how people solve analogies, working in a space of mammal names built from human similarity judgments (apple : tree :: grape : vine is the illustration SLP3 uses for it). Turney and Littman showed in 2005 that sparse count vectors could solve SAT-style analogies, and Mikolov and colleagues brought it to dense neural embeddings in 2013. Their NAACL paper, titled Linguistic Regularities in Continuous Space Word Representations, built a syntactic test set of 8,000 questions and found its recurrent network vectors answered almost 40% correctly. The slide labels it "Mikolov et al. 2013b" and SLP3's bibliography labels the same paper 2013c, so cite it by title.
Recall
Write the parallelogram method for a : a* :: b : b*, and state the two things you must change on slide 99's version.
Quick check
For man : woman :: king : ?, which point does the parallelogram method search around?
Plot man, woman, uncle, aunt, king and queen from Mikolov, Yih and Zweig's recurrent-network language model (the precursor of Word2vec) in two dimensions and draw an arrow from each male word to its female partner. The three arrows come out roughly parallel and roughly the same length. In a second projection of the same space, king to kings and queen to queens are parallel to each other too, and that plural direction cuts across the gender direction.
This is what it means for a relation to be linear in an Embedding space: the offset vector between the two words of a pair is nearly the same for every pair that stands in that relation. Mikolov and colleagues found such offsets for gender (man to woman), verb tense (walking to walked) and country to capital (Spain to Madrid), and their larger 2013 test set organized analogy questions by exactly these semantic and syntactic families. The same picture holds for GloVe: the GloVe project page shows man to woman, sir to madam, heir to heiress, king to queen, uncle to aunt, nephew to niece, brother to sister, earl to countess, duke to duchess and emperor to empress as ten segments that all tilt the same way, and Pennington and colleagues report that offsets also capture comparative and superlative forms.
| Relation | Example pair | What the offset means | Slides |
|---|---|---|---|
| Gender | man → woman | Male form to female form | 100 to 102 |
| Number | king → kings | Singular to plural | 100 |
| Tense | walking → walked | Progressive to past | 101 |
| Capital | Spain → Madrid | Country to its capital city | 101 |
| Title | earl → countess | Male noble title to female counterpart | 102 |
One word can take part in many relations at once. King is the male member of a gender pair, the singular of a number pair, and a royal term next to throne and crown. A 300-dimensional space has room for all of these as different directions, and that is the sentence Mikolov and colleagues put under their figure: in high-dimensional space, multiple relations can be embedded for a single word. Any 2D picture is one projection chosen to show one of those directions, which is why the slides need two panels to show gender and number for the same words.
Recall
How can king lie on a gender direction and a number direction at the same time?
Change the toy space from the first concept so the gender offset is small: man (1, 1), woman (1.4, 1.3), king (4, 1), queen (4.6, 1.9). Now ask man : woman :: king : ?and let every word compete.
Worked example
The exclusion trap
Build the target
t = (1.4, 1.3) − (1, 1) + (4, 1) = (4.4, 1.3).Rank with input words allowed
Distance to king is 0.5, distance to queen is 0.632. The method answers king, the word you gave it.Exclude a, a* and b
Remove man, woman and king from the pool. The nearest remaining word is queen at 0.632.Result
The right answer appears only after exclusion. With a small offset the target barely moves away from b, so b itself is the nearest point.
This is not a toy artefact. Linzen (2016) ran the standard Word2vec analogy benchmark without excluding the inputs: the nearest neighbour of a* − a + b was b in 93% of cases, a* in 5%, and never a. SLP3 makes the same point with cherry : red :: potato : x, which returns potato or potatoes instead of brown unless those are forbidden. Every published analogy accuracy therefore depends on the exclusion rule, and part of the credit belongs to b*'s simply being b's nearest neighbour. Linzen's baselines make this concrete: a method that ignores a, or even both a and a*, and just returns the neighbour of b, scores very high on plurals.
Where the method works and where it does not
- It works for frequent words, for pairs where b* already sits close to b (SLP3's "small distances", which is also why b wins unless it is excluded), and for certain relations (SLP3): country to capital, and inflections such as plural and tense.
- It does poorly on many lexicographic and derivational relations. The BATS set of Gladkova and colleagues has 99,200 questions in 40 categories, against only 15 relations in the Google set, and accuracy varies widely across them.
- Reversing an analogy uses the same offset with the sign flipped, yet Linzen found accuracy dropped in most categories (mean −0.11): US cities fell from .69 to .17 and common capitals from .9 to .53.
- As a model of human analogy making, the parallelogram is too simple: Peterson, Chen and Griffiths (2020) show it cannot account for how people form even simple analogies.
3CosAdd versus 3CosMul
Levy and Goldberg (2014) rewrote the cosine objective with unit vectors and saw it as a balance: two attractors (b* should resemble b and a*) and one repeller (b* should not resemble a). Added together, one large similarity can swamp the others. Their multiplicative version keeps each term in check and generally does better.
Recall
What does the offset method return most often if a, a* and b are allowed as answers?
Quick check
If the input words stay in the candidate pool, what does the offset method usually return?
Follow three words through two centuries of books. In the 1900s gay sits near daft, flaunting, sweet and cheerful; by the 1990s its neighbours are homosexual and lesbian. In the 1850s broadcast sits near sow and seed, a farmer scattering grain; by the 1990s it sits near newspapers, radio and bbc. In the 1850s awful sits near majestic, awe and solemn, full of awe; by the 1900s it sits near terrible and appalling, and by the 1990s near weird and wonderful. That slide from praise to blame is called pejoration.
Hamilton, Leskovec and Jurafsky (2016) produced these pictures with diachronic embeddings. The recipe has three steps. First, train a separate embedding for each decade of text. They compared Positive PMI, SVD and Skip-gram with negative sampling on six historical corpora in four languages, with a window of 4 and 300 dimensions; the English Google Books corpus alone has 8.5 × 1011 tokens covering 1800 to 1999, and COHA has 4.1 × 108 tokens covering 1810 to 2009. Second, align the decades so their axes mean the same thing. Third, measure how far each word moved between aligned decades, and read its old and new neighbours.
Why alignment is needed
SVD and SGNS only care about dot products between vectors, and any rotation of the whole space preserves every dot product. So the 1900 run and the 1990 run can come out rotated relative to each other for no linguistic reason, and comparing the raw coordinates of gay in the two runs measures that arbitrary rotation. Hamilton and colleagues fix this with orthogonal Procrustes: find the orthogonal matrix that best maps one decade's matrix onto the next. Because the matrix is orthogonal it is a rotation (possibly with a reflection), so cosines within each decade are unchanged. PPMI vectors need no alignment, since their dimensions are context words that mean the same thing in every decade.
Two statistical laws
Measuring displacement for thousands of words let Hamilton and colleagues state two laws. The law of conformity: the rate of semantic change scales with an inverse power of word frequency, so frequent words change slowly. The law of innovation: holding frequency fixed, words with more senses (higher Polysemy) change faster.
Recall
Why must decade-specific embeddings be aligned before you measure semantic change, and how?
Recall
State the law of conformity and the law of innovation.
Bolukbasi and colleagues ran the Parallelogram method on Word2vec trained on Google News (3 million words and phrases, 300 dimensions). Asked Paris : France :: Tokyo : x, it answers Japan. Asked father : doctor :: mother : x, it answers nurse. Asked man : computer programmer :: woman : x, it answers homemaker. The same machinery that captured capitals captured stereotypes.
The reason is unsurprising once said aloud. An Embedding is a compressed summary of co-occurrence statistics, so if the training text talks about women and men in different contexts, that difference becomes geometry. Bolukbasi and colleagues found that gender bias is largely captured by a single direction, roughly the she minus he offset, onto which occupation words project unevenly.
Two kinds of harm
| Harm | Definition | Example |
|---|---|---|
| Allocational | A system distributes a resource or opportunity (jobs, loans, search exposure) unfairly across groups. | A resume search that ranks documents by embedding similarity to programmer pushes women's resumes down. |
| Representational | A system demeans, stereotypes or erases a group, whether or not any resource is at stake. | Caliskan et al. find African American names closer to unpleasant words than European American names. |
Embeddings do not just mirror the bias in their text; they can amplify it, exaggerating an association beyond its strength in the corpus or in the world (Zhao et al. 2017, Ethayarajh et al. 2019, Jia et al. 2020, as summarized in SLP3). Caliskan, Bryson and Narayanan (2017) built the Word Embedding Association Test (WEAT) and reproduced classic Implicit Association Test results with GloVe, including the finding that African American names sit closer to unpleasant words. Debiasing methods such as Bolukbasi's neutralize and equalize steps remove the component of gender-neutral words along the gender direction. They reduce measured bias, but Gonen and Goldberg (2019) show that the stereotyped words still cluster together afterwards: the bias is hidden, not removed.
Measuring a century of stereotypes
Garg, Schiebinger, Jurafsky and Zou (2018) turned diachronic embeddings into a tool for social history. Using the decade embeddings from Hamilton and colleagues, they built a group vector for women (the average of words like she, her, woman) and one for men, and scored each neutral word (an adjective or an occupation) by its relative norm difference: its average distance to the men vector minus its average distance to the women vector. A negative score means the word sits closer to men.
Worked example
Relative norm difference in 2D
Place the vectors
Women group vector (0, 2), men group vector (2, 0), adjective smart (1.6, 0.6).Measure both distances
‖smart − men‖ = √(0.4² + 0.6²) = 0.721 and ‖smart − women‖ = √(1.6² + 1.4²) = 2.126.Result
0.721 − 2.126 = −1.405. A negative value means smart is closer to the men vector.
Run over each decade, the measure tells a story. Competence and intelligence adjectives (smart, wise, thoughtful, logical) were biased toward men, and that bias has been decreasing since the 1960s. Words used to describe outsiders (barbaric, monstrous, hateful, bizarre) were most associated with Asian last names before 1950 and declined steadily afterwards. The embeddings reproduce a 1933 survey of ethnic stereotypes. Occupation bias in the Google News embeddings tracks the 2015 share of women in each occupation (r² = .46), and the decade-by-decade trend in the historical embeddings follows US Census data. Embeddings, in other words, are a usable instrument for measuring culture.
| Query | Constrained answer | Unconstrained answer |
|---|---|---|
| man : doctor :: woman : x | gynecologist | doctor |
| man : computer programmer :: woman : x | homemaker | computer programmer |
Recall
Give one allocational harm and one representational harm caused by embeddings.
Recall
Compute Garg's relative norm difference for w = (1, 1), women = (0, 2), men = (2, 0), and interpret it.
Before testing an embedding, look at it. Slide 107 shows a hierarchical clustering of noun embeddings (from Rohde et al., reproduced in SLP3). Read it bottom-up. Body parts merge early: wrist with ankle, then shoulder, arm and leg. Animals form their own branch: dog with cat, then puppy and kitten. Places split off near the top. And look closely at the places: Tokyo sits with Chicago and the other US cities, while Moscow and Hawaii sit with countries and continents.
The height of each join is the dissimilarity at the moment the two clusters merged, so low joins mean close vectors. The odd placements are a lesson in what an embedding encodes: the tree reflects the contexts words appear in, not a geography ontology. Other ways to look include the nearest-neighbour lists you saw for awful, and 2D projections such as t-SNE (van der Maaten and Hinton 2008), which keep local neighbourhoods but distort global distances.
Intrinsic evaluation: test the vectors directly
Intrinsic evaluation scores the vectors on a small task designed to probe them, without building a full system. The most common form compares model similarity with human judgments.
- WordSim-353 (Finkelstein et al. 2002) asks people to rate 353 word pairs from 0 (totally unrelated) to 10 (very much related or identical); plane and car get 5.77. Because the instruction is about relatedness, cup and coffee can score as high as cup and mug. That is Word relatedness, not Word similarity.
- SimLex-999 (Hill, Reichart and Korhonen 2015) was built to fix that: 999 adjective, noun and verb pairs rated for genuine similarity, so cup and mug score high and cup and coffee score low.
- The TOEFL synonym test has 80 questions with 4 choices each: which word is closest to levied? The model picks the choice with the highest cosine. Latent semantic analysis scored 64.4% (Landauer and Dumais 1997), almost exactly the 64.5% average of non-native college applicants in the US.
- Analogy sets, with all the caveats of the earlier concept.
For a similarity dataset the score is the Spearman rank correlation between the model's cosines and the human ratings. Spearman compares orders, not values, so it does not matter that cosines live in [−1, 1] and ratings in [0, 10].
Worked example
Spearman correlation on four WordSim pairs
Rank both columns
Human ratings are the real WordSim-353 values; the model cosines are illustrative.Pair Human rating Model cosine Human rank Model rank d drink, ear 1.31 0.08 1 1 0 plane, car 5.77 0.42 2 2 0 drink, eat 6.87 0.61 3 4 1 planet, star 8.45 0.55 4 3 1 Sum the squared rank differences
Σd² = 0 + 0 + 1 + 1 = 2, with n = 4.Result
ρ = 1 − (6 · 2) / (4 · 15) = 1 − 0.2 = 0.8. The model orders the pairs almost like people do; it only swaps drink-eat and planet-star.
Worked example
Answering a TOEFL item
The question
Which word is closest in meaning to levied: imposed, believed, requested or correlated?Score each choice
Compute cos(levied, choice) for all four choices.Result
Return the argmax. A good space puts imposed first, because both words appear around taxes, fines and duties.
Extrinsic evaluation: test inside a real task
Extrinsic evaluation plugs each candidate embedding into an actual system (named entity recognition, machine translation, coreference), trains it, and compares the task metric. SLP3 calls this the most important evaluation for vector models, because the task is what you care about. It is also slow and noisy, which is why intrinsic tests remain popular for quick comparisons.
| Evaluation | Type | Data | Score |
|---|---|---|---|
| WordSim-353 | Intrinsic | 353 pairs, 0 to 10, relatedness | Spearman ρ between cosine and human ratings |
| SimLex-999 | Intrinsic | 999 pairs, similarity only | Spearman ρ; penalizes cup and coffee scoring like cup and mug |
| TOEFL synonyms | Intrinsic | 80 items, 4 choices each | Accuracy of argmax cosine over the choices |
| Analogy sets (Google, BATS) | Intrinsic | a : a* :: b : ? questions | Accuracy of the parallelogram method |
| NER, MT, coreference | Extrinsic | A full task with its own labelled data | Task metric (F1, BLEU) with each embedding plugged in |
Recall
Intrinsic or extrinsic: (a) Spearman ρ with SimLex-999, (b) NER F1 with GloVe versus word2vec inputs, (c) TOEFL synonym accuracy.
Quick check
A team reports Spearman correlation with SimLex-999 ratings for new embeddings. What evaluation is this?
Feed "I saw a cat." into a neural network. Each token passes through an Embedding layer, a lookup table from word to vector, and the network sits on top. Where should that table come from? There are three answers: copy it from Word2vec or GloVe and freeze it; copy it and keep updating it on your task; or start it from random numbers and learn it together with the network.
The slides give the rule. When there is not enough labelled data, or the task is simple, use embeddings pretrained on another task. Pretrained vectors bring knowledge distilled by Self-supervision from billions of unlabelled tokens, which your few thousand labelled examples could never teach. When there is enough data and the task is hard, such as language modelling or machine translation, train the embeddings with the model, because the task itself supplies the signal and task-specific vectors fit it better. Fine-tuning pretrained vectors is the middle path.
| Option | Where vectors come from | Updated during training? | When to use | Example |
|---|---|---|---|---|
| Frozen pretrained | word2vec or GloVe trained on another corpus | No | Little labelled data, simple task | Kim's CNN-static |
| Fine-tuned pretrained | Copied from word2vec or GloVe | Yes, starting from the pretrained values | Moderate data, want task-specific nuance | Kim's CNN-non-static |
| Joint from random | Learned from scratch with the network | Yes, from random values | Large data and a hard task (LM, MT) | Kim's CNN-rand; large LMs and MT systems |
Two studies give the evidence. Kim (2014) trained a CNN sentence classifier on small benchmarks three ways. CNN-rand, with random embeddings learned jointly, did poorly. CNN-static, with frozen word2vec vectors, "performs remarkably well", and CNN-non-static, which fine-tunes them, improves further; Kim concludes that pretrained vectors are good, universal feature extractors. Qi and colleagues (2018) asked the same question for neural machine translation and found gains of up to 20 BLEU in the most favourable setting. The gains have a sweet spot: they are largest when training data is scarce, but not so scarce that the system cannot be trained at all (baseline BLEU around 3 to 4), and they shrink as parallel data grows.
Everything in this lecture has been a Static embedding: one fixed vector per word type, whatever the sentence. Later lectures replace it with a Contextual embedding, where the vector for bank depends on its sentence. In those language models the embedding layer is trained jointly with the network, exactly the right-hand side of slide 110.
Recall
You have 500k parallel sentences for a hard MT task. Pretrained or joint, and why?
Quick check
You build a dialect sentiment classifier from 2,000 labeled tweets. How should its embeddings start?
Recap
If you remember nothing else
- Parallelogram method: b̂* = argmin over x of distance(x, a* − a + b), or argmax of cosine, with a, a* and b excluded. Slide 99's argmax of distance is a typo.
- Relations such as gender, number, tense and country to capital appear as roughly constant offsets; one word can sit on several such directions at once.
- Without exclusion the method returns b 93% of the time. It works best for frequent words, answers that already sit near b, and inflectional or capital relations; 3CosMul improves on 3CosAdd.
- Decade-specific embeddings, aligned by orthogonal Procrustes, show gay, broadcast and awful shifting. Frequent words change slowly; polysemous words change fast.
- Embeddings reproduce and amplify stereotypes (father : doctor :: mother : nurse), causing allocational and representational harms. Debiasing hides more than it removes.
- Garg et al. track bias over a century with relative norm differences: competence adjectives lean male (weakening since the 1960s), outsider words were tied to Asian names before 1950, and occupation bias tracks census data.
- Visualize with dendrograms, nearest-neighbour lists or t-SNE; all of them show usage, not an ontology.
- Intrinsic evaluation: WordSim-353, SimLex-999, TOEFL, analogies. Extrinsic evaluation is a downstream task and the more important one.
- Use pretrained vectors with little data or a simple task; train embeddings jointly with a large dataset and a hard task; fine-tuning sits between.
Sources
- Speech and Language Processing (3rd ed. draft), ch. 5 EmbeddingsBookJurafsky and Martin, StanfordSections 5.6 to 5.9: visualizing, the parallelogram, bias and evaluation(opens in a new tab)
- Linguistic Regularities in Continuous Space Word RepresentationsPaperMikolov, Yih and Zweig, NAACL 2013(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov et al., 2013(opens in a new tab)
- GloVe: Global Vectors for Word RepresentationPaperPennington, Socher and Manning, EMNLP 2014(opens in a new tab)
- GloVe project page (linear substructures)DocsStanford NLP Group(opens in a new tab)
- A model for analogical reasoningPaperRumelhart and Abrahamson, Cognitive Psychology 5(1), 1973(opens in a new tab)
- Linguistic Regularities in Sparse and Explicit Word RepresentationsPaperLevy and Goldberg, CoNLL 2014(opens in a new tab)
- Issues in evaluating semantic spaces using word analogiesPaperLinzen, RepEval 2016(opens in a new tab)
- Analogy-based detection of morphological and semantic relations with word embeddings (BATS)PaperGladkova, Drozd and Matsuoka, NAACL SRW 2016(opens in a new tab)
- Towards Understanding Linear Word AnalogiesPaperEthayarajh, Duvenaud and Hirst, ACL 2019(opens in a new tab)
- Parallelograms revisited: Exploring the limitations of vector space models for simple analogiesPaperPeterson, Chen and Griffiths, Cognition 205, 2020(opens in a new tab)
- Diachronic Word Embeddings Reveal Statistical Laws of Semantic ChangePaperHamilton, Leskovec and Jurafsky, ACL 2016(opens in a new tab)
- Syntactic Annotations for the Google Books Ngram CorpusPaperLin et al., ACL 2012 demo(opens in a new tab)
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word EmbeddingsPaperBolukbasi et al., NIPS 2016(opens in a new tab)
- Word embeddings quantify 100 years of gender and ethnic stereotypesPaperGarg, Schiebinger, Jurafsky and Zou, PNAS 2018(opens in a new tab)
- Semantics derived automatically from language corpora contain human-like biasesPaperCaliskan, Bryson and Narayanan, Science 2017(opens in a new tab)
- Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove ThemPaperGonen and Goldberg, NAACL 2019(opens in a new tab)
- Fair Is Better than Sensational: Man Is to Doctor as Woman Is to DoctorPaperNissim, van Noord and van der Goot, Computational Linguistics 46(2), 2020(opens in a new tab)
- The WordSimilarity-353 Test CollectionDocsFinkelstein et al., 2002(opens in a new tab)
- SimLex-999: Evaluating Semantic Models With (Genuine) Similarity EstimationPaperHill, Reichart and Korhonen, Computational Linguistics 41(4), 2015(opens in a new tab)
- Problems With Evaluation of Word Embeddings Using Word Similarity TasksPaperFaruqui et al., RepEval 2016(opens in a new tab)
- TOEFL Synonym Questions (State of the art)DocsACL Wiki(opens in a new tab)
- Visualizing Data using t-SNEPapervan der Maaten and Hinton, JMLR 9, 2008(opens in a new tab)
- Convolutional Neural Networks for Sentence ClassificationPaperKim, EMNLP 2014(opens in a new tab)
- When and Why Are Pre-Trained Word Embeddings Useful for Neural Machine Translation?PaperQi, Sachan, Felix, Padmanabhan and Neubig, NAACL 2018(opens in a new tab)