ICS 582Lecture 04Part 02
Similarity, relatedness, antonymy and connotation
The graded relations between word senses (similarity rated by humans, relatedness through semantic fields, antonymy as opposition on one feature) and the affective meaning captured by valence, arousal and dominance.
- Concepts
- 6
- Slides
- 11-17
- Reading
- 36 min
Why this part matters
Before we build a single vector, we need to know what a good vector is supposed to capture. This part names the meaning relations that every embedding model in the rest of the lecture is judged against.
Similarity is the target of SimLex-style intrinsic evaluation. Relatedness is what co-occurrence counts actually pick up, whether you wanted it or not. Antonymy is the classic failure case of distributional models. Connotation is the raw material of sentiment and affect lexicons. Exams like to hand you a list of word pairs and ask which relation each one shows, and a research project that reports a score on a word benchmark must know which of these relations that benchmark measures. Part 01 gave us lemmas, senses and synonymy. Here we add the graded, messier relations that sit around them.
By the end you can
- Explain word similarity as a graded, human-rated relation and read SimLex-999 scores.
- Distinguish similarity from relatedness, and identify a semantic field from examples.
- Define antonymy, tell scale or binary opposites from reversives, and explain why antonyms look similar to distributional models.
- Describe connotation and evaluation, and give word sets that differ only in connotation.
- Define valence, arousal and dominance, and interpret NRC VAD scores.
- Label a word pair with the right relation: synonym, similar, related, antonym, or a connotation contrast.
Read these pairs and give each one a number from 0 to 10 for how alike the two meanings are: vanish and disappear, behave and obey, belief and impression, muscle and bone, modest and flexible, hole and agreement. You probably gave the first pair close to 10, the last close to 0, and found the middle harder but not impossible. Hundreds of people did exactly this, and their averages are the numbers below.
| Pair | SimLex similarity (0 to 10) | USF association | POS |
|---|---|---|---|
| vanish / disappear | 9.8 | 2.76 | verb |
| behave / obey | 7.3 | 0.21 | verb |
| belief / impression | 5.95 | 0.10 | noun |
| muscle / bone | 3.65 | 0.13 | noun |
| modest / flexible | 0.98 | 0 | adjective |
| hole / agreement | 0.3 | 0 | noun |
The association column comes from the University of South Florida free-association norms: how often people answer the second word when given the first as a cue. It measures connection, not shared features.
The scores fall away smoothly. There is no point where the pairs stop being similar and start being dissimilar, which is the first thing to notice. Synonymy, from part 01, is close to a yes or no question about two senses. Word similarity is a matter of degree: two words are similar when their meanings share features, and they can share many features, a few, or none. vanish and disappear share nearly all of them. muscle and bone share a few (body tissue, anatomy) and differ on the rest. hole and agreement share essentially nothing.
The second thing to notice is that similarity is a relation between words, not senses. This is the car and bicycle idea from the end of part 01. To say whether two senses are synonyms you need a sense inventory; to ask a person how similar two words feel, you do not. That makes word similarity cheap to collect and directly comparable to anything that produces one number per word pair, which is exactly what a vector model does.
Where the numbers come from
The table is SimLex-999 (Hill, Reichart and Korhonen 2015). It contains 999 pairs: 666 noun pairs, 222 verb pairs and 111 adjective pairs, mixing concrete and abstract words. About 500 Mechanical Turk workers rated them on an integer slider from 0 to 6, and the mean ratings were then linearly rescaled to 0 to 10. The slide shows the rescaled values, which match the released data file exactly. Individual raters disagree a fair amount: the average Spearman correlation between two raters is 0.67, and between one rater and the mean of the others it is 0.78. The average is what is stable.
Why should you care about these particular numbers? Because later in this lecture they become the gold standard for intrinsic evaluation. You compute the cosine between the two vectors of every SimLex pair, rank the pairs by cosine, and report the Spearman rank correlation with the human ranking. A model that agrees with people about which pairs are more alike scores high.
Recall
What scale does SimLex-999 use, and what does 0.3 mean for hole and agreement?
Quick check
Which pair would SimLex-999 annotators rate as most similar?
Take hot and cold. Both are adjectives. Both describe temperature. Both slot into the same frames: hot coffee and cold coffee, hot weather and cold weather, it is too hot today and it is too cold today. Line up everything you know about the two words and they agree on almost every point. They disagree on exactly one: which end of the temperature scale they name.
That is the definition of antonymy. Antonyms are senses that are opposite with respect to only one feature of meaning and otherwise very similar. It sounds paradoxical that opposites are mostly alike, but you cannot be opposite to something unless you are first comparable to it. Hot is not the opposite of Tuesday.
Kinds of opposition
The single opposed feature can be of different types. In the first group the two words name the two values of a binary choice or the two ends of a scale: long and short, fast and slow, big and little. In the second group, the reversives, the two words describe change or movement in opposite directions: rise and fall, up and down. Mohammad, Dorr, Hirst and Turney refine this further (antipodals, complementaries, gradable opposites) and note that many contrasting pairs, such as warm and cold, are not strict opposites at all.
| Kind | Pairs | What is opposed |
|---|---|---|
| Opposite ends of a scale | long / short, fast / slow, hot / cold | A position on one graded dimension (length, speed, temperature) |
| Binary opposition | in / out | Two values with no middle ground |
| Reversive | rise / fall, up / down | The direction of a change or movement |
What humans say, and why models struggle
Because antonyms differ on a feature people care about, human raters call them dissimilar. Because they share everything else, they are among the most strongly associated pairs in the language. SimLex records both, and the gap is striking. Hill and colleagues conclude that antonyms are the most strongly associated word pairs among the finer-grained relations they examined.
| Pair | SimLex similarity | USF association |
|---|---|---|
| night / day | 1.88 | 8.19 |
| old / new | 1.58 | 7.25 |
| short / long | 1.23 | 5.36 |
| bottom / top | 0.70 | 6.96 |
| large / big (synonyms, for contrast) | 9.55 | 0.68 |
Now look ahead. The distributional hypothesis that drives the rest of this lecture says that words in similar contexts have similar meanings. hot and cold occur in almost identical contexts, so a model built on contexts will put them close together. Opposites even co-occur in the same sentence more often than chance would predict (Charles and Miller, cited by Mohammad and colleagues), which pulls them closer still. SLP3 is blunt about the result: automatically distinguishing synonyms from antonyms can be difficult.
Recall
What do antonyms have in common, and why does that matter for distributional models?
Recall
Name the two kinds of antonymy on slide 14, with an example of each.
Quick check
Why do distributional models often place hot close to cold?
A museum shop sells a replica of an ancient vase. A street stall sells a knockoff. Both objects are copies of a real thing, and a description of either would read much the same. But the first word is close to praise and the second is an accusation.
The part of meaning that differs here is connotation: the aspects of a word's meaning tied to a writer's or reader's emotions, sentiment, opinions or evaluations. Some words exist mainly to evaluate: great and love are positive, terrible and hate are negative. Others, like replica and knockoff, describe the same thing while carrying different attitudes toward it. Positive or negative evaluation in language is called sentiment, and connotation is what sentiment analysis, stance detection, and NLP work on political language and consumer reviews all exploit.
Connotation can be measured. Affect lexicons give each word a score, and the one we meet in the next concept, the NRC VAD Lexicon, gives a valence (pleasantness) between 0 and 1. Here is the SLP3 example worked through with its real numbers.
Worked example
Two sets of copies, one difference in feeling
Look up each word's valence
Negative set: fake 0.073, knockoff 0.350, forgery 0.235. Positive set: copy 0.460, replica 0.480, reproduction 0.800.Average each set
Negative mean: (0.073 + 0.350 + 0.235) / 3 ≈ 0.219. Positive mean: (0.460 + 0.480 + 0.800) / 3 = 0.580.Compare, and read the numbers critically
The positive set is about 0.36 higher. But positive is relative here: copy, at 0.460, sits just below the neutral midpoint. And reproduction scores high partly because it is polysemous: its biological sense (having children) is pleasant, and a lexicon with one score per word averages over all senses.Result
Words that refer to nearly the same thing can sit far apart on valence. Other pairs show the same pattern: innocent 0.729 against naive 0.406, great 0.958 against terrible 0.061, love 1.000 against hate 0.031.
Recall
Give two words with nearly the same reference but different connotation, and say how you would measure the difference.
Compare napping and toxic. napping is pleasant and calm, and it puts you in no particular position of control. toxic is unpleasant and agitating. A single positive or negative score would capture the first difference, but not the second, and not the question of who has the power.
Osgood and colleagues (1957) found that people's ratings of words consistently varied along three affective dimensions, now called valence, arousal and dominance:
- Valence is the pleasantness of the stimulus. napping 0.765, toxic 0.008.
- Arousal is the intensity of emotion the stimulus provokes. napping 0.046, toxic 0.885.
- Dominance is the degree of control the stimulus exerts. napping 0.306, toxic 0.492.
Three numbers per word means each word becomes a point in a three-dimensional space. SLP3 calls this the first expression of the idea behind vector semantics: on the 1 to 9 scales of Warriner and colleagues (2013), heartbreak sits at [2.45, 5.65, 3.58]. Part 03 takes that idea and scales it from three hand-chosen dimensions to hundreds learned from text.
How the NRC VAD Lexicon was built
The numbers come from the NRC VAD Lexicon (Mohammad 2018). Version 1 covers about 20,000 English words; the released file has 19,971 entries. Asking people to rate a word on a slider is unreliable, because everyone uses the slider differently. Instead Mohammad used best-worst scaling. An annotator sees four words and picks the one highest on the dimension (say, most pleasant) and the one lowest. Over many such four-word sets, each word is scored by how often it won minus how often it lost.
Comparative judgements are much more consistent than absolute ones. Splitting the annotators into two random halves and correlating the scores each half produces gives split-half reliability of r = 0.95 for valence, 0.90 for arousal and 0.91 for dominance. Version 2, released in March 2025, extends the lexicon to over 55,000 terms (about 10,000 of them multiword phrases) on a -1 to 1 scale, with split-half Spearman of 0.98, 0.97 and 0.96.
| Word | Valence | Arousal | Dominance |
|---|---|---|---|
| love | 1.000 | 0.519 | 0.673 |
| happy | 1.000 | 0.735 | 0.772 |
| toxic | 0.008 | 0.885 | 0.492 |
| nightmare | 0.005 | 0.810 | 0.436 |
| elated | 0.792 | 0.960 | 0.725 |
| frenzy | 0.610 | 0.965 | 0.682 |
| mellow | 0.633 | 0.069 | 0.265 |
| napping | 0.765 | 0.046 | 0.306 |
| calm | 0.875 | 0.100 | 0.282 |
| excited | 0.908 | 0.931 | 0.709 |
| powerful | 0.865 | 0.830 | 0.991 |
| leadership | 0.870 | 0.690 | 0.983 |
| controlling | 0.490 | 0.441 | 0.885 |
| weak | 0.180 | 0.241 | 0.045 |
| empty | 0.188 | 0.183 | 0.081 |
Recall
Define valence, arousal and dominance, and place napping on each.
Recall
How is an NRC VAD score produced, and how reliable is it?
Quick check
In NRC VAD v1, napping scores 0.046 on arousal. What does that tell you?
Step back and look at what we now have. One word can map to many senses: mouse is a rodent or a pointing device. One sense can map to many words: couch and sofa. The mapping between words and concepts is many to many, and on top of it sits a set of relations, some between senses and some between whole words.
| Relation | Level | Graded? | Example | Typical evidence |
|---|---|---|---|---|
| Synonymy | Sense | Rarely exact | couch / sofa | Thesaurus or WordNet synsets |
| Antonymy | Sense | Comes in kinds | hot / cold | WordNet antonym links |
| Similarity | Word | Graded | vanish / disappear | SimLex-999 ratings |
| Relatedness | Word | Graded | coffee / cup | Association norms or WordSim-353 |
| Connotation | Word | Graded | replica / knockoff | NRC VAD Lexicon |
A lemma groups its senses, and polysemy is the fact that it has several. Synonymy and antonymy are relations between senses. Similarity, relatedness and connotation are graded and can be asked of whole words, which is what makes them measurable with ratings and lexicons.
This table is a list of desiderata. Any representation of word meaning we build next should put similar words near each other, keep related words in recognisable neighbourhoods, encode affect in some consistent direction, and ideally tell synonyms from antonyms. Part 03 introduces vectors as that representation. A vector space captures relatedness readily and affect reasonably well; separating true similarity from relatedness is harder, and telling synonyms from antonyms is where it struggles most, as concepts 2 and 3 warned.
Quick check
replica and knockoff differ mainly in which relation?
Recap
If you remember nothing else
- Similarity is graded and word level. SimLex-999 (999 pairs, 0 to 10) runs from vanish/disappear 9.8 down to hole/agreement 0.3.
- Relatedness (association) is broader: coffee/cup are related through a shared event but not similar. Similarity is a special case of relatedness.
- A semantic field is a set of words covering one domain with structured relations: the hospital, restaurant and house fields. Topic models induce fields from text.
- Antonyms are opposite on one feature and alike on the rest. The kinds are binary or scalar opposites (long/short) and reversives (rise/fall).
- Antonyms score low on SimLex similarity but high on association (night/day 1.88 against 8.19), so distributional models tend to put them close together.
- Connotation is affective meaning. Near-synonyms can differ sharply: replica 0.480 against fake 0.073 valence.
- Osgood's three affective dimensions are valence (pleasantness), arousal (intensity) and dominance (control).
- NRC VAD v1 has about 20k words scored 0 to 1 by best-worst scaling. v2 (2025) has over 55k terms on -1 to 1.
- Every relation here is a test that later vector representations must pass.
Sources
- Speech and Language Processing, 3rd edition draft, chapter 5: EmbeddingsBookJurafsky and Martin, Stanford UniversitySection 5.1 lexical semantics: similarity, relatedness, semantic fields, connotation examples, VAD as the first vector semantics(opens in a new tab)
- Speech and Language Processing, 3rd edition draft, appendix I: Word Senses and WordNetBookJurafsky and Martin, Stanford UniversityAntonymy: binary and scalar opposites, reversives, and the difficulty of separating synonyms from antonyms(opens in a new tab)
- Speech and Language Processing, 3rd edition draft, chapter 23: Lexicons for Sentiment, Affect, and ConnotationBookJurafsky and Martin, Stanford UniversityBest-worst scaling procedure, score formula and split-half reliability(opens in a new tab)
- SimLex-999: Evaluating Semantic Models With (Genuine) Similarity EstimationPaperHill, Reichart and Korhonen, Computational Linguistics 41(4), 2015Dataset design, the 0 to 6 slider rescaled to 0 to 10, and why WordSim-353 measures association(opens in a new tab)
- SimLex-999 homepageDocsFelix HillPair counts, inter-annotator agreement; its clothes-closet figure conflicts with the data file, which gives 3.27(opens in a new tab)
- Evaluating WordNet-based Measures of Lexical Semantic RelatednessPaperBudanitsky and Hirst, Computational Linguistics 32(1), 2006Similarity as a special case of relatedness, and Resnik's cars, gasoline and bicycles example(opens in a new tab)
- Computing Lexical ContrastPaperMohammad, Dorr, Hirst and Turney, Computational Linguistics 39(3), 2013Kinds of opposites and contrasting non-opposites(opens in a new tab)
- Integrating Distributional Lexical Contrast into Word Embeddings for Antonym-Synonym DistinctionPaperNguyen, Schulte im Walde and Vu, ACL 2016Distributional models retrieve both synonyms and antonyms as related words(opens in a new tab)
- Obtaining Reliable Human Ratings of Valence, Arousal, and Dominance for 20,000 English WordsPaperSaif M. Mohammad, ACL 2018NRC VAD v1 construction and Table 2 extremes(opens in a new tab)
- NRC Valence, Arousal, and Dominance LexiconDocsNational Research Council CanadaRelease history, v1 and v2 coverage and scales(opens in a new tab)
- NRC VAD Lexicon v2: Norms for Valence, Arousal, and Dominance for over 55k English TermsPaperSaif M. Mohammad, 2025The -1 to 1 scale and split-half reliability of v2(opens in a new tab)