Majid Al-RaimiSimilarity, relatedness, antonymy and connotation

ICS 582Lecture 04Part 02

Similarity, relatedness, antonymy and connotation

The graded relations between word senses (similarity rated by humans, relatedness through semantic fields, antonymy as opposition on one feature) and the affective meaning captured by valence, arousal and dominance.

Concepts
6
Slides
11-17
Reading
36 min
Understood
0/6 concepts

Why this part matters

Before we build a single vector, we need to know what a good vector is supposed to capture. This part names the meaning relations that every embedding model in the rest of the lecture is judged against.

Similarity is the target of SimLex-style intrinsic evaluation. Relatedness is what co-occurrence counts actually pick up, whether you wanted it or not. Antonymy is the classic failure case of distributional models. Connotation is the raw material of sentiment and affect lexicons. Exams like to hand you a list of word pairs and ask which relation each one shows, and a research project that reports a score on a word benchmark must know which of these relations that benchmark measures. Part 01 gave us lemmas, senses and synonymy. Here we add the graded, messier relations that sit around them.

By the end you can

  1. Explain word similarity as a graded, human-rated relation and read SimLex-999 scores.
  2. Distinguish similarity from relatedness, and identify a semantic field from examples.
  3. Define antonymy, tell scale or binary opposites from reversives, and explain why antonyms look similar to distributional models.
  4. Describe connotation and evaluation, and give word sets that differ only in connotation.
  5. Define valence, arousal and dominance, and interpret NRC VAD scores.
  6. Label a word pair with the right relation: synonym, similar, related, antonym, or a connotation contrast.

Read these pairs and give each one a number from 0 to 10 for how alike the two meanings are: vanish and disappear, behave and obey, belief and impression, muscle and bone, modest and flexible, hole and agreement. You probably gave the first pair close to 10, the last close to 0, and found the middle harder but not impossible. Hundreds of people did exactly this, and their averages are the numbers below.

PairSimLex similarity (0 to 10)USF associationPOS
vanish / disappear9.82.76verb
behave / obey7.30.21verb
belief / impression5.950.10noun
muscle / bone3.650.13noun
modest / flexible0.980adjective
hole / agreement0.30noun
Slide 11 pairs with their SimLex-999 similarity, plus the USF free-association strength recorded in the same file

The association column comes from the University of South Florida free-association norms: how often people answer the second word when given the first as a cue. It measures connection, not shared features.

The scores fall away smoothly. There is no point where the pairs stop being similar and start being dissimilar, which is the first thing to notice. Synonymy, from part 01, is close to a yes or no question about two senses. Word similarity is a matter of degree: two words are similar when their meanings share features, and they can share many features, a few, or none. vanish and disappear share nearly all of them. muscle and bone share a few (body tissue, anatomy) and differ on the rest. hole and agreement share essentially nothing.

The second thing to notice is that similarity is a relation between words, not senses. This is the car and bicycle idea from the end of part 01. To say whether two senses are synonyms you need a sense inventory; to ask a person how similar two words feel, you do not. That makes word similarity cheap to collect and directly comparable to anything that produces one number per word pair, which is exactly what a vector model does.

Where the numbers come from

The table is SimLex-999 (Hill, Reichart and Korhonen 2015). It contains 999 pairs: 666 noun pairs, 222 verb pairs and 111 adjective pairs, mixing concrete and abstract words. About 500 Mechanical Turk workers rated them on an integer slider from 0 to 6, and the mean ratings were then linearly rescaled to 0 to 10. The slide shows the rescaled values, which match the released data file exactly. Individual raters disagree a fair amount: the average Spearman correlation between two raters is 0.67, and between one rater and the mean of the others it is 0.78. The average is what is stable.

Why should you care about these particular numbers? Because later in this lecture they become the gold standard for intrinsic evaluation. You compute the cosine between the two vectors of every SimLex pair, rank the pairs by cosine, and report the Spearman rank correlation with the human ranking. A model that agrees with people about which pairs are more alike scores high.

Recall

What scale does SimLex-999 use, and what does 0.3 mean for hole and agreement?

0 to 10, rescaled from a 0 to 6 slider. A score of 0.3 means raters judged the pair to share almost no meaning.

Quick check

Which pair would SimLex-999 annotators rate as most similar?

Relatedness: words that belong to the same scene

Compare two pairs: coffee and tea, then coffee and cup. Coffee and tea are both hot drinks made by steeping or brewing a plant, both contain caffeine, both are served in the morning. They share features, so they aresimilar. Coffee and cup share practically no features: one is a drink from a plant, the other is a manufactured container. Yet nobody would call them unconnected. They take part together in one everyday event, drinking coffee from a cup.

That second kind of connection is word relatedness, also called association. Two words are related when they are connected in any way at all: by shared features, by a shared event, by part and whole (car and wheel), by function (pencil and paper), even by opposition (hot and cold). Budanitsky and Hirst put the hierarchy plainly: similarity is a special case of relatedness. Every similar pair is related, but many related pairs are not similar. They quote Resnik's example: cars and gasoline are more closely related than cars and bicycles, but the latter pair are certainly more similar.

coffee at the centre. tea sits close on a short solid edge because the two share features. cup sits far away on a long dashed edge: it shares almost no features with coffee, but the two belong to the same scene.
PairShares features?Same scene?Label
coffee / teaYes: both are hot drinksOftenSimilar (and related)
coffee / cupAlmost none: drink against objectYes: drinking coffeeRelated, not similar
surgeon / scalpelNo: person against toolYes: an operationRelated, not similar
car / bicycleSome: wheeled vehiclesRarelySimilar, weakly related
car / gasolineNo: vehicle against fuelYes: driving, refuellingRelated, not similar
hole / agreementNoNoNeither
Two separate questions: do the words share features, and do they appear in the same scene?

The SimLex file records both quantities, so you can see the split in real numbers. clothes and closet have similarity 3.27 but association 4.83: related more than they are alike. car and bicycle have similarity 3.47 and association only 0.41: alike more than they are associated.

Semantic fields

When you collect all the words that keep turning up in the same scene, you get a semantic field: a set of words that cover one domain and bear structured relations to each other. The hospital field holds surgeon, scalpel, nurse, anaesthetic and hospital. The restaurant field holds waiter, menu, plate, food and chef. The house field holds door, roof, kitchen, family and bed. Inside a field the relations are varied: a surgeon uses a scalpel, a nurse assists a surgeon, the anaesthetic is given in the hospital. What unites them is the domain, not shared features.

Three fields, five words each. Each field lights in turn and its words join up: the edges inside a field are relatedness, not similarity.

Fields are not only a linguist's classification. Topic models such as Latent Dirichlet Allocation read a large collection of documents with no labels and discover clusters of words that tend to occur together, and the clusters they find look very much like semantic fields. That is a first hint of the theme of this whole lecture: the company a word keeps carries information about its meaning.

Recall

Why are coffee and tea similar, but coffee and cup only related?

Tea shares features with coffee: both are hot, caffeinated drinks. A cup shares almost none, but it takes part in the same event, drinking coffee, so the pair is connected through a scene or semantic field rather than through shared meaning.

Recall

Which semantic field do waiter, menu, plate and chef belong to, and what holds them together?

The restaurant field. They share almost no features of meaning; what unites them is the domain, and the varied relations inside it (a waiter brings the menu, a chef prepares the food on the plate).

Quick check

A model gives coffee and cup a high score. What has it most likely captured?

Take hot and cold. Both are adjectives. Both describe temperature. Both slot into the same frames: hot coffee and cold coffee, hot weather and cold weather, it is too hot today and it is too cold today. Line up everything you know about the two words and they agree on almost every point. They disagree on exactly one: which end of the temperature scale they name.

That is the definition of antonymy. Antonyms are senses that are opposite with respect to only one feature of meaning and otherwise very similar. It sounds paradoxical that opposites are mostly alike, but you cannot be opposite to something unless you are first comparable to it. Hot is not the opposite of Tuesday.

hot and cold share their part of speech, their dimension and the nouns they modify, so the first three rows light identically. Only the scale position flips to opposite ends. Below, a reversive pair: rise and fall move in opposite directions.

Kinds of opposition

The single opposed feature can be of different types. In the first group the two words name the two values of a binary choice or the two ends of a scale: long and short, fast and slow, big and little. In the second group, the reversives, the two words describe change or movement in opposite directions: rise and fall, up and down. Mohammad, Dorr, Hirst and Turney refine this further (antipodals, complementaries, gradable opposites) and note that many contrasting pairs, such as warm and cold, are not strict opposites at all.

KindPairsWhat is opposed
Opposite ends of a scalelong / short, fast / slow, hot / coldA position on one graded dimension (length, speed, temperature)
Binary oppositionin / outTwo values with no middle ground
Reversiverise / fall, up / downThe direction of a change or movement
The two families on slide 14, with the scale case split from the binary one

What humans say, and why models struggle

Because antonyms differ on a feature people care about, human raters call them dissimilar. Because they share everything else, they are among the most strongly associated pairs in the language. SimLex records both, and the gap is striking. Hill and colleagues conclude that antonyms are the most strongly associated word pairs among the finer-grained relations they examined.

PairSimLex similarityUSF association
night / day1.888.19
old / new1.587.25
short / long1.235.36
bottom / top0.706.96
large / big (synonyms, for contrast)9.550.68
SimLex-999 similarity against USF association for antonym pairs, with a synonym pair for contrast

Now look ahead. The distributional hypothesis that drives the rest of this lecture says that words in similar contexts have similar meanings. hot and cold occur in almost identical contexts, so a model built on contexts will put them close together. Opposites even co-occur in the same sentence more often than chance would predict (Charles and Miller, cited by Mohammad and colleagues), which pulls them closer still. SLP3 is blunt about the result: automatically distinguishing synonyms from antonyms can be difficult.

Recall

What do antonyms have in common, and why does that matter for distributional models?

They are opposite on one feature only and alike on everything else, so they occur in the same contexts and end up close together in a distributional space. Synonyms and antonyms are therefore hard to separate.

Recall

Name the two kinds of antonymy on slide 14, with an example of each.

Binary opposition or opposite ends of a scale (long/short, fast/slow), and reversives, which describe change or movement in opposite directions (rise/fall, up/down).

Quick check

Why do distributional models often place hot close to cold?

Connotation: the feeling a word carries

A museum shop sells a replica of an ancient vase. A street stall sells a knockoff. Both objects are copies of a real thing, and a description of either would read much the same. But the first word is close to praise and the second is an accusation.

The part of meaning that differs here is connotation: the aspects of a word's meaning tied to a writer's or reader's emotions, sentiment, opinions or evaluations. Some words exist mainly to evaluate: great and love are positive, terrible and hate are negative. Others, like replica and knockoff, describe the same thing while carrying different attitudes toward it. Positive or negative evaluation in language is called sentiment, and connotation is what sentiment analysis, stance detection, and NLP work on political language and consumer reviews all exploit.

Connotation can be measured. Affect lexicons give each word a score, and the one we meet in the next concept, the NRC VAD Lexicon, gives a valence (pleasantness) between 0 and 1. Here is the SLP3 example worked through with its real numbers.

Worked example

Two sets of copies, one difference in feeling

  1. Look up each word's valence

    Negative set: fake 0.073, knockoff 0.350, forgery 0.235. Positive set: copy 0.460, replica 0.480, reproduction 0.800.
  2. Average each set

    Negative mean: (0.073 + 0.350 + 0.235) / 3 ≈ 0.219. Positive mean: (0.460 + 0.480 + 0.800) / 3 = 0.580.
  3. Compare, and read the numbers critically

    The positive set is about 0.36 higher. But positive is relative here: copy, at 0.460, sits just below the neutral midpoint. And reproduction scores high partly because it is polysemous: its biological sense (having children) is pleasant, and a lexicon with one score per word averages over all senses.
  4. Result

    Words that refer to nearly the same thing can sit far apart on valence. Other pairs show the same pattern: innocent 0.729 against naive 0.406, great 0.958 against terrible 0.061, love 1.000 against hate 0.031.

Recall

Give two words with nearly the same reference but different connotation, and say how you would measure the difference.

replica and fake (or knockoff). Look up their valence in the NRC VAD Lexicon: replica 0.480, fake 0.073.

Compare napping and toxic. napping is pleasant and calm, and it puts you in no particular position of control. toxic is unpleasant and agitating. A single positive or negative score would capture the first difference, but not the second, and not the question of who has the power.

Osgood and colleagues (1957) found that people's ratings of words consistently varied along three affective dimensions, now called valence, arousal and dominance:

  • Valence is the pleasantness of the stimulus. napping 0.765, toxic 0.008.
  • Arousal is the intensity of emotion the stimulus provokes. napping 0.046, toxic 0.885.
  • Dominance is the degree of control the stimulus exerts. napping 0.306, toxic 0.492.

Three numbers per word means each word becomes a point in a three-dimensional space. SLP3 calls this the first expression of the idea behind vector semantics: on the 1 to 9 scales of Warriner and colleagues (2013), heartbreak sits at [2.45, 5.65, 3.58]. Part 03 takes that idea and scales it from three hand-chosen dimensions to hundreds learned from text.

A V, A, D frame. love, toxic, napping and powerful start bunched at the centre, then slide to their real NRC VAD coordinates, with droplines to the valence-dominance floor showing how high each sits on arousal.

How the NRC VAD Lexicon was built

The numbers come from the NRC VAD Lexicon (Mohammad 2018). Version 1 covers about 20,000 English words; the released file has 19,971 entries. Asking people to rate a word on a slider is unreliable, because everyone uses the slider differently. Instead Mohammad used best-worst scaling. An annotator sees four words and picks the one highest on the dimension (say, most pleasant) and the one lowest. Over many such four-word sets, each word is scored by how often it won minus how often it lost.

score(w)=#best(w)#seen(w)−#worst(w)#seen(w)\text{score}(w) = \frac{\#\text{best}(w)}{\#\text{seen}(w)} - \frac{\#\text{worst}(w)}{\#\text{seen}(w)}
Best-worst score, which runs from -1 to 1 before rescaling to 0 to 1 in v1

Comparative judgements are much more consistent than absolute ones. Splitting the annotators into two random halves and correlating the scores each half produces gives split-half reliability of r = 0.95 for valence, 0.90 for arousal and 0.91 for dominance. Version 2, released in March 2025, extends the lexicon to over 55,000 terms (about 10,000 of them multiword phrases) on a -1 to 1 scale, with split-half Spearman of 0.98, 0.97 and 0.96.

WordValenceArousalDominance
love1.0000.5190.673
happy1.0000.7350.772
toxic0.0080.8850.492
nightmare0.0050.8100.436
elated0.7920.9600.725
frenzy0.6100.9650.682
mellow0.6330.0690.265
napping0.7650.0460.306
calm0.8750.1000.282
excited0.9080.9310.709
powerful0.8650.8300.991
leadership0.8700.6900.983
controlling0.4900.4410.885
weak0.1800.2410.045
empty0.1880.1830.081
NRC VAD v1 scores (0 to 1). Sort mentally by each column and the order changes: the dimensions are independent.

Recall

Define valence, arousal and dominance, and place napping on each.

Valence is pleasantness, arousal is intensity of emotion, dominance is degree of control. napping has valence 0.765 (pleasant), arousal 0.046 (very calm) and dominance 0.306 (low control).

Recall

How is an NRC VAD score produced, and how reliable is it?

Best-worst scaling over four-word sets: the proportion of times a word is chosen best minus the proportion chosen worst, rescaled to 0 to 1 in v1. Split-half reliability is r = 0.95, 0.90 and 0.91 for valence, arousal and dominance.

Quick check

In NRC VAD v1, napping scores 0.046 on arousal. What does that tell you?

Step back and look at what we now have. One word can map to many senses: mouse is a rodent or a pointing device. One sense can map to many words: couch and sofa. The mapping between words and concepts is many to many, and on top of it sits a set of relations, some between senses and some between whole words.

RelationLevelGraded?ExampleTypical evidence
SynonymySenseRarely exactcouch / sofaThesaurus or WordNet synsets
AntonymySenseComes in kindshot / coldWordNet antonym links
SimilarityWordGradedvanish / disappearSimLex-999 ratings
RelatednessWordGradedcoffee / cupAssociation norms or WordSim-353
ConnotationWordGradedreplica / knockoffNRC VAD Lexicon
The five relations that a representation of word meaning should reproduce

A lemma groups its senses, and polysemy is the fact that it has several. Synonymy and antonymy are relations between senses. Similarity, relatedness and connotation are graded and can be asked of whole words, which is what makes them measurable with ratings and lexicons.

This table is a list of desiderata. Any representation of word meaning we build next should put similar words near each other, keep related words in recognisable neighbourhoods, encode affect in some consistent direction, and ideally tell synonyms from antonyms. Part 03 introduces vectors as that representation. A vector space captures relatedness readily and affect reasonably well; separating true similarity from relatedness is harder, and telling synonyms from antonyms is where it struggles most, as concepts 2 and 3 warned.

Quick check

replica and knockoff differ mainly in which relation?

Recap

If you remember nothing else

  • Similarity is graded and word level. SimLex-999 (999 pairs, 0 to 10) runs from vanish/disappear 9.8 down to hole/agreement 0.3.
  • Relatedness (association) is broader: coffee/cup are related through a shared event but not similar. Similarity is a special case of relatedness.
  • A semantic field is a set of words covering one domain with structured relations: the hospital, restaurant and house fields. Topic models induce fields from text.
  • Antonyms are opposite on one feature and alike on the rest. The kinds are binary or scalar opposites (long/short) and reversives (rise/fall).
  • Antonyms score low on SimLex similarity but high on association (night/day 1.88 against 8.19), so distributional models tend to put them close together.
  • Connotation is affective meaning. Near-synonyms can differ sharply: replica 0.480 against fake 0.073 valence.
  • Osgood's three affective dimensions are valence (pleasantness), arousal (intensity) and dominance (control).
  • NRC VAD v1 has about 20k words scored 0 to 1 by best-worst scaling. v2 (2025) has over 55k terms on -1 to 1.
  • Every relation here is a test that later vector representations must pass.

Sources