ICS 582Lecture 04Glossary

Glossary

Every term in Word embeddings, defined once and used the same way in every part. Each entry links to the slides where the idea appears.

Terms
51
Letters
14

A

Alpha-weighted context probability

Raising context counts to alpha = 0.75 before normalizing, so rare contexts gain probability and their inflated PMI falls; word2vec draws its noise words from the same flattened distribution.

Antonymy

Senses that are opposite on only one feature of meaning, as binary or scalar opposites or reversives.

C

Collection frequency

The total count of a term across all documents in the collection, as opposed to its document frequency.

Connotation

The affective meaning of a word, positive or negative, including evaluation.

Context window

The span of words on either side of a target word, such as +/- 2, whose words count as its context.

Contextual embedding

A separate embedding computed for each token of a word in its context, as in ELMo or BERT.

Continuous bag of words

A word2vec variant that predicts the center word from the average (some presentations use the sum) of its context word vectors.

Cosine similarity

The dot product divided by the product of the two vector lengths, ranging from -1 to 1 (0 to 1 for counts).

D

Dense vector

A short vector (50 to 1000 dimensions) whose elements are mostly non-zero.

Diachronic embeddings

Embeddings trained separately on text from different decades, used to see word meanings shift over time.

Distributional hypothesis

The idea that words with almost identical environments (neighboring words) have similar meanings.

Document frequency

The number of documents a term occurs in, distinct from its total collection frequency.

Dot product

The sum of products of corresponding vector entries, high when both vectors have large values in the same dimensions.

E

Embedding

A vector representation of a word, so called because the word is embedded into a space.

Extrinsic evaluation

Evaluating embeddings by their effect on a real downstream task such as MT or NER.

F

FastText

A word2vec extension that represents a word as a bag of character n-grams, giving it subword information.

G

GloVe

Global Vectors: embeddings fit to log co-occurrence counts with a weighting function that limits rare and very frequent pairs.

I

Intrinsic evaluation

Evaluating embeddings on an intermediate task, such as correlation with human similarity ratings or TOEFL questions.

Inverse document frequency

log10(N / df_t), which gives low weight to terms such as the that occur in nearly every document.

L

Learning rate

The factor eta that scales each gradient descent step.

Lemma

The dictionary form of a word (for example mouse (N)) that groups its senses.

Lexical semantics

The linguistic study of word meaning.

N

Negative sampling

Drawing k random noise words from the unigram distribution raised to 0.75 as negative examples for each positive pair, instead of a softmax over the whole vocabulary.

P

Parallelogram method

Solving a:a*::b:? by finding the word closest to a* - a + b, with a, a* and b excluded from the candidates.

Pointwise mutual information

log2 of P(x,y) / (P(x)P(y)), measuring whether two words co-occur more than if they were independent.

Polysemy

The property of a lemma having multiple senses.

Positive PMI

PMI with negative values replaced by 0.

Principle of contrast

The idea that any difference in linguistic form signals some difference in meaning.

S

Self-supervision

Using naturally occurring neighbors in a corpus as gold labels, so no human labels are needed.

Semantic field

A set of words that cover one semantic domain and bear structured relations to each other.

Sense

A meaning component of a word, also called a concept; a lemma can have several.

Sigmoid

1 / (1 + exp(-x)), used to turn a dot product into the probability that c is a neighbor of w.

Singular value decomposition

A matrix factorization used to get short dense vectors from a count matrix; LSA (latent semantic analysis) is a special case.

Skip-gram

A word2vec model that predicts whether a candidate word is a neighbor of the target word.

Skip-gram with negative sampling

Skip-gram trained to separate true (target, context) pairs from k randomly sampled noise pairs.

Skip-thought vectors

A sentence-level analogue of skip-gram that encodes a sentence to predict its neighboring sentences.

Sparse vector

A long vector (length |V|) whose elements are mostly zero, such as tf-idf or PMI vectors.

Static embedding

One fixed vector per word type, as in word2vec or GloVe, unlike contextual embeddings.

Stochastic gradient descent

Learning by repeatedly moving the weights opposite the gradient of the loss, scaled by the learning rate, to make positive pairs more likely and negative pairs less likely.

Synonymy

A relation between senses that have the same meaning in some or all contexts.

T

Target and context matrices

The two embedding sets W and C learned by SGNS, often summed as w_i + c_i.

Term frequency

The count of a term in a document, usually squashed as log10(count + 1).

Term-context matrix

A word-word matrix counting how often each target word occurs near each context word.

Term-document matrix

A matrix with one row per word and one column per document, each cell a count.

tf-idf

A weight equal to term frequency times inverse document frequency.

V

Valence, arousal, dominance

Osgood's three affective dimensions: pleasantness, intensity of emotion, and degree of control.

Vector length

The square root of the sum of a vector's squared entries.

Vector semantics

Representing a word's meaning as a point in a multidimensional space built from its distribution.

W

Word relatedness

Association between words related in any way, for example through a semantic field, such as coffee and cup.

Word similarity

Words that are not synonyms but share some element of meaning, such as car and bicycle.

Word2vec

An embedding method that learns dense vectors by predicting rather than counting nearby words.