ICS 582Lecture 04Glossary
Glossary
Every term in Word embeddings, defined once and used the same way in every part. Each entry links to the slides where the idea appears.
- Terms
- 51
- Letters
- 14
A
- Alpha-weighted context probability
Raising context counts to alpha = 0.75 before normalizing, so rare contexts gain probability and their inflated PMI falls; word2vec draws its noise words from the same flattened distribution.
C
- Collection frequency
The total count of a term across all documents in the collection, as opposed to its document frequency.
- Connotation
The affective meaning of a word, positive or negative, including evaluation.
- Context window
The span of words on either side of a target word, such as +/- 2, whose words count as its context.
- Contextual embedding
A separate embedding computed for each token of a word in its context, as in ELMo or BERT.
- Continuous bag of words
A word2vec variant that predicts the center word from the average (some presentations use the sum) of its context word vectors.
- Cosine similarity
The dot product divided by the product of the two vector lengths, ranging from -1 to 1 (0 to 1 for counts).
D
- Dense vector
A short vector (50 to 1000 dimensions) whose elements are mostly non-zero.
- Diachronic embeddings
Embeddings trained separately on text from different decades, used to see word meanings shift over time.
- Distributional hypothesis
The idea that words with almost identical environments (neighboring words) have similar meanings.
- Document frequency
The number of documents a term occurs in, distinct from its total collection frequency.
- Dot product
The sum of products of corresponding vector entries, high when both vectors have large values in the same dimensions.
E
- Embedding
A vector representation of a word, so called because the word is embedded into a space.
- Extrinsic evaluation
Evaluating embeddings by their effect on a real downstream task such as MT or NER.
F
G
I
- Intrinsic evaluation
Evaluating embeddings on an intermediate task, such as correlation with human similarity ratings or TOEFL questions.
- Inverse document frequency
log10(N / df_t), which gives low weight to terms such as the that occur in nearly every document.
L
- Learning rate
The factor eta that scales each gradient descent step.
- Lexical semantics
The linguistic study of word meaning.
N
- Negative sampling
Drawing k random noise words from the unigram distribution raised to 0.75 as negative examples for each positive pair, instead of a softmax over the whole vocabulary.
P
- Parallelogram method
Solving a:a*::b:? by finding the word closest to a* - a + b, with a, a* and b excluded from the candidates.
- Pointwise mutual information
log2 of P(x,y) / (P(x)P(y)), measuring whether two words co-occur more than if they were independent.
- Positive PMI
PMI with negative values replaced by 0.
- Principle of contrast
The idea that any difference in linguistic form signals some difference in meaning.
S
- Self-supervision
Using naturally occurring neighbors in a corpus as gold labels, so no human labels are needed.
- Semantic field
A set of words that cover one semantic domain and bear structured relations to each other.
- Sigmoid
1 / (1 + exp(-x)), used to turn a dot product into the probability that c is a neighbor of w.
- Singular value decomposition
A matrix factorization used to get short dense vectors from a count matrix; LSA (latent semantic analysis) is a special case.
- Skip-gram
A word2vec model that predicts whether a candidate word is a neighbor of the target word.
- Skip-gram with negative sampling
Skip-gram trained to separate true (target, context) pairs from k randomly sampled noise pairs.
- Skip-thought vectors
A sentence-level analogue of skip-gram that encodes a sentence to predict its neighboring sentences.
- Sparse vector
A long vector (length |V|) whose elements are mostly zero, such as tf-idf or PMI vectors.
- Static embedding
One fixed vector per word type, as in word2vec or GloVe, unlike contextual embeddings.
- Stochastic gradient descent
Learning by repeatedly moving the weights opposite the gradient of the loss, scaled by the learning rate, to make positive pairs more likely and negative pairs less likely.
T
- Target and context matrices
The two embedding sets W and C learned by SGNS, often summed as w_i + c_i.
- Term frequency
The count of a term in a document, usually squashed as log10(count + 1).
- Term-context matrix
A word-word matrix counting how often each target word occurs near each context word.
- Term-document matrix
A matrix with one row per word and one column per document, each cell a count.
V
- Valence, arousal, dominance
Osgood's three affective dimensions: pleasantness, intensity of emotion, and degree of control.
- Vector length
The square root of the sum of a vector's squared entries.
- Vector semantics
Representing a word's meaning as a point in a multidimensional space built from its distribution.
W
- Word relatedness
Association between words related in any way, for example through a semantic field, such as coffee and cup.
- Word similarity
Words that are not synonyms but share some element of meaning, such as car and bicycle.