Majid Al-RaimiHyperparameters, word2vec variants and window size

ICS 582Lecture 04Part 09

Hyperparameters, word2vec variants and window size

Standard SGNS settings, the CBOW, GloVe, FastText and skip-thought alternatives, and how the context window size decides whether neighbors are syntactic or topical.

Concepts
6
Slides
89-97
Reading
36 min
Understood
0/6 concepts

Why this part matters

Picking k, d and the window size is the first thing you do when you train embeddings on a thesis corpus or an Arabic system, and the defaults you inherit from a library are not always the ones the papers recommend. CBOW, GloVe, FastText and skip-thought are the standard comparisons on an exam, and the window size quietly decides whether your vectors find synonyms or topics.

The part opens with the settings that make skip-gram work in practice and what each one trades. It then walks through four relatives of skip-gram, each changing one design decision: CBOW flips the prediction direction, GloVe replaces the sliding window with a global count matrix, FastText breaks words into character pieces, and skip-thought moves the whole idea from words to sentences. It closes with the one hyperparameter that changes not how good the vectors are but what kind of similarity they encode.

By the end you can

  1. State the standard SGNS hyperparameters and explain what raising k, d or the window costs and buys.
  2. Contrast CBOW with skip-gram by input, output, speed and rare-word quality.
  3. Write the GloVe objective and explain each part of its weighting function f.
  4. Decompose a word into FastText n-grams and explain how unseen words get vectors.
  5. Explain skip-thought vectors as skip-gram lifted to sentences.
  6. Predict whether a small or large window yields functional or topical neighbors.

Two students train skip-gram on the same afternoon. One has a clinical corpus of 5 million tokens; the other has 6 billion tokens of web text. Should they use the same number of negative samples? No. The small corpus gives each word only a few true context pairs, so each of those precious pairs has to do more work: contrasting it against many noise words, k = 15 to 20 (the paper allows 5 to 20), squeezes more signal out of it. On the huge corpus every word already sees thousands of real contexts, and k = 2 to 5 is enough while costing a fraction of the compute.

That is the pattern for every knob in skip-gram with negative sampling: each one trades quality against compute or memory, and the right value depends on how much data you have. The lecture gives a "somewhat standard" setting: the model is SGNS, 15 to 20 negative samples for smaller datasets and 2 to 5 for the huge datasets that are usually used, a dense vector of 300 dimensions (100 or 50 also work), and a sliding context window of 5 to 10.

The SGNS knobs, their usual values, and what raising each one buys and costs

Model
Skip-gram with negative sampling (SGNS). The alternatives in this part (CBOW, GloVe, FastText) are judged against it.
Negatives k
15 to 20 on small corpora per the slide (5 to 20 per Mikolov et al.), 2 to 5 on very large corpora. Raising k sharpens the contrast for each true pair and costs one more dot product per positive.
Dimension d
300 is the common choice; 100 or 50 also work. Raising d gives room for more distinctions but grows memory and every dot product linearly, with small gains past 300.
Window half-width m
5 to 10 on the slide, read as words per side, matching word2vec's and Gensim's window parameter (SLP3 says 1 to 10 per side). A ±m window gives 2m context words, which is part 07's L. Raising it adds more pairs per position and shifts neighbors from functional to topical (last concept of this part).
Noise distribution
Unigram counts raised to 3/4, the α of weighted PPMI, so rare words are drawn as negatives a little more often than their raw frequency.
Subsampling threshold t
Around 10^-5 in Mikolov et al.: very frequent words such as the are randomly dropped, which speeds training and helps rare words.

What each knob costs

The cost of negatives is easy to count. For every true (target, context) pair, the skip-gram classifier computes one dot product for the positive and one for each of the k noise words, and each dot product sends a gradient into one row of the context matrix:

dot products per positive pair=1+k\text{dot products per positive pair} = 1 + k

Going from k = 5 to k = 20 therefore makes training about 3.5 times slower (21 / 6). The dimension sets the memory. The model holds two matrices, the target and context matrices W and C, each with one row of length d per vocabulary word:

parameters=2 ∣V∣ d\text{parameters} = 2\,|V|\,d
A 100,000-word vocabulary at d = 300 needs 60 million parameters

Two smaller settings travel with these. Negatives are not drawn by raw frequency but from the unigram distribution raised to 3/4, the same α = 0.75 trick that weighted PPMI uses, which gives rare words a slightly better chance of being picked as noise. And very frequent words are randomly discarded with a threshold around 10^-5, so that the model does not spend most of its updates on pairs like (cat, the) (Mikolov et al. 2013b).

The defaults you meet in practice differ from the slide, which matters when you compare your results to a paper. Gensim, the usual Python library, defaults to CBOW (the variant explained in the next concept), not skip-gram, with 100 dimensions.

SettingdWindowkModel
Mikolov et al. 2013b experiments30055 to 20 (small data)skip-gram
Gensim Word2Vec10055CBOW (sg=0)
fastText10055skip-gram
This slide300 (or 100, 50)5 to 1015 to 20 (small), 2 to 5 (huge)SGNS
Where the numbers come from

Quick check

According to the lecture, which range of k suits a small training corpus?

Recall

State the standard SGNS hyperparameters from the lecture.

Model SGNS. Negatives 15 to 20 on small data (Mikolov says 5 to 20) and 2 to 5 on huge data. Dimension 300, with 100 or 50 also possible. Window 5 to 10.

Take the sentence "I saw a cute grey cat playing in the garden" with a ±2 window around cat. Skip-gram, the model of the last two parts, turns this position into four separate training pairs: (cat, cute), (cat, grey), (cat, playing), (cat, in). The continuous bag of words model, the other half of word2vec, runs the arrow the other way. It looks up the four context vectors for cute, grey, playing and in, combines them into a single hidden vector h, scores every vocabulary word against h, and is trained so that the highest score goes to cat. One position, one prediction.

Same five words, opposite arrows. CBOW draws the four context words into one prediction of cat; skip-gram sends cat out to predict each of the four context words separately.

The combination step is what gives CBOW its name. The context vectors are pooled into one, which throws away their order (a bag), and they are dense real-valued vectors rather than counts (continuous). In Mikolov et al.'s original description the projection layer is shared so that "all words get projected into the same position (their vectors are averaged)", and "the order of words in the history does not influence the projection":

h=12m∑−m≤j≤m, j≠0wt+j\mathbf{h}=\tfrac{1}{2m}\sum_{-m\le j\le m,\,j\ne 0}\mathbf{w}_{t+j}
The CBOW hidden vector: the mean of the 2m context vectors around position t, for a ±m window

Worked example

One sentence, two models

  1. Pick the window

    Target position: cat. With m = 2, the context words are cute, grey (left) and playing, in (right).

  2. CBOW builds one training instance

    h = (w_cute + w_grey + w_playing + w_in) / 4. The model scores h against the output vectors and the loss pushes the score of cat up. Four vectors go in, one gradient signal comes out, and it is shared equally by the four context words.

  3. Skip-gram builds four

    The pairs (cat, cute), (cat, grey), (cat, playing) and (cat, in) are each scored and updated separately, each with its own k negatives. The center vector of cat receives four updates.

  4. Result

    CBOW does one prediction per position; skip-gram does 2m = 4. That factor is why CBOW trains faster and why skip-gram gives each word, including rare ones, more direct updates.

The consequences follow from that count. With negative sampling, CBOW does 1 + k dot products per position and skip-gram 2m × (1 + k), so skip-gram is about 2m times slower (Mikolov et al. 2013a report the same gap with hierarchical softmax). CBOW's averaging also smooths: a rare word that appears as context is blended with its frequent neighbors before any prediction is made, so its own vector gets a diluted signal. Skip-gram gives every occurrence of a rare word its own pairs, which is why it handles rare words better and in Mikolov et al.'s experiments did better on semantic analogies. The fastText tutorial adds a practical note: skip-gram "works better with subword information than cbow".

CBOWSkip-gram
InputThe bag of 2m context vectors, averaged into one hOne center vector
OutputA score for the center wordA score for each context word, one pair at a time
Predictions per position12m
Dot products per position (SGNS)1 + k2m × (1 + k)
Word order inside the windowIgnoredIgnored (each pair is independent)
Rare wordsSmoothed away by averaging with frequent neighborsEach occurrence gives its own updates; better
StrengthSpeed on large corpora, frequent wordsRare words, semantic analogies, subword extensions
CBOW against skip-gram

Quick check

In CBOW, what is the input and what is the prediction target?

Recall

In one sentence each, what do CBOW and skip-gram predict, and which is better for rare words?

CBOW predicts the center word from the averaged (or summed) bag of context vectors. Skip-gram predicts each context word from the center word. Skip-gram is better for rare words, because each occurrence gets its own updates instead of being averaged with frequent neighbors.

Consider three cells of a word-word co-occurrence matrix. The pair (ice, solid) was seen 10 times, (the, of) 1000 times, and (ice, fashion) never. GloVe wants a dot product for each pair that matches the log of its count: about log 10 ≈ 2.30 for (ice, solid) and log 1000 ≈ 6.91 for (the, of). It does not trust every cell equally. With the published settings, the ice and solid pair gets weight 0.178, the pair the and of gets the maximum weight 1 and no more, and the ice and fashion pair gets weight 0, so its undefined log 0 never enters the loss.

That is the whole method. GloVe, Global Vectors (Pennington, Socher and Manning 2014), first counts the term-context matrix N(w, c) over the corpus, the same matrix that PPMI reweights. It then fits a Dot product of a word vector and a context vector, plus two learned biases, to the log count in every cell, as a weighted least-squares regression:

J(θ)=∑w,cf(N(w,c)) (uc⊤vw+bc+bˉw−log⁡N(w,c))2\begin{aligned} J(\theta)=\sum_{w,c} f\big(N(w,c)\big)\,\big(&\mathbf{u}_c^\top\mathbf{v}_w \\ &+b_c+\bar b_w \\ &-\log N(w,c)\big)^2 \end{aligned}
The GloVe objective as on the slide: context vector u_c, word vector v_w, biases b_c and b̄_w

The weighting function f is where the design lives. Pennington et al. ask three things of it. It must vanish at zero, because log 0 is undefined and because zero cells are 75 to 95 percent of the matrix. It must not decrease, "so that rare co-occurrences are not overweighted", since a pair seen once is mostly noise. And it must stay "relatively small for large values of x, so that frequent co-occurrences are not overweighted". Their choice is a power curve that flattens into a cap:

f(x)={(x/xmax⁡)αx<xmax⁡1otherwisexmax⁡=100, α=34\begin{gathered} f(x)=\begin{cases}(x/x_{\max})^{\alpha} & x<x_{\max}\\ 1 & \text{otherwise}\end{cases} \\ x_{\max}=100,\ \alpha=\tfrac{3}{4} \end{gathered}
f(x) = (x/100)^0.75 rises to the dashed x_max marker and stays flat at 1. The bars show the weights of pairs seen 1, 10, 50 and 1000 times: 0.03, 0.18, 0.59 and 1.
Count xWeight f(x)Example
10.0316A single co-occurrence: barely trusted
50.1057
100.1778(ice, solid)
500.5946
1001x = x_max: full weight reached
10001(the, of): capped, cannot dominate
f(x) at the published settings

Worked example

Three pairs through the loss

  1. (ice, solid), 10 co-occurrences

    Target log 10 ≈ 2.30. Weight (10/100)^0.75 ≈ 0.178. If the current prediction u·v + b + b̄ is 1.30, the term contributes 0.178 × 1.0² = 0.178.

  2. (the, of), 1000 co-occurrences

    Target log 1000 ≈ 6.91. The count is above x_max, so the weight is capped at 1. Without the cap, the same curve would give (1000/100)^0.75 ≈ 5.6, and this one function-word pair would count as much as about 32 pairs like (ice, solid).

  3. (ice, fashion), 0 co-occurrences

    Weight f(0) = 0. The term vanishes, so the undefined log 0 is never evaluated and the optimizer only visits the nonzero cells.

  4. Result

    Rare pairs count a little, mid-frequency pairs count a lot, and very frequent pairs are capped. The slide's sum over all w, c ∈ V silently relies on f(0) = 0 to skip the empty cells.

GloVe sits between the two families in this lecture. Like PPMI it is built on global matrix statistics collected once, and like word2vec it learns dense vectors whose dot products carry the meaning. Jurafsky and Martin describe it as "based on ratios of probabilities from the word-word co-occurrence matrix", capturing global corpus statistics as count methods do while learning dense vectors as word2vec does. Like SGNS it ends with two vectors per word, and the authors use the sum W + W̃ as the final embedding, the same trick as adding w and c in skip-gram. The slide's b̄_w is just the second bias set, written b̃_j in the paper. Pennington et al. also chose 300 dimensions for their main results, and they report that α = 3/4 gave "a modest improvement over a linear version with α = 1".

SGNSGloVe
Data it readsA stream of (target, context) pairs from sliding windowsThe global co-occurrence matrix, counted once
ObjectiveLogistic loss: true pairs up, sampled noise pairs downWeighted squared error between u·v + biases and log N(w,c)
Negativesk sampled per positive pairNone: zero cells get weight f(0) = 0 and drop out
Frequency controlNoise drawn from U(w)^(3/4), subsampling of frequent wordsWeight f(x) = (x/x_max)^(3/4), capped at 1
Final vectorsTarget matrix W, or W + CW + W̃ (target plus context)
SGNS against GloVe

Recall

Why does GloVe's weighting function need f(0) = 0, and what does capping f at 1 above x_max do?

log 0 is undefined, and zero cells are 75 to 95 percent of the matrix, so f(0) = 0 drops them. The cap stops very frequent pairs such as (the, of) from dominating the loss. Below x_max, rare noisy pairs get small weights.

Take the word where. FastText first wraps it in boundary symbols, <where>, so that prefixes and suffixes can be told apart from the middle of a word. With n = 3 it then slides a three-character window across: <wh, whe, her, ere, re>. It also keeps the whole word <where> as one more unit. Each of those six units has its own vector, and the vector of where is their sum.

A three-character window slides across <where>, emitting <wh, whe, her, ere and re>; the whole-word unit <where> joins them, and the six vectors sum into one word vector.

Bojanowski et al. 2017 built this on the skip-gram model. Write 𝒢_w for the set of n-grams of word w, including the word itself, and z_g for the vector of n-gram g. The word's vector and its skip-gram score with a context word c become:

uw=∑g∈Gwzg\mathbf{u}_w=\sum_{g\in\mathcal{G}_w}\mathbf{z}_g
s(w,c)=∑g∈Gwzg⊤vcs(w,c)=\sum_{g\in\mathcal{G}_w}\mathbf{z}_g^\top\mathbf{v}_c
The skip-gram score with subword information: every n-gram of the target takes part in every dot product

Worked example

All the n-grams of where at the real settings

  1. Pad

    where becomes <where>, 7 characters.

  2. n = 3 (5 n-grams)

    <wh whe her ere re>

  3. n = 4, 5 and 6 (4 + 3 + 2 n-grams)

    <whe wher here ere>, then <wher where here>, then <where where>.

  4. Add the special whole-word sequence

    <where> is added as its own unit, distinct from the n-gram where found inside it.

  5. Result

    The paper extracts "all the n-grams for n greater or equal to 3 and smaller or equal to 6", which gives 14 n-grams here, plus the whole word: 15 vectors summed into one.

The boundary symbols matter more than they look. The paper's own example: "the sequence <her>, corresponding to the word her, is different from the tri-gram her from the word where". So the pronoun and the piece of where get separate vectors, while the prefix <wh is shared by where, when, what and which.

What the pieces buy

Plain word2vec gives each word type its own row, so a word missing from training has no vector at all. FastText builds a vector for an unseen word such as wherever by summing the vectors of its n-grams, most of which (<wh, whe, her, ere) were learned from other words. The same sharing helps rare inflections in morphologically rich languages such as Arabic, Turkish and German: a rarely seen form borrows statistics from frequent forms that share its stem and affixes. Pretrained FastText vectors exist for 157 languages.

The price is computation, which the slide calls "a lot of additional computation": every update now touches about fifteen vectors for the target instead of one. Memory is kept bounded by hashing all n-grams into a fixed table of 2,000,000 buckets (the bucket default), so collisions are allowed rather than storing every possible substring.

Quick check

FastText meets the unseen word 'wherever' at test time. Which vector does it return?

Recall

List the FastText 3-gram units for 'where' and explain how FastText builds a vector for an unseen word.

<wh, whe, her, ere, re>, plus the whole-word token <where>. In practice n runs from 3 to 6. An unseen word gets the sum of the vectors of its known n-grams. The word <her> differs from the trigram her inside where.

Take three consecutive sentences from a novel: "I got back home. I could see the cat on the steps. This was strange." A skip-thought model reads the middle sentence and compresses it into one vector. Two decoders then have to regenerate the neighbors from that vector alone, word by word: one writes "I got back home", the other writes "This was strange".

Prev decoder

Generates the previous sentence, "I got back home <eos>", conditioned on the vector.

sentence vector
Encoder (GRU)

Reads "I could see the cat on the steps" and outputs one sentence vector.

sentence vector
Next decoder

Generates the next sentence, "This was strange <eos>", conditioned on the vector.

Skip-thought on the triplet from Kiros et al. 2015: one encoder, two decoders, the previous and next sentences as targets; both decoders read the same vector in parallel

The idea is skip-gram moved up one level. Skip-gram uses a word to predict its neighboring words; Kiros et al. 2015 use a sentence to predict its neighboring sentences. The distributional hypothesis carries over: sentences that appear in similar discourse surroundings, with similar sentences before and after them, are pushed toward similar vectors. In the paper's words, "the sentence s_i is encoded and tries to reconstruct the previous sentences_i−1 and next sentence s_i+1".

Three things change on the way up. A sentence is not one vocabulary item that can be looked up, so the encoder is a recurrent network with GRU units that reads the words in order, which makes the vector order-sensitive. The targets are whole sentences, so the decoders generate text instead of classifying pairs, and there is no negative sampling. And the vectors are big: the model was trained on the BookCorpus (11,038 books, 74,004,228 sentences), the unidirectional encoder gives 2400 dimensions and the combined model 4800. The authors froze these vectors and trained only linear classifiers on top of them for 8 evaluation tasks.

Skip-gramSkip-thought
UnitA wordA sentence
ContextWords within ±mThe previous and the next sentence
EncoderA lookup in W (one row per word)A GRU recurrent network over the words, order-sensitive
ObjectiveClassify (target, context) pairs as real or noiseGenerate each neighbor sentence word by word
Negativesk sampled per positiveNone: decoders use a softmax over the vocabulary
Output sized, typically 100 to 3002400 (uni-skip), 4800 (combine-skip)
Skip-gram against skip-thought

Recall

What does a skip-thought model encode and what does it predict?

A GRU encoder turns a sentence into one vector. Two decoders regenerate the previous and the next sentence from that vector. The result is a 2400-dimensional sentence vector, used as frozen features.

Properties of embeddings start with the window. Levy and Goldberg 2014 trained skip-gram on English Wikipedia twice, changing only the window, and looked up the nearest neighbors of Hogwarts. With a ±2 window they were evernight, sunnydale, garderobe, blandings and collinwood: other fictional schools and houses, words that fill the same slot in a sentence. With a ±5 window they were dumbledore, hallows, half-blood, malfoy and snape: the world of Harry Potter. Same corpus, same model, a different idea of what "similar" means.

Two models trained with different windows. With ±2 Hogwarts' nearest neighbors are Sunnydale, Evernight, Garderobe and Blandings; retrained with ±5 they become Dumbledore, Malfoy, half-blood and Snape. The frames stand for the training window, not distance in the vector space.

The reason is what each window can see. Two words on each side of a noun are mostly its syntactic frame: the determiner before it, the preposition, the verb that takes it as an object ("students at Hogwarts", "returned to Sunnydale"). Words that share those frames are words of the same type and function, so a short window rewards similarity. Linguists call this a paradigmatic, or second-order, association: the two words rarely appear together but appear in the same surroundings. Five words on each side reach past the frame into the topic of the passage, so a long window rewards relatedness, words from the same semantic field. That is a syntagmatic, or first-order, association: the words appear near each other.

The slides show the same split with Voita's examples. With larger windows, dog groups with bark and leash, and walking with walked and run: words about the same activity. With smaller windows, Poodle groups with Pitbull and Rottweiler, and walking with running and approaching: words that could replace one another. Jurafsky and Martin summarize it as shorter windows giving representations that are "a bit more syntactic", with neighbors that are "semantically similar words with the same parts of speech", while longer windows give words that are "topically related but not similar".

Small windowLarge window
Window±2 (small)±5 or more (large)
What the window seesThe syntactic frame: determiners, prepositions, the verb that takes the wordThe whole topic of the passage
Neighbor kindFunctional similarity: same slot, same part of speechTopical relatedness: same semantic field
Hogwarts (Levy and Goldberg)sunnydale, evernight, garderobe, blandings, collinwooddumbledore, hallows, half-blood, malfoy, snape
Dog example (Voita)Poodle, Pitbull, Rottweilerdog, bark, leash
Verb example (Voita)walking, running, approachingwalking, walked, run
Good forSynonym finding, POS-like features, slot fillingTopic modelling, retrieval, query expansion
Small against large windows

The same logic explains the window in the term-context matrix of the count methods: it is one hyperparameter shared by PPMI, SGNS and GloVe, and it changes what all of them learn. Levy and Goldberg push it one step further by replacing the window with syntactic dependency contexts. Hogwarts' neighbors then become sunnydale, collinwood, calarts, greendale and millfield: even more purely functional, all schools and fictional places.

Quick check

Skip-gram is trained with a context window of plus or minus 2. Which neighbors does Hogwarts get?

Recall

With a ±2 versus a ±5 window, what are Hogwarts' nearest neighbors, and why?

±2 gives other fictional schools (Sunnydale, Evernight, Blandings), because narrow windows capture syntactic slot or function. ±5 gives the Harry Potter world (Dumbledore, half-blood, Malfoy), because wide windows capture topic.

Recap

If you remember nothing else

  • Standard SGNS: k = 15 to 20 on small data (5 to 20 in Mikolov), 2 to 5 on huge data, d = 300 (or 100, 50), window 5 to 10.
  • Each extra negative adds one dot product per positive pair; d sets the 2|V|d parameter count.
  • CBOW predicts the center from the averaged context bag: faster, smoother. Skip-gram predicts the context from the center: better for rare words.
  • GloVe fits u_c·v_w + b_c + b_w to log N(w,c) by weighted least squares; f(x) = (x/100)^0.75 below 100, then 1, and f(0) = 0 skips zero cells.
  • FastText adds < and > boundaries, n-grams of 3 to 6 characters and the whole word; a word is the sum of its n-gram vectors, so unseen words still get vectors.
  • Skip-thought encodes a sentence with a GRU and decodes the previous and next sentences, giving 2400-dimensional sentence vectors.
  • A ±2 window yields functional look-alikes (Hogwarts near Sunnydale); a ±5 window yields topic-mates (Hogwarts near Dumbledore).

Sources