ICS 582Lecture 04Part 09
Hyperparameters, word2vec variants and window size
Standard SGNS settings, the CBOW, GloVe, FastText and skip-thought alternatives, and how the context window size decides whether neighbors are syntactic or topical.
- Concepts
- 6
- Slides
- 89-97
- Reading
- 36 min
Why this part matters
Picking k, d and the window size is the first thing you do when you train embeddings on a thesis corpus or an Arabic system, and the defaults you inherit from a library are not always the ones the papers recommend. CBOW, GloVe, FastText and skip-thought are the standard comparisons on an exam, and the window size quietly decides whether your vectors find synonyms or topics.
The part opens with the settings that make skip-gram work in practice and what each one trades. It then walks through four relatives of skip-gram, each changing one design decision: CBOW flips the prediction direction, GloVe replaces the sliding window with a global count matrix, FastText breaks words into character pieces, and skip-thought moves the whole idea from words to sentences. It closes with the one hyperparameter that changes not how good the vectors are but what kind of similarity they encode.
By the end you can
- State the standard SGNS hyperparameters and explain what raising k, d or the window costs and buys.
- Contrast CBOW with skip-gram by input, output, speed and rare-word quality.
- Write the GloVe objective and explain each part of its weighting function f.
- Decompose a word into FastText n-grams and explain how unseen words get vectors.
- Explain skip-thought vectors as skip-gram lifted to sentences.
- Predict whether a small or large window yields functional or topical neighbors.
Two students train skip-gram on the same afternoon. One has a clinical corpus of 5 million tokens; the other has 6 billion tokens of web text. Should they use the same number of negative samples? No. The small corpus gives each word only a few true context pairs, so each of those precious pairs has to do more work: contrasting it against many noise words, k = 15 to 20 (the paper allows 5 to 20), squeezes more signal out of it. On the huge corpus every word already sees thousands of real contexts, and k = 2 to 5 is enough while costing a fraction of the compute.
That is the pattern for every knob in skip-gram with negative sampling: each one trades quality against compute or memory, and the right value depends on how much data you have. The lecture gives a "somewhat standard" setting: the model is SGNS, 15 to 20 negative samples for smaller datasets and 2 to 5 for the huge datasets that are usually used, a dense vector of 300 dimensions (100 or 50 also work), and a sliding context window of 5 to 10.
The SGNS knobs, their usual values, and what raising each one buys and costs
- Model
- Skip-gram with negative sampling (SGNS). The alternatives in this part (CBOW, GloVe, FastText) are judged against it.
- Negatives k
- 15 to 20 on small corpora per the slide (5 to 20 per Mikolov et al.), 2 to 5 on very large corpora. Raising k sharpens the contrast for each true pair and costs one more dot product per positive.
- Dimension d
- 300 is the common choice; 100 or 50 also work. Raising d gives room for more distinctions but grows memory and every dot product linearly, with small gains past 300.
- Window half-width m
- 5 to 10 on the slide, read as words per side, matching word2vec's and Gensim's window parameter (SLP3 says 1 to 10 per side). A ±m window gives 2m context words, which is part 07's L. Raising it adds more pairs per position and shifts neighbors from functional to topical (last concept of this part).
- Noise distribution
- Unigram counts raised to 3/4, the α of weighted PPMI, so rare words are drawn as negatives a little more often than their raw frequency.
- Subsampling threshold t
- Around 10^-5 in Mikolov et al.: very frequent words such as the are randomly dropped, which speeds training and helps rare words.
What each knob costs
The cost of negatives is easy to count. For every true (target, context) pair, the skip-gram classifier computes one dot product for the positive and one for each of the k noise words, and each dot product sends a gradient into one row of the context matrix:
Going from k = 5 to k = 20 therefore makes training about 3.5 times slower (21 / 6). The dimension sets the memory. The model holds two matrices, the target and context matrices W and C, each with one row of length d per vocabulary word:
Two smaller settings travel with these. Negatives are not drawn by raw frequency but from the unigram distribution raised to 3/4, the same α = 0.75 trick that weighted PPMI uses, which gives rare words a slightly better chance of being picked as noise. And very frequent words are randomly discarded with a threshold around 10^-5, so that the model does not spend most of its updates on pairs like (cat, the) (Mikolov et al. 2013b).
The defaults you meet in practice differ from the slide, which matters when you compare your results to a paper. Gensim, the usual Python library, defaults to CBOW (the variant explained in the next concept), not skip-gram, with 100 dimensions.
| Setting | d | Window | k | Model |
|---|---|---|---|---|
| Mikolov et al. 2013b experiments | 300 | 5 | 5 to 20 (small data) | skip-gram |
| Gensim Word2Vec | 100 | 5 | 5 | CBOW (sg=0) |
| fastText | 100 | 5 | 5 | skip-gram |
| This slide | 300 (or 100, 50) | 5 to 10 | 15 to 20 (small), 2 to 5 (huge) | SGNS |
Quick check
According to the lecture, which range of k suits a small training corpus?
Recall
State the standard SGNS hyperparameters from the lecture.
Take the sentence "I saw a cute grey cat playing in the garden" with a ±2 window around cat. Skip-gram, the model of the last two parts, turns this position into four separate training pairs: (cat, cute), (cat, grey), (cat, playing), (cat, in). The continuous bag of words model, the other half of word2vec, runs the arrow the other way. It looks up the four context vectors for cute, grey, playing and in, combines them into a single hidden vector h, scores every vocabulary word against h, and is trained so that the highest score goes to cat. One position, one prediction.
The combination step is what gives CBOW its name. The context vectors are pooled into one, which throws away their order (a bag), and they are dense real-valued vectors rather than counts (continuous). In Mikolov et al.'s original description the projection layer is shared so that "all words get projected into the same position (their vectors are averaged)", and "the order of words in the history does not influence the projection":
Worked example
One sentence, two models
Pick the window
Target position: cat. With m = 2, the context words are cute, grey (left) and playing, in (right).
CBOW builds one training instance
h = (w_cute + w_grey + w_playing + w_in) / 4. The model scores h against the output vectors and the loss pushes the score of cat up. Four vectors go in, one gradient signal comes out, and it is shared equally by the four context words.
Skip-gram builds four
The pairs (cat, cute), (cat, grey), (cat, playing) and (cat, in) are each scored and updated separately, each with its own k negatives. The center vector of cat receives four updates.
Result
CBOW does one prediction per position; skip-gram does 2m = 4. That factor is why CBOW trains faster and why skip-gram gives each word, including rare ones, more direct updates.
The consequences follow from that count. With negative sampling, CBOW does 1 + k dot products per position and skip-gram 2m × (1 + k), so skip-gram is about 2m times slower (Mikolov et al. 2013a report the same gap with hierarchical softmax). CBOW's averaging also smooths: a rare word that appears as context is blended with its frequent neighbors before any prediction is made, so its own vector gets a diluted signal. Skip-gram gives every occurrence of a rare word its own pairs, which is why it handles rare words better and in Mikolov et al.'s experiments did better on semantic analogies. The fastText tutorial adds a practical note: skip-gram "works better with subword information than cbow".
| CBOW | Skip-gram | |
|---|---|---|
| Input | The bag of 2m context vectors, averaged into one h | One center vector |
| Output | A score for the center word | A score for each context word, one pair at a time |
| Predictions per position | 1 | 2m |
| Dot products per position (SGNS) | 1 + k | 2m × (1 + k) |
| Word order inside the window | Ignored | Ignored (each pair is independent) |
| Rare words | Smoothed away by averaging with frequent neighbors | Each occurrence gives its own updates; better |
| Strength | Speed on large corpora, frequent words | Rare words, semantic analogies, subword extensions |
Quick check
In CBOW, what is the input and what is the prediction target?
Recall
In one sentence each, what do CBOW and skip-gram predict, and which is better for rare words?
Consider three cells of a word-word co-occurrence matrix. The pair (ice, solid) was seen 10 times, (the, of) 1000 times, and (ice, fashion) never. GloVe wants a dot product for each pair that matches the log of its count: about log 10 ≈ 2.30 for (ice, solid) and log 1000 ≈ 6.91 for (the, of). It does not trust every cell equally. With the published settings, the ice and solid pair gets weight 0.178, the pair the and of gets the maximum weight 1 and no more, and the ice and fashion pair gets weight 0, so its undefined log 0 never enters the loss.
That is the whole method. GloVe, Global Vectors (Pennington, Socher and Manning 2014), first counts the term-context matrix N(w, c) over the corpus, the same matrix that PPMI reweights. It then fits a Dot product of a word vector and a context vector, plus two learned biases, to the log count in every cell, as a weighted least-squares regression:
The weighting function f is where the design lives. Pennington et al. ask three things of it. It must vanish at zero, because log 0 is undefined and because zero cells are 75 to 95 percent of the matrix. It must not decrease, "so that rare co-occurrences are not overweighted", since a pair seen once is mostly noise. And it must stay "relatively small for large values of x, so that frequent co-occurrences are not overweighted". Their choice is a power curve that flattens into a cap:
| Count x | Weight f(x) | Example |
|---|---|---|
| 1 | 0.0316 | A single co-occurrence: barely trusted |
| 5 | 0.1057 | |
| 10 | 0.1778 | (ice, solid) |
| 50 | 0.5946 | |
| 100 | 1 | x = x_max: full weight reached |
| 1000 | 1 | (the, of): capped, cannot dominate |
Worked example
Three pairs through the loss
(ice, solid), 10 co-occurrences
Target log 10 ≈ 2.30. Weight (10/100)^0.75 ≈ 0.178. If the current prediction u·v + b + b̄ is 1.30, the term contributes 0.178 × 1.0² = 0.178.
(the, of), 1000 co-occurrences
Target log 1000 ≈ 6.91. The count is above x_max, so the weight is capped at 1. Without the cap, the same curve would give (1000/100)^0.75 ≈ 5.6, and this one function-word pair would count as much as about 32 pairs like (ice, solid).
(ice, fashion), 0 co-occurrences
Weight f(0) = 0. The term vanishes, so the undefined log 0 is never evaluated and the optimizer only visits the nonzero cells.
Result
Rare pairs count a little, mid-frequency pairs count a lot, and very frequent pairs are capped. The slide's sum over all w, c ∈ V silently relies on f(0) = 0 to skip the empty cells.
GloVe sits between the two families in this lecture. Like PPMI it is built on global matrix statistics collected once, and like word2vec it learns dense vectors whose dot products carry the meaning. Jurafsky and Martin describe it as "based on ratios of probabilities from the word-word co-occurrence matrix", capturing global corpus statistics as count methods do while learning dense vectors as word2vec does. Like SGNS it ends with two vectors per word, and the authors use the sum W + W̃ as the final embedding, the same trick as adding w and c in skip-gram. The slide's b̄_w is just the second bias set, written b̃_j in the paper. Pennington et al. also chose 300 dimensions for their main results, and they report that α = 3/4 gave "a modest improvement over a linear version with α = 1".
| SGNS | GloVe | |
|---|---|---|
| Data it reads | A stream of (target, context) pairs from sliding windows | The global co-occurrence matrix, counted once |
| Objective | Logistic loss: true pairs up, sampled noise pairs down | Weighted squared error between u·v + biases and log N(w,c) |
| Negatives | k sampled per positive pair | None: zero cells get weight f(0) = 0 and drop out |
| Frequency control | Noise drawn from U(w)^(3/4), subsampling of frequent words | Weight f(x) = (x/x_max)^(3/4), capped at 1 |
| Final vectors | Target matrix W, or W + C | W + W̃ (target plus context) |
Recall
Why does GloVe's weighting function need f(0) = 0, and what does capping f at 1 above x_max do?
Take the word where. FastText first wraps it in boundary symbols, <where>, so that prefixes and suffixes can be told apart from the middle of a word. With n = 3 it then slides a three-character window across: <wh, whe, her, ere, re>. It also keeps the whole word <where> as one more unit. Each of those six units has its own vector, and the vector of where is their sum.
Bojanowski et al. 2017 built this on the skip-gram model. Write 𝒢_w for the set of n-grams of word w, including the word itself, and z_g for the vector of n-gram g. The word's vector and its skip-gram score with a context word c become:
Worked example
All the n-grams of where at the real settings
Pad
where becomes <where>, 7 characters.
n = 3 (5 n-grams)
<wh whe her ere re>
n = 4, 5 and 6 (4 + 3 + 2 n-grams)
<whe wher here ere>, then <wher where here>, then <where where>.
Add the special whole-word sequence
<where> is added as its own unit, distinct from the n-gram where found inside it.
Result
The paper extracts "all the n-grams for n greater or equal to 3 and smaller or equal to 6", which gives 14 n-grams here, plus the whole word: 15 vectors summed into one.
The boundary symbols matter more than they look. The paper's own example: "the sequence <her>, corresponding to the word her, is different from the tri-gram her from the word where". So the pronoun and the piece of where get separate vectors, while the prefix <wh is shared by where, when, what and which.
What the pieces buy
Plain word2vec gives each word type its own row, so a word missing from training has no vector at all. FastText builds a vector for an unseen word such as wherever by summing the vectors of its n-grams, most of which (<wh, whe, her, ere) were learned from other words. The same sharing helps rare inflections in morphologically rich languages such as Arabic, Turkish and German: a rarely seen form borrows statistics from frequent forms that share its stem and affixes. Pretrained FastText vectors exist for 157 languages.
The price is computation, which the slide calls "a lot of additional computation": every update now touches about fifteen vectors for the target instead of one. Memory is kept bounded by hashing all n-grams into a fixed table of 2,000,000 buckets (the bucket default), so collisions are allowed rather than storing every possible substring.
Quick check
FastText meets the unseen word 'wherever' at test time. Which vector does it return?
Recall
List the FastText 3-gram units for 'where' and explain how FastText builds a vector for an unseen word.
Take three consecutive sentences from a novel: "I got back home. I could see the cat on the steps. This was strange." A skip-thought model reads the middle sentence and compresses it into one vector. Two decoders then have to regenerate the neighbors from that vector alone, word by word: one writes "I got back home", the other writes "This was strange".
Generates the previous sentence, "I got back home <eos>", conditioned on the vector.
Reads "I could see the cat on the steps" and outputs one sentence vector.
Generates the next sentence, "This was strange <eos>", conditioned on the vector.
The idea is skip-gram moved up one level. Skip-gram uses a word to predict its neighboring words; Kiros et al. 2015 use a sentence to predict its neighboring sentences. The distributional hypothesis carries over: sentences that appear in similar discourse surroundings, with similar sentences before and after them, are pushed toward similar vectors. In the paper's words, "the sentence s_i is encoded and tries to reconstruct the previous sentences_i−1 and next sentence s_i+1".
Three things change on the way up. A sentence is not one vocabulary item that can be looked up, so the encoder is a recurrent network with GRU units that reads the words in order, which makes the vector order-sensitive. The targets are whole sentences, so the decoders generate text instead of classifying pairs, and there is no negative sampling. And the vectors are big: the model was trained on the BookCorpus (11,038 books, 74,004,228 sentences), the unidirectional encoder gives 2400 dimensions and the combined model 4800. The authors froze these vectors and trained only linear classifiers on top of them for 8 evaluation tasks.
| Skip-gram | Skip-thought | |
|---|---|---|
| Unit | A word | A sentence |
| Context | Words within ±m | The previous and the next sentence |
| Encoder | A lookup in W (one row per word) | A GRU recurrent network over the words, order-sensitive |
| Objective | Classify (target, context) pairs as real or noise | Generate each neighbor sentence word by word |
| Negatives | k sampled per positive | None: decoders use a softmax over the vocabulary |
| Output size | d, typically 100 to 300 | 2400 (uni-skip), 4800 (combine-skip) |
Recall
What does a skip-thought model encode and what does it predict?
Properties of embeddings start with the window. Levy and Goldberg 2014 trained skip-gram on English Wikipedia twice, changing only the window, and looked up the nearest neighbors of Hogwarts. With a ±2 window they were evernight, sunnydale, garderobe, blandings and collinwood: other fictional schools and houses, words that fill the same slot in a sentence. With a ±5 window they were dumbledore, hallows, half-blood, malfoy and snape: the world of Harry Potter. Same corpus, same model, a different idea of what "similar" means.
The reason is what each window can see. Two words on each side of a noun are mostly its syntactic frame: the determiner before it, the preposition, the verb that takes it as an object ("students at Hogwarts", "returned to Sunnydale"). Words that share those frames are words of the same type and function, so a short window rewards similarity. Linguists call this a paradigmatic, or second-order, association: the two words rarely appear together but appear in the same surroundings. Five words on each side reach past the frame into the topic of the passage, so a long window rewards relatedness, words from the same semantic field. That is a syntagmatic, or first-order, association: the words appear near each other.
The slides show the same split with Voita's examples. With larger windows, dog groups with bark and leash, and walking with walked and run: words about the same activity. With smaller windows, Poodle groups with Pitbull and Rottweiler, and walking with running and approaching: words that could replace one another. Jurafsky and Martin summarize it as shorter windows giving representations that are "a bit more syntactic", with neighbors that are "semantically similar words with the same parts of speech", while longer windows give words that are "topically related but not similar".
| Small window | Large window | |
|---|---|---|
| Window | ±2 (small) | ±5 or more (large) |
| What the window sees | The syntactic frame: determiners, prepositions, the verb that takes the word | The whole topic of the passage |
| Neighbor kind | Functional similarity: same slot, same part of speech | Topical relatedness: same semantic field |
| Hogwarts (Levy and Goldberg) | sunnydale, evernight, garderobe, blandings, collinwood | dumbledore, hallows, half-blood, malfoy, snape |
| Dog example (Voita) | Poodle, Pitbull, Rottweiler | dog, bark, leash |
| Verb example (Voita) | walking, running, approaching | walking, walked, run |
| Good for | Synonym finding, POS-like features, slot filling | Topic modelling, retrieval, query expansion |
The same logic explains the window in the term-context matrix of the count methods: it is one hyperparameter shared by PPMI, SGNS and GloVe, and it changes what all of them learn. Levy and Goldberg push it one step further by replacing the window with syntactic dependency contexts. Hogwarts' neighbors then become sunnydale, collinwood, calarts, greendale and millfield: even more purely functional, all schools and fictional places.
Quick check
Skip-gram is trained with a context window of plus or minus 2. Which neighbors does Hogwarts get?
Recall
With a ±2 versus a ±5 window, what are Hogwarts' nearest neighbors, and why?
Recap
If you remember nothing else
- Standard SGNS: k = 15 to 20 on small data (5 to 20 in Mikolov), 2 to 5 on huge data, d = 300 (or 100, 50), window 5 to 10.
- Each extra negative adds one dot product per positive pair; d sets the 2|V|d parameter count.
- CBOW predicts the center from the averaged context bag: faster, smoother. Skip-gram predicts the context from the center: better for rare words.
- GloVe fits u_c·v_w + b_c + b_w to log N(w,c) by weighted least squares; f(x) = (x/100)^0.75 below 100, then 1, and f(0) = 0 skips zero cells.
- FastText adds < and > boundaries, n-grams of 3 to 6 characters and the whole word; a word is the sum of its n-gram vectors, so unseen words still get vectors.
- Skip-thought encodes a sentence with a GRU and decodes the previous and next sentences, giving 2400-dimensional sentence vectors.
- A ±2 window yields functional look-alikes (Hogwarts near Sunnydale); a ±5 window yields topic-mates (Hogwarts near Dumbledore).
Sources
- Speech and Language Processing, chapter 5: EmbeddingsBookJurafsky and Martin, draft of August 2026Window of 1 to 10 per side, short versus long windows, the Hogwarts example, first- and second-order co-occurrence, GloVe and FastText summaries(opens in a new tab)
- Efficient Estimation of Word Representations in Vector SpacePaperMikolov, Chen, Corrado and Dean, 2013CBOW averages context vectors and ignores order; training complexity of CBOW and skip-gram; sampling distant words less(opens in a new tab)
- Distributed Representations of Words and Phrases and their CompositionalityPaperMikolov, Sutskever, Chen, Corrado and Dean, NeurIPS 2013k of 5 to 20 for small data and 2 to 5 for large, U(w)^(3/4) noise, subsampling threshold near 10^-5, d = 300(opens in a new tab)
- word2vec Parameter Learning ExplainedPaperXin Rong, 2014The CBOW network figure on the slide; CBOW takes the average of the context vectors(opens in a new tab)
- GloVe: Global Vectors for Word RepresentationPaperPennington, Socher and Manning, EMNLP 2014Equations 8 and 9, x_max = 100, α = 3/4, zero entries are 75 to 95 percent of X, final vectors W + W̃(opens in a new tab)
- GloVe project pageDocsStanford NLP GroupPretrained GloVe vectors(opens in a new tab)
- Enriching Word Vectors with Subword InformationPaperBojanowski, Grave, Joulin and Mikolov, TACL 2017Boundary symbols, n from 3 to 6, the whole word included, <her> versus her, sum of n-gram vectors, OOV words(opens in a new tab)
- Word representations tutorialDocsfastTextDefault dimension 100, subwords of 3 to 6 characters, skip-gram works better with subwords than CBOW(opens in a new tab)
- fastText source: getWordVectorDocsFacebook ResearchgetWordVector divides the summed n-gram vectors by their count; Model::computeHidden in src/model.cc averages the input rows during training(opens in a new tab)
- Skip-Thought VectorsPaperKiros, Zhu, Salakhutdinov, Zemel, Torralba, Urtasun and Fidler, NeurIPS 2015GRU encoder, two decoders, BookCorpus, 2400 and 4800 dimensional vectors, 8 evaluation tasks(opens in a new tab)
- Dependency-Based Word EmbeddingsPaperLevy and Goldberg, ACL 2014Table 1 Hogwarts neighbors for BoW2, BoW5 and dependency contexts; uniform window sampling(opens in a new tab)
- models.word2vec: Word2vec embeddingsDocsGensimDefaults vector_size 100, window 5, negative 5, sg 0 (CBOW), cbow_mean 1, ns_exponent 0.75(opens in a new tab)
- NLP Course For You: Word EmbeddingsArticleLena VoitaSource of the slide 89 wording, the CBOW sum figure and the window size examples on slide 97(opens in a new tab)