Negative sampling and contrastive learning — from word2vec to CLIP
The textbook version of word2vec relies on a softmax output layer to turn the model’s raw scores into a probability for every word in the vocabulary — and that’s where the scaling problem comes from. In training, the loss only ever uses one probability per pair, P(target | center). But softmax defines it as a share of a vocabulary-wide total:
P(target) = exp(score_target) / Σ_w exp(score_w)The denominator sums over all V words, so to get the one number you actually want, you must compute all V scores — not because you want them, but because the normalizer needs them. Drop any single score and the normalization breaks.
With V = 10⁶ vocab words and embedding dim d = 300, scoring the center against every word is a V × d matmul — 300 million multiply-adds — plus a million exponentiations for the denominator, per training pair.
To feel the scale, compare per-example compute to MNIST:
INPUT HIDDEN OUTPUT OUTPUT-LAYER MATMUL
MNIST 784 128 10 128 × 10 = 1,280 ops
word2vec (V = 10⁶) V → 1 300 V 300 × 10⁶ = 300,000,000 opsMNIST is a 10-class classification problem; word2vec is a million-class one — one output neuron per vocab word. Every cost in the output layer — weights, matmul, softmax — scales linearly with that class count.
The 2013 follow-up paper introduced negative sampling as an alternative to hierarchical softmax. It replaces vocabulary prediction with binary classification: distinguish observed word–context pairs from pairs drawn from a noise distribution. This changes the objective, rather than merely approximating the softmax denominator.
One positive and k negative samples require k+1 dot products, costing O((k+1)d) instead of O(Vd). Only sampled output embeddings receive gradients; repeated samples count repeatedly. Here we store output vectors as rows of Ep, with shape (V, d) — the transpose of the E' layout in the word2vec article.
The comparison between observed and sampled pairs connects negative sampling to contrastive learning. CLIP and dense retrievers use related ideas, but their usual softmax losses differ from word2vec’s independent sigmoid terms.
From pairs to negatives
word2vec’s training data is (center, context) pairs from a sliding window over the corpus: the cat sat on the mat produces (sat, cat), (sat, on), (sat, the), and so on — billions of genuine co-occurrences.
Negative sampling keeps the pairs identical, but instead of just maximizing the likelihood of the real (center, context) match, it adds k random pairs per step as counter-balance — adjusting the weights so that their likelihood is also minimized.
Both kinds of pairs share the same center but pick the accompanying word (partner) differently:
Positive pairs come from context windows. Negative samples draw a partner independently of the center, usually with probability proportional to count(w)^0.75. A sampled word can also be a genuine context word: “negative” describes how it was sampled, not a claim that the pair never occurs. These are potential false negatives.
The lookup and dot products are unchanged. The difference is which words are scored and how those scores enter the loss.
A raw dot product v_c · v'_w can be any real number — positive, negative, large, small — but the loss needs a probability between 0 and 1: “how likely is this pair real?” The sigmoid function σ is what gets us there.
Once we have the scores, we can turn each into a probability with sigmoid. σ squashes any real-valued score into (0, 1):
dot product v_c · v'_w | σ | model says |
|---|---|---|
| large positive | ~1 | ”real pair” |
| ~0 | 0.5 | unsure |
| large negative | ~0 | ”random pair” |
For each real pair we want σ to climb toward 1; for each random pair, toward 0 — training pushes the dot products in those directions.
On the same five-word toy from the word2vec article, the widget below runs that score step. Pick a (center, target) training pair, toggle which words are sampled as the negatives, then step through the k+1 dot products — score[w] = v_c · v'_w, expanded term by term, then through its sigmoid. Only the target and the negatives get scored; the other rows of Ep stay greyed, never read.
| word | role | score | σ(score) | direction |
|---|---|---|---|---|
| on | positive | +0.394 | 0.597 | → pushed toward 1 |
| cat | negative | -0.187 | 0.453 | → pushed toward 0 |
| mat | negative | +0.260 | 0.565 | → pushed toward 0 |
Each sigmoid estimates a data-versus-noise label under the chosen sampling scheme. It is not P(context word | center), and it depends on both the noise distribution and the number of negatives. The values need not sum to one across words.
How are negatives sampled?
There are two useful ways to obtain negatives: sample words from a noise distribution, as in word2vec, or reuse other examples in a batch, as in CLIP and DPR.
Raising counts to a power below one reduces the dominance of frequent words while preserving their ranking. The amount of probability assigned to stop words depends on the corpus; there is no universal percentage.
The negative-sampling paper reports that exponent 0.75 outperformed unigram and uniform sampling in its experiments:
P(w) is the noise-sampling probability. The denominator normalizes the powered counts; it is computed from corpus frequencies, not the model’s current scores.
With the distribution in hand, sampling is straightforward: compute P(w) for every word in the vocab using the formula above, then for each positive (center, context) pair draw k random words from this distribution — a weighted dice-roll over the vocabulary, repeated k times, where words with higher P(w) get picked more often.
With in-batch negatives, other examples in the current batch supply the contrasting candidates.
In a batch of N labeled pairs (Qi, Pi), use Pi as the positive for Qi and the other candidates as negatives. This assigns training labels; it does not establish that every off-diagonal pair is semantically unrelated:
batch: (Q1, P1) (Q2, P2) (Q3, P3) (Q4, P4)
for Q1: positive = P1, negatives = {P2, P3, P4}
for Q2: positive = P2, negatives = {P1, P3, P4}
for Q3: positive = P3, negatives = {P1, P2, P4}
for Q4: positive = P4, negatives = {P1, P2, P3}The encoders already computed all N candidate embeddings. Reusing them avoids extra encoder passes, but the N×N score matrix still costs computation and memory. A batch of 256 supplies 255 candidate negatives per query.
Some other candidates may also be relevant to the query, creating false negatives. Batch composition, duplicate handling, and hard-negative selection affect training; in-batch negatives are not automatically better negatives.
The loss
Loss. The per-pair loss is a sum of k+1 log-sigmoid terms — one for the positive, one for each negative — in place of full softmax’s −log P(target | center):
loss = − log σ(v_c · v'_t) − Σ log σ(−v_c · v'_n)
───────────────── ───────────────────────
true (positive) pair k sampled negativesWhere v_c is the center word’s input embedding, v'_t is the true target’s output embedding, v'_n is a sampled negative word’s output embedding, and σ is the sigmoid function. The first term pushes the true pair’s dot product up (toward σ(·) = 1); the second term pushes each negative’s dot product down (toward σ(·) = 0).
The negative term uses σ(−v_c · v'_n) — the negative of the dot product — which works because of the identity σ(−x) = 1 − σ(x). So −log σ(−v_c · v'_n) is just −log(1 − σ(v_c · v'_n)): the standard “wrong class” half of cross-entropy, applied to the “this isn’t a real pair” direction. Each term in the loss is binary cross-entropy (BCE) applied to one (center, w) pair — label 1 for the positive, label 0 for each negative. The total loss is k+1 BCEs added together.
The loss grows when the model assigns low probability to the observed label. Its derivative with respect to the logit is σ(z) − y, bounded between −1 and 1. A large loss therefore does not imply an unbounded logit gradient.
Each scored word contributes one term — −log σ for the positive, −log(1 − σ) for each negative:
| word | role | σ | term | value |
|---|---|---|---|---|
on | positive | 0.5973 | −log(0.5973) | 0.5153 |
cat | negative | 0.4534 | −log(1 − 0.4534) | 0.6041 |
mat | negative | 0.5647 | −log(1 − 0.5647) | 0.8318 |
Total: L ≈ 1.95. mat contributes most — its σ (0.56) is the furthest from where a negative should be (0).
We model the sampled labels as conditionally independent Bernoulli observations given their scores. Multiplying their likelihoods and taking the negative logarithm gives this sum. The terms still share parameters through the center embedding.
P(all right) = P(positive right) × P(neg₁ right) × … × P(neg_k right)−log turns a product into a sum:
−log P(all right) = −log P(positive) + −log P(neg₁) + … + −log P(neg_k)Softmax models the context-word category; negative sampling models the data-versus-noise label. They optimize different likelihoods, so normalized softmax probabilities and negative-sampling sigmoids answer different questions. Neither is automatically calibrated on unseen data.
The gradient
For plain SGD without regularization, only the center’s input row and sampled output rows receive updates. If a word is sampled repeatedly, sum its contributions. Compute every gradient from the same pre-update parameter values.
To minimize L we need its gradient with respect to every parameter that touched the forward pass: v_c (the center’s row in E), v'_t (the target’s row in Ep), and each v'_n (one row per negative). With one calculus fact —
d/dz [ −log σ(z) ] = σ(z) − 1— the chain rule gives all three:
∂L / ∂v'_t = (σ_t − 1) · v_c ← target's output row
∂L / ∂v'_n = σ_n · v_c ← each negative's output row
∂L / ∂v_c = (σ_t − 1) · v'_t + Σ_n σ_n · v'_n ← center's input rowwhere σ_t = σ(v_c · v'_t) and σ_n = σ(v_c · v'_n) — exactly the numbers from loss above.
Notice the symmetry: every output-row gradient (∂L/∂v'_t, ∂L/∂v'_n) is a scalar times v_c, and the center’s gradient is a weighted sum of the output rows it scored against. That’s a direct consequence of the dot product being symmetric in its arguments — differentiating any f(v_c · v'_w) w.r.t. v'_w always yields something proportional to v_c, and vice versa.
The positive output update adds a multiple of v_c; a negative output update subtracts one. This increases or decreases the dot product when v_c is held fixed. “Pull” and “push” describe that score change, not a guaranteed reduction or increase in Euclidean distance.
Using the unrounded sigmoid values from the score widget, the center gradient is:
∂L/∂v_c = (σ_on − 1) · v'_on + σ_cat · v'_cat + σ_mat · v'_mat
≈ [−0.0435, 0.5269, 0.0717]The positive updates share the factor η(1 − σ_t), but their lengths are η(1 − σ_t)‖v_c‖ and η(1 − σ_t)‖v'_t‖. They are equal only when the vector norms are equal. The center also receives all negative contributions.
The update
Gradient descent with learning rate η:
v'_t ← v'_t + η · (1 − σ_t) · v_c ← step toward v_c
v'_n ← v'_n − η · σ_n · v_c ← step away from v_c
v_c ← v_c + η · (1 − σ_t) · v'_t − η · Σ_n σ_n · v'_n
← toward v'_t, away from each v'_nA larger sigmoid error gives a larger coefficient, but the vector norm also affects the update length. Different pairs can produce conflicting gradients, and a large learning rate can overshoot. Training loss need not decrease at every step.
Plugging in the gradient from above with η = 0.1 nudges the center:
v_c = [0.33, −0.27, 0.84]
v_c_new = v_c − 0.1 · ∂L/∂v_c ≈ [0.3344, −0.3227, 0.8328]A small step, but in the direction the loss demands. v'_on, v'_cat, and v'_mat get their own updates at the same time using the formulas above; we focus on v_c here to keep the trace short.
Verifying the step
First update only v_c, holding all output vectors fixed, to isolate the center’s contribution:
| word | before | after | direction |
|---|---|---|---|
v_c · v'_on | 0.39 | 0.41 | up — positive more aligned ✓ |
v_c · v'_cat | −0.19 | −0.22 | down — negative pushed apart ✓ |
v_c · v'_mat | 0.26 | 0.25 | down — negative pushed apart ✓ |
For this center-only update, the loss decreases from 1.9512 to 1.9229. Updating the output vectors simultaneously from the same old values gives 1.8626. These are results for this example and learning rate, not a guarantee for every training step.
Why negative sampling learns useful embeddings
Negative sampling is not an unbiased estimator of the full-softmax gradient. Both objectives reward observed associations, but weight the competing words differently. Negative sampling can learn useful representations because distinguishing co-occurrence from noise captures statistical relationships between words; it need not reproduce the softmax solution.
This 2D example repeatedly trains one positive pair and three fixed negatives. Watch the sigmoid scores and loss as both input and output vectors change. It illustrates one objective; it is not a trained semantic map of these five words.
The next widget trains on multiple pairs and plots how the input embeddings change:
This browser demo trains on a deliberately structured synthetic corpus. Each step processes one pair and draws five negatives with replacement from the smoothed frequency distribution, including possible collisions with the positive. The plot shows input embeddings only; output embeddings are trained separately. Colors label groups for the reader and are not training inputs. The model is restricted to two dimensions for display, so clean semantic clusters are not guaranteed.
The NumPy step below snapshots all required vectors before updating either table. np.add.at accumulates repeated output indices correctly, including a word drawn as both positive and negative. E and Ep are separate floating-point arrays of shape (V, d).
import numpy as np
def sgns_step(E, Ep, c, t, negatives, lr=0.1):
rows = np.r_[t, np.asarray(negatives, dtype=int)]
labels = np.zeros(len(rows))
labels[0] = 1
center = E[c].copy()
outputs = Ep[rows].copy()
logits = outputs @ center
z = np.exp(-np.abs(logits))
probabilities = np.where(logits >= 0, 1 / (1 + z), z / (1 + z))
errors = probabilities - labels
grad_center = errors @ outputs
grad_outputs = errors[:, None] * center
loss = np.sum(np.logaddexp(0, logits) - labels * logits)
np.add.at(Ep, rows, -lr * grad_outputs)
E[c] -= lr * grad_center
return float(loss) # loss before the update
# rng is a NumPy Generator; neg_dist sums to 1 over V words.
for c, t in pairs:
negatives = rng.choice(len(E), size=k, replace=True, p=neg_dist)
loss = sgns_step(E, Ep, c, t, negatives, lr)For an unconstrained score and k independent negatives from distribution q, the population optimum is s*(c,w) = log[P_data(w|c) / (k q(w))]. If q(w) = P_data(w), this becomes PMI(c,w) − log(k). Using q(w) ∝ count(w)^0.75 changes the correction term. Finite-dimensional embeddings can only approximate this matrix of ideal scores.
This connection is the basis of Levy and Goldberg’s shifted-PMI analysis. It explains what SGNS scores capture; it does not make SGNS, SVD, and full softmax interchangeable objectives.
When to use what
Use full softmax when the task needs a normalized distribution over a fixed set of alternatives and computing it is affordable. Use a sampled objective when representation quality and training cost justify that choice. Neither objective guarantees better accuracy for every task.
BERT’s masked-token prediction and autoregressive language modeling commonly use vocabulary softmax. Retrieval systems often compare a relevant item against sampled or in-batch candidates. The choice of candidate set is separate from the choice of sigmoid or softmax loss.
DPR learns query–passage scores using a softmax over candidate passages. Preference reward models instead learn relative response scores, often through −log σ(r_preferred − r_rejected); that is a ranking objective, not word2vec’s sampled-noise objective. CLIP provides another softmax-based contrastive example.
CLIP: contrasting images and captions
CLIP was trained on 400 million image–text pairs. It uses in-batch mismatches as negatives, but its loss is a symmetric softmax cross-entropy, not the independent binary losses used by word2vec negative sampling.
The original CLIP experiments used a text transformer paired with either a ResNet or a vision transformer. Projection layers map both outputs into a shared embedding space; its dimension depends on the model variant.
For each training batch of N pairs (I_1, T_1), ..., (I_N, T_N):
L2-normalize both sets of embeddings and form S[i,j] = exp(t) × dot(image[i], text[j]), where t is a learned log-scale. The dot product is now cosine similarity. Apply cross-entropy to each row with target i, and to each column with target i; average the two mean losses. The diagonal is labeled positive, although off-diagonal pairs can contain false negatives. See the CLIP implementation.
The original training batch contained 32,768 pairs, giving each image 32,767 candidate negative captions. Encoder outputs are reused, but computing the pairwise scores and synchronizing embeddings across devices still has a cost.
What you get: an embedding space where semantically related images and texts land close, unrelated ones land far. That’s why CLIP can do zero-shot image classification — compute text embeddings for class names (“a photo of a dog”, “a photo of a cat”, etc.), then classify an image by which class embedding it’s closest to. The geometry the contrastive loss carved out already encodes meaning across the two modalities; no labeled classifier needed.
A raw dot product is not a cosine: SGNS can change vector norms as well as angles. CLIP normalizes vectors and learns a score scale. These different geometries do not imply a universal target angle for negatives, nor do they by themselves explain the batch size a model needs.
The common idea is to learn from comparisons. The details matter: word2vec SGNS classifies sampled pairs independently, while CLIP makes candidates compete through softmax. Choose the sampling scheme and the loss together, according to what the resulting scores need to represent.