Skip to main content

Skip-gram model and negative sampling

In the previous post, we have seen the 3 word2vec models: skip-gram, CBOW and GloVe. Now let's have a look at negative sampling and what it is used to make training skip-gram faster.
The idea is originated from this paper:
"Distributed Representations of Words and Phrases and their Compositionality” (Mikolov et al. 2013)
In the previous example, we have seen that if we have a vocabulary of size 10K, and we want to train word vectors of size 300. Then the number of parameters we have to estimate in each layer is 10Kx300. This number is big and makes training prone to over-fitting and gives too much focus on words that appear often, and less focus on rare words.

Subsampling of frequent words

So the idea of subsampling is that: we try to maximize the probability that "real outside word" appears, and minimize the probability that "random words" appear around center word. Real outside words are words that characterize the meaning of the center word, while "random words" tend to occur very often and go together with many other different words (e.g., "the", "an", "a").
So each word will have a probability of being "kept", where more frequent words have lower probability of being "kept" and vice versa.

Negative Sampling

In our model, we have a very big number of weights. Everytime we process a new training sample, we will have to go back and update all our weights in the model, which makes this process very slow.
We use negative sampling to address this problem: instead of modifying all of the weights, we modify only a small percentage of them.
In particular, we will take some random negative words (i.e., words that are not in the context) and update weights for our positive words (i.e., words that are in the context).
So how many "random negative words" should we draw? In the paper, they suggest 2-5 words for large dataset and 5-20 words for small dataset.


Comments

Popular posts from this blog

Random variables

A random variable is a mapping from a sample space to real numbers $\Omega \rightarrow \mathrm{R}$ At a certain point in most probability courses, we don't see the sample space, but it's always there, lurking in the background. For example: Let $\Omega = \{(x,y); x^2 + y^2 \leq 1\}$ be the unit disc. Consider drawing a point "at random" from $\Omega$. Outcome: $\omega = (x,y)$. Examples of random variables: $X(\omega) = x$, $X(\omega) = y$, $Z(\omega) = x + y$

SAXParser: too many exceptions for invalid XML character..

I'm working on my Similarity Search project, in which I have to implement the Tree Edit Distance and Traversal String Edit Distance. Trees are all represented in XML format and I'm using SAXParser to parse those XML files in java. I've used it a lot of times before but still, I don't quite like. So my first step is to create a valid XML database. However, "valid" to be parsed using SAXParser is complicated!! Here is what I get again and again: File Read Error: org.xml.sax.SAXParseException : The content of elements must consist of well-formed character data or markup. org.xml.sax.SAXParseException: The content of elements must consist of well-formed character data or markup. The reasons can be different, like: - Tags cannot contain number (e.g., is an invalid tag) - Tags cannot contain some symbols, like {, ., ?, etc. ("_" or "-" is fine) - Tags cannot be empty However, in my database, all of the tag are numbers.. To make it a valid XML fi...