Skip to main content

openNLP: getting started, code example

openNLP is an interesting tool for Natural Language Processing. Today, it took me a while to get started. So I want to write again that next time (or anyone) who wants to try it out, it will take less time for you.

Here we go:

1. Download and build:
This is the main website: http://opennlp.sourceforge.net/

You can either download the package there and build to the .jar file (where you have to set your $JAVA_HOME environment - see below). Or you can directly download the .jar file from this link.

This step took me a while since I didn't know how to set my $JAVA_HOME, and didn't find out that there's already a .jar file to download.


2. Some code


So now, you want to start with some code. Here is some sample code for doing Sentence Detection and Tokenization.

Note that you can either download the models from the previous website or have the training dataset yourself.

In this example, I used 2 models of openNLP (EnglishSD.bin.gz and EnglishTok.bin.gz).

//This is the path to your model files
SentenceDetector sendet = new SentenceDetector("opennlp-tools-1.4.3/Models/EnglishSD.bin.gz");
Tokenizer tok = new Tokenizer("opennlp-tools-1.4.3/Models/EnglishTok.bin.gz");

//Sentence detection
String[] sens = sendet.sentDetect("This is sentence one. This is sentence two.");

for (int i=0; i<sens.length; i++)

{
System.out.println("Sentence " + i + ": ");
String[] tokens = tok.tokenize(sens[i]);
for (int j=0; j
<sens.length; j++)
System.out.print(tokens[j] + " - ");
System.out.println();
}



3. Other notes
If you got some exception when running the above code, it's probably that you didn't include the .jar files (e.g., maxent.jar and trove.jar) in the /lib folder.

Good luck!

Comments

Popular posts from this blog

Quick text files merging, data preparation

It's very often that in natural language processing, you will have to re-format your data to take as inputs to different systems. In this case, these simple linux commands will help you do it much quicker without having to write a script. 1. Merging two files to one file with two column Input f1 looks like this: 1 2 3 4 Input f2 looks like this: a b c d Output f3 will look like this: 1  a 2  b 3  c 4  d Command: paste f1 f2 > f3  The delimiter by default is a tab. You can also define it (for example, separated by a comma) as follows: paste -d ',' f1 f2 > f3 2.  Create a line number to each line of a text file Assume that you want to create an index to each line in a text file, i.e. inserting a line number and then a tab before the content of each line: Input f1: a b c d Output f2: 1  a 2  b 3  c 4  d Command: nl f1 > f2 3. Joining two files with a common field Input f1: 1   aaa...

Random variables

A random variable is a mapping from a sample space to real numbers $\Omega \rightarrow \mathrm{R}$ At a certain point in most probability courses, we don't see the sample space, but it's always there, lurking in the background. For example: Let $\Omega = \{(x,y); x^2 + y^2 \leq 1\}$ be the unit disc. Consider drawing a point "at random" from $\Omega$. Outcome: $\omega = (x,y)$. Examples of random variables: $X(\omega) = x$, $X(\omega) = y$, $Z(\omega) = x + y$