Skip to main content

Probability Theory

Following "A concise Course in Statistical Inference", I will start summarizing the main ideas.

Data analysis, pattern recognition, machine learning, data mining are general terms that we have seen a lot, which indeed refer to special cases of statistical inference. Particular problems are classification, prediction, clustering, estimation.

- In Probability, we study: Given a data generating process, what are the properties of the outcome?
- In Statistical Inference, we study: Given the outcomes, what can we say about the process that generated the data?

This first entry will devote for "Probability".

Probability theory, or probability distribution/density is a function that describes the probability of a random variable taking certain value. There are two kinds:

- Discrete probability distribution (or probability on finite sample spaces): characterized by a probability mass function
uniform probability distribution: if the sample space is finite and each outcome is equally likely.

- Continuous probability distribution, i.e., a distribution of a continuous random variable X, is a probability distribution that has a probability density function. For example: normal, uniform, chi-squared distribution.

Comments

Popular posts from this blog

Quick text files merging, data preparation

It's very often that in natural language processing, you will have to re-format your data to take as inputs to different systems. In this case, these simple linux commands will help you do it much quicker without having to write a script. 1. Merging two files to one file with two column Input f1 looks like this: 1 2 3 4 Input f2 looks like this: a b c d Output f3 will look like this: 1  a 2  b 3  c 4  d Command: paste f1 f2 > f3  The delimiter by default is a tab. You can also define it (for example, separated by a comma) as follows: paste -d ',' f1 f2 > f3 2.  Create a line number to each line of a text file Assume that you want to create an index to each line in a text file, i.e. inserting a line number and then a tab before the content of each line: Input f1: a b c d Output f2: 1  a 2  b 3  c 4  d Command: nl f1 > f2 3. Joining two files with a common field Input f1: 1   aaa...

Random variables

A random variable is a mapping from a sample space to real numbers $\Omega \rightarrow \mathrm{R}$ At a certain point in most probability courses, we don't see the sample space, but it's always there, lurking in the background. For example: Let $\Omega = \{(x,y); x^2 + y^2 \leq 1\}$ be the unit disc. Consider drawing a point "at random" from $\Omega$. Outcome: $\omega = (x,y)$. Examples of random variables: $X(\omega) = x$, $X(\omega) = y$, $Z(\omega) = x + y$