Skip to main content

Probability Theory

Following "A concise Course in Statistical Inference", I will start summarizing the main ideas.

Data analysis, pattern recognition, machine learning, data mining are general terms that we have seen a lot, which indeed refer to special cases of statistical inference. Particular problems are classification, prediction, clustering, estimation.

- In Probability, we study: Given a data generating process, what are the properties of the outcome?
- In Statistical Inference, we study: Given the outcomes, what can we say about the process that generated the data?

This first entry will devote for "Probability".

Probability theory, or probability distribution/density is a function that describes the probability of a random variable taking certain value. There are two kinds:

- Discrete probability distribution (or probability on finite sample spaces): characterized by a probability mass function
uniform probability distribution: if the sample space is finite and each outcome is equally likely.

- Continuous probability distribution, i.e., a distribution of a continuous random variable X, is a probability distribution that has a probability density function. For example: normal, uniform, chi-squared distribution.

Comments

Popular posts from this blog

Random variables

A random variable is a mapping from a sample space to real numbers $\Omega \rightarrow \mathrm{R}$ At a certain point in most probability courses, we don't see the sample space, but it's always there, lurking in the background. For example: Let $\Omega = \{(x,y); x^2 + y^2 \leq 1\}$ be the unit disc. Consider drawing a point "at random" from $\Omega$. Outcome: $\omega = (x,y)$. Examples of random variables: $X(\omega) = x$, $X(\omega) = y$, $Z(\omega) = x + y$

SAXParser: too many exceptions for invalid XML character..

I'm working on my Similarity Search project, in which I have to implement the Tree Edit Distance and Traversal String Edit Distance. Trees are all represented in XML format and I'm using SAXParser to parse those XML files in java. I've used it a lot of times before but still, I don't quite like. So my first step is to create a valid XML database. However, "valid" to be parsed using SAXParser is complicated!! Here is what I get again and again: File Read Error: org.xml.sax.SAXParseException : The content of elements must consist of well-formed character data or markup. org.xml.sax.SAXParseException: The content of elements must consist of well-formed character data or markup. The reasons can be different, like: - Tags cannot contain number (e.g., is an invalid tag) - Tags cannot contain some symbols, like {, ., ?, etc. ("_" or "-" is fine) - Tags cannot be empty However, in my database, all of the tag are numbers.. To make it a valid XML fi...