Professional Documents
Culture Documents
Part 3 - Mathematical Background Statistics and Linear Regression
Part 3 - Mathematical Background Statistics and Linear Regression
Linear regression
Part III
Mathematical Background:
Statistics and Linear Regression
Linear regression
Context
Table of contents
Linear regression
Linear regression
Linear regression
Linear regression
Probability: Example
Consider as an example the precipitation on a given day, and to make
the definitions precise let h denote precipitation in mm.
Sample set: e.g. = {dry (p = 0), drizzle (0 < p 2), rain (2 <
p 10), downpour (p > 10)}, with the outcomes taking any of
these values.
Event A: any one of the outcomes, e.g. A = {drizzle}, and in
addition any union of outcomes, such as
A = {drizzle} {rain} {downpour}; call A = wet.
Then, one example of probability measure is P({dry}) = 0.5,
P({drizzle}) = 0.2, P({rain}) = 0.2, P({downpour}) = 0.1, and we
use condition 3 to generate the probabilities of combined events; e.g.
P(wet) = 0.2 + 0.2 + 0.1 = 0.5. Note that conditions 1 and 2
(P() = 1) are satisfied.
Linear regression
Probability: Independence
The joint probability of two events A and B is defined as
P(A, B) := P(A B).
Definition
Two events A and B are called independent if P(A, B) = P(A)P(B).
Examples:
The event of rolling a 6 with a dice is independent of the event of
getting a 6 at the previous roll (or, indeed, any other value and
any previous roll).
The event of rolling two consecutive 6-s, however, is not
independent of the previous roll!
(Incidentally, failing to understand the first example leads to the
so-called gamblers fallacy. Having just had a long sequence of
bador goodgames at the casino, means nothing for the next game.
Take this into account the next time youre in Vegas!)
Linear regression
Random variable
Definition
A random variable is a function X : X defined on the sample set
, taking values in some arbitrary space X .
Intuitively, random variables associate interesting quantities to the
outcomes. A specific (deterministic) value of X is denoted x. Such a
value is called a realization of X .
The probability of X taking value x is the probability of all outcomes
leading to x:
P(X = x) = P({ | X () = x })
where the first, shorter notation is used for convenience.
Linear regression
Linear regression
Linear regression
For the mathematically inclined: the same may appear true for infinitely many
discrete values, but there we can get around the problem by using sequences
and limits. That no longer helps when the values are continuous
mathematical term: uncountable.
Linear regression
dF (x)
dx
Remarks:
The PDF is the correspondent to the PMF from the discrete case.
R
For any set ZP X , P(X Z ) = xZ f (x) (in the discrete case,
P(X Z ) = xZ P(x)).
Linear regression
Example: Gaussian
(x )2
fG (x) =
exp
2 2
2 2
1
Table of contents
Linear regression
Linear regression
Linear regression
Practical probabilities
Linear regression
Expected value
Definition
(P
p(x)x
E {X } = R xX
f (x)x
xX
Linear regression
Linear regression
Variance
Definition
Var {X } = E (X E {X })2 = E X 2 (E {X })2
Intuition: the spread of the random values around the expectation.
(P
p(x)(x E {X })2 if discrete
Var {X } = R xX
f (x)(x E {X })2
if continuous
xX
(P
p(x)x 2 (E {X })2 if discrete
= R xX
f (x)x 2 (E {X })2
if continuous
xX
Examples:
For a fair dice, Var {X } = 16 12 + 16 22 + . . . + 16 62 (7/2)2 = 35/12.
If X is distributed with PDF f (x) = fG (x), Gaussian, then
Var {X } = 2 .
Linear regression
Notation
Linear regression
Covariance
Definition
Cov {X , Y } = E {(X E {X })(Y E {Y })} = E {(X X )(Y Y )}
where X , Y denote the means (expected values) of the two random
variables.
Intuition: how much the two variables change together (positive if
they change in the same way, negative if they change in opposite
ways).
Remark: Var {X } = Cov {X , X }.
Linear regression
Uncorrelated variables
Definition
Random variables X and Y are uncorrelated if Cov {X , Y } = 0.
Otherwise, they are correlated.
Examples:
The education level of a person is correlated with their income.
Hair color may be uncorrelated with income (at least in an ideal
world).
Remarks:
If X and Y are independent, they are uncorrelated.
But the reverse is not necessarily true!
Linear regression
>
Cov {X } :=
Linear regression
(2)N
1
p
det()
>
exp (x )1 (x )
Linear regression
Stochastic process
Definition
A stochastic process X is a sequence of random variables
X = (X1 , . . . , Xk , . . . , XN ).
It is in fact just a vector of random variables, with the additional
structure that the index in the vector has the meaning of time step k.
In system identification, signals (e.g., inputs, outputs) will often be
stochastic processes evolving over discrete time steps k .
Linear regression
Definition
The stochastic process X is zero-mean white noise if: k, E {Xk } = 0
(zero-mean), and k, k 0 6= k , Cov {Xk , Xk 0 } = 0 (the values at different
steps are uncorrelated). In addition, the variance Var {Xk } must be
finite k.
Stated concisely using vector notation: mean E {X } = 0 RN and
covariance matrix = Cov {X } is diagonal (with positive finite
elements on the diagonal).
In system identification, noise processes (such as zero-mean white
noise) often affect signal measurements.
Linear regression
Stationary process
Table of contents
Linear regression
Regression problem and solution
Examples
Analysis and discussion
Linear regression
Linear regression
Regression problem
Problem elements:
A collection of known samples y (k) R, indexed by
k = 1, . . . , N: y is the regressed variable.
For each k, a known vector (k) Rn : contains the regressors
i (k ), i = 1, . . . , n.
An unknown parameter vector Rn .
Objective: identify the behavior of the regressed variable from the
data, using the linear model:
y(k) = > (k )
Linear regression is classical and very common, e.g. it was used by
Gauss to compute the orbits of the planets in 1809.
Linear regression
unknown g(x)
b(x)
g
Linear regression
Linear regression
= 1 k k 2 . . . k n1
3
. . .
n
= > (k )
Linear regression
d
2
X
(x
c
)
j
j
(d-dim)
= exp
2
b
j
j=1
Linear regression
Example 3: Interpolation
Suitable for function approximation.
d-dimensional grid of points in the input space.
Linear interpolation between the points.
Equivalent with pyramidal basis functions (triangular in 1-dim)
Linear regression
Linear system
Writing the model for each of the N data points, we get a linear
system of equations:
y (1) = 1 (1)1 + 2 (1)2 + . . . n (1)n
y (2) = 1 (2)1 + 2 (2)2 + . . . n (2)n
y (1)
1 (1) 2 (1) . . .
y (2)
1 (2) 2 (2) . . .
.. =
(N)
(N)
.
..
1
2
y(N)
1
n (1)
2
n (2)
3
. . .
n (N)
n
Y =
with newly introduced variables Y RN and RNn .
Linear regression
Least-squares problem
If N = n, the system can be solved with equality.
In practice, it is a good idea to use N > n, due e.g. to noise and
disturbances. In this case, the system can no longer be solved with
equality, but only in an approximate sense.
Error at k : (k ) = y(k) > (k ),
>
error vector = [(1), (2), . . . , (N)] .
Objective function to be minimized:
N
V () =
1X
1
(k )2 = >
2
2
k=1
Least-squares problem
Find the parameter vector b that minimizes the objective function:
b = arg min V ()
Linear regression
Formal solution
Linear regression
Alternative expression
> =
N
X
(k)> (k ), > Y =
k=1
N
X
> (k )y (k )
k=1
N
X
k=1
#1 "
>
(k) (k)
N
X
#
>
(k)y(k)
k =1
Table of contents
Linear regression
Regression problem and solution
Examples
Analysis and discussion
Linear regression
Linear regression
y (N) = (N) = 1 b
In matrix form:
y(1)
1
.. ..
. = .
y (N)
Y =
Linear regression
b = (> )1 > Y
1
1
.
1
= 1 1 ..
= N 1 1
y(1)
.
1 ..
y (N)
y (1)
1 ...
y(N)
1
(y (1) + . . . + y (N))
N
Linear regression
Linear regression
Linear regression
Linear regression
Linear regression
Due to the better flexibility of the neural network, the results are better
than with linear regression.
Table of contents
Linear regression
Regression problem and solution
Examples
Analysis and discussion
Linear regression
Linear regression
Geometrical interpretation
Linear regression
Linear regression
Intuition: Part 1 says that the solution makes (statistical) sense, while
Part 2 can be interpreted as measuring the confidence in the solution.
E.g., smaller errors e(k) have smaller variance 2 , which means the
covariances are smaller i.e. better confidence that b is close to the
true value 0 .
Remark: 2 is unknown, but can be estimated as
b is available).
expression for V ()
b
2V ()
Nn
Linear regression
Model choice
Assume that given a model size n, we have a way to generate
regressors (k ) that make the model more expressive (e.g., basis
functions on a finer grid). Then we expect the objective function to
behave as follows:
Linear regression
Linear regression
Multivariable case