Lecture 2: Properties of the OLS estimator
Economics 527: Econometric Methods
Model and assumptions
Where the first lecture ended
The first lecture produced the estimator and stopped there: \hat\beta = (X^\top X)^{-1} X^\top Y = \left(\sum_{i=1}^n X_i X_i^\top\right)^{-1}\sum_{i=1}^n X_i Y_i, the unique minimiser of \sum_{i=1}^n (Y_i - X_i^\top b)^2.
Nothing so far says whether it is a good estimator.
\hat\beta is a function of the sample, so it is random, and in general it will not equal \beta.
Three questions, taken in this order: whether \hat\beta is right on average, how far it moves from sample to sample, and whether a better estimator of the same kind exists.
Model in matrix form
Stack the n observations, as in the first lecture: \underset{n\times1}{Y} = \begin{pmatrix} Y_1 \\ \vdots \\ Y_n \end{pmatrix},\qquad \underset{n\times k}{X} = \begin{pmatrix} X_1^\top \\ \vdots \\ X_n^\top \end{pmatrix},\qquad \underset{n\times1}{U} = \begin{pmatrix} U_1 \\ \vdots \\ U_n \end{pmatrix}.
The n scalar equations Y_i = X_i^\top\beta + U_i become one: \underset{n\times1}{Y} = \underset{n\times k}{X}\ \underset{k\times1}{\beta} + \underset{n\times1}{U}.
\beta is fixed and unknown. X, Y and U are random, and U is not observed.
Four assumptions
Assumption 1. Y = X\beta + U for some \beta \in \mathbb{R}^k.
Assumption 2. \mathrm{E}\left[U \mid X\right] = 0.
Assumption 3. \mathrm{Var}\left(U \mid X\right) = \sigma^2 I_n for some \sigma^2 > 0.
Assumption 4. \operatorname{rank}(X) = k.
Unbiasedness
What unbiased means
\hat\beta is a function of the sample, hence a random vector: \hat\beta(\omega) = \left(\sum_{i=1}^n X_i(\omega) X_i(\omega)^\top\right)^{-1}\sum_{i=1}^n X_i(\omega) Y_i(\omega).
With a continuously distributed error, \mathrm{P}(\hat\beta = \beta) = 0. Hitting \beta exactly is not the standard.
Definition. An estimator \hat\theta of \theta is unbiased if \mathrm{E}\left[\hat\theta\right] = \theta, and conditionally unbiased given X if \mathrm{E}\left[\hat\theta \mid X\right] = \theta, whatever the true value of \theta: the estimator is a rule that does not know \theta, and the equality must hold for every value \theta could take.
The conditional statement is the stronger one, and it is the one this lecture proves.
Sampling error
Substitute Assumption 1 into the estimator; Assumption 4 is what makes (X^\top X)^{-1} exist: \begin{aligned} \hat\beta &= (X^\top X)^{-1} X^\top Y \\ &= (X^\top X)^{-1} X^\top (X\beta + U) \\ &= \underbrace{(X^\top X)^{-1} X^\top X}_{=\, I_k}\, \beta + (X^\top X)^{-1} X^\top U. \end{aligned}
The first term produces \beta, and we have signal+noise decomposition: \boxed{\ \hat\beta = \underbrace{\beta}_{\text{signal}} + \underbrace{{\color{blue}(X^\top X)^{-1} X^\top U}}_{\text{noise}}\ }
Everything in this section and the next is a statement about the noise term.
Taking expectations, first attempt
Take expectations of the sampling error: \mathrm{E}\left[\hat\beta\right] = \beta + \mathrm{E}\left[{\color{blue}(X^\top X)^{-1} X^\top U}\right].
Suppose only that \mathrm{E}\left[X_i U_i\right] = 0 for every i. Then \mathrm{E}\left[X^\top U\right] = \mathrm{E}\left[\sum_{i=1}^n X_i U_i\right] = \sum_{i=1}^n \mathrm{E}\left[X_i U_i\right] = 0.
That is not enough, because in general \mathrm{E}\left[(X^\top X)^{-1} X^\top U\right] \;\text{ need not equal }\; \mathrm{E}\left[(X^\top X)^{-1}\right]\,\mathrm{E}\left[X^\top U\right].
The two factors are built from the same X, and once \mathrm{E}\left[U \mid X\right] is not zero, the expectation of their product need not factor.
Expectation of a ratio
The one-regressor case makes the failure visible: \mathrm{E}\left[\frac{\sum_i X_i U_i}{\sum_i X_i^2}\right] \;\text{ need not equal }\; \frac{\mathrm{E}\left[\sum_i X_i U_i\right]}{\mathrm{E}\left[\sum_i X_i^2\right]} = 0.
In general the expectation of a ratio is not the ratio of expectations. For any two scalar random variables Z and S, \mathrm{E}\left[\frac{Z}{S}\right] = \iint \frac{z}{s}\, f_{Z,S}(z,s)\, dz\, ds \;\text{ need not equal }\; \frac{\int z f_Z(z)\, dz}{\int s f_S(s)\, ds} = \frac{\mathrm{E}\left[Z\right]}{\mathrm{E}\left[S\right]}.
The way out is to condition on X, which turns (X^\top X)^{-1}X^\top into a known matrix.
Conditional unbiasedness: the claim
Claim. Under Assumptions 1, 2 and 4, \mathrm{E}\left[\hat\beta \mid X\right] = \beta.
Assumption 3 is not used. Unbiasedness does not depend on the variance of the errors.
Proof of conditional unbiasedness
Condition the sampling error on X: \begin{aligned} \mathrm{E}\left[\hat\beta \mid X\right] &= \mathrm{E}\left[\beta + (X^\top X)^{-1}X^\top U \mid X\right] \\ &= \beta + \mathrm{E}\left[(X^\top X)^{-1}X^\top U \mid X\right] \\ &= \beta + (X^\top X)^{-1}X^\top \underbrace{\mathrm{E}\left[U \mid X\right]}_{=\,0 \text{ by Assumption 2}} \\ &= \beta. \end{aligned}
The third line takes out what is known, as in the first lecture. Given X, the matrix (X^\top X)^{-1}X^\top is known and comes out of the expectation.
From conditional to unconditional
The law of iterated expectations, from the first lecture, finishes the proof: \mathrm{E}\left[\hat\beta\right] = \mathrm{E}\left[\mathrm{E}\left[\hat\beta \mid X\right]\right] = \mathrm{E}\left[\beta\right] = \beta.
Conditional unbiasedness implies unbiasedness whenever the unconditional mean exists. The converse does not hold.
A weaker condition
Replace Assumption 2 by \mathrm{E}\left[X_i U_i\right] = 0 \quad \text{for all } i, which, together with \mathrm{E}\left[U_i\right] = 0, says that each regressor is uncorrelated with the error.
\beta is still identified under this condition.
It is not enough for unbiasedness: \mathrm{E}\left[\hat\beta\right] = \beta + \underbrace{\mathrm{E}\left[(X^\top X)^{-1}X^\top U\right]}_{\text{need not be } 0}.
So \hat\beta need not be unbiased. The noise term is a nonlinear function of X and U, and once \mathrm{E}\left[U \mid X\right] is not zero, nothing forces its mean to vanish.
Variance of a random vector
Definition
For a random vector V that is n \times 1 with finite second moments, and with \mathrm{E}\left[V\right] = (\mathrm{E}\left[V_1\right], \ldots, \mathrm{E}\left[V_n\right])^\top, \mathrm{Var}\left(V\right) = \mathrm{E}\left[\underset{n\times 1}{(V - \mathrm{E}\left[V\right])}\ \underset{1\times n}{(V - \mathrm{E}\left[V\right])^\top}\right] \qquad (n \times n).
The product inside is an outer product, not an inner product, so the result is a matrix rather than a number.
Entry by entry
Entry (i,j) is \mathrm{E}\left[(V_i - \mathrm{E}\left[V_i\right])(V_j - \mathrm{E}\left[V_j\right])\right], so \mathrm{Var}\left(V\right) = \begin{pmatrix} \mathrm{Var}\left(V_1\right) & \mathrm{Cov}\left(V_1, V_2\right) & \cdots & \mathrm{Cov}\left(V_1, V_n\right) \\ \mathrm{Cov}\left(V_2, V_1\right) & \mathrm{Var}\left(V_2\right) & & \vdots \\ \vdots & & \ddots & \\ \mathrm{Cov}\left(V_n, V_1\right) & \cdots & & \mathrm{Var}\left(V_n\right) \end{pmatrix}.
The variances sit on the diagonal and the covariances off it.
Symmetry
\mathrm{Var}\left(V\right)^\top = \mathrm{Var}\left(V\right), because \mathrm{Cov}\left(V_i, V_j\right) = \mathrm{Cov}\left(V_j, V_i\right).
The same statement from the definition, using that the transpose passes through the expectation, and then (AB)^\top = B^\top A^\top: \mathrm{Var}\left(V\right)^\top = \mathrm{E}\left[\big((V - \mathrm{E}\left[V\right])(V - \mathrm{E}\left[V\right])^\top\big)^\top\right] = \mathrm{E}\left[(V - \mathrm{E}\left[V\right])(V - \mathrm{E}\left[V\right])^\top\right] = \mathrm{Var}\left(V\right).
Positive semidefiniteness
Claim. \mathrm{Var}\left(V\right) is positive semidefinite.
Proof. Take any non-random a \in \mathbb{R}^n and set W = a^\top (V - \mathrm{E}\left[V\right]), a scalar. Then \begin{aligned} a^\top \mathrm{Var}\left(V\right) a &= a^\top \mathrm{E}\left[(V - \mathrm{E}\left[V\right])(V - \mathrm{E}\left[V\right])^\top\right] a \\ &= \mathrm{E}\left[\underbrace{a^\top (V - \mathrm{E}\left[V\right])}_{W}\ \underbrace{(V - \mathrm{E}\left[V\right])^\top a}_{W}\right] \\ &= \mathrm{E}\left[W^2\right] \;\ge\; 0. \end{aligned}
The constant a moves inside the expectation because expectation is linear.
This is the argument the first lecture used for \mathrm{E}\left[X_i X_i^\top\right], with W_i = X_i^\top c.
When it is positive definite
For a \neq 0, with W = a^\top (V - \mathrm{E}\left[V\right]) as above, track the equality case: a^\top \mathrm{Var}\left(V\right) a = 0 \iff \mathrm{E}\left[W^2\right] = 0 \iff a^\top (V - \mathrm{E}\left[V\right]) = 0, that is, a_1 (V_1 - \mathrm{E}\left[V_1\right]) + \cdots + a_n (V_n - \mathrm{E}\left[V_n\right]) = 0.
So \mathrm{Var}\left(V\right) is positive definite exactly when no non-zero linear combination of V_1, \ldots, V_n is constant, that is, when no exact linear relation holds among the centred variables V_i - \mathrm{E}\left[V_i\right].
Variance under a linear map
Claim. For a non-random m \times n matrix \Gamma, \boxed{\ \mathrm{Var}\left(\Gamma V\right) = \Gamma \mathrm{Var}\left(V\right) \Gamma^\top \ } \qquad (m \times m).
Proof. \mathrm{E}\left[\Gamma V\right] = \Gamma \mathrm{E}\left[V\right], so \begin{aligned} \mathrm{Var}\left(\Gamma V\right) &= \mathrm{E}\left[\big(\Gamma V - \Gamma\mathrm{E}\left[V\right]\big)\big(\Gamma V - \Gamma\mathrm{E}\left[V\right]\big)^\top\right] \\ &= \mathrm{E}\left[\Gamma (V - \mathrm{E}\left[V\right]) (V - \mathrm{E}\left[V\right])^\top \Gamma^\top\right] \\ &= \Gamma \mathrm{E}\left[(V - \mathrm{E}\left[V\right])(V - \mathrm{E}\left[V\right])^\top\right] \Gamma^\top = \Gamma \mathrm{Var}\left(V\right) \Gamma^\top, \end{aligned} using (AB)^\top = B^\top A^\top on the second factor, and linearity to move the non-random \Gamma and \Gamma^\top outside the expectation.
Check against the scalar case: with m = n = 1 this is \mathrm{Var}\left(\gamma V\right) = \gamma^2 \mathrm{Var}\left(V\right).
Adding a constant
For a non-random \alpha that is m \times 1, \mathrm{Var}\left(\alpha + \Gamma V\right) = \mathrm{Var}\left(\Gamma V\right) = \Gamma \mathrm{Var}\left(V\right) \Gamma^\top.
The constant shifts the mean, so the deviations from the mean are unchanged and it drops out.
This is why \beta will drop out of the variance of the sampling error.
Conditional version
The same definition with conditional expectations throughout: \mathrm{Var}\left(V \mid X\right) = \mathrm{E}\left[\big(V - \mathrm{E}\left[V \mid X\right]\big)\big(V - \mathrm{E}\left[V \mid X\right]\big)^\top \,\Big|\, X\right].
Everything above holds conditionally, with \Gamma allowed to be a function of X. Given X, such a \Gamma is known and comes out of the conditional expectation, by taking out what is known, as in the proof of unbiasedness.
When the conditional mean is zero, the definition simplifies. Under Assumption 2, \mathrm{Var}\left(U \mid X\right) = \mathrm{E}\left[U U^\top \mid X\right].
Variance of the OLS estimator
Variance matrix of \hat\beta
\hat\beta is k \times 1, so \mathrm{Var}\left(\hat\beta \mid X\right) is k \times k: the conditional variance of each estimated coefficient on the diagonal, the conditional covariances between them off it.
Conditioning on X again, for the same reason it worked for unbiasedness: it makes (X^\top X)^{-1}X^\top a known matrix.
Sandwich formula
Start from the sampling error and name the matrix that multiplies U: \hat\beta = \beta + \underbrace{(X^\top X)^{-1} X^\top}_{\Gamma_X}\, U.
\beta is a constant and drops out, and \Gamma_X is known given X, so the rule for a linear map applies: \begin{aligned} \mathrm{Var}\left(\hat\beta \mid X\right) &= \mathrm{Var}\left(\beta + \Gamma_X U \mid X\right) \\ &= \mathrm{Var}\left(\Gamma_X U \mid X\right) \\ &= \Gamma_X \mathrm{Var}\left(U \mid X\right) \Gamma_X^\top. \end{aligned}
Sandwich formula written out
Substituting \Gamma_X = (X^\top X)^{-1}X^\top, with \Gamma_X^\top = X (X^\top X)^{-1} because X^\top X and its inverse are symmetric: \boxed{\ \mathrm{Var}\left(\hat\beta \mid X\right) = (X^\top X)^{-1} X^\top\ \mathrm{Var}\left(U \mid X\right)\ X (X^\top X)^{-1}\ }
The box uses Assumptions 1 and 4, and needs the errors to have finite conditional second moments, so that \mathrm{Var}\left(U \mid X\right) exists. Assumption 2 enters only to write the middle factor as \mathrm{E}\left[U U^\top \mid X\right].
Apart from requiring \mathrm{Var}\left(U \mid X\right) to exist, nothing yet has been assumed about the errors’ conditional variances or covariances.
Assumption 3: homoskedasticity and no autocorrelation
Assumption 3. \mathrm{Var}\left(U \mid X\right) = \sigma^2 I_n for some \sigma^2 > 0, that is, \mathrm{Var}\left(U \mid X\right) = \begin{pmatrix} \sigma^2 & & 0 \\ & \ddots & \\ 0 & & \sigma^2 \end{pmatrix}.
Read the diagonal: every error has the same conditional variance, \mathrm{Var}\left(U_1 \mid X\right) = \cdots = \mathrm{Var}\left(U_n \mid X\right) = \sigma^2.
This is homoskedasticity. It rules out errors that are more variable at some values of the regressors than at others.
Read the off-diagonal entries: for i \neq j, with the first equality by Assumption 2, \mathrm{Cov}\left(U_i, U_j \mid X\right) = \mathrm{E}\left[U_i U_j \mid X\right] = 0.
The errors are conditionally uncorrelated across observations. This is no autocorrelation, also called no serial correlation.
Assumption 3 is these two statements together, and nothing more. It does not say the errors are independent, and it does not say they are normal.
Variance of OLS under Assumption 3
Put \mathrm{Var}\left(U \mid X\right) = \sigma^2 I_n into the sandwich: \begin{aligned} \mathrm{Var}\left(\hat\beta \mid X\right) &= (X^\top X)^{-1} X^\top\ \sigma^2 I_n\ X (X^\top X)^{-1} \\ &= \sigma^2 (X^\top X)^{-1} \underbrace{X^\top X (X^\top X)^{-1}}_{=\, I_k} \\ &= \sigma^2 (X^\top X)^{-1}. \end{aligned}
Under Assumptions 1 to 4: \boxed{\ \mathrm{E}\left[\hat\beta \mid X\right] = \beta, \qquad \mathrm{Var}\left(\hat\beta \mid X\right) = \sigma^2 (X^\top X)^{-1}\ }
This pair is the central result of the lecture.
What \sigma^2 is
From Assumptions 2 and 3 together, \sigma^2 = \mathrm{Var}\left(U_i \mid X\right) = \mathrm{E}\left[U_i^2 \mid X\right], and the law of iterated expectations gives \sigma^2 = \mathrm{E}\left[U_i^2\right] = \mathrm{Var}\left(U_i\right), the last step because \mathrm{E}\left[U_i\right] = \mathrm{E}\left[\mathrm{E}\left[U_i \mid X\right]\right] = 0 by Assumption 2.
\sigma^2 is one unknown number, the same for every observation. Estimating it is a question for the next lecture.
One regressor
Take k = 1 and no intercept, so Y_i = \beta X_i + U_i and X is n \times 1. Then X^\top X = \sum_i X_i^2 and \mathrm{Var}\left(\hat\beta \mid X\right) = \frac{\sigma^2}{\sum_{i=1}^n X_i^2}.
Two things move the variance. Noisier errors raise it. A larger \sum_i X_i^2 lowers it: with no intercept this is the sum of squares of the regressor about zero. With an intercept the same formula gives the variance of the slope, with \sum_i (X_i - \bar X)^2, the spread of the regressor about its mean, in place of \sum_i X_i^2.
Why spread in the regressor helps
With the regressor bunched together, many quite different slopes fit the data about equally well, so the estimate moves a lot from sample to sample.
With the regressor spread out, the slope is pinned down.
Without homoskedasticity
If the errors are heteroskedastic but still conditionally uncorrelated across observations, so that \mathrm{Var}\left(U \mid X\right) is diagonal with entries that may depend on X, \mathrm{Var}\left(U \mid X\right) = \Omega = \begin{pmatrix} \sigma_1^2 & & 0 \\ & \ddots & \\ 0 & & \sigma_n^2 \end{pmatrix}, the sandwich does not collapse. The middle factor becomes X^\top \Omega X = \begin{pmatrix} X_1 & \cdots & X_n \end{pmatrix} \Omega \begin{pmatrix} X_1^\top \\ \vdots \\ X_n^\top \end{pmatrix} = \sum_{i=1}^n \sigma_i^2 X_i X_i^\top, so \mathrm{Var}\left(\hat\beta \mid X\right) = (X^\top X)^{-1} \left(\sum_{i=1}^n \sigma_i^2 X_i X_i^\top\right) (X^\top X)^{-1}.
Unbiasedness is untouched. Only the variance formula changes.
Where each assumption was used
| Assumptions | What they deliver |
|---|---|
| 1, 2, 4 | \mathrm{E}\left[\hat\beta \mid X\right] = \beta, and \mathrm{E}\left[\hat\beta\right] = \beta whenever that mean exists |
| 1, 4, finite conditional second moments | \mathrm{Var}\left(\hat\beta \mid X\right) = \Gamma_X \mathrm{Var}\left(U \mid X\right) \Gamma_X^\top, with \Gamma_X = (X^\top X)^{-1}X^\top |
| 1, 3, 4 | \mathrm{Var}\left(\hat\beta \mid X\right) = \sigma^2 (X^\top X)^{-1} |
Gauss-Markov theorem
Linear estimators
Definition. An estimator of \beta is linear if it can be written \tilde\beta = A Y for a matrix A that is k \times n and depends on X alone, not on Y.
OLS is linear, with A = (X^\top X)^{-1} X^\top.
Given X, each entry of a linear estimator is a fixed linear combination of the observations on the dependent variable.
Two more linear estimators
The two-point slope. With k = 1 and X_1 \neq X_n, \tilde\beta = \frac{Y_n - Y_1}{X_n - X_1} = -\frac{1}{X_n - X_1} Y_1 + \frac{1}{X_n - X_1} Y_n. It uses two observations and throws the rest away.
Weighted least squares. For a known n \times n matrix W that is symmetric and positive definite, and that may depend on X but not on Y, \tilde\beta = (X^\top W X)^{-1} X^\top W\, Y. OLS is the case W = I_n, but for other choices of W the estimator is different.
Both are unbiased
The two-point slope, assume Y_i = \beta X_i + U_i and \mathrm{E}\left[U_i \mid X\right] = 0: \tilde\beta = \beta + \frac{U_n - U_1}{X_n - X_1}, \qquad \mathrm{E}\left[\tilde\beta \mid X\right] = \beta.
Weighted least squares, assume Y = X\beta + U and \mathrm{E}\left[U \mid X\right] = 0: \tilde\beta = \beta + (X^\top W X)^{-1} X^\top W\, U, \qquad \mathrm{E}\left[\tilde\beta \mid X\right] = \beta.
Being linear and unbiased does not single OLS out. There are many conditionally unbiased linear estimators.
One can show that under Assumptions 1-4, \mathrm{Var}\left(\tilde\beta \mid X\right) = \sigma^2 (X^\top W X)^{-1} X^\top W\, W X (X^\top W X)^{-1}, which reduces to \sigma^2 (X^\top X)^{-1} when W = I_n.
Which of these matrices is smallest, and what smallest means for matrices, are the next questions.
Comparison to settle
For every linear conditionally unbiased \tilde\beta, compare the two conditional variance matrices: \underset{k\times k}{\mathrm{Var}\left(\tilde\beta \mid X\right)} - \underset{k\times k}{\mathrm{Var}\left(\hat\beta \mid X\right)} \ \overset{?}{\ge}\ 0.
Here \ge 0 means positive semidefinite.
Reading the matrix ordering
If a k \times k matrix M is positive semidefinite, then c^\top M c \ge 0 for every c \in \mathbb{R}^k.
Take c to be the jth coordinate vector e_j, the vector with a 1 in position j and a 0 everywhere else: e_j = (0, \ldots, 0, 1, 0, \ldots, 0)^\top, \qquad \text{the } 1 \text{ sitting in position } j.
Writing the quadratic form out entry by entry, and using (e_j)_l = 0 for every l \neq j and (e_j)_j = 1, e_j^\top M e_j = \sum_{l=1}^k \sum_{m=1}^k (e_j)_l\, M_{lm}\, (e_j)_m = M_{jj}.
So every diagonal entry of a positive semidefinite matrix is non-negative.
Suppose M = \mathrm{Var}\left(\tilde\beta \mid X\right) - \mathrm{Var}\left(\hat\beta \mid X\right)\ge 0 in the positive semidefinite sense. Then: \mathrm{Var}\left(\tilde\beta_j \mid X\right) \ \ge\ \mathrm{Var}\left(\hat\beta_j \mid X\right), \qquad j = 1, \ldots, k.
Gauss-Markov theorem
Theorem. Suppose Assumptions 1 to 4 hold:
- Y = X\beta + U for some \beta \in \mathbb{R}^k;
- \mathrm{E}\left[U \mid X\right] = 0;
- \mathrm{Var}\left(U \mid X\right) = \sigma^2 I_n for some \sigma^2 > 0;
- \operatorname{rank}(X) = k.
Let \hat\beta = (X^\top X)^{-1} X^\top Y be the OLS estimator.
Let \tilde\beta be any estimator that is
- linear: \tilde\beta = A Y, with A a k \times n matrix depending on X alone;
- conditionally unbiased: \mathrm{E}\left[\tilde\beta \mid X\right] = \beta for every value of \beta.
Then \mathrm{Var}\left(\tilde\beta \mid X\right) - \mathrm{Var}\left(\hat\beta \mid X\right)\ge 0 \ \text{(positive semidefinite).}
Meaning: OLS is the best linear unbiased estimator (BLUE): best in the sense above, linear in Y, conditionally unbiased, and an estimator of \beta.
Proof, step 1: linearity and unbiasedness force A X = I_k
Take conditional expectations of \tilde\beta = A Y: \begin{aligned} \beta = \mathrm{E}\left[\tilde\beta \mid X\right] &= \mathrm{E}\left[A(X\beta + U) \mid X\right] \\ &= A X \beta + A \underbrace{\mathrm{E}\left[U \mid X\right]}_{=\,0 \text{ by Assumption 2}} \\ &= A X \beta. \end{aligned}
This must hold for every value of \beta and all realizations of X, so \boxed{\ A X = I_k\ }
Linearity plus conditional unbiasedness is exactly this one condition.
In our examples:
- OLS: A = (X^\top X)^{-1} X^\top, so: A X = (X^\top X)^{-1} X^\top X = I_k.
- Weighted least squares: A = (X^\top W X)^{-1} X^\top W, so: A X = (X^\top W X)^{-1} X^\top W X = I_k.
Proof, step 2: the covariance between \tilde\beta and \hat\beta
Using A X = I_k for the general estimator: \tilde\beta = A X \beta + A U = \beta + A U.
Therefore: \begin{align*} \tilde\beta - \beta &= A U. \end{align*}
For OLS: \hat\beta - \beta = (X^\top X)^{-1} X^\top U.
Consider the covariance between \tilde\beta and \hat\beta: \begin{aligned} \mathrm{Cov}\left(\tilde\beta, \hat\beta \mid X\right) &= \mathrm{E}\left[\big(\tilde\beta - \beta\big)\big(\hat\beta - \beta\big)^\top \,\Big|\, X\right] \\ &= \mathrm{E}\left[\underbrace{A U}_{\tilde\beta - \beta}\, \underbrace{U^\top X (X^\top X)^{-1}}_{(\hat\beta - \beta)^\top} \,\big|\, X\right] \\ &= A\, \underbrace{\mathrm{E}\left[U U^\top \mid X\right]}_{=\,\sigma^2 I_n \text{ by Assumptions 2 and 3}}\, X (X^\top X)^{-1} \\ &= \sigma^2 \underbrace{A X}_{=\, I_k} (X^\top X)^{-1} \\ &= \sigma^2 (X^\top X)^{-1}. \end{aligned}
The right-hand side is \mathrm{Var}\left(\hat\beta \mid X\right), so \boxed{\ \mathrm{Cov}\left(\tilde\beta, \hat\beta \mid X\right) = \mathrm{Var}\left(\hat\beta \mid X\right)\ }
Every linear conditionally unbiased estimator has the same covariance with OLS, namely the variance of OLS itself.
Proof, step 3: the difference of the variances
Expand the variance of the difference: \begin{aligned} \mathrm{Var}\left(\tilde\beta - \hat\beta \mid X\right) &= \mathrm{Var}\left(\tilde\beta \mid X\right) + {\color{blue}\mathrm{Var}\left(\hat\beta \mid X\right)} \\ &\quad- {\color{blue}\mathrm{Cov}\left(\tilde\beta, \hat\beta \mid X\right)} - {\color{blue}\mathrm{Cov}\left(\hat\beta, \tilde\beta \mid X\right)}\\ &= \mathrm{Var}\left(\tilde\beta \mid X\right) + {\color{blue}\mathrm{Var}\left(\hat\beta \mid X\right)} \\ &\quad - \underbrace{\color{blue}\mathrm{Var}\left(\hat\beta \mid X\right)}_{=\mathrm{Cov}\left(\tilde\beta, \hat\beta \mid X\right)} - \underbrace{\color{blue}\mathrm{Var}\left(\hat\beta \mid X\right)}_{=\mathrm{Cov}\left(\hat\beta, \tilde\beta \mid X\right)^\top} \\ &= \mathrm{Var}\left(\tilde\beta \mid X\right) - \mathrm{Var}\left(\hat\beta \mid X\right). \end{aligned}
Therefore: 0\leq \mathrm{Var}\left(\tilde\beta - \hat\beta \mid X\right) = \mathrm{Var}\left(\tilde\beta \mid X\right) - \mathrm{Var}\left(\hat\beta \mid X\right).
Uniqueness
Suppose some linear conditionally unbiased \tilde\beta attains the same variance matrix. By the identity of step 3, \mathrm{Var}\left(\tilde\beta \mid X\right) = \mathrm{Var}\left(\hat\beta \mid X\right) \quad\Longrightarrow\quad \mathrm{Var}\left(\tilde\beta - \hat\beta \mid X\right) = 0.
A random vector with zero conditional variance equals its conditional mean (the equality case of positive semidefiniteness, one coordinate at a time), so \tilde\beta - \hat\beta = \delta(X) for some function of X alone.
Both are conditionally unbiased, so taking conditional expectations, \underbrace{\mathrm{E}\left[\tilde\beta \mid X\right]}_{=\,\beta} = \underbrace{\mathrm{E}\left[\hat\beta \mid X\right]}_{=\,\beta} + \delta(X) \quad\Longrightarrow\quad \delta(X) = 0.
Therefore \tilde\beta = \hat\beta.
Where Assumption 3 was used
Assumption 3 entered in step 2, where \mathrm{E}\left[U U^\top \mid X\right] was replaced by \sigma^2 I_n, and again when \sigma^2 (X^\top X)^{-1} was recognised as \mathrm{Var}\left(\hat\beta \mid X\right). Without it the cancellation that produced \sigma^2 A X fails, and the covariance need not equal the variance of OLS.
If Assumption 3 fails, OLS is still unbiased but need not be best: Assumption 3 is sufficient for efficiency, not necessary.
If Assumption 3 fails: generalised least squares
Suppose now that \mathrm{Var}\left(U \mid X\right) = \Omega for a known positive definite \Omega, not necessarily diagonal, with \Omega \neq \sigma^2 I_n for every \sigma^2 > 0.
Let \Omega^{-1/2} be the symmetric positive definite matrix with \Omega^{-1/2}\Omega^{-1/2} = \Omega^{-1}.
- It is built from the eigenvalue decomposition \Omega^{-1/2} = C \Lambda^{-1/2} C^\top.
Multiply the model through by \Omega^{-1/2}: \begin{align*} &\underbrace{\Omega^{-1/2} Y}_{Y^*} = \underbrace{\Omega^{-1/2} X}_{X^*} \beta + \underbrace{\Omega^{-1/2} U}_{U^*}\\ &Y^* = X^* \beta + U^*,\\ & \mathrm{E}\left[U^* \mid X\right] = 0,\\ & \mathrm{Var}\left(U^* \mid X\right) = \Omega^{-1/2} \Omega \Omega^{-1/2} = I_n. \end{align*}
The transformed model satisfies all four assumptions, so Gauss-Markov applies to it.
OLS on the transformed model is the best linear unbiased estimator: \hat\beta^* = (X^{*\top} X^*)^{-1} X^{*\top} Y^* = (X^\top \Omega^{-1} X)^{-1} X^\top \Omega^{-1} Y, which is weighted least squares with W = \Omega^{-1}. This is the GLS estimator: generalized least squares.
Suppose \mathrm{E}\left[U_i^2 \mid X\right] = \sigma_i^2 and \mathrm{E}\left[U_i U_j \mid X\right] = 0 for i \neq j, so that \Omega = \begin{pmatrix} \sigma_1^2 & & 0 \\ & \ddots & \\ 0 & & \sigma_n^2 \end{pmatrix}.
Then the GLS estimator is: \hat\beta^* = (X^\top \Omega^{-1} X)^{-1} X^\top \Omega^{-1} Y = \left(\sum_{i=1}^n \frac{X_i X_i^\top}{\sigma_i^2}\right)^{-1} \sum_{i=1}^n \frac{X_i Y_i}{\sigma_i^2}. Observations with larger \sigma_i^2 get less weight, and those with smaller \sigma_i^2 get more weight.
This is equivalent transforming the data by scaling each observation i by its \sigma_i:
\frac{Y_i}{\sigma_i} = \frac{X_i^\top}{\sigma_i} \beta + \frac{U_i}{\sigma_i}, \qquad i = 1, \ldots, n, so that \mathrm{Var}\left(U_i / \sigma_i \mid X\right) = 1 for every i.
In practice \Omega is not known.
Summary
Summary: unbiasedness and variance
Substituting the model into the estimator gives the sampling error \hat\beta = \beta + (X^\top X)^{-1} X^\top U. Everything follows from it.
Unbiasedness. Under Assumptions 1, 2 and 4, conditioning on X makes (X^\top X)^{-1}X^\top known, so \mathrm{E}\left[\hat\beta \mid X\right] = \beta, and the law of iterated expectations gives \mathrm{E}\left[\hat\beta\right] = \beta whenever that mean exists. The weaker condition \mathrm{E}\left[X_i U_i\right] = 0 is not enough, because once \mathrm{E}\left[U \mid X\right] is not zero, the expectation of the noise term need not vanish, and the parabola shows that zero correlation does not make the conditional mean zero.
Variance. Under Assumptions 1 and 4, given finite conditional second moments, \mathrm{Var}\left(\hat\beta \mid X\right) = (X^\top X)^{-1} X^\top \mathrm{Var}\left(U \mid X\right) X (X^\top X)^{-1}, and adding Assumption 3, that \mathrm{Var}\left(U \mid X\right) = \sigma^2 I_n, gives \mathrm{Var}\left(\hat\beta \mid X\right) = \sigma^2 (X^\top X)^{-1}. With one regressor and no intercept this is \sigma^2 / \sum_i X_i^2: noisier errors raise it; a larger \sum_i X_i^2 lowers it.
Summary: efficiency
Gauss-Markov. Under Assumptions 1 to 4, OLS has the smallest conditional variance among all estimators that are linear in Y and conditionally unbiased, whatever the value of \beta. For every such \tilde\beta, the difference of the two conditional variance matrices is positive semidefinite: \mathrm{Var}\left(\tilde\beta \mid X\right) - \mathrm{Var}\left(\hat\beta \mid X\right) \ \ge\ 0, so in particular \mathrm{Var}\left(\tilde\beta_j \mid X\right) \ge \mathrm{Var}\left(\hat\beta_j \mid X\right) for every coefficient.
The proof. Linearity and conditional unbiasedness force A X = I_k. Assumption 3 then makes \mathrm{Cov}\left(\tilde\beta, \hat\beta \mid X\right) = \mathrm{Var}\left(\hat\beta \mid X\right). Expanding \mathrm{Var}\left(\tilde\beta - \hat\beta \mid X\right), which is positive semidefinite, leaves exactly the difference of the two variances.
Uniqueness. Equality of the variances forces \tilde\beta = \hat\beta.
Assumption 3 is what buys efficiency.
Assumptions 1 and 2 give unbiasedness.
Assumption 4 gives existence of the estimator.