Lecture 5: Hypothesis testing

Economics 527: Econometric Methods

Author

Vadim Marmer, UBC

Hypotheses and errors

Estimator and parameter set

  • Normally distributed estimator with known variance \omega^2: \underset{1\times1}{\hat\theta} \sim N\big(\underset{1\times1}{\theta},\ \underset{1\times1}{\omega^2}\big), \qquad \frac{\hat\theta - \theta}{\omega} = Z \sim N(0,1)

  • Regression example, conditional on X: \theta = \beta_1,\, \hat\theta = \hat\beta_1, \, \omega^2 = \sigma^2 / (X_1^\top M_2 X_1).

  • Parameter set, for example \Theta = \mathbb{R}: \theta \in \Theta

  • Two disjoint parts: \begin{aligned} \Theta &= \Theta_0 \cup \Theta_1 \\ \Theta_0 \cap \Theta_1 &= \emptyset \end{aligned}

  • Null and alternative hypotheses: H_0 : \theta \in \Theta_0 \qquad \text{vs} \qquad H_1 : \theta \in \Theta_1

Examples of \Theta_0 and \Theta_1

  • Point null, \Theta = \mathbb{R}: \begin{aligned} \Theta_0 &= \{0\} \\ \Theta_1 &= \mathbb{R} \setminus \{0\} \\ &\Longleftrightarrow\quad H_0 : \theta = 0 \quad \text{vs} \quad H_1 : \theta \neq 0 \end{aligned}

  • More generally, for a known \theta_0: \begin{aligned} \Theta_0 &= \{\theta_0\} \\ \Theta_1 &= \mathbb{R} \setminus \{\theta_0\} \\ &\Longleftrightarrow\quad H_0 : \theta = \theta_0 \quad \text{vs} \quad H_1 : \theta \neq \theta_0 \end{aligned}

  • One-sided null: \begin{aligned} \Theta_0 &= [\,\theta_0,\ +\infty) \\ \Theta_1 &= (-\infty,\ \theta_0) \\ &\Longleftrightarrow\quad H_0 : \theta \ge \theta_0 \quad \text{vs} \quad H_1 : \theta < \theta_0 \end{aligned}

  • Definition. Simple hypothesis: a single-point set; otherwise composite: \Theta_0 = \{\theta_0\} : \ \text{simple}, \qquad \Theta_0 = [\,0,\ +\infty) : \ \text{composite}

Test statistic and decision rule

  • Test statistic T: a function of the data, random, with range S: \begin{aligned} T &\in S \\ S &= S_A \cup S_R, \qquad S_A \cap S_R = \emptyset \end{aligned}

  • S_A: acceptance region; S_R: rejection region.

  • Decision rule: \begin{aligned} T \in S_A \quad &\Longrightarrow \quad \text{decide in favor of } \theta \in \Theta_0 \qquad (\text{fail to reject } H_0) \\ T \in S_R \quad &\Longrightarrow \quad \text{decide in favor of } \theta \in \Theta_1 \qquad (\text{reject } H_0) \end{aligned}

Type I and Type II errors

Truth against decision:

Truth Decision H_0 Decision H_1
H_0 correct Type I error
H_1 Type II error correct
  • Probabilities of the errors: \begin{aligned} \mathrm{P}(\text{Type I error}) &= \mathrm{P}(T \in S_R \mid \theta \in \Theta_0) \\ \mathrm{P}(\text{Type II error}) &= \mathrm{P}(T \notin S_R \mid \theta \in \Theta_1) \\ &= \mathrm{P}(T \in S_A \mid \theta \in \Theta_1) \end{aligned}

  • To make \mathrm{P}(\text{Type I error}) smaller: \begin{aligned} & S_R \downarrow \\ \Longrightarrow\quad & S_A \uparrow \\ \Longrightarrow\quad & \mathrm{P}(\text{Type II error}) \uparrow \end{aligned}

  • We can’t make both error probabilities small at the same time.

Valid test, size, significance level

  • Definition. Valid test, with \alpha chosen in advance: \begin{aligned} \mathrm{P}(\text{Type I error}) &= \mathrm{P}(\text{falsely deciding in favour of } H_1 \mid H_0 \text{ true}) \\ &\le \alpha \end{aligned}

  • Composite \Theta_0, supremum of the Type I error probabilities: \sup_{\theta \in \Theta_0} \mathrm{P}(T \in S_R \mid \theta) \le \alpha

  • Size \alpha test: \sup_{\theta \in \Theta_0} \mathrm{P}(T \in S_R \mid \theta) = \alpha

  • Significance level \alpha: small, for example 0.01, 0.05, 0.1.

Two-sided test

Point null against a two-sided alternative

  • Hypotheses, for example \theta_0 = 0: \underbrace{H_0 : \theta = \theta_0}_{\text{simple}} \qquad \text{vs} \qquad \underbrace{H_1 : \theta \neq \theta_0}_{\text{composite}}

  • Three objects: \begin{aligned} \theta &: \ \text{true value, unknown, fixed} \\ \hat\theta &: \ \text{estimator, data, random} \\ \theta_0 &: \ \text{chosen by you, known, fixed} \end{aligned}

  • Regression example, \theta = \beta_1 and \theta_0 = 0: H_0 : \beta_1 = 0 \qquad \text{vs} \qquad H_1 : \beta_1 \neq 0

Test statistic T(\theta_0)

  • Distance between the data and H_0: T(\theta_0) = \left|\frac{\hat\theta - \theta_0}{\omega}\right|

  • Reject H_0 when: \begin{aligned} & T(\theta_0) \ \text{ is large} \\ \Longleftrightarrow\quad & T(\theta_0) > c \end{aligned}

  • Regions, c > 0: \begin{aligned} S_R &= (c,\ +\infty) \\ S_A &= [\,0,\ c\,] \\ S &= S_A \cup S_R = [\,0,\ +\infty) \end{aligned}

  • Choose c for size control: \mathrm{P}(\text{Type I error}) \le \alpha

Choice of c

  • Size \alpha: set the Type I error probability at \theta = \theta_0 to \alpha: \begin{aligned} \alpha &= \mathrm{P}(\text{Type I error}) \\ &= \mathrm{P}(\text{reject } H_0 \mid H_0 \text{ true}) \\ &= \mathrm{P}\big(T(\theta_0) > c \mid \theta = \theta_0\big) \\ &= \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > c \;\Big|\; \theta = \theta_0\Big) \\ &\overset{\theta = \theta_0}{=} \mathrm{P}\Big(\Big|\underbrace{\frac{\hat\theta - \theta}{\omega}}_{=\,Z \sim N(0,1)}\Big| > c\Big) \\ &= \mathrm{P}(|Z| > c) \\ &= \mathrm{P}\big(\underbrace{\{Z > c\}}_{A} \ \text{ or } \ \underbrace{\{-Z > c\}}_{B}\big) \\ &= \mathrm{P}(Z > c) + \mathrm{P}(Z < -c) \\ &= 2\,\mathrm{P}(Z > c) \\ \Longrightarrow\quad \mathrm{P}(Z > c) &= \frac{\alpha}{2} \\ \Longrightarrow\quad \mathrm{P}(Z \le c) &= 1 - \frac{\alpha}{2} \\ \Longrightarrow\quad c &= z_{1-\alpha/2} \qquad \big(\mathrm{P}(Z \le z_\tau) = \tau\big) \end{aligned}

Size \alpha two-sided test

  • Size \alpha test of H_0 : \theta = \theta_0 against H_1 : \theta \neq \theta_0: \boxed{\ \text{Reject } H_0 \ \text{ if } \ \left|\frac{\hat\theta - \theta_0}{\omega}\right| > z_{1-\alpha/2}\ }

  • For \alpha = 0.05: \begin{aligned} & z_{0.975} \approx 1.96 \approx 2 \\ \Longrightarrow\quad & \text{reject if } \ \Big|\frac{\hat\theta - \theta_0}{\omega}\Big| \gtrsim 2 \\ \Longleftrightarrow\quad & \text{reject if } \ |\hat\theta - \theta_0| \gtrsim 2\,\underbrace{\omega}_{\text{std. dev.}} \end{aligned}

  • Compare with: Reject if |\hat\theta - \theta_0| > 5\,\omega: \mathrm{P}(Z > 5) \approx 3 \cdot 10^{-7}, \qquad \alpha = \mathrm{P}(|Z| > 5) \approx 6 \cdot 10^{-7}

Significance at level \alpha

  • Definition. Significance at level \alpha, by a size \alpha test: \begin{aligned} \Big|\frac{\hat\theta}{\omega}\Big| > z_{1-\alpha/2} \quad&\Longrightarrow\quad H_0 : \theta = 0 \ \text{ rejected in favour of } \ H_1 : \theta \neq 0 \\ &\Longrightarrow\quad \hat\theta \ \text{ significant at level } \alpha \end{aligned}

Test and confidence interval

  • Solve for \theta_0: \begin{aligned} H_0 \text{ rejected} \quad&\Longleftrightarrow\quad \left|\frac{\hat\theta - \theta_0}{\omega}\right| > z_{1-\alpha/2} \\ H_0 \text{ not rejected} \quad&\Longleftrightarrow\quad \left|\frac{\hat\theta - \theta_0}{\omega}\right| \le z_{1-\alpha/2} \\ &\Longleftrightarrow\quad -z_{1-\alpha/2} \le \frac{\hat\theta - \theta_0}{\omega} \le z_{1-\alpha/2} \\ &\Longleftrightarrow\quad -z_{1-\alpha/2}\,\omega \le \hat\theta - \theta_0 \le z_{1-\alpha/2}\,\omega \\ &\Longleftrightarrow\quad z_{1-\alpha/2}\,\omega \ge \theta_0 - \hat\theta \ge -z_{1-\alpha/2}\,\omega \\ &\Longleftrightarrow\quad \hat\theta + z_{1-\alpha/2}\,\omega \ge \theta_0 \ge \hat\theta - z_{1-\alpha/2}\,\omega \\ &\Longleftrightarrow\quad \hat\theta - z_{1-\alpha/2}\,\omega \le \theta_0 \le \hat\theta + z_{1-\alpha/2}\,\omega \\ &\Longleftrightarrow\quad \theta_0 \in CI_{1-\alpha} = \big[\hat\theta - z_{1-\alpha/2}\,\omega,\ \hat\theta + z_{1-\alpha/2}\,\omega\big] \end{aligned}

  • Equivalent statement of the test: \boxed{\ \text{Reject } H_0 : \theta = \theta_0 \ \text{ if } \ \theta_0 \notin CI_{1-\alpha}\ }

CI_{1-\alpha} as a set of null values

  • All \theta_0 the test does not reject: \begin{aligned} CI_{1-\alpha} &= \big\{\theta_0 : T(\theta_0) \le z_{1-\alpha/2}\big\} \\ &= \Big\{\theta_0 : \Big|\frac{\hat\theta - \theta_0}{\omega}\Big| \le z_{1-\alpha/2}\Big\} \\ &= \big\{\theta_0 : \text{a size } \alpha \text{ test does not reject } H_0 : \theta = \theta_0\big\} \end{aligned}

  • Suppose in your data: CI_{0.95} = [\,0.04,\ 0.06\,]

Power

Noise and signal

  • Type II error: \mathrm{P}(\text{Type II error}) = \mathrm{P}(\text{fail to reject } H_0 \mid \theta \neq \theta_0)

  • Split (\hat\theta - \theta_0)/\omega, under the true \theta: \begin{aligned} \underbrace{\frac{\hat\theta - \theta_0}{\omega}}_{\text{data} \,-\, H_0} &= \underbrace{\frac{\hat\theta - \theta}{\omega}}_{=\,Z \sim N(0,1),\ \text{noise}} + \underbrace{\frac{\theta - \theta_0}{\omega}}_{\text{truth} \,-\, H_0,\ \text{signal}} \\ &= Z + \frac{\theta - \theta_0}{\omega} \\ &\sim N\Big(\frac{\theta - \theta_0}{\omega},\ 1\Big) \end{aligned}

Power function \pi(\theta)

  • Probability of rejecting H_0, as a function of the true \theta: \begin{aligned} \pi(\theta) &= \mathrm{P}(\text{reject } H_0 \mid \theta) \\ &= \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2} \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta}{\omega} + \frac{\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2} \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(\Big|\underbrace{Z}_{N(0,1)} + \frac{\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2}\Big) \end{aligned}

  • Known: \theta_0, \omega, z_{1-\alpha/2}, the distribution of Z.

  • Unknown: \theta \in \mathbb{R}.

Type II error probability

  • Type II error probability depends on \theta: For \theta \neq \theta_0, \begin{aligned} \mathrm{P}(\text{Type II error} \mid \theta) &= \mathrm{P}(\text{fail to reject } H_0 \mid \theta) \\ &= 1-\mathrm{P}(\text{reject } H_0 \mid \theta) \\ &= 1 - \pi(\theta). \end{aligned}

Distribution of (\hat\theta - \theta_0)/\omega under H_0 and H_1

  • Density of (\hat\theta - \theta_0)/\omega, with \theta_0 < \theta_1 < \theta_2:

  • Under \theta \neq \theta_0, more mass in the rejection region: \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2} \;\Big|\; \theta\Big) > \alpha

Power curve

  • \pi(\theta) against \theta: minimum \alpha at \theta_0, approaches 1 far from \theta_0:

  • Largest Type II error probability, as \theta \to \theta_0: \begin{aligned} \sup_{\theta \neq \theta_0} \mathrm{P}(\text{Type II error} \mid \theta) &= \sup_{\theta \neq \theta_0} \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| \le z_{1-\alpha/2} \;\Big|\; \theta\Big) \\ &= \sup_{\theta \neq \theta_0} \mathrm{P}\Big(\Big|Z + \frac{\theta - \theta_0}{\omega}\Big| \le z_{1-\alpha/2}\Big) \\ &= 1 - \alpha \end{aligned}

Maximum against supremum

  • Closed interval: \max_{x \in [0,1]} x = 1

  • Half-open interval: \begin{aligned} \max_{x \in [0,1)} x \ &\ \text{does not exist} \\ \sup_{x \in [0,1)} x \ &= 1 \end{aligned}

  • Definition. Supremum: \sup_{x \in \mathcal{X}} f(x) = \text{smallest upper bound of } \{f(x) : x \in \mathcal{X}\}

  • \mathrm{P}(\text{Type II error} \mid \theta) over \theta \neq \theta_0: no maximum.

Example: \hat\theta \sim N(\theta,\ 9)

  • H_0 : \theta = 0 vs H_1 : \theta \neq 0, \alpha = 0.05, \omega = 3: \text{Reject } H_0 \ \text{ if } \ \Big|\frac{\hat\theta}{3}\Big| > 1.96

  • Power function: \begin{aligned} \pi(\theta) &= \mathrm{P}\Big(\Big|\frac{\hat\theta}{3}\Big| > 1.96 \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta}{3} + \frac{\theta}{3}\Big| > 1.96\Big) \\ &= \mathrm{P}\Big(\Big|Z + \frac{\theta}{3}\Big| > 1.96\Big) \\ &= \mathrm{P}\Big(Z + \frac{\theta}{3} > 1.96 \ \text{ or } \ Z + \frac{\theta}{3} < -1.96\Big) \\ &= \mathrm{P}\Big(Z > 1.96 - \frac{\theta}{3}\Big) + \mathrm{P}\Big(Z < -1.96 - \frac{\theta}{3}\Big) \end{aligned}

Example, power at \theta = 3

  • Substitute \theta = 3: \begin{aligned} \pi(3) &= \mathrm{P}\Big(Z > 1.96 - \frac{3}{3}\Big) + \mathrm{P}\Big(Z < -1.96 - \frac{3}{3}\Big) \\ &= \mathrm{P}(Z > 0.96) + \mathrm{P}(Z < -2.96) \\ &\approx \underbrace{\mathrm{P}(Z \le -0.96)}_{\approx\,0.16853} + 0.00154 \\ &\approx 0.17 \end{aligned}

  • \pi(\theta) at four values of \theta: \pi(\theta) \approx \begin{cases} 0.063 & \theta = -1 \\ 0.05 & \theta = 0 \\ 0.063 & \theta = 1 \\ 0.17 & \theta = 3 \end{cases}

Failing to reject \beta_1 = 0

  • Regression, \theta = \beta_1, \omega^2 = \sigma^2 / (X_1^\top M_2 X_1): Y = \beta_1 X_1 + X_2 \beta_2 + U, \qquad H_0 : \beta_1 = 0 \quad \text{vs} \quad H_1 : \beta_1 \neq 0

  • Suppose H_0 is not rejected: \Big|\frac{\hat\beta_1}{\omega}\Big| \le z_{1-\alpha/2} \qquad \overset{?}{\Longrightarrow} \qquad \beta_1 = 0

  • Type II error probability near 1 - \alpha for \theta near \theta_0.

  • T(\theta_0) > z_{1-\alpha/2}: strong evidence that \theta \neq \theta_0.

  • Claim. T(\theta_0) \le z_{1-\alpha/2}: not strong evidence in favour of \theta_0.

Minimum detectable deviation

  • Detectable \theta: \pi(\theta) \ge \beta, for example \beta = 0.85 or 0.5.

  • Smallest detectable \bar\theta > \theta_0 solves: \mathrm{P}\Big(\Big|Z + \frac{\bar\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2}\Big) = \beta, \qquad \theta_0,\ \omega,\ z_{1-\alpha/2},\ \beta \ \text{ known}

  • |\bar\theta - \theta_0|: the minimum detectable deviation.

More precise estimator, more power

  • Smaller variance: \begin{aligned} \hat\theta^* &\sim N(\theta,\ \omega_*^2) \\ \omega_*^2 &< \omega^2 \\ \Longrightarrow\quad \omega_* &< \omega \end{aligned}

  • Both tests valid: \underbrace{\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2}}_{\text{test with } \hat\theta} \qquad \text{or} \qquad \underbrace{\Big|\frac{\hat\theta^* - \theta_0}{\omega_*}\Big| > z_{1-\alpha/2}}_{\text{test with } \hat\theta^*}

  • Suppose \theta > \theta_0, larger signal with \hat\theta^*: \begin{aligned} \frac{\hat\theta - \theta_0}{\omega} &= Z + \frac{\theta - \theta_0}{\omega} \sim N\Big(\frac{\theta - \theta_0}{\omega},\ 1\Big) \\ \frac{\hat\theta^* - \theta_0}{\omega_*} &\sim N\Big(\frac{\theta - \theta_0}{\omega_*},\ 1\Big) \\ \omega_* < \omega \quad &\Longrightarrow\quad \frac{\theta - \theta_0}{\omega} < \frac{\theta - \theta_0}{\omega_*} \end{aligned}

  • Power functions: \begin{aligned} \pi(\theta) &= \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2} \;\Big|\; \theta\Big) \\ \pi^*(\theta) &= \mathrm{P}\Big(\Big|\frac{\hat\theta^* - \theta_0}{\omega_*}\Big| > z_{1-\alpha/2} \;\Big|\; \theta\Big) \end{aligned}

  • Test with \hat\theta^*: more powerful, \pi^*(\theta) \ge \pi(\theta).

Test that ignores the data

  • V \sim \text{Uniform}(0,1), density and CDF: \begin{aligned} f(v) &= \mathbf{1}\{0 < v < 1\} \\ &= \begin{cases} 1 & 0 < v < 1 \\ 0 & \text{otherwise} \end{cases} \\ F(v) &= \begin{cases} 0 & v \le 0 \\ v & 0 < v < 1 \\ 1 & v \ge 1 \end{cases} \end{aligned}

  • For 0 < x < 1: \begin{aligned} \mathrm{P}(V < x) &= \mathrm{P}(V \le x) \qquad (\text{continuous}) \\ &= F(x) \\ &= \int_0^x f(v)\, dv \\ &= \int_0^x 1 \, dv \\ &= x \end{aligned}

  • Test:

    1. Throw \hat\theta away.
    2. Draw V \sim \text{Uniform}(0,1).
    3. Reject H_0 : \theta = \theta_0 if V < \alpha.
  • V does not depend on \theta: \begin{aligned} \mathrm{P}(\text{Type I error} \mid \theta = \theta_0) &= \mathrm{P}(V < \alpha) \\ &= \alpha \\ \tilde\pi(\theta) &= \mathrm{P}(\text{reject } H_0 \mid \theta) \\ &= \mathrm{P}(V < \alpha) \\ &= \alpha \qquad \text{for every } \theta \end{aligned}

  • Power \alpha: trivial.

One-sided tests

One-sided hypotheses

  • One-sided null and alternative, for example \theta_0 = 0: \underbrace{H_0 : \theta \le \theta_0}_{\text{composite}} \qquad \text{vs} \qquad \underbrace{H_1 : \theta > \theta_0}_{\text{composite}}

  • Test: reject H_0 in favour of H_1 if: \frac{\hat\theta - \theta_0}{\omega} > c

  • Type I error probability for each \theta \in \Theta_0, and its supremum: \begin{aligned} \mathrm{P}(\text{Type I error} \mid \theta) &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > c \;\Big|\; \theta\Big), \qquad \theta \le \theta_0 \\ \sup_{\theta \le \theta_0} \mathrm{P}(\text{Type I error} \mid \theta) &\le \alpha \end{aligned}

  • Split (\hat\theta - \theta_0)/\omega: \frac{\hat\theta - \theta_0}{\omega} = \underbrace{\frac{\hat\theta - \theta}{\omega}}_{=\,Z} + \underbrace{\frac{\theta - \theta_0}{\omega}}_{\le\, 0 \ \text{ for } \theta \le \theta_0}

Choice of c, one-sided

  • Supremum over \Theta_0, attained at \theta = \theta_0: \begin{aligned} \alpha &= \sup_{\theta \le \theta_0} \mathrm{P}(\text{Type I error} \mid \theta) \\ &= \sup_{\theta \le \theta_0} \mathrm{P}(\text{reject } H_0 \mid \theta) \\ &= \sup_{\theta \le \theta_0} \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > c \;\Big|\; \theta\Big) \\ &= \sup_{\theta \le \theta_0} \mathrm{P}\Big(\underbrace{\frac{\hat\theta - \theta}{\omega}}_{=\,Z \sim N(0,1)} + \frac{\theta - \theta_0}{\omega} > c \;\Big|\; \theta\Big) \\ &= \sup_{\theta \le \theta_0} \mathrm{P}\Big(Z + \underbrace{\frac{\theta - \theta_0}{\omega}}_{\le\,0,\ \text{largest at } \theta = \theta_0} > c\Big) \\ &= \mathrm{P}(Z > c) \\ \Longrightarrow\quad \mathrm{P}(Z \le c) &= 1 - \alpha \\ \Longrightarrow\quad c &= z_{1-\alpha} \end{aligned}

One-sided test and interval

  • Size \alpha test of H_0 : \theta \le \theta_0 against H_1 : \theta > \theta_0: \boxed{\ \text{Reject } H_0 \ \text{ if } \ \frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha}\ }

  • One-sided interval, all \theta_0 the test does not reject: \begin{aligned} CI^{1}_{1-\alpha} &= \Big\{\theta_0 : \frac{\hat\theta - \theta_0}{\omega} \le z_{1-\alpha}\Big\} \\ &= \big\{\theta_0 : \theta_0 \ge \hat\theta - z_{1-\alpha}\,\omega\big\} \\ &= \big[\hat\theta - z_{1-\alpha}\,\omega,\ +\infty\big) \end{aligned}

Mirror case: H_0 : \theta \ge \theta_0

  • Reject H_0 : \theta \ge \theta_0 in favour of H_1 : \theta < \theta_0 if, for some c > 0: \frac{\hat\theta - \theta_0}{\omega} < -c

  • Choose c from the supremum over \theta \ge \theta_0: \begin{aligned} \alpha &\ge \mathrm{P}(\text{Type I error} \mid \theta) \qquad (\theta \ge \theta_0) \\ &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} < -c \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} < -c\Big) \\ \alpha &= \sup_{\theta \ge \theta_0} \mathrm{P}\Big(Z + \underbrace{\frac{\theta - \theta_0}{\omega}}_{\ge\,0} < -c\Big) \qquad (\text{size } \alpha) \\ &= \mathrm{P}(Z < -c) \qquad (\theta = \theta_0) \\ \Longrightarrow\quad -c &= z_{\alpha} \\ \Longrightarrow\quad c &= z_{1-\alpha} \end{aligned}

  • Test: \boxed{\ \text{Reject } H_0 : \theta \ge \theta_0 \ \text{ if } \ \frac{\hat\theta - \theta_0}{\omega} < -z_{1-\alpha}\ }

Power through \Phi

  • \Phi(x) = \mathrm{P}(Z \le x), the CDF of N(0,1).

  • Left-tailed test of H_0 : \theta \ge \theta_0: \begin{aligned} \pi(\theta) &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} < -z_{1-\alpha} \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} < -z_{1-\alpha}\Big) \\ &= \mathrm{P}\Big(Z < -z_{1-\alpha} - \frac{\theta - \theta_0}{\omega}\Big) \\ &= \Phi\Big(\frac{\theta_0 - \theta}{\omega} - z_{1-\alpha}\Big) \end{aligned}

Two tests for H_0 : \theta = \theta_0

  • Against H_1 : \theta \neq \theta_0, both of size \alpha: \begin{aligned} \text{Test 1 (one-sided)} &: \quad \text{reject } H_0 \ \text{ if } \ \frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha} \\ \text{Test 2 (two-sided)} &: \quad \text{reject } H_0 \ \text{ if } \ \Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2} \end{aligned}

  • Type I error of Test 1: \begin{aligned} \mathrm{P}(\text{Type I error of Test 1}) &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha} \;\Big|\; \theta = \theta_0\Big) \\ &\overset{\theta = \theta_0}{=} \mathrm{P}\Big(\frac{\hat\theta - \theta}{\omega} > z_{1-\alpha}\Big) \\ &= \mathrm{P}(Z > z_{1-\alpha}) \\ &= \alpha \end{aligned}

  • Critical values: \begin{aligned} z_{1-\alpha} &< z_{1-\alpha/2} \\ \Longrightarrow\quad -z_{1-\alpha/2} &< -z_{1-\alpha} \end{aligned}

Power of Test 1 and Test 2

  • Power of Test 1: \begin{aligned} \pi_1(\theta) &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha} \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} > z_{1-\alpha}\Big) \end{aligned}

  • Power of Test 2: \begin{aligned} \pi_2(\theta) &= \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2} \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(\Big|Z + \frac{\theta - \theta_0}{\omega}\Big| > z_{1-\alpha/2}\Big) \\ &= \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} < -z_{1-\alpha/2}\Big) + \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} > z_{1-\alpha/2}\Big) \end{aligned}

One-sided against two-sided power

  • \pi_1(\theta) increasing, \pi_2(\theta) U-shaped:

  • For \theta > \theta_0, Test 1 is more powerful; compare the right tails: \begin{aligned} \Big|\frac{\hat\theta - \theta_0}{\omega}\Big| &= \frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha/2} \qquad \Big(\frac{\hat\theta - \theta_0}{\omega} > 0\Big) \\ \Longrightarrow\quad \frac{\hat\theta - \theta_0}{\omega} &> z_{1-\alpha} \end{aligned}

  • For \theta < \theta_0, Test 2 is more powerful: \begin{aligned} \frac{\hat\theta - \theta_0}{\omega} &= Z + \underbrace{\frac{\theta - \theta_0}{\omega}}_{<\,0} > z_{1-\alpha} \qquad (\text{Test 1 rejects}) \\ \pi_1(\theta) &= \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} > z_{1-\alpha}\Big) \\ &= \mathrm{P}\Big(Z > z_{1-\alpha} - \frac{\theta - \theta_0}{\omega}\Big) \\ &< \mathrm{P}(Z > z_{1-\alpha}) \\ &= \alpha \end{aligned}

  • No uniformly most powerful test against H_1 : \theta \neq \theta_0.

  • Test 1 biased against H_1 : \theta \neq \theta_0.

  • Largest Type II error probability of Test 1: \begin{aligned} \text{against } H_1 : \theta > \theta_0 &: \quad \sup_{\theta > \theta_0} \big(1 - \pi_1(\theta)\big) = 1 - \alpha \\ \text{against } H_1 : \theta \neq \theta_0 &: \quad \sup_{\theta \neq \theta_0} \big(1 - \pi_1(\theta)\big) = 1 \end{aligned}

Two-sided test with unequal tails

  • Test 3 of H_0 : \theta = \theta_0 against H_1 : \theta \neq \theta_0, reject H_0 when: \frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha/3} \qquad \text{or} \qquad \frac{\hat\theta - \theta_0}{\omega} < -z_{1-2\alpha/3}

  • Type I error probability, with z_\tau = -z_{1-\tau}: \begin{aligned} \mathrm{P}(\text{Type I error}) &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha/3} \ \text{ or } \ \frac{\hat\theta - \theta_0}{\omega} < -z_{1-2\alpha/3} \;\Big|\; \theta = \theta_0\Big) \\ &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha/3} \;\Big|\; \theta = \theta_0\Big) \\ &\qquad + \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} < -z_{1-2\alpha/3} \;\Big|\; \theta = \theta_0\Big) \\ &\overset{\theta = \theta_0}{=} \mathrm{P}\Big(\frac{\hat\theta - \theta}{\omega} > z_{1-\alpha/3}\Big) + \mathrm{P}\Big(\frac{\hat\theta - \theta}{\omega} < -z_{1-2\alpha/3}\Big) \\ &= \underbrace{\mathrm{P}(Z > z_{1-\alpha/3})}_{=\,\alpha/3} + \underbrace{\mathrm{P}(Z < -z_{1-2\alpha/3})}_{=\,\mathrm{P}(Z < z_{2\alpha/3})\,=\,2\alpha/3} \\ &= \alpha \end{aligned}

Power with unequal tails

  • Power of Test 3: \begin{aligned} \pi_3(\theta) &= \mathrm{P}(\text{reject } H_0 \mid \theta) \\ &= \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha/3} \;\Big|\; \theta\Big) + \mathrm{P}\Big(\frac{\hat\theta - \theta_0}{\omega} < -z_{1-2\alpha/3} \;\Big|\; \theta\Big) \\ &= \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} > z_{1-\alpha/3}\Big) + \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} < -z_{1-2\alpha/3}\Big) \\ &= \mathrm{P}\Big(Z > z_{1-\alpha/3} - \frac{\theta - \theta_0}{\omega}\Big) + \mathrm{P}\Big(Z < -z_{1-2\alpha/3} - \frac{\theta - \theta_0}{\omega}\Big) \end{aligned}

  • At \theta = \theta_0: \begin{aligned} \pi_3(\theta_0) &= \mathrm{P}(Z > z_{1-\alpha/3}) + \mathrm{P}(Z < -z_{1-2\alpha/3}) \\ &= \frac{\alpha}{3} + \frac{2\alpha}{3} \\ &= \alpha \end{aligned}

  • For \theta > \theta_0, (\theta - \theta_0)/\omega > 0: \pi_3(\theta) = \underbrace{\mathrm{P}\Big(Z > z_{1-\alpha/3} - \frac{\theta - \theta_0}{\omega}\Big)}_{>\,\alpha/3} + \underbrace{\mathrm{P}\Big(Z < -z_{1-2\alpha/3} - \frac{\theta - \theta_0}{\omega}\Big)}_{<\,2\alpha/3}

  • \pi_3(\theta) against \pi_2(\theta), both \alpha at \theta_0:

  • \pi_3(\theta) < \alpha just right of \theta_0: Test 3 biased against H_1 : \theta \neq \theta_0.

Strict inequality in the null

  • H_0 : \theta < \theta_0 vs H_1 : \theta \ge \theta_0, same one-sided test: \begin{aligned} \sup_{\theta < \theta_0} \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} > z_{1-\alpha}\Big) &= \sup_{\theta \in (-\infty,\ \theta_0)} \mathrm{P}\Big(Z + \frac{\theta - \theta_0}{\omega} > z_{1-\alpha}\Big) \\ &= \mathrm{P}(Z > z_{1-\alpha}) \\ &= \alpha \end{aligned}

  • \theta_0 \in \Theta_1: no power above \alpha at \theta_0.

Null \theta \neq \theta_0: not testable

  • H_0 : \theta \neq \theta_0 vs H_1 : \theta = \theta_0; size control: \sup_{\theta \neq \theta_0} \mathrm{P}(\text{reject } H_0 \mid \theta) \le \alpha

  • Reject if |(\hat\theta - \theta_0)/\omega| < c: \begin{aligned} \alpha &\ge \sup_{\theta \neq \theta_0} \mathrm{P}\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big| < c \;\Big|\; \theta\Big) \\ &= \sup_{\theta \neq \theta_0} \mathrm{P}\Big(\Big|Z + \frac{\theta - \theta_0}{\omega}\Big| < c\Big) \\ &= \mathrm{P}\big(|Z| < c\big) \qquad (\theta \to \theta_0) \\ &= \mathrm{P}(\text{reject } H_0 \mid \theta = \theta_0) \end{aligned}

  • Size \alpha: no power above \alpha at \theta_0, not testable.

p-values

Question

  • Question. What is the p-value?
    1. \mathrm{P}(\text{Type I error})
    2. Probability of drawing \hat\theta from the distribution centred at \theta
    3. Probability of rejecting
    4. Probability of drawing another \hat\theta further than the one in your sample

Two-sided p-value

  • Two tails beyond \pm T(\theta_0):

  • p-value: \begin{aligned} \text{p-value} &= 2\big(1 - \Phi(T(\theta_0))\big) \\ &= \boxed{\ 2\Big(1 - \Phi\Big(\Big|\frac{\hat\theta - \theta_0}{\omega}\Big|\Big)\Big)\ } \\ &= 2\,\Phi\Bigl(-\Big|\frac{\hat\theta - \theta_0}{\omega}\Big|\Bigr) \qquad \big(1 - \Phi(x) = \Phi(-x)\big) \end{aligned}

  • p-value: random before the data are plugged in, a statistic.

  • Test: \text{Reject } H_0 : \theta = \theta_0 \ \text{ if } \ \text{p-value} < \alpha

  • Claim. Size \alpha: \begin{aligned} \mathrm{P}(\text{Type I error}) &= \mathrm{P}(\text{p-value} < \alpha \mid \theta = \theta_0) \\ &= \mathrm{P}\big(2(1 - \Phi(|Z|)) < \alpha\big) \qquad \Big(\tfrac{\hat\theta - \theta_0}{\omega} = Z \text{ under } \theta = \theta_0\Big) \\ &= \alpha \end{aligned}

Two-sided p-value under \theta = \theta_0

  • CDF of the p-value, for t \in (0,1): \begin{aligned} \mathrm{P}(\text{p-value} \le t \mid \theta = \theta_0) &= \mathrm{P}\Big(2\Big(1 - \Phi\Big(\Big|\underbrace{\frac{\hat\theta - \theta_0}{\omega}}_{=\,Z \sim N(0,1)}\Big|\Big)\Big) \le t \;\Big|\; \theta = \theta_0\Big) \\ &= \mathrm{P}\big(2(1 - \Phi(|Z|)) \le t\big) \\ &= \mathrm{P}\Big(1 - \Phi(|Z|) \le \frac{t}{2}\Big) \\ &= \mathrm{P}\Big(\Phi(|Z|) \ge 1 - \frac{t}{2}\Big) \\ &= \mathrm{P}\Big(|Z| \ge \Phi^{-1}\Big(1 - \frac{t}{2}\Big)\Big) \\ &= \mathrm{P}\big(|Z| \ge z_{1-t/2}\big) \\ &= \mathrm{P}\big(Z \ge z_{1-t/2}\big) + \mathrm{P}\big(Z \le \underbrace{-z_{1-t/2}}_{=\,z_{t/2}}\big) \\ &= \frac{t}{2} + \frac{t}{2} \\ &= t \end{aligned}

  • Under \theta = \theta_0, for all t \in (0,1): \begin{aligned} & \mathrm{P}(\text{p-value} \le t \mid \theta = \theta_0) = t \\ \Longrightarrow\quad & \boxed{\ \text{p-value} \sim \text{Uniform}(0,1)\ } \end{aligned}

One-sided p-value, right tail

  • Right tail beyond (\hat\theta - \theta_0)/\omega, for H_0 : \theta \le \theta_0:

  • Right-tailed test: \boxed{\ \text{p-value} = 1 - \Phi\Big(\frac{\hat\theta - \theta_0}{\omega}\Big)\ }

  • The same, as a probability over Z^* \sim N(0,1) independent of the data: \begin{aligned} \text{p-value} &= \mathrm{P}^*\Big(Z^* > \frac{\hat\theta - \theta_0}{\omega} \;\Big|\; \hat\theta\Big) \\ &= 1 - \Phi\Big(\frac{\hat\theta - \theta_0}{\omega}\Big) \end{aligned}

  • Reject H_0 if p-value < \alpha.

  • Same decision as the one-sided test: \begin{aligned} & 1 - \Phi\Big(\frac{\hat\theta - \theta_0}{\omega}\Big) < \alpha \\ \Longleftrightarrow\quad & \frac{\hat\theta - \theta_0}{\omega} > z_{1-\alpha} \end{aligned}

One-sided p-value, left tail

  • Left tail below (\hat\theta - \theta_0)/\omega, for H_0 : \theta \ge \theta_0:

  • Left-tailed test: reject if (\hat\theta - \theta_0)/\omega < -z_{1-\alpha}.

  • p-value: \boxed{\ \text{p-value} = \Phi\Big(\frac{\hat\theta - \theta_0}{\omega}\Big)\ }

  • Reject H_0 if p-value < \alpha.

Size of the right-tailed p-value test

  • We will use: \Phi^{-1}(1 - \alpha) = z_{1-\alpha}.

  • Supremum over \Theta_0, attained at \theta = \theta_0: \begin{aligned} \sup_{\theta \le \theta_0} \mathrm{P}(\text{p-value} < \alpha \mid \theta) &= \sup_{\theta \le \theta_0} \mathrm{P}\Big(1 - \Phi\Big(\frac{\hat\theta - \theta_0}{\omega}\Big) < \alpha \;\Big|\; \theta\Big) \\ &= \sup_{\theta \le \theta_0} \mathrm{P}\Big(\Phi\Big(\frac{\hat\theta - \theta_0}{\omega}\Big) > 1 - \alpha \;\Big|\; \theta\Big) \\ &= \sup_{\theta \le \theta_0} \mathrm{P}\Big(\Phi\Big(Z + \underbrace{\frac{\theta - \theta_0}{\omega}}_{\le\,0}\Big) > 1 - \alpha\Big) \\ &= \mathrm{P}\big(\Phi(Z) > 1 - \alpha\big) \\ &= \mathrm{P}\big(Z > \Phi^{-1}(1 - \alpha)\big) \\ &= \mathrm{P}(Z > z_{1-\alpha}) \\ &= \alpha \end{aligned}

Distribution of \Phi(Z)

  • New random variable: W = \Phi(Z), \qquad W \in (\,0, 1\,)

  • CDF of W: \mathrm{P}(W \le x) = \begin{cases} 0 & x \le 0 \\ ? & 0 < x < 1 \\ 1 & x \ge 1 \end{cases}

  • We will use: a \le b \Longleftrightarrow g(a) \le g(b) for strictly increasing g.

  • For 0 < x < 1, apply \Phi^{-1} inside: \begin{aligned} \mathrm{P}(W \le x) &= \mathrm{P}\big(\Phi(Z) \le x\big) \\ &= \mathrm{P}\big(\Phi^{-1}(\Phi(Z)) \le \Phi^{-1}(x)\big) \\ &= \mathrm{P}\big(Z \le \underbrace{\Phi^{-1}(x)}_{=\,z_x}\big) \\ &= \Phi\big(\Phi^{-1}(x)\big) \\ &= x \end{aligned}

  • Hence: \mathrm{P}\big(\Phi(Z) \le x\big) = x \quad\Longrightarrow\quad \boxed{\ \Phi(Z) \sim \text{Uniform}(0,1)\ }

Distribution of 1 - \Phi(Z)

  • CDF of 1 - \Phi(Z), for 0 < u < 1: \begin{aligned} F_{1-W}(u) &= \mathrm{P}(1 - W \le u) \\ &= \mathrm{P}\big(1 - \Phi(Z) \le u\big) \\ &= \mathrm{P}\big(\Phi(Z) \ge 1 - u\big) \\ &= \mathrm{P}\big(Z \ge \Phi^{-1}(1 - u)\big) \\ &= 1 - \Phi\big(\Phi^{-1}(1 - u)\big) \\ &= 1 - (1 - u) \\ &= u \end{aligned}

  • Hence: \mathrm{P}\big(1 - \Phi(Z) \le u\big) = u \quad\Longrightarrow\quad \boxed{\ 1 - \Phi(Z) \sim \text{Uniform}(0,1)\ }

p-values under \theta = \theta_0

  • (\hat\theta - \theta_0)/\omega = Z: \begin{aligned} 1 - \Phi(Z) &\sim \text{Uniform}(0,1) \\ \Phi(Z) &\sim \text{Uniform}(0,1) \\ 2\big(1 - \Phi(|Z|)\big) &\sim \text{Uniform}(0,1) \end{aligned}

  • (\hat\theta - \theta_0)/\omega \sim N(0,1): critical values z_{1-\alpha}, z_{1-\alpha/2} from N(0,1).

  • p-value \sim \text{Uniform}(0,1): critical value \alpha from \text{Uniform}(0,1).

  • For every p-value: \begin{aligned} \mathrm{P}(\text{p-value} < \alpha \mid \theta = \theta_0) &= \mathrm{P}(\text{p-value} \le \alpha \mid \theta = \theta_0) \\ &= \text{CDF of the p-value at } \alpha, \ \text{ under } \theta = \theta_0 \\ &= \mathrm{P}(V \le \alpha), \quad V \sim \text{Uniform}(0,1) \\ &= \alpha \end{aligned}

  • In general: X \sim F, F continuous and strictly increasing \Longrightarrow F(X) \sim \text{Uniform}(0,1).

Summary

Hypotheses, errors, size

  • \hat\theta \sim N(\theta, \omega^2), \omega^2 known; \theta \in \Theta = \Theta_0 \cup \Theta_1, \Theta_0 \cap \Theta_1 = \emptyset; H_0 : \theta \in \Theta_0 vs H_1 : \theta \in \Theta_1.

  • Test statistic T \in S = S_A \cup S_R; reject H_0 if T \in S_R.

  • Type I error: reject a true H_0; Type II error: fail to reject a false H_0.

  • S_R \downarrow \ \Longrightarrow\ \mathrm{P}(\text{Type I error}) \downarrow \ \Longrightarrow\ \mathrm{P}(\text{Type II error}) \uparrow.

  • Valid test: \sup_{\theta \in \Theta_0} \mathrm{P}(T \in S_R \mid \theta) \le \alpha; size \alpha test: equality.

Two-sided test and power

  • Reject H_0 : \theta = \theta_0 if |(\hat\theta - \theta_0)/\omega| > z_{1-\alpha/2}; at \alpha = 0.05, z_{0.975} \approx 1.96.

  • |\hat\theta / \omega| > z_{1-\alpha/2}: \hat\theta significant at level \alpha.

  • Fail to reject \Longleftrightarrow \theta_0 \in CI_{1-\alpha}; CI_{1-\alpha} = \{\theta_0 : T(\theta_0) \le z_{1-\alpha/2}\}.

  • (\hat\theta - \theta_0)/\omega = Z + (\theta - \theta_0)/\omega \sim N\big((\theta - \theta_0)/\omega,\ 1\big).

  • \pi(\theta) = \mathrm{P}(|Z + (\theta - \theta_0)/\omega| > z_{1-\alpha/2}), \pi(\theta_0) = \alpha.

  • \sup_{\theta \neq \theta_0} \mathrm{P}(\text{Type II error} \mid \theta) = \sup_{\theta \neq \theta_0} \big(1 - \pi(\theta)\big) = 1 - \alpha.

  • Smaller \omega: more power.

  • Test that ignores the data: size \alpha, power \alpha.

One-sided tests and p-values

  • H_0 : \theta \le \theta_0: reject if (\hat\theta - \theta_0)/\omega > z_{1-\alpha}; z_{1-\alpha} < z_{1-\alpha/2}.

  • \sup_{\theta \le \theta_0} \mathrm{P}(\text{reject } H_0 \mid \theta) attained at \theta = \theta_0.

  • CI^{1}_{1-\alpha} = [\hat\theta - z_{1-\alpha}\,\omega,\ +\infty).

  • Against H_1 : \theta \neq \theta_0: one-sided test biased; no uniformly most powerful test.

  • Unequal tails, \alpha/3 right and 2\alpha/3 left: size \alpha, biased against H_1 : \theta \neq \theta_0.

  • p-values; reject if p-value < \alpha: \begin{aligned} \text{two-sided:}&\ \ 2\big(1 - \Phi(|(\hat\theta - \theta_0)/\omega|)\big) \\ \text{right-tailed:}&\ \ 1 - \Phi\big((\hat\theta - \theta_0)/\omega\big) \\ \text{left-tailed:}&\ \ \Phi\big((\hat\theta - \theta_0)/\omega\big) \end{aligned}

  • Under \theta = \theta_0: every p-value \sim \text{Uniform}(0,1); \mathrm{P}(\text{p-value} < \alpha \mid \theta = \theta_0) = \alpha.