Showing posts with label Neyman-Pearson theory. Show all posts
Showing posts with label Neyman-Pearson theory. Show all posts

Saturday, May 16, 2015

Neyman and Pearson: "On the Problem of the Most Efficient Tests of Statistical Hypotheses" (1933)

The Wikipedia page on the Neyman-Pearson lemma provides a statement and proof of the theorem which deviates quite a lot from the corresponding formulations in the article in which it originally appeared (pp. 300–301). I've been thinking a bit about how best to state and prove this theorem, and here's what I've come up with.

Dichotomies and Tests

Suppose two probability measures $P_0$ and $P_1$ on a set $\Omega$ are given, and that we are given a data set $X=x$ drawn from $\Omega$ according to one of these two distributions. Our goal is now to define a function $T: \Omega\rightarrow \{0,1\}$ which will return, for any data set, a guess at which of the two hypotheses the data came from.

Such a test will thus itself be a random variable $T=T(X)$. The quality of this test can be measured in terms of two error probabilities,
$$
P_0(T=1) \qquad \textrm{and} \qquad P_1(T=0).
$$Minimizing these two sources of error will generally be conflicting goals, and can be seen by considering the constant functions $T=0$ and $T=1$.

Likelihood Ratio Tests and the Neyman-Pearson Lemma

One family of tests we always have at our disposal are the likelihood ratio tests. A likelihood ratio test at threshold $r$ is the indicator variable that returns the value 1 when the evidence supports hypothesis 1 by a factor of more than $r$:
$$
R(x) \; = \; \mathbb{I}\left( \frac{P_1(x)}{P_0(x)} \geq r \right).
$$The content of the Neyman-Pearson lemma is that these likelihood ratio tests are the only optimal ones, in a certain sense.

Specifically, suppose any other test $T$ is given. The lemma then states that if
$$
P_0(T=1) \; \leq \; P_0(R=1) \qquad \Longrightarrow \qquad P_1(T=0) \; \geq \; P_1(R=0).
$$In other words, when the probability of an error of the first kind is held below a fixed level, we achieve the lowest rate of errors of the other kind by choosing a likelihood ration test. You cannot achieve a lower error rate under $P_0$ without increasing your error rate under $P_1$.

From $1\times 2$ to $2 \times 2$

The best way of proving this is to directly compare the regions of $\Omega$ on which the two tests disagree. There are two of these regions, $T=1 \wedge R=0$ and $T=0 \wedge R=1$. The difference in the measure of these two regions under the two hypotheses determine the difference in error rates for the two tests.

A paraphrase of the lemma could then be that if
$$
P_0(T=1 \wedge R=0)  \;  \leq  \; P_0(T=0 \wedge R=1),
$$then we also have
$$
P_1(T=0 \wedge R=1)  \;  \geq  \; P_1(T=1 \wedge R=0).
$$This alternative formulation can be translated back into the original form by adding back in the corner regions $T=1 \wedge R=1$ and $T=0 \wedge R=0$ on which the two tests give the same result.

Ratio Conversions

Using this formulation, a proof strategy suggests itself: Looking at cases according to the value of $R$. Specifically, if an event $A\subseteq \Omega$ satisfies the condition
$$
A \; \subseteq \; \{x:\;R(x)=0\} \; = \; \left\{ x:\; \frac{P_1(x)}{P_0(x)}  \; <  \;  r\right\},
$$then $\frac{1}{r}P_1 < P_0 $ on $A$. We can therefore obtain a lower bound on the $P_0$-measure of $A$ by performing the integration using the smaller measure $\frac{1}{r}P_1$ instead:
$$
\frac{1}{r} P_1(A) \; <  \; P_0(A).
$$Similarly, if
$$
A \; \subseteq \;  \{ R=1 \} \; = \; \left\{ \frac{P_1}{P_0}  \; \geq  \;  r\right\},
$$then $\frac{1}{r}P_1 \geq P_0$ on $A$, and
$$
P_0(A) \; \leq \; \frac{1}{r} P_1(A).
$$Since the two sets $R=0$ and $R=1$ form a partition of $\Omega$, we can split any set $A$ up according to its overlap with these cells and thus translate bounds on $P_0$ into bounds on $P_1$.

The Extended Sandwich

Now, by applying these considerations to the condition
$$
P_0(T=1 \wedge R=0)  \;  \leq  \; P_0(T=0 \wedge R=1),
$$we obtain the result that
$$
\frac{1}{r} P_1(T=1 \wedge R=0)  \;  \leq  \; \frac{1}{r} P_1(T=0 \wedge R=1).
$$By cancelling the $\frac{1}{r}$, we thus get the result.

I like this formulation of the lemma, because it makes it vivid what it is that's going on: When somebody splits up $\Omega$ into $\{T=0\}$ and $\{T=1\}$, we use our ratio test $R$ to further split these up into two subregions each. This allows us to use the translation between the two measures, and thus to compare the $P_0$ error rates with the $P_1$ error rates.

It also gives better intuitions about the potential differences between the tests $R$ and $T$: The more the set $T=1$ overlaps with the set $R=1$, the more the two tests are alike, and the more squeezed will the sandwich of inequalities be.

Monday, December 8, 2014

Edwards: Likelihood (1972)

Edwards (from his Cambridge site)
The geneticist A. F. W. Edwards is a (now retired) professor of biometry who was massively influenced by Ronald Fisher in his scientific writings. His books Likelihood argues that the likelihood concept is the only sound basis for scientific inference, but it reads at times almost like one long rant against Bayesian statistics (particularly ch. 4) and Neyman-Pearson theory (particularly ch. 9).

Don't Do Probs

As an alternative to these approaches to statistics, Edwards proposes that we limit ourselves to makes assertions only in terms of likelihood, "support" (log-likelihood, p.12), and likelihood ratios. In the brief epilogue of the book, he states that this
… allows us to do most of the things which we want to do, whilst restraining us from doing some things which, perhaps, we should not do. (p. 212)
In particular, this approach emphatically prohibits the comparison of hypotheses in probabilitistic terms. The kind of uncertainty we have about scientific theories is simply not, Edwards states, of a nature that can be quantified in terms of probabilities: "The beliefs are of a different kind," and they are "not commensurate" (p. 53)

The Difference Between Bad and Worse

He briefly mentions Ramsey and his Dutch book-style argument for the calculus of probability, and then goes on to speculate that, had not died so young,
… perhaps he would have argued that his demonstration that absolute degrees of belief in propositions must, for consistency's sake, obey the law of probability, did not compel anyone to apply such a theory to scientific hypotheses. Should they decline to do so (as I do), then they might consider a theory of relative degrees of belief, such as likelihood supplies. (p. 28)
In other words, it might be true that you cannot assign numbers to propositions in any other way than according to the calculus of probabilities, but you can always reject to have a quantitative opinion in the first place (or not make a bet).

Nulls Only

Consistently with Fisher's approach to statistics, Edwards finds it important to distinguish between null and not-null hypotheses: That is, in opposition to Neyman-Pearson theory, he refuses to explicitly formulate the alternative hypothesis against which a chance hypothesis is tested.

Here as elsewhere, this is a serious limitation with quite profound consequences:
It should be noted that the class of hypotheses we call 'statistical' is not necessarily closed with respect to the logical operations of alternation ('or') and negation ('not'). For a hypothesis resulting from either of these operations is likely to be composite, and composite hypotheses do not have well-defined statistical consequences, because the probabilities of occurrence of the component simple hypotheses are undefined. For example, if $p$ is the parameter of a binomial model, about which inferences are to be made from some particular binomial results, '$p=\frac{1}{2}$' is a statistical hypothesis because its consequences are well-defined in probability terms, but its negation, '$p\neq\frac{1}{2}$', is not a statistical hypothesis, its consequences being ill-defined. Similarly, '$p=\frac{1}{4}$ or $p=\frac{1}{2}$' is not a statistical hypothesis, except in the trivial case of each simple hypothesis having identical consequences. (p. 5)
This should also be contrasted with Jeffreys' approach, in which the alternative hypothesis has a free parameter and thus is allowed to 'learn', while the null has the parameter fixed at a certain value.

Scientists With Attitude

At several points, in the book, Edwards uses the concerns of the working scientist as an argument in favor of a likelihood-based reasoning calculus. He thus faults Bayesian statistics for "fail[ing] to answer questions of the type many scientists ask" (p. 54).

This question, I presume, is "What does the data tell my about my hypotheses?" This is distinct from "What should I do?" or "Which of these hypotheses is correct?" in that it only supplies the objective, quantitative measure of support, not the conclusion:
The scientist must be the judge of his own hypotheses, not the statistician. The perpetual sniping which statisticians suffer at the hands of practising scientists is largely due to their collective arrogance in presuming to direct the scientists in his consideration of hypotheses; the best contribution they can make is to provide some measure of 'support', and the failure of all but a few to admit the weaknesses of the conventional approaches has not improved the scientists' opinion. (p. 34)
In brief form, this leads to the following tirade against Bayesian statistics:
Inverse probability, in its various forms, is considered and rejected on the grounds of logic (concerning the representation of ignorance), utility (it does not allow answers in the form desired), oversimplicity (in problems involving the treatment of frequency probabilities) and inconsistency (in the allocation of prior probability distributions). (p. 67–68)