Showing posts with label frequentism. Show all posts
Showing posts with label frequentism. Show all posts

Tuesday, April 14, 2015

Chrystal: "On Some Fundamental Principles in the Theory of Probability" (1891)

I've been trying to get my hands on the following paper:
George Chrystal (University of Edinburgh): "On Some Fundamental Principles in the Theory of Probability." Transactions of the Actuarial Society of Edinburgh, Volume 2, January 1891, pages 420–439.
So far, no luck. The Cambridge Journals database has a copy, but it's behind a paywall, and my university doesn't have a subscription.

However, a number of other sources quote extensively from the paper, so I've been able to piece together an understanding of what it looks like.

Posterior Frequencies

It seems that Chrystal's main beef is with the use of Bayes' rule to update the probability of a certain set of hypotheses whose long-term frequencies are already given in advance. His reasoning seems to be that this conflates our subjective degree of belief (which may indeed change) with the objective frequency (which, by assumption, cannot).

This philosophical distinction is nicely presented in the following quote. It comes from an 1894 review (also available in book form on the Internet Archive) by somebody called G. F. Hardy, not to be confused with G. H. Hardy.
"There is," says Professor Chrystal, "in Laplace's view, a confusion between two senses of the word 'Probability', which although distinct are often more or less associated in point of fact. In common speech we say that a single event is more or less 'probable', and by this word we indicate our own mental attitude towards the event, an attitude that may be well or ill justified by facts. When an actuary says that the probability that a man of 20 will live to be 60 is $\frac{59}{97}$, he is not, strictly speaking, referring to any one event at all, but merely making an assertion to the effect that out of any considerable number of men of 20 years of age about $\frac{59}{97}$ will reach the age of 60. No one knows better than an actuary that this statement is a fact, established (under certain circumstances, and with certain limitations), by experience, and that it has nothing whatever to do with the mental attitude of anyone. Everyone will admit that we could never arrive at this result by analyzing the event of a man of 20 reaching or not reaching the age of 60 into cases regarding each of which we should be equally undecided,—mentally suspended, as it were, like Buridan's ass between the equal bundles of hay." (Hardy, p. 316)
It's not clear whether the emphases are in the original, but I'm guessing not.

Hardy does not give a page reference, but proper reference seems to be page 423. I got that figure from the manuscript of a 1893 presentation by a certain John Govan, F.F.A. (whose name, location, and date fit the proselytizing businessman John George Govan).

Govan's rendition of the quote occurs on his page 212. He doesn't use the emphases.

The Burial of Bayes

Several sources also report that Chrystal summarizes his discussion with the following tirade:
… both from the point of view of practical common-sense, and from the point of view of logic, the so-called laws of Inverse Probability are a useless appendage to the first principles of probability, if indeed they be not a flat contradiction of those very principles.
This is cited by a number of authors, including E. T. Whittaker, F.R.S. (in a footnote to a 1920 presentation, p. 165), Andrew I. Dale (1999, p. 485) and Sharon McGrayne (2011, p. 37).

Dale reports that this quote is found on page 438. Whittaker apparently reports it as page 421 (but that would be at odds with Dale's description of the quote as a occurring near the conclusion of the essay). McGrayne doesn't give a page number.

According to Whitaker, the quote continues as follows:
The laws of Inverse Probability being dead, should be decently buried out of sight, and not embalmed in text-books and examination papers.
McGrayne further reports the following conclusion:
The indiscretions of great men should be quietly allowed to be forgotten.
Checking the relevant footnote of McGrayne's book (note 9 of ch. 3, p. 260), this turns out to be a recycled quote from Anders Hald's A History of Mathematical Statistics, page 275. That book doesn't have a Google Books preview, and it's not in my library.

Sort-of-Long-Term Frequencies

Hardy's review continues to quote Chrystal's discussion of how probability is to be defined:
"The notion of probability is always attached to a class or series of events, which usually have more or less of other attributes in common, but are always distinguished by this mark, that certain phases of them, although not predicable with the smallest certainty in any individual case, are predicable with more or less uniformity in a certain proportion of cases in the long run. The fundamental features of this series are statistical uniformity combined with irregularity of every conceivable kind in the individual instance. The number of the events in the series must be large. Its extension both as to space and time is arbitrary, and in certain ideal cases infinite. It is in this last respect alone that probability has anything to do with our mental attitude; we may choose our standpoint, and this determines the probability to which our knowledge may make a better or worse approximation. As the series is varied the probability alters. . . .  We are thus led to the following abstract definition of the probability or chance or an event. If, on taking any very large number, N out of a series of cases in which an event A is in question, A happens on pN occasions, the probability of the event A is said to be p." (Hardy, p. 317)
Again, no page number is given.

Posterior Priors

The examples in Chrystal's paper seem all to be of the same kind: He describes a set-up in which certain a priori frequencies are given, and he then tells us that no amount of evidence should be able to change those frequencies; the only mental operation we can perform is to exclude logically impossible cases, not to compute posterior probabilities.

Govan thus quotes him as discussing a situation in which you draw two white balls from a bag of black and white balls. Then:
"Any one," says Professor Chrystal, "who knows the definition of mathematical probability, and who considers this question apart from the Inverse Rule, will not hesitate for a moment to say that the chance is $\frac{1}{2}$; that is to say, that the third ball is just as likely to be white as black. For there are four possible constitutions of the bag . . . each of which we are told occurs equally often in the long run, and among those cases there are two . . . in which there are two white balls, and among these the case in which there are three white occurs in the long run, just as often as the case in which there are only two." (Govan, p. 208)
According to the text of Govan's discussion, this quote must be on or around page 435 of Chrystal's text.

Another very similar example is attributed to Chrystal's page 437:
"A bag contains five balls which are known to be either all black or all white—and both these are equally probable. A white ball is dropped into the bag, and then a ball is drawn out at random and found to be white. What is now the chance that the original balls were all white?" Professor Chrystal asserts that the chance is precisely what it was before, viz. $\frac{1}{2}$. (Govan, p. 208–209)
"The ball drawn out," says Professor Chrystal, "may have been the one we put in, it may not; and this is all that any one can say." (Govan, p. 209)
Note that this is quite upside-down compared to how we usually think about frequentism: Here, Chrystal tells us to ignore the likelihoods and put all our confidence in the priors. We are used to thinking about frequentists as doing the exact opposite.

The Essential Tension

Govan objects to Chrystal's principle of rejecting the evidence:
… let us say we have two bags before us, one containing six white balls, the other five black balls and one white. There is nothing to indicate which is which. We draw from one of the bags chosen at random a ball which proves to be white. It is difficult to believe that any man in the possession of his faculties, say if his life depended on his guessing aright from which bag the ball had come, would hesitate to guess the former. Even according to Professor Chrystal he would be right 6 times out of 7 in the long run. Yet, again according to Professor Chrystal he would would be just as likely to be wrong as to be right. (Govan, p. 209)
Although they are talking past each other, this is certainly the core of the issue: The distinction between optimal, adaptive gambling behavior and fixed, objective frequencies.

Monday, December 8, 2014

Edwards: Likelihood (1972)

Edwards (from his Cambridge site)
The geneticist A. F. W. Edwards is a (now retired) professor of biometry who was massively influenced by Ronald Fisher in his scientific writings. His books Likelihood argues that the likelihood concept is the only sound basis for scientific inference, but it reads at times almost like one long rant against Bayesian statistics (particularly ch. 4) and Neyman-Pearson theory (particularly ch. 9).

Don't Do Probs

As an alternative to these approaches to statistics, Edwards proposes that we limit ourselves to makes assertions only in terms of likelihood, "support" (log-likelihood, p.12), and likelihood ratios. In the brief epilogue of the book, he states that this
… allows us to do most of the things which we want to do, whilst restraining us from doing some things which, perhaps, we should not do. (p. 212)
In particular, this approach emphatically prohibits the comparison of hypotheses in probabilitistic terms. The kind of uncertainty we have about scientific theories is simply not, Edwards states, of a nature that can be quantified in terms of probabilities: "The beliefs are of a different kind," and they are "not commensurate" (p. 53)

The Difference Between Bad and Worse

He briefly mentions Ramsey and his Dutch book-style argument for the calculus of probability, and then goes on to speculate that, had not died so young,
… perhaps he would have argued that his demonstration that absolute degrees of belief in propositions must, for consistency's sake, obey the law of probability, did not compel anyone to apply such a theory to scientific hypotheses. Should they decline to do so (as I do), then they might consider a theory of relative degrees of belief, such as likelihood supplies. (p. 28)
In other words, it might be true that you cannot assign numbers to propositions in any other way than according to the calculus of probabilities, but you can always reject to have a quantitative opinion in the first place (or not make a bet).

Nulls Only

Consistently with Fisher's approach to statistics, Edwards finds it important to distinguish between null and not-null hypotheses: That is, in opposition to Neyman-Pearson theory, he refuses to explicitly formulate the alternative hypothesis against which a chance hypothesis is tested.

Here as elsewhere, this is a serious limitation with quite profound consequences:
It should be noted that the class of hypotheses we call 'statistical' is not necessarily closed with respect to the logical operations of alternation ('or') and negation ('not'). For a hypothesis resulting from either of these operations is likely to be composite, and composite hypotheses do not have well-defined statistical consequences, because the probabilities of occurrence of the component simple hypotheses are undefined. For example, if $p$ is the parameter of a binomial model, about which inferences are to be made from some particular binomial results, '$p=\frac{1}{2}$' is a statistical hypothesis because its consequences are well-defined in probability terms, but its negation, '$p\neq\frac{1}{2}$', is not a statistical hypothesis, its consequences being ill-defined. Similarly, '$p=\frac{1}{4}$ or $p=\frac{1}{2}$' is not a statistical hypothesis, except in the trivial case of each simple hypothesis having identical consequences. (p. 5)
This should also be contrasted with Jeffreys' approach, in which the alternative hypothesis has a free parameter and thus is allowed to 'learn', while the null has the parameter fixed at a certain value.

Scientists With Attitude

At several points, in the book, Edwards uses the concerns of the working scientist as an argument in favor of a likelihood-based reasoning calculus. He thus faults Bayesian statistics for "fail[ing] to answer questions of the type many scientists ask" (p. 54).

This question, I presume, is "What does the data tell my about my hypotheses?" This is distinct from "What should I do?" or "Which of these hypotheses is correct?" in that it only supplies the objective, quantitative measure of support, not the conclusion:
The scientist must be the judge of his own hypotheses, not the statistician. The perpetual sniping which statisticians suffer at the hands of practising scientists is largely due to their collective arrogance in presuming to direct the scientists in his consideration of hypotheses; the best contribution they can make is to provide some measure of 'support', and the failure of all but a few to admit the weaknesses of the conventional approaches has not improved the scientists' opinion. (p. 34)
In brief form, this leads to the following tirade against Bayesian statistics:
Inverse probability, in its various forms, is considered and rejected on the grounds of logic (concerning the representation of ignorance), utility (it does not allow answers in the form desired), oversimplicity (in problems involving the treatment of frequency probabilities) and inconsistency (in the allocation of prior probability distributions). (p. 67–68)

Fisher: The Design of Experiments (4th ed., 1947), Chapter I

Fisher; from Judea Pearl's website.
Now, here's a revealing turn of phrase:
In the foregoing paragraphs the subject-matter of this book has been regarded from the point of view of an experimenter, who wishes to carry out his work competently, and having done so wishes to safeguard his results, so far as they are validly established, from ignorant criticism by different sorts of superior persons. (p. 3)
You could hardly spell out more explicitly the philosophy that lies behind Fisher's concept of statistics: It's a strategic ritual, not designed to ensure a result, but to protect against criticism.

Perfectly Rigorous and Unequivocal

Such protection only goes as far as the mathematical consensus on the validity of the logic. But Fisher goes on to state that "rigorous deductive argument" is possible even in the context of random events, citing gambling as a proof of concept:
The mere fact that inductive inferences are uncertain cannot, therefore, by accepted as precluding perfectly rigorous and unequivocal inference. (p. 4)
This seems to confuse the issues of probability and statistics, unless his argument here really only amounts to saying that distributions are non-stochastic entities.

Useless for Scientific Purposes

This leads him to a discussion of "inverse probability," which he gives three reasons for rejecting: First,
… advocates of inverse probability seem forced to regard mathematical probability, not as an objective quantity measured by observed frequencies, but as measuring merely psychological tendencies, theorems respecting which are useless for scientific purposes. (p. 6–7)
Second, Bayes' axiom (about the flat prior for a coin flip) is not self-evident, that is, the choice of prior is not unequivocal (p. 7).

Ever Since the Dawn of Man…

And third,
… inverse probability has been only very rarely used in the justification of conclusions from experimental facts, although the theory has been widely taught, and is widespread in the literature of probability. Whatever the reasons are which could give experimenters confidence that they can draw valid conclusions from their results, they seem to act just as powerfully whether the experimenter has heard of the theory of inverse probability or not. (p. 7)
That's a funny sociological proof, given that he has just rejected Bayesian statistics for its psychologism. But he himself sometimes seems to think that his statistics is a kind of theory of learning, whatever that means:
Men have always been capable of some mental processes of the kind we call "learning by experience." … Experimental observations are only experience carefully planned in advance, and designed to form a secure basis of new knowledge; (p. 8)

Saturday, December 6, 2014

Chernoff: "A career in statistics" (2014)

Chernoff makes some interesting remarks about the philosophy of statistics in his recent autobiographical essay.

First, an anecdote about the three classical decision criteria considered in decision theory:
I had always been interested in the philosophical issues in statistics, and Jimmie Savage claimed to have resolved one. Wald had proposed the minimax criterion for deciding how to select one among the many “admissible” strategies. Some students at Columbia had wondered why Wald was so tentative in proposing this criterion. The criterion made a good deal of sense in dealing with two-person zero-sum games, but the rationalization seemed weak for games against nature. In fact, a naive use of this criterion would suggest suicide if there was a possibility of a horrible death otherwise. Savage pointed out that in all the examples Wald used, his loss was not an absolute loss, but a regret for not doing the best possible under the actual state of nature. He proposed that minimax regret would resolve the problem. At first I bought his claim, but later discovered a simple example where minimax regret had a similar problem to that of minimax expected loss. For another example the criterion led to selecting the strategy A, but if B was forbidden, it led to C and not A. This was one of the characteristics forbidden in Arrow’s thesis.
Savage tried to defend his method, but soon gave in with the remark that perhaps we should examine the work of de Finetti on the Bayesian approach to inference. He later became a sort of high priest in the ensuing controversy between the Bayesians and the misnamed frequentists. (pp. 32–33)
He immediately moves on to one of his own more dismal conclusions about the issue:
I posed a list of properties that an objective scientist should require of a criterion for decision theory problems. There was no criterion satisfying that list in a problem with a finite number of states of nature, unless we canceled one of the requirements. In that case the only criterion was one of all states being equally likely. To me that meant that there could be no objective way of doing science. I held back publishing those results for a few years hoping that time would resolve the issue (Chernoff, 1954). (p. 33)
I haven't read the paper he is referring to here, but it seems like the text has jumbled up the conclusions: I think what he meant to say is that there is no single good decision function when we have infinitely many states, since the criteria essentially require us to use a uniform distribution. But I would have to check the details.

Finally, he moves on to a more on-record explication of his position:
In the controversy, I remained a frequentist. My main objection to Bayesian philosophy and practice was based on the choice of the prior probability. In principle, it should come from the initial belief. Does that come from birth? If we use instead a non-informative prior, the choice of one may carry hidden assumptions in complicated problems. Besides, the necessary calculation was very forbidding at that time. The fact that randomized strategies are not needed for Bayes procedures is disconcerting, considering the important role of random sampling. On the other hand, frequentist criteria lead to the contradiction of the reasonable criteria of rationality demanded by the derivation of Bayesian theory, and thus statisticians have to be very careful about the use of frequentist methods. 
In recent years, my reasoning has been that one does not understand a problem unless it can be stated in terms of a Bayesian decision problem. If one does not understand the problem, the attempts to solve it are like shooting in the dark. If one understands the problem, it is not necessary to attack it using Bayesian analysis. My thoughts on inference have not grown much since then in spite of my initial attraction to statistics that came from the philosophical impact of Neyman–Pearson and decision theory. (p. 33)

Tuesday, October 28, 2014

Blackwell and Girshick: Theory of Games and Statistical Decisions (1954), Ch. 4

There's an interesting representation theorem in Blackwell and Girshick's textbook in statistics: It provides a set of sufficient conditions for a preference ordering over lotteries to be expressible as a prior probability distribution (Th. 4.3.1, p. 118).

I assume the theorem comes from either Wald, Savage, or de Finetti, but no reference is given.

Well-Behaved Preferences

A lottery can here be defined as a function from the sample space to the real numbers. The conditions in the theorem are then the following:
  • The ordering of two lotteries $f$ and $g$ cannot depend on the availability of other lotteries.
  • If a lottery $f$ provides a higher payoff than a lottery $g$ at all points in the sample space $\Omega$, then $f$ must be preferred to $g$.
  • If $f$ is preferred to $g$, then $f+h$ must be preferred to $g+h$.
An inspection of the proof also shows that they should have included a continuity condition:
  • If $f_1, f_2, \ldots$ is a series of lotteries converging to a limit $f$, and if $f_i$ is preferred to $g$ for all $i$, then $f$ must also be preferred to $g$.
When these conditions are met, the preference ordering over the lotteries can be expressed as a distribution over the sample space, unless it categorizes all lotteries as equally good.

The Large and the Good

The proof of the theorem uses the fact that two convex, disjoint, open sets can be separated by a hyperplane. Here's a sketch:

If we let $e$ be the lottery that pays zero in all situations $\omega \in \Omega,$ then we can define the following sets:
\begin{eqnarray}
F[>] &=& \{f\ |\ \forall \omega \in \Omega: f(\omega) > 0\},
\\
F[\gtrsim] &=& \{f\ |\ f\gtrsim e\},
\end{eqnarray}
and we can further define, in the usual way, $F[>] + F[\gtrsim]$ to be the sums of lotteries from those two sets.

Now, $F[>]$ is an open set, and hence $F[>] + F[\gtrsim]$ is, too. Further, $F[>]$ and $F[\gtrsim]$ are both closed under addition (due to the third assumption), and hence their sum is, too. They are also both closed under multiplication with a scalar, and again, so is their sum — but this latter argument requires a bit more spelling out.

Rational and Real Convexity

Suppose a lottery $f$ is preferred to the zero lottery, that is, $f \in F[\gtrsim]$. The third assumption of the theorem then tells us that
$$
e \;\lesssim\; f \;\lesssim\; f + f \;\lesssim\; f + f + f \;\lesssim\; \ldots \;\lesssim\; nf.
$$
By further adding multiple copies of the zero lottery to both sides of this preference ineqality, we can see that
$$
me \;\lesssim\; (m-1)e + nf \;=\; nf,
$$
where we have selectively used the fact that $e$ is the zero lottery. Putting these facts together, we then have the preference inequality
$$
e \;\lesssim\; \left(\frac{n}{m}\right)f.
$$
By using a positive sequence of rational approximations $(n/m) \rightarrow \lambda$, we can use this fact along with the continuity assumption to conclude that $F[\gtrsim]$ is closed under multiplication with any positive, real scalar $\lambda$.

I don't think there's a way around this last technicality. It is, incidentally, the same proof technique used to prove that the logarithmic functions are the only continuous functions that turn products into sums.

Cutting the Cake

At any rate, $F[>] + F[\gtrsim]$ is an open and convex set separated from the singleton set $\{e\}$. We can therefore conclude that there is a lottery (or vector) $p$ which defines the hyperplane $\{f\ |\ f\cdot p=0\}$ separating $F[>] + F[\gtrsim]$ from $\{e\}$. The set $F[>] + F[\gtrsim]$ is thus a subset of the half-space $\{f\ |\ f\cdot p \geq 0\}$.

This vector $p$ must have nonnegative coordinates, since the set $F[>]$ is unbounded in all positive directions. If $p$ had a negative coordinate, $p(\omega) \leq 0$, we could choose a lottery for which the corresponding coordinate, $f(\omega)$, was so large that $f\cdot p < 0$. This would violate the definition of $p$, and $p$ hence has to be a nonnegative vector which can be interpreted as a probability distribution.

It would also have the property that $f\cdot p \geq g\cdot p$ if and only if $f \gtrsim g$. This follows from the fact that $f\cdot p \geq g\cdot p$ if and only if $(f - g) \in F[>] + F[\gtrsim]$, which holds if and only if $f$ can be expressed as the sum of $g$ and some lottery preferable to the zero lottery.

Saturday, July 26, 2014

Wald: Sequential Analysis (1947)

I've always thought that Shannon's insights in the 1940s papers seemed like pure magic, but this book suggests that at parts of his way of thinking were already in the air at the time: From his own frequentist perspective, Wald comes oddly close to defining a version of information theory.

The central question of the book is:
How can we devise decision procedures that map observations into {Accept, Reject, Continue} in such a way that (1) the probability of wrongly choosing Accept or Reject is low, and (2) the expected number of Continue decisions is low?
Throughout the book, the answer that he proposes is to use likelihood ratio tests. This puts him strangely close to the Bayesian tradition, including for instance Chapter 5 of Jeffreys' Theory of Probability.

A sequential binomial test ending in rejection (p. 94)

In particular, the coin flipping example that Jeffreys considers in his Chapter 5.1 is very close in spirit to the sequential binomial test that Wald considers in Chapter 5 of Sequential Analysis. However, some differences are:
  • Wald compares two given parameter values p0 and p1, while Jeffreys compares a model with a free parameter to one with a fixed value for that parameter.
  • Jeffreys assigns prior probabilities to everything; but Wald only uses the likelihoods given the two parameters. From his perspective, the statistical test will thus have different characteristics depending on what the underlying situation is, and he performs no averaging over these values.
This last point also means that Wald is barred from actually inventing the notion of mutual information, although he comes very close. Since he cannot take a single average over the log-likelihood ratio, he cannot compute any single statistic, but always has to bet on two horses simultanously.

Friday, May 30, 2014

Fisher: "On the Mathematical Foundations of Theoretical Statistics" (1921)

I haven't had time to study this paper in detail yet, but based on a quick skim, it seems that Fisher
I'll read the whole thing later. But for now, a few quotes.

First, a centerpiece in Bayes' original paper was the postulate that the uncertainty about the bias of a coin should be represented by means of a uniform distribution. Fisher comments:
The postulate would, if true, be of great importance in bringing an immense variety of questions within the domain of probability. It is, however, evidently extremely arbitrary. Apart from evolving a vitally important piece of knowledge, that of the exact form of the distribution of values of p, out of an assumption of complete ignorance, it is not even a unique solution. (p. 325)
Second Bayesian topic is ratio tests: That is, assigning probabilities to two exclusive and exhaustive hypotheses X and Y based on the ratio between how well they explain the data set A, that is,
Fisher in 1931; image from the National Portrait Gallery.
Pr(A | X) / Pr(A | Y).
Fisher comments:
This amounts to assuming that before A was observed, it was known that our universe had been selected at random for [= from] an infinite population in which X was true in one half, and Y true in the other half. Clearly such an assumption is entirely arbitrary, nor has any method been put forward by which such assumptions can be made even with consistent uniqueness. (p. 326)
The introduction of the likelihood concept:
There would be no need to emphasise the baseless character of the assumptions made under the titles of inverse probability and BAYES' Theorem in view of the decisive criticism to which they have been exposed at the hands of BOOLE, VENN, and CHRYSTAL, were it not for the fact that the older writers, such as LAPLACE and POISSON, who accepted these assumptions, also laid the foundations of the modern theory of statistics, and have introduced into their discussions of this subject ideas of a similar character. I must indeed plead guilty in my original statement of the Method of the Maximum Likelihood (9) to having based my argument upon the principle of inverse probability; in the same paper, it is true, I emphasised the fact that such inverse probabilities were relative only. That is to say, that while we might speak of one value of p as having an inverse probability three times that of another value of p, we might on no account introduce the differential element dp, so as to be able to say that it was three times as probable that p should lie in one rather than the other of two equal elements. Upon consideration, therefore, I perceive that the word probability is wrongly used in such a connection: probability is a ratio of frequencies, and about the frequencies of such values we can know nothing whatever. We must return to the actual fact that one value of p, of the frequency of which we know nothing, would yield the observed result three times as frequently as would another value of p. If we need a word to characterise this relative property of different values of p, I suggest that we may speak without confusion of the likelihood of one value of p being thrice the likelihood of another, bearing always in mind that likelihood is not here used loosely as a synonym of probability, but simply to express the relative frequencies with which such values of the hypothetical quantity p would in fact yield the observed sample. (p. 326)
In the conclusion, he says that likelihood and probability are "two radically distinct concepts, both of importance in influencing our judgment," but "confused under the single name of probability" (p. 367) Note that these concepts are "influencing our judgment" — that is, they are not just computational methods for making a decision, but rather a kind of model of a rational mind.

Friday, May 9, 2014

Zabell: "R. A. Fisher and the Fiducial Argument" (1992)

Chapter 3.3 of Fisher's 1956 book is dedicated to his so-called "Fiducial Argument."

I was extremely confused by his presentation and looked around for some secondary literature. This brought me to this wonderful paper by Sandy Zabell, which explains how Fisher's ideas about fiducial inference were indeed quite confused and changed a lot over time. It also explains the core of his argument better than he did himself (to my mind, at least).

As I now understand Fisher's argument, this is the idea: When you have a statistical model with a flat prior, the posterior probability of a specific parameter setting is proportional to the likelihood of the data under that parameter setting,

Pr(p | x) ∝ Pr(x | p)

However, for many unbounded parameter spaces, the likelihood Pr(X = x | p) does not have a finite integral when considered as a function of p. In such cases, a straightforward use of Bayesian inference with a flat prior is not an option.

But, Fisher says, consider instead the cumulative likelihood Pr(X < x | p). This is a function of x which always lies between 0 and 1, and it is 0 at negative infinity and 1 at positive infinity.

Pr(X < x | p) for uniform distributions with right end-points p = 3, p = 5, and p = 7.

The trick now is to view this cumulative likelihood as a function of p instead of a function of x. In many but not all cases, this function will be 1 when the parameter is at negative infinity and 0 when it is at positive infinity.

Cumulative likelihood at x = 1.5, x = 2.5, and p = 3.5 as a function of the parameter.

For instance, if p is the mean of a normal distribution, the cumulative likelihood Pr(X < x | p) decreases in this way. The reason is that the upper bound x stays where it is, while the expected value of the variable increases.

Uniform likelihoods Pr(X < x | p) with varying right end-points p.

In such cases, we can thus interpret the cumulative likelihood as the complement of a CDF for the parameter,
G(p) = 1 – Pr(X < x | p).
If this function G is differentiable, we can further interpret G' as a PDF for the parameter p given observation x.

As an example, suppose (as on the pictures) that a number X is drawn from a uniform distribution on the interval [0, p]. The cumulative likelihood is then
Pr(X < x | p) = x/p    (0 < x < p),
and 0 and 1 below and above the interval, respectively.

Fiducial PDFs given the observations x = 1.5, x = 2.5, and x = 3.5.

Considering the complement of this function, 1 – x/p, as a CDF for the parameter p, we can differentiate it in order to get the density
G'(p) = x/p2    (x < p),
when p > x and 0 otherwise. We have thus obtained a posterior probability distribution for the parameter without assuming anything about the prior.

It should be noted that
  • This method does not always work; consider for instance the case in which the cumulative likelihood oscillates between a unimodal normal and a bimodal normal distribution as the parameters runs along the real number line.
  • The method can also give inconsistent results; for instance, the fiducial distribution of X2 is, as far as I understand, not necessarily the distribution you would get by finding the fiducial distribution of X and then deriving a distribution for X2.
  • The method has no single, natural extension to the multi-parameter case, and there are some serious obstacles to constructing such an extension.
It is also interesting that in the example above, the fiducial distribution corresponds to the posterior you get if you assume the improper prior 1/p. It can thus not be rationalized as a posterior inference using only ordinary probability distributions, but it can if we allow ourselves crazy, unnormalizable ones.

Tuesday, May 6, 2014

de Finetti: Probability, Induction, and Statistics (1972)

In Chapters 8 and 9 of this anthology, Bruno de Finetti reiterates his reasons for espousing Bayesian probability theory as the unique optimal calculus of reasoning. This brings him into a discussion of several controversies surrounding the two paradigms of statistics.

Bruno de Finetti and a computer; image from www.moebiusonline.eu.

No Unknown Unknowns

According to de Finetti, the ordinary meaning of the word "probability" is "a degree of belief" (p. 148), and he rejects any attempt to define it in terms of frequency:
… we reject the idea that the ostensible notion of identical events or trials gives a suitable basis for an empirical formulation of a frequentist theory of probability or for some objectivistic form of the "law of large numbers". (p. 154)
Consequently:
The probability of an event conditional on, or in the light of, a specified result is a different probability, not a better evaluation of the original probability. (p. 149)
There is thus no such things as an "unknown probability." You always know your own uncertainty:
Any assertion concerning probabilities of events is merely the expression of somebody's opinion and not itself an event. There is no meaning, therefore, in asking whether such an assertion is true or false or more or less probable. (p. 189)
Thus, "speaking of unknown probabilities must be forbidden as meaningless" (p. 190) and in fact rejected as a "superstition" (p. 154–55).

But of course we do have problems assigning numbers of things, so de Finetti has some explaining to do. He thus invokes the analogy of choosing a price for a commodity:
A personal probability is, in effect, a quantitative decision closely akin to deciding on a price. In seeking to fix such a number with precision the person will sooner or later encounter difficulties that evoke the expressions "vagueness", "insecurity", or "vacillation". Analysis of this omnipresent phenomenon has given rise to misunderstandings. Thus, attempts to say that the exact probabilities are "meaningless" or "non-existent" pose more severe problems than they are intended to resolve, similarly for replacements of individual probabilities by intervals or by second-order probabilities. […] Sight should not be lost of the the fact that a person may find himself in an economic situation that entails acting in accordance with a sharply defined probability, whether the person chooses his act with security or not. (p. 145)
In spite of this seeming pluralism about personal opinion, he still maintains that the mathematical concept of probability is an idealization:
The (subjectivistic) theory of probability is a normative theory (p. 151).
But of course, the latter refers only to the mechanics of the calculus, not the choice of priors.

Rants Against Frequentism

De Finetti hates frequentist statistics. In his brief historical sketch, he says that the frequentist theory is a set of "substitutes" for Bayesian reasoning which were supposed to fill the "void" left after the analysis by Bayes was rejected (p. 161).

He adds:
The method pursued in the construction of such substitutes consists in general of adopting or imitating some case where the correct method reduces to a simple form based on summarizing parameters, however substituting for the true formulation and justification some incomplete and fragmentary justification or even no justification at all, as comes to seem legitimate when each notion is interpreted as something autonomous and arbitrary. For each isolated problem it appeared thus legitimate to devise as many ad hoc expedients as desired, and in fact it often happens that several are devised, proposed, and applied, to a single problem. (p. 161)
Shortly after, another rant follows:
In this manner, any notion of a systematic and meaningful interpretation of the problem of statistical inference is abandoned for the position of devising, case by case, "tests" of hypotheses or methods of "estimating" parameters. This means formulating, as an autonomous and largely arbitrary question, the problem of extracting from experience something that is apparently to be employed as though it were a conclusion or conviction, while asserting that it is neither one nor the other. (p. 162)
He is specifically angry about the "grossly inconsistent" notion of tests and hypothesis rejections, which he finds to be perverse distortions of the proper use of Bayes' rule (p. 163):
The severest of these mutilations is that of the oversimplified criteria according to which a probability P(E | H) us taken as a basis for rejecting the isolated hypothesis H if this probability, for the observation E, is small. (p. 163)
Such hypothesis rejection are, namely, ambiguous about what event E the observed data actually testifies to, as in the problem of choosing between one-sided and two-sided tests:
If, for example, as is often the case, E consists in having observed the exact value x of a random number X, such a deviation, the probability of that exact value is ordinarily zero. In order to eliminate the evident meaninglessness of this criterion that rejects the hypothesis no matter what value x may have, some other is substituted for it, such as observation of a value equal to or greater than x in absolute value, or equal or greater in absolute value and of the same sign. But all these variables are arbitrary, at least in the framework of so crudely mutilated a formulation. (p. 163)
On the following page, he also gives the example of having to decide whether a point on a target was hit by a particular marksman. He gives various examples of sets that such a point can belong to: The singleton set containing only the point itself, a circle having the point as a center, a slice of the target containing the point, a circle having the center of the target as its center, etc.

Various ways of construing the acceptance region for a test.

He continues to say that "One might say that all the deficiencies of objectivistic statistics stem from insistence on using only what appears to be soundly based" (p. 165). This, he says, is like setting a price according to the things that are easiest to measure rather than the things that are most relevant.

The issue of building a statistical enterprise on likelihoods alone is, he contends, like a systematic attempt to find P(E | H) when you are looking for P(H | E). In an example he attributes to Halphen:
We need a cement that will not be harmed by water. The merchant advises us to buy a certain kind that, he assures us, will not harm water. He does not try to cheat us by saying that the two things are equivalent but he want to convince us not to insist on asking for what we need (p. 173).
This is apparently a commentary on related example used by Neyman.

De Finetti on Wald

In a series of papers from the 1940s and 50s, Abraham Wald developed a theory of "admissible decision functions" for decision problems with uncertainty (see, e.g., here). His idea was to consider a decision admissible if it minimized the maximal damage that could obtain in the given situation. This correspond to the solution of a two-person zero-sum game against a malevolent nature.

In his discussion of Wald's theory, de Finetti helpfully "completes" the specification of a decision problem by putting a prior probability on the various hypotheses. Having provided these marginal probabilities, he comments:
Of course, these marginal elements do not appear in Wald's formulation; their absence there is just what prevents the problem of decision from having the solution that is obvious when the table is thus completed. Namely, choose the decision corresponding to the minimal(mean) loss, or equivalently to the maximal(mean) gain or the maximal (mean) utility. Here we have always put "mean" between parentheses but from now on shall suppress the word altogether; for value and utility in an uncertain situation is, by definition, the mathematical expectation of the values of utilities. (p. 179)
In his own work, Wald concluded that the admissible strategies are the mixed strategies whose support consists of pure strategies that are optimal for some parameter setting. But these are also the ones that can be rationalized by some prior probability distribution, so de Finetti happily concludes that
… the admissible decisions are the Bayesian ones; that is, those that minimize the loss with respect to some evaluation of the [prior probabilities]. (p. 181).
Abraham Wald; image from Wikimedia.
Having thus turned Wald into a closet Bayesian, de Finetti only needs to object a bit to the distribution-free worst-case reasoning that Wald applied in order to reach his conclusion:
Wald did not explicitly recognize the rule of the probability evaluation in induction and, even more, he seemed inclined to emphasize everywhere the application of the minimax principle, which is reasonable only in strategic situations (like the zero-sum-two-person case in the theory of games) or under such a superstition as that of a "malevolent nature". In spite of its shortcomings, Wald's formulation avoids the narrow interpretation of decisions as acceptance of hypotheses, and offers freedom to choose the proper decision according to a not yet openly recognized prior opinion. (p. 183)
He also later criticizes the minimax solutions on the grounds that "their initial assumptions seem rather arbitrary and artificial" (p. 198). He thus notes:
If the subjectivistic formulation were to lead to conclusion diverging from the objectivistic ones, opposition would be understandable; but the conclusions are the same. Among the admissible rules, the objectivistic theory requires that one be chosen arbitrarily, and it cannot give any criterion of preference; the subjectivistic theory does the same but explains each possible choice as corresponding to a suitable initial opinion. Why then reject this compelling unification? (p. 185)
This is, I think, quite crude, and also misses the essential concern about statistical consistency which plays such a large role in frequentist reasoning, and which has no place in Bayesian reasoning, where all priors are considered equal. Another way of saying this is that Wald would have worried as much about the admissible priors as he worried about the admissible decisions if he had turned Bayesian. A foundation for statistical reasoning cannot itself be statistical.

Tuesday, April 15, 2014

Fisher: "Statistical Methods and Scientific Induction" (1955)

Ronald Fisher; image from Wikimedia Commons.
In this brief paper, Sir Ronald Fisher militates against what he sees as wrong and absurd interpretations of the notion of a statistical test.

The Ideology of Statistics

The core of his argument is that a test only gives positive information when yields a significant difference and thus warrants the rejection of a hypothesis — an absence of a significant difference does not mean "accept." He contends that
… this difference in point of view originated when Neyman, thinking that he was correcting and improving my own early work on tests of significance, … in fact reinterpreted them in terms of that technological and commercial apparatus which is known as an acceptance procedure. (p. 69)
And although acceptance procedures might be good enough for commerce, they have no place in science:
I am casting no contempt on acceptance procedures, and I am thankful, whenever I travel by air, that the high level of precision and reliability required can really be achieved by such means. But the logical differences between such an operation and the work of scientific discovery by physical or biological experimentation seem to me so wide that the analogy between them is not helpful, and the identification of the two sorts of operation is decidedly misleading. (pp. 69–70)
Then comes the juicy part:
I shall hope to bring out some of the logical differences more distinctly, but there is also, I fancy, in the background an ideological difference. Russians are made familiar with the ideal that research in pure science can and should be geared to technological performance, in the comprehensive organized effort of a five-year plan for the nation. How far, within such a system, personal and individual inferences from observed facts are permissible we do not know, but it may be safer, and even, in such a political atmosphere, more agreeable, to regard one's scientific work simply as a contributary element in a great machine, and to conceal rather than to advertise the selfish and perhaps heretical aim of understanding for oneself the scientific situation. In the U.S. also the great importance of organized technology has I think made it easy to confuse the process appropriate for drawing correct conclusions, with those aimed rather at, let us say, speeding production, or saving money. There is therefore something to be gained by at least being being able to think of our scientific problems in a language distinct from that of technological efficiency. (p. 70)
So there you have it: In the technological regime of either of the two Cold War superpowers, "learning," "inference," and private, inner thought are taboo, according to Fisher. Presumably we are to contrast this with the aims of British science going back to Newton.

The Three Issues

Fisher singles out three phrases that he finds particularly offensive in scientific statistics:
  1. "Repeated sampling from the same distribution"
  2. Errors of the "second kind"
  3. "Inductive behaviour"
I'll discuss these one by one.

1. "Repeated sampling from the same distribution"

The issue with the first one is not completely clear to me, but here is what I make of his discussion (pp. 71–72): Suppose you are performing a test to see whether the mean of some population has a specific value; suppose further that the standard deviation of that population is unknown, but that you have estimated it based on the available sample.

The problem then is, if I understand Fisher correctly, that the test depends on the standard deviation being constant and known, but in reality, it is an unknown quantity that you have estimated by a maximum likelihood method. This is, strictly speaking, illegitimate, since any estimate should be based on a numerous and representative sample; but since the standard deviation is a property of samples of size N, you should really have M samples of a sample of size N in order to have some data to estimate from. But clearly, this sets a far too high standard for the amount of data required.

It's a convoluted argument, but I think it makes sense from a rigorously frequentist standpoint: If parameters are consistently interpreted as frequencies, then the only legitimate statistical procedure for learning about an unknown quantity t is to obtain a large number of samples dependent on t and then wait for the law of large numbers to kick in.

Strictly speaking, this means that the amount of data points you need in order to estimate all the parameters in a model will grow exponentially in the number of parameters. That sounds sort of crazy, but if you do not allow yourself to have any model in the absence of data, you really have to wait for the data to overwhelm your initial ignorance before you can say that you have a model of the situation. That takes time.

2. Errors of the "Second Kind"

Errors of the first kind are false negatives: Cases in which, for instance, a population in fact has mean m, but nevertheless exhibits a sample average so far away from m that the hypothesis is rejected. Such errors have a frequentist interpretation, because the likelihoods given m are well-defined even in the absence of a prior distribution over m.

Errors of the second kind are false positives: Some other mean m' different from m produces a sample average so close to m that the false hypothesis of a mean of m is confirmed. This kind of error has no frequentist interpretation, because it requires the alternative hypotheses m' to have prior probabilities, and because it requires that there be a loss function associated with accepting the hypothesis of m when the true mean m' is close to m.


Jerzy Neyman in the classroom, 1973; image from Wikimedia Commons.

Fisher is not willing to assume any of those two instruments. He writes:
It was only when the relation between a test of significance and its corresponding null hypothesis was confused with an acceptance procedure that it seemed suitable to distinguish errors in which the hypothesis is rejected wrongly, from errors in which it is "accepted wrongly" as the phrase does. (p. 73)
Such language is not just scientifically irresponsible, he thinks — it also misunderstands the private states of mind present in the head of a scientist:
The fashion of speaking of a null hypothesis as "accepted when false", whenever a test of significance gives us no strong reason for rejecting it, and when in fact it is in some way imperfect, shows real ignorance of the research worker's attitude, by suggesting that in such a case he has come to an irreversible decision. (p. 73; Fisher's emphasis)
Of course, neither positive nor negative decisions are immune to revision as more data comes in (cf. p. 76), so Fisher prefers to depict the scientist's attitude as one of cautious learning in the face of data. This contrasts with the forced-choice nature of acceptance procedures:
In an acceptance procedure, on the other hand, acceptance is irreversible, whether the evidence for it was strong or weak. It is the result of applying mechanically rules laid down in advance; no thought is given to the particular case, and the tester's state of mind, or his capacity for learning, is inoperative.
By contrast, conclusions drawn by a scientific worker from a test of significance are provisional, and involve an intelligent attempt to understand the experimental situation. (pp. 73–74; Fisher's emphasis).
Note again the insistence on private states of mind as the hallmark of scientific rationality.

3. "Inductive Behaviour"

The last issue Fisher has with Neyman's brand of statistics is shelves under the heading above, but it is really about an issue of linguistics: Neyman contends (according to Fisher's summary — there is no direct reference) that statements like
There is 5% probability that the sample average deviates strongly from the mean
have a meaningful and well-defined interpretation (in terms of likelihood). On the other hand,
There is 5% probability that the mean deviates strongly from the sample average
is meaningless, because the mean is not a random variable.

Fisher disagrees, not because he is a fan of prior probability distributions on the parameters, but because he thinks that such statements could only ever refer to likelihoods. To make this point vivid, he considers (I am changing the example a bit here) a statement of the form
Pr(m < x) = 5%,
where m is a parameter and x is an observation, and he contrasts this with
Pr(m < 17) = 5%.
If one of these statements has a meaning, he says, clearly the other one must have a meaning too, unless we want to "deny the syllogistic process of making a substitution" (p. 75). But Neyman contends that the probability of a statement of the second kind should be "necessarily either 0 or 1" (p. 75), so that only the former probability (the likelihood given the mean) is well-defined.

Fisher comments:
The paradox is rather childish, for it requires that we should wilfully misinterpret the probability statement so as to pretend that the population to which it refers is not defined by our observations and their precision, but is absolutely independent of them. (p. 75)
By this he means that the reference class (the "population") is defined arbitrarily by our experimental set-up. And as he says about populations earlier in the paper, "no one of them has objective reality, all being products of the statistician's imagination" (p. 71).

An Englishman's Duty

In the conclusion, Fisher comes back to the ethical standards of statistics:
As an act of construction the hypothesis is not altogether impersonal, for the scientist's personal capacity for theorizing comes into it; moreover, the criteria by which it is approved require a certain honesty, or integrity, in their application. (p. 75)
Again, he explains that decision-theoretic methods (such as Bayesian statistics) have no business in scientific inference, since the goal is not optimal decisions, but the attainment of truth:
Finally, in inductive inference we introduce no cost functions for faulty judgments … In fact, scientific research is not geared to maximize the profits of any particular organization, but is rather an attempt to improve public knowledge undertaken as an act of faith to the effect that, as more becomes known, or more surely known, the intelligent pursuit of a great variety of aims, by a great variety of men, and groups of men, will be facilitated. We make no attempt to evaluate these consequences, and do not assume that they are capable of evaluation in any sort of currency.
… We aim, in fact, at methods of inference which should be equally convincing to all rational minds, irrespective of any intentions they may have in utilizing the knowledge inferred.
We have the duty of formulating, of summarizing, and of communicating our conclusions, in intelligible form, in recognition of the right or other free minds to utilize them in making their own decisions. (p. 77)
We could hardly have it more explicit: The difference in statistical paradigm is one of ethics.

Saturday, March 1, 2014

Attneave: Applications of Information Theory to Psychology (1959)

Fred Attneave's book on information theory and psychology is a sober and careful overview of the various ways in which information theory had been applied to psychology (by people like George Miller) by 1959.

Attneave explicitly tries to stay clear of the information theory craze which followed the publication of Shannon's 1948 paper:
Applications of Information Theory to Psychology (cover)
Book cover; from Amazon.
Thus presented with a shiny new tool kit and a somewhat esoteric new vocabulary to go with it, more than a  few psychologists reacted with an excess of enthusiasm. During the early fifties some of the attempts to apply informational techniques to psychological problems were successful and illuminating, some were pointless, and some were downright bizarre. At present two generalizations may be stated with considerable confidence:
(1) Information theory is not going to provide a ready-made solution to all psychological problems; (2) Employed with intelligence, flexibility, and critical insight, information theory can have great value both in the formulation or certain psychological problems and in the analysis of certain psychological data (pp. v–vi)
Or in other words: Information theory can provide the descriptive statistics, but there is no hiding from the fact that you and you alone are responsible for your model.

Language Only, Please

Chapter 2 of the book is about entropy rates, and about the entropy of English in particular. Attneave talks about various estimation methods, and he discusses Shannon's guessing game and a couple of related studies.

As he sums up the various mathematical estimation tricks, he notes that predictions from statistical tables tend to be more reliable than predictions from human subjects with respect to the first couple of letters of a text. This means that estimates from human predictions will tend to overestimate the unpredictability of the first few letters of a string.

He then comments:
What we are concerned with above is the obvious possibility that calculated values (or rather, brackets) of HN [= the entropy of letter N given letter 1 through N – 1] will be too high because of the subject's incomplete appreciation of statistical regularities which are objectively present. On the other hand, there is the less obvious possibility that a subject's guesses may, in a certain sense, be too good. Shannon's intent is presumably to study statistical restraints which pertain to language. But a subject given a long sequence of letters which he has probably never encountered before, in that exact pattern, may be expected to base his prediction of the next letter not only upon language statistics, but also upon his general knowledge [p. 40] of the world to which language refers. A possible reply to this criticism is that all but the lowest orders of sequential dependency in language are in any case attributable to natural connections among the referents of words, and that it is entirely legitimate for a human predictor to take advantage of of such natural connections to estimate transitional probabilities of language, even when no empirical frequencies corresponding to the probabilities exist. It is nevertheless important to realize that a human predictor  is conceivably superior to a hypothetical "ideal predictor" who knows none of the connections between words and their referents, but who (with unlimited computational facilities) has analyzed all the English ever written and discovered all the statistical regularities residing therein. (pp. 39–40; emphases in original)
I'm not sure that was "Shannon's intent." Attneave seems to rely crucially on an objective interpretation of probability as well as an a priori belief in language as an autonomous object.

Just like Laplace's philosophical commitments became obvious when he starting talking in hypothetical terms, it is also the "ideal predictor" in this quote which reveals the philosophy of language that informs Attneave's perspective.