Showing posts with label Abraham Wald. Show all posts
Showing posts with label Abraham Wald. Show all posts

Friday, May 8, 2015

Jaynes: Probability (2003), ch. 3

While discussing the probability of drawing the same ball twice when sampling with replacement from an urn, Jaynes breaks off into a little "sermon" (pp. 64–66):
In probability theory there is a very clever trick for handling a problem that becomes too difficult. We just solve it anyway by:
  1. Making it still harder;
  2. Redefining what we mean by “solving” it, so that it becomes something we can do;
  3. Inventing a dignified and technical-sounding word to describe this procedure, which has the psychological effect of concealing the real nature of what we have done, and making it appear respectable.
In the case of sampling with replacement, we apply this strategy by
  1. Supposing that after tossing the ball in, we shake up the urn. However complicated the problem was initially, it now becomes many orders of magnitude more complicated, because the solution now depends on every detail of the precise way we shake it, in addition to all the factors mentioned above;
  2. Asserting that the shaking has somehow made all these details irrelevant, so that the problem reverts back to the simple one where the Bernoulli Urn Rule applies;
  3. Inventing the dignified-sounding word randomization to describe what we have done. This term is, evidently, a euphemism whose real meaning is: deliberately throwing away relevant information when it becomes too complicated for us to handle. (p. 64)
Just to be sure, he adds:
Shaking does not make the result “random,” because that term is basically meaningless as an attribute of the real world; it has no clear definition applicable in the real world. (p. 65)
I think I would agree with Jaynes that randomization in field experiments is akin to a protective ritual. To the extent that such a tirual is intended to protect us against a malicious Nature out to get us, it is absurd.

But like the proponents of randomized trials, Jaynes here fails to see that the evil adversary might in fact be a dishonest experimenter rather than a malicious Nature. Both of these characters are in principle vulnerable to strategies like randomization, but unlike Evil Nature, Evil Scientist actually exists.

Saturday, December 6, 2014

Chernoff: "A career in statistics" (2014)

Chernoff makes some interesting remarks about the philosophy of statistics in his recent autobiographical essay.

First, an anecdote about the three classical decision criteria considered in decision theory:
I had always been interested in the philosophical issues in statistics, and Jimmie Savage claimed to have resolved one. Wald had proposed the minimax criterion for deciding how to select one among the many “admissible” strategies. Some students at Columbia had wondered why Wald was so tentative in proposing this criterion. The criterion made a good deal of sense in dealing with two-person zero-sum games, but the rationalization seemed weak for games against nature. In fact, a naive use of this criterion would suggest suicide if there was a possibility of a horrible death otherwise. Savage pointed out that in all the examples Wald used, his loss was not an absolute loss, but a regret for not doing the best possible under the actual state of nature. He proposed that minimax regret would resolve the problem. At first I bought his claim, but later discovered a simple example where minimax regret had a similar problem to that of minimax expected loss. For another example the criterion led to selecting the strategy A, but if B was forbidden, it led to C and not A. This was one of the characteristics forbidden in Arrow’s thesis.
Savage tried to defend his method, but soon gave in with the remark that perhaps we should examine the work of de Finetti on the Bayesian approach to inference. He later became a sort of high priest in the ensuing controversy between the Bayesians and the misnamed frequentists. (pp. 32–33)
He immediately moves on to one of his own more dismal conclusions about the issue:
I posed a list of properties that an objective scientist should require of a criterion for decision theory problems. There was no criterion satisfying that list in a problem with a finite number of states of nature, unless we canceled one of the requirements. In that case the only criterion was one of all states being equally likely. To me that meant that there could be no objective way of doing science. I held back publishing those results for a few years hoping that time would resolve the issue (Chernoff, 1954). (p. 33)
I haven't read the paper he is referring to here, but it seems like the text has jumbled up the conclusions: I think what he meant to say is that there is no single good decision function when we have infinitely many states, since the criteria essentially require us to use a uniform distribution. But I would have to check the details.

Finally, he moves on to a more on-record explication of his position:
In the controversy, I remained a frequentist. My main objection to Bayesian philosophy and practice was based on the choice of the prior probability. In principle, it should come from the initial belief. Does that come from birth? If we use instead a non-informative prior, the choice of one may carry hidden assumptions in complicated problems. Besides, the necessary calculation was very forbidding at that time. The fact that randomized strategies are not needed for Bayes procedures is disconcerting, considering the important role of random sampling. On the other hand, frequentist criteria lead to the contradiction of the reasonable criteria of rationality demanded by the derivation of Bayesian theory, and thus statisticians have to be very careful about the use of frequentist methods. 
In recent years, my reasoning has been that one does not understand a problem unless it can be stated in terms of a Bayesian decision problem. If one does not understand the problem, the attempts to solve it are like shooting in the dark. If one understands the problem, it is not necessary to attack it using Bayesian analysis. My thoughts on inference have not grown much since then in spite of my initial attraction to statistics that came from the philosophical impact of Neyman–Pearson and decision theory. (p. 33)

Tuesday, October 28, 2014

Blackwell and Girshick: Theory of Games and Statistical Decisions (1954), Ch. 4

There's an interesting representation theorem in Blackwell and Girshick's textbook in statistics: It provides a set of sufficient conditions for a preference ordering over lotteries to be expressible as a prior probability distribution (Th. 4.3.1, p. 118).

I assume the theorem comes from either Wald, Savage, or de Finetti, but no reference is given.

Well-Behaved Preferences

A lottery can here be defined as a function from the sample space to the real numbers. The conditions in the theorem are then the following:
  • The ordering of two lotteries $f$ and $g$ cannot depend on the availability of other lotteries.
  • If a lottery $f$ provides a higher payoff than a lottery $g$ at all points in the sample space $\Omega$, then $f$ must be preferred to $g$.
  • If $f$ is preferred to $g$, then $f+h$ must be preferred to $g+h$.
An inspection of the proof also shows that they should have included a continuity condition:
  • If $f_1, f_2, \ldots$ is a series of lotteries converging to a limit $f$, and if $f_i$ is preferred to $g$ for all $i$, then $f$ must also be preferred to $g$.
When these conditions are met, the preference ordering over the lotteries can be expressed as a distribution over the sample space, unless it categorizes all lotteries as equally good.

The Large and the Good

The proof of the theorem uses the fact that two convex, disjoint, open sets can be separated by a hyperplane. Here's a sketch:

If we let $e$ be the lottery that pays zero in all situations $\omega \in \Omega,$ then we can define the following sets:
\begin{eqnarray}
F[>] &=& \{f\ |\ \forall \omega \in \Omega: f(\omega) > 0\},
\\
F[\gtrsim] &=& \{f\ |\ f\gtrsim e\},
\end{eqnarray}
and we can further define, in the usual way, $F[>] + F[\gtrsim]$ to be the sums of lotteries from those two sets.

Now, $F[>]$ is an open set, and hence $F[>] + F[\gtrsim]$ is, too. Further, $F[>]$ and $F[\gtrsim]$ are both closed under addition (due to the third assumption), and hence their sum is, too. They are also both closed under multiplication with a scalar, and again, so is their sum — but this latter argument requires a bit more spelling out.

Rational and Real Convexity

Suppose a lottery $f$ is preferred to the zero lottery, that is, $f \in F[\gtrsim]$. The third assumption of the theorem then tells us that
$$
e \;\lesssim\; f \;\lesssim\; f + f \;\lesssim\; f + f + f \;\lesssim\; \ldots \;\lesssim\; nf.
$$
By further adding multiple copies of the zero lottery to both sides of this preference ineqality, we can see that
$$
me \;\lesssim\; (m-1)e + nf \;=\; nf,
$$
where we have selectively used the fact that $e$ is the zero lottery. Putting these facts together, we then have the preference inequality
$$
e \;\lesssim\; \left(\frac{n}{m}\right)f.
$$
By using a positive sequence of rational approximations $(n/m) \rightarrow \lambda$, we can use this fact along with the continuity assumption to conclude that $F[\gtrsim]$ is closed under multiplication with any positive, real scalar $\lambda$.

I don't think there's a way around this last technicality. It is, incidentally, the same proof technique used to prove that the logarithmic functions are the only continuous functions that turn products into sums.

Cutting the Cake

At any rate, $F[>] + F[\gtrsim]$ is an open and convex set separated from the singleton set $\{e\}$. We can therefore conclude that there is a lottery (or vector) $p$ which defines the hyperplane $\{f\ |\ f\cdot p=0\}$ separating $F[>] + F[\gtrsim]$ from $\{e\}$. The set $F[>] + F[\gtrsim]$ is thus a subset of the half-space $\{f\ |\ f\cdot p \geq 0\}$.

This vector $p$ must have nonnegative coordinates, since the set $F[>]$ is unbounded in all positive directions. If $p$ had a negative coordinate, $p(\omega) \leq 0$, we could choose a lottery for which the corresponding coordinate, $f(\omega)$, was so large that $f\cdot p < 0$. This would violate the definition of $p$, and $p$ hence has to be a nonnegative vector which can be interpreted as a probability distribution.

It would also have the property that $f\cdot p \geq g\cdot p$ if and only if $f \gtrsim g$. This follows from the fact that $f\cdot p \geq g\cdot p$ if and only if $(f - g) \in F[>] + F[\gtrsim]$, which holds if and only if $f$ can be expressed as the sum of $g$ and some lottery preferable to the zero lottery.

Saturday, July 26, 2014

Wald: Sequential Analysis (1947)

I've always thought that Shannon's insights in the 1940s papers seemed like pure magic, but this book suggests that at parts of his way of thinking were already in the air at the time: From his own frequentist perspective, Wald comes oddly close to defining a version of information theory.

The central question of the book is:
How can we devise decision procedures that map observations into {Accept, Reject, Continue} in such a way that (1) the probability of wrongly choosing Accept or Reject is low, and (2) the expected number of Continue decisions is low?
Throughout the book, the answer that he proposes is to use likelihood ratio tests. This puts him strangely close to the Bayesian tradition, including for instance Chapter 5 of Jeffreys' Theory of Probability.

A sequential binomial test ending in rejection (p. 94)

In particular, the coin flipping example that Jeffreys considers in his Chapter 5.1 is very close in spirit to the sequential binomial test that Wald considers in Chapter 5 of Sequential Analysis. However, some differences are:
  • Wald compares two given parameter values p0 and p1, while Jeffreys compares a model with a free parameter to one with a fixed value for that parameter.
  • Jeffreys assigns prior probabilities to everything; but Wald only uses the likelihoods given the two parameters. From his perspective, the statistical test will thus have different characteristics depending on what the underlying situation is, and he performs no averaging over these values.
This last point also means that Wald is barred from actually inventing the notion of mutual information, although he comes very close. Since he cannot take a single average over the log-likelihood ratio, he cannot compute any single statistic, but always has to bet on two horses simultanously.

Tuesday, May 6, 2014

de Finetti: Probability, Induction, and Statistics (1972)

In Chapters 8 and 9 of this anthology, Bruno de Finetti reiterates his reasons for espousing Bayesian probability theory as the unique optimal calculus of reasoning. This brings him into a discussion of several controversies surrounding the two paradigms of statistics.

Bruno de Finetti and a computer; image from www.moebiusonline.eu.

No Unknown Unknowns

According to de Finetti, the ordinary meaning of the word "probability" is "a degree of belief" (p. 148), and he rejects any attempt to define it in terms of frequency:
… we reject the idea that the ostensible notion of identical events or trials gives a suitable basis for an empirical formulation of a frequentist theory of probability or for some objectivistic form of the "law of large numbers". (p. 154)
Consequently:
The probability of an event conditional on, or in the light of, a specified result is a different probability, not a better evaluation of the original probability. (p. 149)
There is thus no such things as an "unknown probability." You always know your own uncertainty:
Any assertion concerning probabilities of events is merely the expression of somebody's opinion and not itself an event. There is no meaning, therefore, in asking whether such an assertion is true or false or more or less probable. (p. 189)
Thus, "speaking of unknown probabilities must be forbidden as meaningless" (p. 190) and in fact rejected as a "superstition" (p. 154–55).

But of course we do have problems assigning numbers of things, so de Finetti has some explaining to do. He thus invokes the analogy of choosing a price for a commodity:
A personal probability is, in effect, a quantitative decision closely akin to deciding on a price. In seeking to fix such a number with precision the person will sooner or later encounter difficulties that evoke the expressions "vagueness", "insecurity", or "vacillation". Analysis of this omnipresent phenomenon has given rise to misunderstandings. Thus, attempts to say that the exact probabilities are "meaningless" or "non-existent" pose more severe problems than they are intended to resolve, similarly for replacements of individual probabilities by intervals or by second-order probabilities. […] Sight should not be lost of the the fact that a person may find himself in an economic situation that entails acting in accordance with a sharply defined probability, whether the person chooses his act with security or not. (p. 145)
In spite of this seeming pluralism about personal opinion, he still maintains that the mathematical concept of probability is an idealization:
The (subjectivistic) theory of probability is a normative theory (p. 151).
But of course, the latter refers only to the mechanics of the calculus, not the choice of priors.

Rants Against Frequentism

De Finetti hates frequentist statistics. In his brief historical sketch, he says that the frequentist theory is a set of "substitutes" for Bayesian reasoning which were supposed to fill the "void" left after the analysis by Bayes was rejected (p. 161).

He adds:
The method pursued in the construction of such substitutes consists in general of adopting or imitating some case where the correct method reduces to a simple form based on summarizing parameters, however substituting for the true formulation and justification some incomplete and fragmentary justification or even no justification at all, as comes to seem legitimate when each notion is interpreted as something autonomous and arbitrary. For each isolated problem it appeared thus legitimate to devise as many ad hoc expedients as desired, and in fact it often happens that several are devised, proposed, and applied, to a single problem. (p. 161)
Shortly after, another rant follows:
In this manner, any notion of a systematic and meaningful interpretation of the problem of statistical inference is abandoned for the position of devising, case by case, "tests" of hypotheses or methods of "estimating" parameters. This means formulating, as an autonomous and largely arbitrary question, the problem of extracting from experience something that is apparently to be employed as though it were a conclusion or conviction, while asserting that it is neither one nor the other. (p. 162)
He is specifically angry about the "grossly inconsistent" notion of tests and hypothesis rejections, which he finds to be perverse distortions of the proper use of Bayes' rule (p. 163):
The severest of these mutilations is that of the oversimplified criteria according to which a probability P(E | H) us taken as a basis for rejecting the isolated hypothesis H if this probability, for the observation E, is small. (p. 163)
Such hypothesis rejection are, namely, ambiguous about what event E the observed data actually testifies to, as in the problem of choosing between one-sided and two-sided tests:
If, for example, as is often the case, E consists in having observed the exact value x of a random number X, such a deviation, the probability of that exact value is ordinarily zero. In order to eliminate the evident meaninglessness of this criterion that rejects the hypothesis no matter what value x may have, some other is substituted for it, such as observation of a value equal to or greater than x in absolute value, or equal or greater in absolute value and of the same sign. But all these variables are arbitrary, at least in the framework of so crudely mutilated a formulation. (p. 163)
On the following page, he also gives the example of having to decide whether a point on a target was hit by a particular marksman. He gives various examples of sets that such a point can belong to: The singleton set containing only the point itself, a circle having the point as a center, a slice of the target containing the point, a circle having the center of the target as its center, etc.

Various ways of construing the acceptance region for a test.

He continues to say that "One might say that all the deficiencies of objectivistic statistics stem from insistence on using only what appears to be soundly based" (p. 165). This, he says, is like setting a price according to the things that are easiest to measure rather than the things that are most relevant.

The issue of building a statistical enterprise on likelihoods alone is, he contends, like a systematic attempt to find P(E | H) when you are looking for P(H | E). In an example he attributes to Halphen:
We need a cement that will not be harmed by water. The merchant advises us to buy a certain kind that, he assures us, will not harm water. He does not try to cheat us by saying that the two things are equivalent but he want to convince us not to insist on asking for what we need (p. 173).
This is apparently a commentary on related example used by Neyman.

De Finetti on Wald

In a series of papers from the 1940s and 50s, Abraham Wald developed a theory of "admissible decision functions" for decision problems with uncertainty (see, e.g., here). His idea was to consider a decision admissible if it minimized the maximal damage that could obtain in the given situation. This correspond to the solution of a two-person zero-sum game against a malevolent nature.

In his discussion of Wald's theory, de Finetti helpfully "completes" the specification of a decision problem by putting a prior probability on the various hypotheses. Having provided these marginal probabilities, he comments:
Of course, these marginal elements do not appear in Wald's formulation; their absence there is just what prevents the problem of decision from having the solution that is obvious when the table is thus completed. Namely, choose the decision corresponding to the minimal(mean) loss, or equivalently to the maximal(mean) gain or the maximal (mean) utility. Here we have always put "mean" between parentheses but from now on shall suppress the word altogether; for value and utility in an uncertain situation is, by definition, the mathematical expectation of the values of utilities. (p. 179)
In his own work, Wald concluded that the admissible strategies are the mixed strategies whose support consists of pure strategies that are optimal for some parameter setting. But these are also the ones that can be rationalized by some prior probability distribution, so de Finetti happily concludes that
… the admissible decisions are the Bayesian ones; that is, those that minimize the loss with respect to some evaluation of the [prior probabilities]. (p. 181).
Abraham Wald; image from Wikimedia.
Having thus turned Wald into a closet Bayesian, de Finetti only needs to object a bit to the distribution-free worst-case reasoning that Wald applied in order to reach his conclusion:
Wald did not explicitly recognize the rule of the probability evaluation in induction and, even more, he seemed inclined to emphasize everywhere the application of the minimax principle, which is reasonable only in strategic situations (like the zero-sum-two-person case in the theory of games) or under such a superstition as that of a "malevolent nature". In spite of its shortcomings, Wald's formulation avoids the narrow interpretation of decisions as acceptance of hypotheses, and offers freedom to choose the proper decision according to a not yet openly recognized prior opinion. (p. 183)
He also later criticizes the minimax solutions on the grounds that "their initial assumptions seem rather arbitrary and artificial" (p. 198). He thus notes:
If the subjectivistic formulation were to lead to conclusion diverging from the objectivistic ones, opposition would be understandable; but the conclusions are the same. Among the admissible rules, the objectivistic theory requires that one be chosen arbitrarily, and it cannot give any criterion of preference; the subjectivistic theory does the same but explains each possible choice as corresponding to a suitable initial opinion. Why then reject this compelling unification? (p. 185)
This is, I think, quite crude, and also misses the essential concern about statistical consistency which plays such a large role in frequentist reasoning, and which has no place in Bayesian reasoning, where all priors are considered equal. Another way of saying this is that Wald would have worried as much about the admissible priors as he worried about the admissible decisions if he had turned Bayesian. A foundation for statistical reasoning cannot itself be statistical.