Showing posts with label objective probability. Show all posts
Showing posts with label objective probability. Show all posts

Wednesday, March 26, 2014

von Mises: Probability, Statistics and Truth (1951), ch. 1

Richard von Mises; from Wikipedia.
Richard von Mises was an important proponent of the frequentist philosophy of probability.

In his book Probability, Statistics and Truth, he militates against the use of the word "probability" for anything other than indefinitely repeatable experiments with converging relative frequencies (pp. 10–12). He also compares probabilities to physical constants like the velocity of a molecule (p. 21) and asserts that the law of large numbers is an empirical generalization comparable to physical laws like the conservation of energy (pp. 16, 22, and 26).

Reference Class Relativism

A consequence of this frequentist notion of probability is that specific events do not have probabilities. Only infinite classes of comparable events can have probabilities.

For instance, when your coin comes up heads at 10 o'clock, that's a different event from the coin coming up heads at 11 o'clock in infinitely many ways. Only because you choose which properties from the situation to select can you identify the two events as equivalent.

As a kind of argument for this reference class relativism, von Mises asserts that a specific person has a different probability of dying depending on the reference class (e.g., people over 40, men over 40, male smokers over 40, etc). We thus have to explicitly select the reference class before we can talk about "the" probability.

He comments:
One might suggest that a correct value of the probability of death for Mr. X may be obtained by restricting the collective to which he belongs as far as possible, by taking into consideration more and more of his individual characteristics. There is, however, no end to this process, and if we go further and further into the selection of the members of the collective, we shall be left finally with this individual alone. (p. 18)
Even from a frequentist perspective, I'm not sure this makes sense. The fact that we have narrowed down our reference class so much that there is only a single real person left in it should not change the fact that we still have an intensional definition of the class. In so far we do, we should be able to apply that definition to the outcome of any sequence of candidates, like an infinite sequence of people or experiments. In reality, it is only data sparsity that keeps use from going "further and further."

So I think von Mises has a theoretical choice to make: Either, he must require that reference classes be actually infinite, or he must merely require that they be potentially infinite.

"Randomness" and Insensitivity to Subsequence Selection

Von Mises spends a large part of the lecture elaborating a notion of "randomness" which is intended to capture the difference between asymptotically i.d.d. sequences and not asymptotically i.d.d. sequences with the same limiting frequencies. He does so by adding the requirement that the limiting frequencies are independent of subsequence selection.

A possibly more intuitive way of stating that definition would be in terms of a Topsøe-style game: A structure-finding player is tasked to pick infinitely many places in a sequence based on past data and is rewarded when the empirical frequencies fails to converge to a given distribution; a structure-hiding player is tasked to select the sequence and is rewarded when the frequencies do converge to the given distribution.

If the structure-hider then introduces any systematic dependence between the experiments, the structure-finder can exploit these regularities to outgamble the structure-hider. Thus, only asymptotically i.d.d. sequences are part of an equilibrium.

I haven't checked the details, but this game seems to be the same as that suggested by Shafer and Vovk, although (if I remember correctly), they only consider fair (that is, maximum-entropy) i.i.d. coins, not arbitrary biases. But at any rate, coin flipping is, like distributions on a finite set, one of the cases in which there is a maximum entropy distribution even in the absence of an externally given mean.

Tuesday, March 18, 2014

Lewis: "Humean Supervenience Debugged" (1994)

In this paper, David Lewis wrings his hands at a phenomenon he calls "undermining." He considers a probabilistic model as "undermined" if it assigns a positive probability to a data set that would cause a rational agent to adopt a different model.

"Contradiction!"

To say this in a vocabulary closer to Lewis', suppose that C(F | E) is the posterior probability a rational believer would assign to the event F in light of the evidence E. Suppose further that we are looking at the specific case in which E reveals the actual parameters of the world (e.g., "This coin has bias 0.3"), and F is a possible future which would produce a different subjective belief in our hypothetical observer (e.g., "The empirical frequency will be 0.4")

The question is then: Given E, does the possible future F have zero probability, or positive probability? Without giving any argument, Lewis asserts that
there is some present chance that events would would go in such a way as to complete a chancemaking pattern that would make the present chances different from what they actually are. (p. 482)
But he contrasts that with the following "argument":
But F is inconsistent with E, so C(F/E) = 0. Contradiction. (p. 483)
The former of these quotes seem to indicate that he is thinking about a finite sample from the model (consistent with the example he gives on p. 488). The latter argument, on the other hand, seems to assume that he is talking about a limiting frequency from an ergodic process or something like that — unless he seriously believes that empirical frequencies cannot differ from parameter values.

The Super-Objectivist

But this way of putting the argument is of course alien to Lewis. He has no concept of a statistical model, and he thinks that the credence of a rational agent is a unique and well-defined concept that doesn't require any assumptions:
Despite appearances and the odd metaphor, this is not epistemology! You're welcome to spot an analogy, but I insist that I am not talking about how evidence determines what's reasonable to believe about laws and chances. Rather, I'm talking about how nature—the Humean arrangement of qualities—determines what's true about laws and chances. Whether there are any believers living in the lawful and chancy world has nothing to do with it. (pp. 481–82)
This is even stronger and more absurd than classical objectivism. Instead of just discarding certain models as inconsistent with the evidence, Lewis assumes that the evidence suggests a single optimal model out of its own accord. For no apparent reason, he also wants "nature" to do this in a retrospective manner even though there is no reason to, given that he has expelled all subjective observers from the universe.

The Big Flip

Lewis' own solution to the "paradox" is to say that credences should be conditioned on "theories" as well as data — but "theory" doesn't quite mean what it sounds like. This is evident from the example he gives towards the end of the paper.

In this example, he assumes that a coin has exhibited a frequency of 2/3 heads in the past, and he assumes that this means that our hypothetical rational agent estimates its bias to be 2/3.

The "theory" T that he wants us to consider is then that the next 10,002 coin flips exhibit a frequency of exactly 2/3 heads, i.e., 6,668 heads and 3,334 tails. This event has the binomial probability
Pr(T) = B(6,668; 10,002, 2/3).
He then asks us to consider a possible future A in which the next four coin flips come up heads. Still using the parameter estimate of 2/3, this has the binomial probability
Pr(A) = B(4; 4, 2/3).
What is the conditional probability Pr(A | T)? Since the "theory" T did not change the parameter estimate 2/3, one might think that it equals the unconditional probability Pr(A). But for no apparent reason, Lewis decides to take the four coin flips in A from the coin flips in T, producing an amputated event T' with 3 fewer heads and 1 fewer tails. Even more oddly, he computes Pr(A, T') as if the two events were independent even though the observation of either would clearly change the parameter estimate used to compute the conditional probability of the other.

So according to his logic, the "joint probability" of A and T' is then
Pr(A, T') = B(4; 4, 2/3) B(6,665; 9,998, 2/3).
By dividing this by Pr(T), he supposedly finds the "conditional probability" of A given T.

This computation is, of course, completely absurd. If the parameter had been 1/3 instead of 2/3, it would have produced a "probability" larger than 1. So I'm afraid the example isn't doing much good.

Saturday, March 1, 2014

Attneave: Applications of Information Theory to Psychology (1959)

Fred Attneave's book on information theory and psychology is a sober and careful overview of the various ways in which information theory had been applied to psychology (by people like George Miller) by 1959.

Attneave explicitly tries to stay clear of the information theory craze which followed the publication of Shannon's 1948 paper:
Applications of Information Theory to Psychology (cover)
Book cover; from Amazon.
Thus presented with a shiny new tool kit and a somewhat esoteric new vocabulary to go with it, more than a  few psychologists reacted with an excess of enthusiasm. During the early fifties some of the attempts to apply informational techniques to psychological problems were successful and illuminating, some were pointless, and some were downright bizarre. At present two generalizations may be stated with considerable confidence:
(1) Information theory is not going to provide a ready-made solution to all psychological problems; (2) Employed with intelligence, flexibility, and critical insight, information theory can have great value both in the formulation or certain psychological problems and in the analysis of certain psychological data (pp. v–vi)
Or in other words: Information theory can provide the descriptive statistics, but there is no hiding from the fact that you and you alone are responsible for your model.

Language Only, Please

Chapter 2 of the book is about entropy rates, and about the entropy of English in particular. Attneave talks about various estimation methods, and he discusses Shannon's guessing game and a couple of related studies.

As he sums up the various mathematical estimation tricks, he notes that predictions from statistical tables tend to be more reliable than predictions from human subjects with respect to the first couple of letters of a text. This means that estimates from human predictions will tend to overestimate the unpredictability of the first few letters of a string.

He then comments:
What we are concerned with above is the obvious possibility that calculated values (or rather, brackets) of HN [= the entropy of letter N given letter 1 through N – 1] will be too high because of the subject's incomplete appreciation of statistical regularities which are objectively present. On the other hand, there is the less obvious possibility that a subject's guesses may, in a certain sense, be too good. Shannon's intent is presumably to study statistical restraints which pertain to language. But a subject given a long sequence of letters which he has probably never encountered before, in that exact pattern, may be expected to base his prediction of the next letter not only upon language statistics, but also upon his general knowledge [p. 40] of the world to which language refers. A possible reply to this criticism is that all but the lowest orders of sequential dependency in language are in any case attributable to natural connections among the referents of words, and that it is entirely legitimate for a human predictor to take advantage of of such natural connections to estimate transitional probabilities of language, even when no empirical frequencies corresponding to the probabilities exist. It is nevertheless important to realize that a human predictor  is conceivably superior to a hypothetical "ideal predictor" who knows none of the connections between words and their referents, but who (with unlimited computational facilities) has analyzed all the English ever written and discovered all the statistical regularities residing therein. (pp. 39–40; emphases in original)
I'm not sure that was "Shannon's intent." Attneave seems to rely crucially on an objective interpretation of probability as well as an a priori belief in language as an autonomous object.

Just like Laplace's philosophical commitments became obvious when he starting talking in hypothetical terms, it is also the "ideal predictor" in this quote which reveals the philosophy of language that informs Attneave's perspective.