Showing posts with label causality. Show all posts
Showing posts with label causality. Show all posts

Wednesday, June 19, 2013

Meier and Robinson: "Why the Sunny Side Is Up" (2004)

This Psychological Science paper by Brian P. Meier and Michael D. Robinson is a report on three experiments which collectively show that
  1. reading positive and negative words facilitates the recognition of letters that are placed high or a low, respectively,
  2. but looking up and down does facilitate the comprehension of positive and negative words.
In the terminology of Judea Pearl, we can thus say that activation of GOOD causes activation of UP, and BAD of DOWN:


Another result that came out of the experiments was that people recognize positive words faster than negative words, regardless of condition. Meier and Robinson do not discuss this effect explicitly, and I'm also not sure what to make of it.

From GOOD to UP, Not from UP to GOOD

The paper opens with a discussion of an experiment which shows that people are faster at judging whether a word is "positive" or "negative" if it's placed in a position on the computer screen which is congruent with the meaning. When the word polite is flashed at the top of the screen, for instance, you're faster at pressing the button corresponding to "positive" than when it flashes at the bottom.

This shows a correlation, but does not disentangle the question of causation. To do this, the two variables need to be manipulated independently of each other. This was achieved in two follow-up experiments which placed a spatial judgment task before or after a valence judgment task.

In the first of these follow-up experiments, the GOOD-to-UP direction was set up by a design looking roughly as follows:


That is, first the subject made a judgment about the valence of a word, and then about the identity of a letter which was presented in either a high or a low position. The last screen showed the word INCORRECT if the subject failed at the second task (as I've assumed in this example). This study found a significant facilitation effect.

The second follow-up experiment set up the UP-to-GOOD direction in a design that may be pictured thus:
 

Here, the subject first had to identify where a spatial cue was, and then to judge the valence of a word. In this task, they found no facilitation effect.

Is This Consistent With Cognitive Metaphor Theory?

In their brief conclusion, Meier and Robinson write:
These findings suggest that, when making evaluations, people automatically assume that objects that are high in visual space are good, whereas objects that are low in visual space are bad. (p. 247)
But there's an issue buried here which is usually evaded in cognitive metaphor theory: In which direction do the connections between source and target domain run? Meier and Robinson's results as well as the notion of "understanding in terms of" suggest that the connections run from target domain to source domain.

This analysis, however, leaves cognitive metaphor theory with little explanatory value with respect to word semantics. Consider for instance this paradigmatic instance of a "conceptual" metaphor:
  • I'm feeling up.
According to the standard analysis, I understand that the word up here means "good" because the domain of spatial position translates into the domain of moods.

But this is exactly the opposite association of what we need. Meier and Robinson's show that we have a completely literal understanding of spatial position, and that looking upwards doesn't associate to feeling good:
People can use their senses to determine whether an object is up or down, white or black. There is no need to borrow metaphor to achieve an understanding of vertical position. Because of these considerations, we view it as unlikely that physical cues, in the absence of an evaluative context, activate evaluations. (p. 245)
There is thus no neurological reason why the word up should activate the meaning "good."

While Meier and Robinson thus offer excellent evidence in favor of a system of analogies in thought, it is not obvious that these should be responsible for the way we understand metaphorical language—only that they can be seen during our comprehension of literal language. To my mind, this creates a huge problem for cognitive metaphor theory, in particular in its more recent brain-talky incarnation.

Thursday, September 13, 2012

Thompson and Mann: "Perceived Necessity Explains the Dissociation Between Logic and Meaning" (1995)

Under which conditions do people think that If A, then B can be paraphrased as A only if B? This paper by Valerie Thompson and Jacqueline Mann is an empirical investigation of the question, checking a couple of relevant parameters.

As it turns out, two factors play a major role: The temporal order between A and B, and whether we perceive A and B to be equivalent in the concrete case at hand.

By contrast, the type of discourse relationship between A and B plays no role. It thus doesn't matter whether the relationship between them is causation, permission, co-occurrence, definition, etc.

Independent Variables

Let me just fix some terminology. What I here call the discourse relationship is what Thompson and Mann call "pragmatic relations." I just dislike this term because it's not quite consistent with the jorgon of linguistics.

The two most important discourse relations that they are dealing with are causation and permission:
  • Butter melts if it's heated. (causation)
  • You may enter if you're over 18. (permission)
They introduce a couple more (p. 1557), but since disourse relationship turns out to have no effect, this is of a minor importance.

Second, when Thompson and Mann talk about "necessity" relationships, they are really talking about condtional perfection. This is the backwards conditional If B, then A that we sometimes infer when we hear the forward one:
  • If water is heated to 100°C, it boils (… and vice versa).
  • If it rains, the pavement will be wet (… but not necessarily vice versa).
The effect is a conflation of implication and bi-implication. This "logical" difference does, unsurprisingly, turn out to have an effect on the acceptability of paraphrases.

Lastly, the notion of temporal succession is the most interesting one, and it interacts in some non-trivial ways with the psychology undergraduates' intuitions about synonymy:
  • If a plant has received enough care, it grows. (A before B)
  • If a plant grows, it has received enough care. (A after B)
In terms of the relationships visible to classical logic, these sentence mean very different things: The first one rules out rules out externalities that could hinder growth even in the event of care; the second one rules out other sufficient causes of growth. However, from an intuitive perspective, the sentences seem to point towards the same underlying causal relationship.

Results

Thompson and Mann's main concern is whether their subjects think that a sentence of the form If A, then B is synonymous with A only if B, and whether it is synonymous with B only if A. As I mentioned above, it turns out that this depends strongly on whether the (inferred, perceived) temporal order of A and B, and the (inferred, perceived) equivalence of A and B.

Thompson and Mann used a super-weird scoring scheme in which their subjects had to assign a 1 to a perfect match and a 7 to a complete mismatch. "For ease of comprehension," they report the transformed score 8 – x instead of x (p. 1557; why didn't they just use the easy one in the first place?).

This gives means between 1 and 7. I've transformed these means into percentages to make it easier to see how far the various means are from the maximal and minimal scores. I did this by computing 100/7 * (y – 1) from the reported y = (8 – x). So let's look at a couple of snapshots from the results of Thompson and Mann's experiment 2b.

First, causal relationships with forward-moving time and no conditional perfection. An example of this is the following:
  • If the car runs out of gas, then it will stall.
    1. The car only runs out of gas if it stalls (13% — equivalent)
    2. The car only stalls if it runs out of gas (50% — not equivalent)
In this case, subjects do not like the actually equivalent form (which suggests a modus tollens inference schema). Note that the percentages are the average scores for this class of sentences, not the specific example.

Now a causal relationship with backward-moving time, but still no conditional perfection:
  • If the car drives, then there is gas in the tank.
    1. The car only drives if there is gas in the tank (79% — equivalent)
    2. There is only gas in the tank if the car drives (21% — not equivalent)
So this reversal of time completely turns the intuitions upside-down: Now, the equivalent paraphrase seems more consistent with the order of terms (STATE only if PRECONDITION), and the non-equivalent seems less natural.

If we put these two sets of statistics together, we get the following chart of acceptabilities:

Lastly, a forward-moving example with conditional perfection:
  • If water is heated to 100°C, it boils.
    1. Water is heated to 100°C only if it boils (26% — equivalent)
    2. Water only boils if it is heated to 100°C (75% — not equivalent)
On a very coarse level, this is the pattern of forward-moving time without conditional perfection; there is an intensity effect, but no reversal of judgments.

The Role of Time

So it seems that the single most predictive factor about intuitions of synonymy and inference is the distinction between forward-moving and backward-moving time. Certain ways of construing a causal situation highlight the potential for following the actual causal direction in your thoughts, and other ways highlight the possibility of following the order of inference rather than the order of events.

If this is true, then it would have some consequences for how difficulty various inference types are, as well as how they errors will occur through "normalization." For instance, denial of the antecedent can be seen as a natural thought to have if we follow the order of events in a case where the literal meaning of the premises requires us to follow the order of inference.

Wednesday, September 12, 2012

Literature on the meaning of "only if"

I've been looking for some empirical studies of A only if B constructions. In the theoretical literature on natural language semantics, there is a number of models, but I want to know more about how they are actually understood. Fortunately, there seem to be some facts about that out there, too.

What Does It Mean, Allegedly?

The problematic issue with the only if construction is that it is supposed to be logically equivalent to a number of related constructions, even though non-logicians sometimes disagree with this. According to the classical convention, the following sentences thus all mean the same:
  • It only thunders if it rains.
  • If it thunders, it rains.
  • If it doesn't rain, it doesn't thunder.
On the other hand, if we reverse the implication, we change the truth conditions:
  • It only rains if it thunders.
  • If it rains, it thunders.
  • It it doesn't thunder, it doesn't rain.
If this was just a mere convention about logical language, all would be fine. The problem is, however, that these sentence forms are not used in the same situations, and they do not integrate equally well into all reasoning patterns in spite of their (alleged) equivalence.

The Performance Problem

One difference between the If A, then B and A only if B forms is that if form is generally more difficult to use in a modus tollens inference than the only if. At least, this is what Carlos Santamaría and Orlando Espino say (Santamaría and Espino 2002, p. 42). They're referring to three studies, including one by Jonathan Evans and M. A. Beck (Evans and Beck 1981).

The problematic case is thus the following inference:
If it thunders, it rains.
It doesn't rain.
–––––––––––––––––
It doesn't thunder.
This (clasically valid) inference should be performed more readily when served in this alternative, and supposedly equivalent formulation:
It only thunders if it rains.
I doesn't rain.
–––––––––––––––––––––
It doesn't thunder.
Cognitively, or perhaps in terms of actual natrual language semantics, this seems to indicate that A only if B works more like the contrapositive If not B, then not A than like its positive translation, If A, then B. Or at least, it seems to issue a conversational warrant closer to it.

It would be interesting to know if this alternative formulation comes with a corresponding decrease—are we trading of willingness to perform the straightforward modus ponens inference for higher rates of modus tollens? This would imply that the following inference generally is less accepted:
It only thunders if it rains.
It thunders.
–––––––––––––––––––––
It rains.
If the only if formulation really does works like a contrapositive, then this inference should appear to us like a modus tollens inference in terms of plausibility and difficulty. I do not know right now whether such an effect can actually be measured or not.

The Issue of Time

Another interesting proposal that Santamaría and Espino cite, also coming from Evans and Beck, is that there is a systematic interaction between our conception of temporal order and the choice of form.

Thus, even though If A, then B, and A only if B are supposedly logically equivalent, we get different patterns of acceptability or naturalness depending on whether A or B happened first. For A preceding B, we then (perhaps) have:
  • If you bought on Tuesday, you're paying on Wednesday.
  • (?) You bought on Tuesday only if you paying on Wednesday.
And for B preceding A:
  • (?) If you're paying on Wednesday, you bought on Tuesday.
  • You're only paying on Wednesday if you bought on Tuesday.
Of course, much clearer intuitions can be produced if we ruffle up the tenses a bit. But this, I think, relatively fair example to start the discussion from.

So, Causality?

Note that the issue of before/after interfaces with the concept of causality, which is notoriously bound up with implication, even if logicians and statisticians hate to admit this fact.

Possibly, the the only way we can really justify an inference from a later effect to a prior cause in the form of If EFFECT, then CAUSE is to objectify the cause and the effect by thinking about the observation of the effect and the deduction of the cause. In this way, we would straighten out the temporal sequence so that EFFECT could in fact precede CAUSE.

If this is true, it has a quite important consequence for the psychology of reasoning: We would then only be able to understand abductive inference by effectivly embedding a cause/effect relationship in a different and larger cause/effect relationship—namely the only in which a real or imagined person reasons from fire (the logical "cause") to smoke (the logical "effect").

Monday, July 30, 2012

Judea Pearl: Causality (2000)

This book serves two puposes: It's a textbook in Bayesian statistics, and it launches a theory of causality which contrasts with Pearl's own earlier position on the topic.

The Meaning of Causation

His new theory introduces a meaningful distinction between causality and correlation by internalizing the concept of an intervention. The idea is that information only propagates downstream after an intervention, while it propagates both upstream and downstream after an observation.

For example, if I observe that the street is wet, I consider both rain and wet shoes more likely. On the other hand, if I make the streets wet (say, by emptying a bucket of water), my subjective probability of rain remains unchanged, while I still consider wet shows more likely.

In section 7.2.1, he applies this idea to the example of price regulation. Price and demand are mutually dependent, but observing a price at a specific level and controlling the price does not have the same effect on the demand.

The Stuff of Semantics

The methods that Pearl discusses have a surprisingly logical flavor, given that the book sells itself as a kind of applied statistics. He is quite keen on presenting his probability calculus as a semantics defined on a set of logical expressions (or queries).

The structures that are used to evaluate sentences in this logic are causal models (arrow-drawings expressing a hypothesis about what affects what) and, more specifically, settings of specific variables in such models. This corresponds roughly to worlds and propositional truth values in modal logic.

Within a causal model (a bunch or arrows connecting the variables in various ways), one can thus be right or wrong about specific probability assignments. But above that, a causal model can also be consistent or inconsistent with a specific probability distribution.

Consider for instance this Bayes net:


With the probability tables below, this is consistent with the distribution P, but not with the distribution Q:
 x  y  z  P(x,y,z)  Q(x,y,z)
0001/41/3
000
000
0001/4
000
0001/41/3
000
0001/41/3

Very frequently, we are in practice interested in answering questions within a fixed causal model rather than saying things that are true for any probability distribution.

The Meaning of Counterfactuals

As I said above, Pearl's logical language allows both for regular conditioning, X | Y, and for intervention, do(Y).

While the former is defined as usual, the latter has an original meaning: The result of updating with do(Y = y) is that we delete all incoming arrows to the node Y and replace it by the value Y. This contrasts with simply conditioning on Y, in which case the causal skeleton remains intact.

Interestingly, Pearl uses his new operator to define the meaning of counterfactual statements. "If X were the case, then Y" is defined in terms of the intervention do(Y), while "Y given X" is defined in terms of conventional conditioning.

For some reason, he requires conditioning to take place before intervention (cf., e.g., p. 206). This makes a mathematical difference in some cases, but I'm not sure whether that difference has any interesting philosophical or linguistic counterpart. An example where there is a difference is the Bayes net below, supposing that X is a coin flip, Y = X, and Z = Y:


In this causal model, we have
  • P( "Given Z, X if Y were true" ) = 1.
However,
  • P( "If Y were true, X given Z" ) = ½.
Maybe that's a desirable quality of a probabilistic logic. I don't know.

Inside and Outside the Model

Counterfactuals have often been criticized in the literature for being empirically meaningless and hence a pollutant in science. However, the semantics introduced by Pearl gives them a specific meaning and allows us in principle to evaluate them. But it's interesting that he notes that they convey information more about the underlying causal model than about empirical values of the variables (see p. 219).

A speaker uttering a counterfactual will thus not communicate something about (irrelevant, non-existent) states of affairs, but about how the world works. In a sense, a sentence like "If you had ... " is then a statement about everything that is not mentioned in the sentence.

This is quite important, also because of the framing problem that Pearl hardly even recognizes the existence of: Any particular evaluation of a counterfactual statement presupposes a causal model, including a split between endogenous variables (whose value is determined by other variables) and exogenous variables (whose values are stochastic).

But it's still quite ambiguous what variables we change and what variable we put in the model in the first place when we understand a sentence like "If I were you, I wouldn't ... " It makes sense that it should convey a model of reality as it does appear a little more cautious than a direct recommendation ("You shouldn't ... "). It's like showing someone your watch rather than telling them what time it is.