Showing posts with label experimental semantics. Show all posts
Showing posts with label experimental semantics. Show all posts

Thursday, October 11, 2012

Bob van Tiel: "Embedded scalars and typicality" (2012)

Bob van Tiel has, as far as I understand, been arguing for a while that the various empirical problems surrounding scalar implicatures can be explained in terms of typicality. So the strangeness of saying that I ate some of the apples if I in fact ate all of them should be compared to the strangeness of saying there's a bird in the garden if there is in fact an ostrich in my garden.

This argument is nicely and succinctly presented in a manuscript archived at the online repository The Semantics Archive. It contains a fair amount of nice empirical data.

A Bibliography

First of all, the paper contains pointers to most of the interesting recent literature on the subject. Let me just liberally snip out a handful of good references that I either have read or should read:
This list should probably also include the following, which I still have to read:

Quantification According to van Teil

In sections 6 and 8 of the paper, van Teil suggests a very particular semantics for the use of some and any, both extracted from "goodness" ratings by 30 American subjects regarding the sentences All the circles are black and Some of the circles are black.

Semantics for All

His suggestion for the semantics of all is, loosely speaking, that the truth value V("all x are F") should be computed as the harmonic mean of the truth values of V("x1 is F"), V("x2 is F"), etc.

This obviously only makes sense for finite sets, but more strangely, it does not make sense if the truth value 0 occurs anywhere (since the harmonic mean involves a division). Consequently, he has to assume that V("x is black") = .1 when then x is white, and = .9 when x is black.

While this is not completely unreasonable, it does introduce yet another degree of freedom in his statistical fit (remember, he already chose the aggregation function himself), and it be a cause for some caution when interpreting his significance levels.

Semantics for Some

With respect to some, his suggestion is that the paradigmatic case of some circles are black is half of the circles are black. He thus sets the truth value V("some x are F") to be 1 minus the squared difference between the actual case and the half-of-the-individuals case. Ideally, this should give rise to truth value computation of the form
T(k) = 1 – (n/2 – k)2.
However, on the graph on page 17 of the paper, we can see that T(5) < 7 (7 being the maximal "goodness" level), so even when exactly half of the circles are black, we do not get maximal truth. This must be due to some additional assumption like the .9 parameter introduced above, but as far as I can see, he doesn't explain this anywhere in the paper.

One assumption he does make explicit is that
this definition is supplemented with penalties for the situations where the target sentence is unequivocally false (i.e., the 0 and 1 situations) (p. 18)
While these seems relatively innocuous as a general move, we should note that the situation in which exactly one circle is black counts as a counterexample to Some of the circles are black. It also seems to postulate to different mechanisms for evaluating a sentence: First comparing it to a prototype example, and then in addition checking whether it is "really" true. This extra postulation makes his typicality model lose a lot of its attraction, since it discreetly smuggles conventional truth-conditional semantics back into the system rather than superseding it.

Van Tiel's Comments on Chemla and Spector

While the rest of the paper is reasonably clear, there is one part that I do not understand. This is the part where van Tiel recreates the results from Chemla and Spector's letter-and-circle judgment task.

Here's what I do get: He says that the sentence used by Chemla and Spector,
Every letter is connected to some of its circles
suggests most strongly a some-but-not-all reading (labeled "Mixed"), less strongly an all reading, and least strongly a none reading. So however a subject rates the seven different pictures given by Chemla and Spector (0 to 6 connections), they should respect this constrain on appropriateness orderings.

But then van Tiel says the following:

Using Excel, I randomly generated 5,000 values for each of the three cases such that every triplet obeyed the constraint [that some suggests Mixed more than All, and All more than None]. For every triplet, I calculated the typicality value for the seven situations. Ultimately, I derived the mean from these values for comparison with the results of Chemla & Spector. The product-moment correlation between the mean typicality values from the Monte Carlo simulation and the mean suitability values found by Chemla & Spector was nearly perfect (r = 0.99, p < .001). This demonstrates that Chemla & Spector’s results can almost entirely be explained as typicality effects. (p. 19)
I don't get what it is that he is simulating here. Since he randomly generates triplets (not 7-tuples), the stochastic part must be the proposed "goodness" intuition of a random subject. But how does he go from those three numbers to assigning ratings to all seven cases? I suppose you could compute backwards from the three values to the parameter settings for the model discussed above, but that doesn't seem to be what he's doing. So what is he doing?

I think it would have made more sense to compute the theoretically expected truth value of Chemla and Spector's sentence directly now that he has just gone through such pains to construct a compositional semantics for some and every.

We have the number of connections for each picture, so we can compute the truth value of, say, The letter A is connected to some of its circles; and we also have, in each condition, the set of pictures, so we could compute the harmonic mean of these values for the six truth values that are presented to the subject. Why not do that instead if we really want to test the model?

Thursday, September 13, 2012

Thompson and Mann: "Perceived Necessity Explains the Dissociation Between Logic and Meaning" (1995)

Under which conditions do people think that If A, then B can be paraphrased as A only if B? This paper by Valerie Thompson and Jacqueline Mann is an empirical investigation of the question, checking a couple of relevant parameters.

As it turns out, two factors play a major role: The temporal order between A and B, and whether we perceive A and B to be equivalent in the concrete case at hand.

By contrast, the type of discourse relationship between A and B plays no role. It thus doesn't matter whether the relationship between them is causation, permission, co-occurrence, definition, etc.

Independent Variables

Let me just fix some terminology. What I here call the discourse relationship is what Thompson and Mann call "pragmatic relations." I just dislike this term because it's not quite consistent with the jorgon of linguistics.

The two most important discourse relations that they are dealing with are causation and permission:
  • Butter melts if it's heated. (causation)
  • You may enter if you're over 18. (permission)
They introduce a couple more (p. 1557), but since disourse relationship turns out to have no effect, this is of a minor importance.

Second, when Thompson and Mann talk about "necessity" relationships, they are really talking about condtional perfection. This is the backwards conditional If B, then A that we sometimes infer when we hear the forward one:
  • If water is heated to 100°C, it boils (… and vice versa).
  • If it rains, the pavement will be wet (… but not necessarily vice versa).
The effect is a conflation of implication and bi-implication. This "logical" difference does, unsurprisingly, turn out to have an effect on the acceptability of paraphrases.

Lastly, the notion of temporal succession is the most interesting one, and it interacts in some non-trivial ways with the psychology undergraduates' intuitions about synonymy:
  • If a plant has received enough care, it grows. (A before B)
  • If a plant grows, it has received enough care. (A after B)
In terms of the relationships visible to classical logic, these sentence mean very different things: The first one rules out rules out externalities that could hinder growth even in the event of care; the second one rules out other sufficient causes of growth. However, from an intuitive perspective, the sentences seem to point towards the same underlying causal relationship.

Results

Thompson and Mann's main concern is whether their subjects think that a sentence of the form If A, then B is synonymous with A only if B, and whether it is synonymous with B only if A. As I mentioned above, it turns out that this depends strongly on whether the (inferred, perceived) temporal order of A and B, and the (inferred, perceived) equivalence of A and B.

Thompson and Mann used a super-weird scoring scheme in which their subjects had to assign a 1 to a perfect match and a 7 to a complete mismatch. "For ease of comprehension," they report the transformed score 8 – x instead of x (p. 1557; why didn't they just use the easy one in the first place?).

This gives means between 1 and 7. I've transformed these means into percentages to make it easier to see how far the various means are from the maximal and minimal scores. I did this by computing 100/7 * (y – 1) from the reported y = (8 – x). So let's look at a couple of snapshots from the results of Thompson and Mann's experiment 2b.

First, causal relationships with forward-moving time and no conditional perfection. An example of this is the following:
  • If the car runs out of gas, then it will stall.
    1. The car only runs out of gas if it stalls (13% — equivalent)
    2. The car only stalls if it runs out of gas (50% — not equivalent)
In this case, subjects do not like the actually equivalent form (which suggests a modus tollens inference schema). Note that the percentages are the average scores for this class of sentences, not the specific example.

Now a causal relationship with backward-moving time, but still no conditional perfection:
  • If the car drives, then there is gas in the tank.
    1. The car only drives if there is gas in the tank (79% — equivalent)
    2. There is only gas in the tank if the car drives (21% — not equivalent)
So this reversal of time completely turns the intuitions upside-down: Now, the equivalent paraphrase seems more consistent with the order of terms (STATE only if PRECONDITION), and the non-equivalent seems less natural.

If we put these two sets of statistics together, we get the following chart of acceptabilities:

Lastly, a forward-moving example with conditional perfection:
  • If water is heated to 100°C, it boils.
    1. Water is heated to 100°C only if it boils (26% — equivalent)
    2. Water only boils if it is heated to 100°C (75% — not equivalent)
On a very coarse level, this is the pattern of forward-moving time without conditional perfection; there is an intensity effect, but no reversal of judgments.

The Role of Time

So it seems that the single most predictive factor about intuitions of synonymy and inference is the distinction between forward-moving and backward-moving time. Certain ways of construing a causal situation highlight the potential for following the actual causal direction in your thoughts, and other ways highlight the possibility of following the order of inference rather than the order of events.

If this is true, then it would have some consequences for how difficulty various inference types are, as well as how they errors will occur through "normalization." For instance, denial of the antecedent can be seen as a natural thought to have if we follow the order of events in a case where the literal meaning of the premises requires us to follow the order of inference.

Wednesday, September 12, 2012

Literature on the meaning of "only if"

I've been looking for some empirical studies of A only if B constructions. In the theoretical literature on natural language semantics, there is a number of models, but I want to know more about how they are actually understood. Fortunately, there seem to be some facts about that out there, too.

What Does It Mean, Allegedly?

The problematic issue with the only if construction is that it is supposed to be logically equivalent to a number of related constructions, even though non-logicians sometimes disagree with this. According to the classical convention, the following sentences thus all mean the same:
  • It only thunders if it rains.
  • If it thunders, it rains.
  • If it doesn't rain, it doesn't thunder.
On the other hand, if we reverse the implication, we change the truth conditions:
  • It only rains if it thunders.
  • If it rains, it thunders.
  • It it doesn't thunder, it doesn't rain.
If this was just a mere convention about logical language, all would be fine. The problem is, however, that these sentence forms are not used in the same situations, and they do not integrate equally well into all reasoning patterns in spite of their (alleged) equivalence.

The Performance Problem

One difference between the If A, then B and A only if B forms is that if form is generally more difficult to use in a modus tollens inference than the only if. At least, this is what Carlos Santamaría and Orlando Espino say (Santamaría and Espino 2002, p. 42). They're referring to three studies, including one by Jonathan Evans and M. A. Beck (Evans and Beck 1981).

The problematic case is thus the following inference:
If it thunders, it rains.
It doesn't rain.
–––––––––––––––––
It doesn't thunder.
This (clasically valid) inference should be performed more readily when served in this alternative, and supposedly equivalent formulation:
It only thunders if it rains.
I doesn't rain.
–––––––––––––––––––––
It doesn't thunder.
Cognitively, or perhaps in terms of actual natrual language semantics, this seems to indicate that A only if B works more like the contrapositive If not B, then not A than like its positive translation, If A, then B. Or at least, it seems to issue a conversational warrant closer to it.

It would be interesting to know if this alternative formulation comes with a corresponding decrease—are we trading of willingness to perform the straightforward modus ponens inference for higher rates of modus tollens? This would imply that the following inference generally is less accepted:
It only thunders if it rains.
It thunders.
–––––––––––––––––––––
It rains.
If the only if formulation really does works like a contrapositive, then this inference should appear to us like a modus tollens inference in terms of plausibility and difficulty. I do not know right now whether such an effect can actually be measured or not.

The Issue of Time

Another interesting proposal that Santamaría and Espino cite, also coming from Evans and Beck, is that there is a systematic interaction between our conception of temporal order and the choice of form.

Thus, even though If A, then B, and A only if B are supposedly logically equivalent, we get different patterns of acceptability or naturalness depending on whether A or B happened first. For A preceding B, we then (perhaps) have:
  • If you bought on Tuesday, you're paying on Wednesday.
  • (?) You bought on Tuesday only if you paying on Wednesday.
And for B preceding A:
  • (?) If you're paying on Wednesday, you bought on Tuesday.
  • You're only paying on Wednesday if you bought on Tuesday.
Of course, much clearer intuitions can be produced if we ruffle up the tenses a bit. But this, I think, relatively fair example to start the discussion from.

So, Causality?

Note that the issue of before/after interfaces with the concept of causality, which is notoriously bound up with implication, even if logicians and statisticians hate to admit this fact.

Possibly, the the only way we can really justify an inference from a later effect to a prior cause in the form of If EFFECT, then CAUSE is to objectify the cause and the effect by thinking about the observation of the effect and the deduction of the cause. In this way, we would straighten out the temporal sequence so that EFFECT could in fact precede CAUSE.

If this is true, it has a quite important consequence for the psychology of reasoning: We would then only be able to understand abductive inference by effectivly embedding a cause/effect relationship in a different and larger cause/effect relationship—namely the only in which a real or imagined person reasons from fire (the logical "cause") to smoke (the logical "effect").

Tuesday, May 15, 2012

Chemla and Spector: "Experimental Evidence for Embedded Scalar Implicatures" (2010)

I remember Benjamin Spector seeing speaking about this experiment at ESSLLI 2010. The paper argues that the sentence
  • Every student solved some problems 
has a so-called "localist" reading. A counterexample to such a localist reading is a student who solved no questions, or a student who solved all questions. The second option is the crucial one, since a strict Gricean model does not predict any implicatures that warrant this reading. Chemla and Spector claim that the localist reading is available, although not the dominant one.

Experimental Set-Up and Subject Responses

The empirical method that they employ in order to prove this involves some sentences and some drawings. The drawings show six letters, each surrounded by six circles. Each letter is connected to some number of the circles surrounding it, as in the following figure:

 

The relevant sentences were then of the following kind:
  • Every letter is connected with some of its circles
In particular, the question was how subjects evaluate such sentences when no letters are completely disconnected, and at least one letter is connected every circle. Will anyone judge the sentence to be false in that scenario?

However, instead of just asking their subjects this question (a methodology that has previously falsified the theory) they asked subjects to click somewhere on a bar connecting the word "Yes" and the word "No." With this methodology, subjects did indeed pick points closer to "No" when the drawing included some letters that were connected to all of their circles.

So it seems that when pushed, subjects do start to doubt whether the word "some" should be taken to mean "at least one" or "at least one and not all." Once they notice this ambiguity, they might become slightly more scared of giving an unequivocal "Yes" and consequently click somewhere lower on the scale.

Or to put it differently, as soon as the subjects start fearing that they are in some kind of polemic language game, they switch to a safer strategy by committing to less. So a localist reading is indeed available, and in some circumstances a possible claim inherent in the sentence.

Relevance Considerations

However, what really caught my eye on this second rereading of the paper was the comments that Chemla and Spector make about relevance while discussing possible experimental set-ups:
the local reading (‘every square is connected with some of the circles and not with all of them’) is relevant typically in a context in which we are interested in knowing, for each square, whether it is connected with some, all, or no circle. Such a context would for instance result from raising the following question: ‘Which squares are connected to which circles?’. (Section 2.2.2, page 365)
The globalist reading, one might add, would be the most relevant answer to the question Which sqaures are connected to some circles? The only counterexamples to the globalist reading consist of sqaures that are not connected to anything, as stated above.

The reason I find this comment interesting is that it connects the meaning of the sentence with the expectation of the subjects. Experimental materials place the subject in some particular role, and this implicitly suggests certain answers to the big question: What does he expect me to do with this question?

Chemla and Spector obviously embrace some kind of objectivist perspective on semantics, with sentences having grammatically determined meanings, and a clear division between grammar and pragmatics. But their sensitivity to the perspective of the subject is very commendable and opens up the possibility of founding the notion of meaning of the notion of social context.