Showing posts with label British National Corpus. Show all posts
Showing posts with label British National Corpus. Show all posts

Tuesday, July 23, 2013

Rodd, Davis, and Johnsrude: "The Neural Mechanisms of Speech Comprehension" (2005)

This paper reports on two fMRI experiments which contrast ambiguous speech to unambiguous speech, and unambiguous speech to speech-like noise.

Blob One and Blob Two

By comparing pictures taken in each of these conditions, extracting significant differences, this design gives an indication of where in the head ambiguity is sorted out. The conclusion is that two areas in particular seemed to be disproportionally active when the experimental subjects listens to ambiguous speech:
The results of two fMRI experiments show that when volunteers listen to sentences that contain semantically ambiguous words, activity increases in both temporal and frontal brain regions. This confirms the involvement of these regions in the semantic aspects of sentence comprehension (i.e. activating, selecting or integrating word meanings). (p. 1266)
Roughly speaking, the areas in question were the bit of the brain behind the ears (on both sides of the head), and the bit behind the eyebrow (on the left side only).


Anatomical drawing of a brain from a 1918 textbook.
The parts of the brain discussed in the text are roughly located behind
the lower part of the temple and the lower left side of the forehead.


Decoding Efforts vs. Selection Efforts

The study did not include any distinctions more fine-grained than ambiguous/unambiguous. In particular, it did not contrasts skewed and balanced ambiguity; this is significant, since reading a word used in one of its less frequent meanings involves inhibition which may require cognitive effort.

Consider for instance the following example sentence from the paper:
  • the cymbals/symbols were making a racket/racquet
There are 67 results for cymbal(s) in the BNC; there are about 3000 for symbol(s). If I read this sentence aloud for you and took a picture of your brain while you listened, I would see a lot of activity in some regions; but this might be best interpreted as a trace of the force you exert in order to suppress the dominant but irrelevant meaning of the sound /ˈsɪmbəɫ/.

This process of suppressing a loud noise may or may not be a different from choosing between two competing alternatives, but we can't say on the basis of this experiment.

Thursday, June 20, 2013

Wilson and Gibbs: "Real and Imagined Body Movement Primes Metaphor Comprehension" (2007)

One of the recurring problems with experimental tests of cognitive metaphor theory is that it is exceedingly difficult to disentangle lexical priming from semantic priming. For instance, the notion of dragging is not only semantically related to "boredom" and "delay," but also discursively related to it: About 4 out of the 10 results for dragging out in the BNC are time-related metaphors.

In an attempt to circumvent this problem, Nicole L. Wilson and Ray Gibbs have performed two experiments in which a non-verbal movement served as the prime for a reading task. Specifically, they had their subjects learn to make certain movements like a stretching motion on cue, and then had them read small phrases like stretch for understanding. It turns out that performing these actions decreases reading time.

Are the Clichés Really Clichés?

The phrases they used were the following (p. 725):
  • Stamp out fear
  • Push the argument
  • Swallow your pride
  • Sniff out the truth
  • Spit out the facts
  • Shake off a feeling
  • Grasp a concept
  • Chew on an idea
  • Stretch for understanding
These phrases are not all equally standard. This is a bit problematic because Wilson and Gibbs explicitly use the data to argue against "phrasal lexicon" and "clichés or dead metaphors" accounts (p. 723). It would thus have been more convincing if they had used actual clichés instead of constructed and not completely natural phrases.

These differences can be quantified by counting co-occurrences. To do so, I've taken all the verb/noun pairs above and looked for cases in which they co-occur in the BNC.

For instance, I took all the forms of the verb shake (shake, shakes, shook, shaken) and paired them with all the forms of the noun feeling (feeling, feelings). I then checked whether any combination of a word form from the first list co-occurred with one from the second list up to 20 words apart, and in any order.

Compiling such counts gives the following table:

v
n
#(n)
#(v)
#(v, n)
P(n | v)
P(v | n)
grasp
concept
280
2485
8988
11.27%
3.12%
chew
idea
72
1116
31876
6.45%
0.23%
swallow
pride
112
2585
2913
4.33%
3.84%
sniff
truth
40
1165
8397
3.43%
0.48%
spit
fact
40
1371
41801
2.92%
0.10%
shake
feeling
228
9109
17559
2.50%
1.30%
stretch
understanding
40
6239
9552
0.64%
0.42%
push
argument
48
10703
12006
0.45%
0.40%
stamp
fear
8
3086
14578
0.26%
0.05%

So it turns out that we quite often grasp concepts, but we rarely if ever stamp out fear.

We should thus expect such a phrase to be experienced as much more "fresh," or alternatively, much more awkward. It would be interesting to check whether these statistics correlate in any way with the priming effect, but there's no way to do so directly, because Wilson and Gibbs do not report reading times for individual reading times in the experiment.

Thursday, March 14, 2013

Sereno, O'Donnell, and Rayner: "Eye Movements and Lexical Ambiguity Resolution" (2006)

In the literature on word comprehension, some studies have found that people usually take quite a long time looking at an ambiguous word if it occurs in a context that strongly favors one of its less frequent meanings.

This paper raises the issue of whether this is mainly because of clash between the high contextual fit and the low frequency, or mainly because of the frequency.

The Needle-in-a-Haystack Effect

A context preceding a word can either be neutral or biased, and a meaning of an ambiguous word can either be dominant (more frequent) or subordinate (less frequent). When a biased context favors the subordinate meaning, it is called a subordinate-biasing context.

The subordinate-bias effect is the phenomenon that people spend more time looking at an ambiguous word in a subordinate-biasing context than they take looking at an unambiguous word in the same context — given that the two words have the same frequency.

For instance, the word port can mean either "harbor" or "sweet wine," but the former is much more frequent than the latter. In this case, the subordinate-biasing effect is that people take longer to read the sentence
  • I decided to drink a glass of port
than the sentence
  • I decided to drink a glass of beer
This is true even though the words port and beer have almost equal frequencies (in the BNC, there are 3691 vs. 3179 occurrences of port vs. beer, respectively).

Balanced Meaning Frequencies = Balanced Reading Time

The question is whether these absolute word frequencies are the right thing to count, and Sereno, O'Donnell, and Rayner argue that they aren't. Instead, they suggest that it would be more fair to compare the sentence
  • I decided to drink a glass of port
to the sentence
  • I decided to drink a glass of rum
This is because port occurs in the meaning "sweet wine" approximately as often as the word rum occurs in absolute terms — i.e., much more rarely than beer. (A casual inspection of the frequencies of the phrases drink port/rum and a glass of port/rum seem to confirm the close match.)

What the Measurements Say

This means that you get three relevant conditions:
  1. one in which the target word is ambiguous, and in which its intended meaning is not the most frequent one;
  2. one in which the target word has the same absolute frequency as the ambiguous word;
  3. and one in which the target word has the same absolute frequency as the intended meaning of the ambiguous word.
Each of these are then associated with an average reading time:


It's not like the effect is overwhelming, but here's what you see: The easiest thing to read is a high-frequent word with only a single meaning (middle row); the most difficult thing to read is a low-frequent word with only a single meaning (top row).

Between these two things in terms of reading time, you find the ambiguous word whose meaning was consistent with the context, but whose absolute frequency was higher.

Why are Ambiguous Words Easier?

In the conclusion of the paper, Sereno, O'Donnell, and Rayner speculate a bit about the possible causes of this "reverse subordinate-biasing effect," but they don't seem to find an explanation they are happy about (p. 345).

It seems to me that one would have to look closer at the sentences to find the correct answer. For instance, consider the following incomplete sentence:
  • She spent hours organizing the information on the computer into a _________
If you had to bet, how much money would you put on table, paper, and graph, respectively? If you would put more money on table than on graph, that probably also means that you were already anticipating seeing the word table in its "figure" meaning when your eyes reached the blank in the end of the sentence.

If people in general have such informed expectations, then that would explain why they are faster at retrieving the correct meaning of the anticipated word than they are at comprehending an unexpected word. But checking whether this is in fact the case would require a more careful information-theoretic study of the materials used in the experiment.

Eviatar and Just: "Brain correlates of discourse processing" (2006)

This paper shows that three different kind of text snippets lead to three different patterns of brain activity. This is interpreted as showing how literal, metaphorical, and ironic language is processed.

However, a closer look at the experimental materials show that these labels should be approached with some caution. The literal statements are not all "literal" by the standards of cognitive metaphor theory, and the metaphorical statements are in many cases not as conventional as they are claimed to be.

Are the Literal Sentences Literal?

Here are some examples of text snippets that Eviatar and Just categorized as "literal," with underscores added by me:
  • Betsy and Mary were on the basketball team. Mary scored a lot of points in the game. Betsy said, “Mary is a great player.”
  • Harry waited in line for 3 h to see the movie. He enjoyed himself. He said, “That was worth waiting for.”
  • George promised to be quiet in the library. He sat in a corner looking at a book. His dad said, “Thanks for keeping your promise.”
  • Laura was out sick for a week. Johnny called her every day. Laura said, “Thanks for worrying about me.”
  • Betty and Laura were in the same class. Laura finished her homework before Betty. Laura said, “You sure are a slow worker.”
Several of these should not be categorized as literal according to the standards of cognitive metaphor theory.

For instance, great is typically taken as an example of a metaphor with size as its source domain. Similarly, worth, keeping, and perhaps worrying are here used in senses that are not their most "basic" ones. Further, slow worker should probably be categorized as a metonymy.

Are the Metaphorical Sentences Conventional?

Here are some examples that they consider to be "frozen" metaphors (p. 2350):
  • In the morning John came to work early. He started to work right away at a fast pace. His boss said, “John is a hurricane.”
  • Mary got straight A’s on her report card. Her parents were proud of her. They said, “You are as sharp as a razor.”
  • Susie helped her mom when her brother got sick. She took good care of him. Her mom said, “You are an angel from heaven.”
  • Donna was always late for everything. Today she made it home on time for supper. Her dad said, “You have turned over a new leaf.”
  • George went to Betty’s birthday party. Fifty people crowded into her small apartment. He said, “I feel like a sardine.”
  • Betty and Laura were in the same class. Laura finished her homework before Betty. Laura said, “You work like a snail.”
No doubt that these examples constitute metaphors; it's only that they are a very different kind of metaphors than unemployment is growing or I'll handle the press. One is very overt and almost cries out for attention, while the other is quiet and discreet.

Some Factors Causing the Metaphors to Grab Attention

We can hardly expect a sentence like
  • You are an angel from heaven
to be processed the same way as
  • The chocolate cake was divine
The resonance of angel with heaven, the quotation marks in the original story, and the syntactic clumsiness of the sentences all contribute to provoke a very vivid mental picture in the reader. This would probably not be the case for the divine chocolate cake.

Other sentences from the materials provoke such images because they are not really conventional, or occur in an abnormal form. For instance, the common "sardine" metaphor virtually always occurs in the plural form like sardines and never in the form like a sardine.

Lastly, because the subject-predicate form in sentences like You are X or John is X is so semantically weak in its subject, it draws a lot of attention to the metaphor which is its topic. Compare for instance
  • He was a tallish man with a mind as sharp as a razor. (BNC)
  • You are a razor.
To my intuitions, the razor in the first sentence seems to recede very much into the background, while in the second sentence, it is being put forward as the explicit topic of the sentence. The reasons are probably both syntactic (information is spread unevenly over constituents) and semantic (the first sentence contains more competing content words).

Are the Metaphors Metaphorical?

One last thing that I should mention is that at least one of the "metaphorical" examples could be interpreted as literal:
  • Ken was worried about having his hair cut. When the barber finished, Ken’s ears stuck out. He said, “You’ve turned me into a clown.”
Without a more specific theory of what "metaphor" means, this is a very problematic borderline case.

So What?

This doesn't mean that the data that Eviatar and Just has collected is useless, or that it should be discarded. But it does mean that, once again, the notion of "metaphor" is so vague and contested that it can't just be transplanted from one field to another without some problems.

In particular, linguists should be careful not to take evidence from brain scans at face value; we need to look carefully into the details of what the neuroscience actually shows, and be explicit about what our own semantic theory actually say.

In this case, that means the following: First, some sentences provoke mental imagery more strongly than others; this is a product of several interacting factors, including word frequencies and resonance effects. And second, it is not at all clear what the relation between mental imagery and meaning is, and this relation cannot be made clear unless we come clean about what we thing linguistic meaning is.

Saturday, September 24, 2011

Mike Thelwall: "Fk yea I swear" (2008)

This is a corpus-based study on swearing on UK Myspace profiles. From my perspective, the article is mostly interesting because it contains some valuable statistics on the uses of swear words like fuck, cunt, twat, and shit.

It was published in the journal Corpora, but a preprint is available Mike Thelwall's website.

Metaphors with taboo source domains
As one part of the study, Thelwall and a helper cateogized 427 swear words from their custom-tailored corpus of Myspace comments and profiles.

They used a category scheme borrowed from a similar study on the British National Corpus. It includes categories like "Predicative negative adjective," "Emphatic adjective," etc. This taxonomy was proposed by Tom McEnery and Richard Xiao in "Swearing in modern British English" (2004).

Thelwall and his helper found, out of the 427 cases, not a single case of metaphorical use of a swear word. Thus, fuck was never used in a sense that extended its sexual meaning, as in, I suppose, I'm going to take this delicious cake back to my room and fuck it. It's even hard to force such a metaphorical reading on this sentence.

Literal use of taboo terms
He did find some literal (sexual, religious, etc.) uses of some swear words, but they only constituted 3% of all cases.

Using my definition of "literal," I would probably have to categorize the emphatic use of curse words as literal, since that was their most frequent use in the corpus. This boils down to saying that when you process a phrase like fucking tired, you do not retrieve or need to retrieve the sexual meaning of fuck.

A word like bloody (used emphatically) might have a slightly higher tendency to evoke the "blood-stained" meaning, since it is less common as an emphatic adjective relative to its "literal" meaning.