Showing posts with label corpus linguistics. Show all posts
Showing posts with label corpus linguistics. Show all posts

Wednesday, May 28, 2014

Herdan: The Advanced Theory of Language as Choice and Chance (1966)

This really odd book is based to some extent on the earlier Language as Choice and Chance (1956), but contains, according to the author, a lot of new material. It discusses a variety of topics and language statistics, often in a rather unsystematic and bizarrely directionless way.

The part about "linguistics duality" (Part IV) is particularly confusion and strange, and I'm not sure quite what to make of it. Herdan seems to want to make some great cosmic connection between quantum physics, natural language semantics, and propositional logic.

But leaving that aside, I really wanted to quote it because he so clearly expresses the deterministic philosophy of language — that speech isn't a random phenomenon, but rather has deep roots in free will.

"Human Willfulness"

He thus explains that language has one region which is outside the control of the speaker, but that once we are familiar with this set of restrictions, we can express ourselves within its bounds:
It leaves the individual free to exercise his choice in the remaining features of language, and insofar language is free. The determination of the extent to which the speaker is bound by the linguistic code he uses, and conversely, the extent to which he is free, and can be original, this is the essence of what I call quantitative linguistics. (pp. 5–6)
Similarly, he comments on a study comparing the letter distribution in two texts as follows:
There can be no doubt about the possibility of two distributions of this kind, in any language, being significantly different, if only for the simple reason that the laws of language are always subject, to some extent at least, to human willfulness or choice. By deliberately using words in one text, which happen to be of rather singular morphological structure, it may well be possible to achieve a significant difference of letter frequencies in the two texts. (p. 59)
And finally, in the chapter on the statistical analysis of style, he goes on record completely:
The deterministic view of language regards language as the deliberate choice of such linguistic units as are required for expressing the idea one has in mind. This may be said to be a definition in accordance with current views. Introspective analysis of linguistic expression would seem to show it is a deterministic process, no part of which is left to chance. A possible exception seems to be that we often use a word or expression because it 'happened' to come into our mind. But the fact that memory may have features of accidental happenings does not mean that our use of linguistic forms is of the same character. Supposing a word just happened to come into our mind while we are in the process of writing, we are still free to use it or not, and we shall do one or the other according to what is needed for expressing what we have in mind. It would seem that the cause and effect principle of physical nature has its parallel in the 'reason and consequence' or 'motive and action' principle of psychological nature, part of which is the linguistic process of giving expression to thought. Our motive for using a particular expression is that it is suited better than any other we could think of for expressing our thought. (p. 70)
He then goes on to say that "style" is a matter of self-imposed constraints, so that we are in fact a bit less free to choose our words that it appears, although we ourselves come up with the restrictions.

"Grammar Load"

In Chapter 3.3, he expands these ideas about determinism and free will in language by suggesting that we can quantify the weight of the grammatical restrictions of a language in terms of a "grammar load" statistic. He suggest that this grammar load can be assessed by counting the number of word forms per token in a sample of text from the language (p. 49). He does discuss entropy later in the book (Part III(C)), but doesn't make the connection to redundancy here.

Inspects an English corpus of 78,633 tokens, he thus finds 53,102 different forms and concludes that English has a "grammar load" of
53,102 / 78,633 = 67.53%.
Implicit in this computation is the idea that the number of new word forms grows linearly with the number of tokens you inspect. This excludes sublinear spawn rates such as
In fact, vocabulary size does seem to grow sublinearly as a function of corpus size. The spawn rate for new word forms in thus not constant in English text.

Word forms in the first N tokens in the brown corpus for N = 0 … 100,000.

However, neither of the growth functions listed above fit the data very well (as far as I can see).

Thursday, June 20, 2013

Wilson and Gibbs: "Real and Imagined Body Movement Primes Metaphor Comprehension" (2007)

One of the recurring problems with experimental tests of cognitive metaphor theory is that it is exceedingly difficult to disentangle lexical priming from semantic priming. For instance, the notion of dragging is not only semantically related to "boredom" and "delay," but also discursively related to it: About 4 out of the 10 results for dragging out in the BNC are time-related metaphors.

In an attempt to circumvent this problem, Nicole L. Wilson and Ray Gibbs have performed two experiments in which a non-verbal movement served as the prime for a reading task. Specifically, they had their subjects learn to make certain movements like a stretching motion on cue, and then had them read small phrases like stretch for understanding. It turns out that performing these actions decreases reading time.

Are the Clichés Really Clichés?

The phrases they used were the following (p. 725):
  • Stamp out fear
  • Push the argument
  • Swallow your pride
  • Sniff out the truth
  • Spit out the facts
  • Shake off a feeling
  • Grasp a concept
  • Chew on an idea
  • Stretch for understanding
These phrases are not all equally standard. This is a bit problematic because Wilson and Gibbs explicitly use the data to argue against "phrasal lexicon" and "clichés or dead metaphors" accounts (p. 723). It would thus have been more convincing if they had used actual clichés instead of constructed and not completely natural phrases.

These differences can be quantified by counting co-occurrences. To do so, I've taken all the verb/noun pairs above and looked for cases in which they co-occur in the BNC.

For instance, I took all the forms of the verb shake (shake, shakes, shook, shaken) and paired them with all the forms of the noun feeling (feeling, feelings). I then checked whether any combination of a word form from the first list co-occurred with one from the second list up to 20 words apart, and in any order.

Compiling such counts gives the following table:

v
n
#(n)
#(v)
#(v, n)
P(n | v)
P(v | n)
grasp
concept
280
2485
8988
11.27%
3.12%
chew
idea
72
1116
31876
6.45%
0.23%
swallow
pride
112
2585
2913
4.33%
3.84%
sniff
truth
40
1165
8397
3.43%
0.48%
spit
fact
40
1371
41801
2.92%
0.10%
shake
feeling
228
9109
17559
2.50%
1.30%
stretch
understanding
40
6239
9552
0.64%
0.42%
push
argument
48
10703
12006
0.45%
0.40%
stamp
fear
8
3086
14578
0.26%
0.05%

So it turns out that we quite often grasp concepts, but we rarely if ever stamp out fear.

We should thus expect such a phrase to be experienced as much more "fresh," or alternatively, much more awkward. It would be interesting to check whether these statistics correlate in any way with the priming effect, but there's no way to do so directly, because Wilson and Gibbs do not report reading times for individual reading times in the experiment.

Thursday, February 14, 2013

Hanks: "Metaphoricity is gradable" (2006)

This is Patrick Hanks' contribution to Corpus-Based Approaches to Metaphor and Metonymy (2006). As the title says, its claim is that expressions can be more or less metaphorical.

Hanks grades a number of expressions involving the word sea according to how metaphorical he finds them, but it is unclear how to generalize his methods. In the course of the text, he gives at least two different answers to the question of how one might measure the metaphoricity of a phrase.

The Frequency Measure of Metaphoricity

First, he suggest the absolute frequency of the phrase as a measure (p. 21). An area of research is thus perceived as relatively literal because the string as a whole is quite frequent. By contrast, an oasis of calm is quite metaphorical because infrequent.

Hanks may have a point here, but it's difficult to tell, because he never submits any evidence that the statistic of interest is the absolute frequencies rather than the relative ones. It is true that area of research only accounts for about 0.1% of the 34,824 occurrences of area in the BNC, while the 6 occurrences of oasis of calm add up to 2.4% of the 248 occurrences of oasis.

But on the other hand, these numbers are seriously sensitive to paraphrase: if you include study and search into the analysis,  the tally jumps from 43 to 168, i.e., a near quadrupling. For the oasis of calm, expanding the search to include silence, peace, quiet, tranquillity, repose, and serenity only yields 9 more hits, i.e. a growth of about 50%. It is thus not at all clear which of those numbers we should trust more, and which will be responsible for how we think.


study work research research concern search analysis enquiry
area of … 87 61 43 43 40 36 5 4


calm peace tranquility serenity silence quiet repose
oasis of … 6 3 2 2 1 1 1

The Resonance Measure of Metaphoricity

Second, Hanks suggests that the degree of metaphoricity is determined by certain quality called "resonance," a concept he has picked up from Max Black (pp. 20, 22). Hanks writes:
In the most metaphorical cases, the the secondary subject [source domain] shares the fewest properties with the primary subject [= target domain]. Therefore, the reader or hearer has to work correspondingly harder to create a relevant interpretation. At the other extreme, the more shared properties there are, the weaker the metaphoricity. (p. 22)
The term "resonance" subsequently seems to be used as the inverse of "semantic distance" (p. 22), and a similarity-based prototype theory is taken for granted. Accordingly, he finds that an Antarctic oasis—an ice-free area on Antarctica—is slightly metaphorical because it deviates from the stereotypical cartoon oasis with palm trees etc. (p. 30–31).

Meaning and Personal Experience

I thought I should just cite the summary of cognitive metaphor theory that Hanks gives, because it succinctly sums up some of the concerns that a number of people have had about it:
Lakoff and Johnson's basic thesis about metaphor is that its function is to enable us to interpret concepts (especially abstract concepts) in terms of familiar, everyday cognitive experiences. This is broadly satisfactory, though we might be tempted to substitute 'perceptual experiences' for 'cognitive experiences', and common sense forces us to acknowledge that the 'everyday experience' in question is that of the language community at large, not each individual. (p. 19–20)
In fact, I don't think a metaphor needs to be grounded in anybody's first-hand experience in the general case; and metaphors that involve dying are good illustrations of this phenomenon.

Friday, September 14, 2012

Barbara Dancygier: "Mental space embeddings, counterfactuality, and the use of unless" (2002)

According to many native English speakers, the word unless has a tendency to work less well in hypothetical contexts:
  • You're not clinically depressed unless you're tired.
  • ?You wouldn't be clinically depressed unless you were tired.
In Barbara Dancygier's paper on the use of unless, she tries to explain this fact by means of some rather extravagant cognitive assumptions about people's use of hypothetical mental spaces. This explanation mainly amounts to describing the different distributions of unless and except if.

Unacceptability: Counterexamples

Although unless is indeed not appropriate in many hypothetical constructions, there are couterexamples. Dancygier gives a number of quite nice corpus examples, including
  • 'Unless I was naked, you'd still worry I was wearing a gun or a wire.' (p. 365)
  • 'I have pulled my tail off,' replied the younger Mouse, 'but as I should still be on the sorcerer's table unless I had, I do not regret it.' (p. 368)
  • [I]f Miss Catherine had the misfortune to marry him, he would not be beyond her control, unless she were extremely and foolishly indulgent (p. 369)
So apparently, some hypothetical uses are OK, whatever the reason is.

Some Acceptability Judgments

Zooming a bit out from this observation, we have, according to Dancygier's intuitions, the following pattern (p. 369; I've abbreviated the sentences a bit):
  • If she married him, he would be under control unless she were extremely indulgent.
  • If she married him, he would be under control except if she were extremely indulgent.
Further, we have the following distribution in the mouse example, again according to Dancygier's intuitions (p. 372):
  • As I should be on the table unless I had pulled my tail off, I do not regret it.
  • *As I should be on the table *except if I had pulled my tail off, I do not regret it.
I suppose it's fair to assume the following acceptability judgments as well, even though Dancygier does not explicitly say so:
  • *I wouldn't have finished unless you had helped me.
  • *I wouldn't have finished except if you had helped me.
If these three cases are representative, then unless is more lax than except if: Whenever except if fits in a frame, unless does, too. Whatever story we tell about these distributions, it thus better be one in which except if requires some kind of higher standard of well-formedness.


Acceptability Patterns: A Possible Explanation

Dancygier's theory has something to do with contrast.

She proposes that except if needs to introduce a condition that stands in contrast with the whole hypothetical context (I think?) all the way up to the actual situation. Unless, on the other hand, only needs to introduce a condition that stands in a contrastive relationship with the immediate intensional context.

What this means is not quite clear, and it requires some further (and quite strong) assumptions about the introduction of new layers of intensional context. Such layers are introduced quite often in Dancygier's theory; not only modal verbs like would and should trigger them, but also future tenses and conditional items like if.

Beyond Control

So, and example: Let's look at the marriage example and show how this example allegedly builds up layers of context according to Dancygier's analysis. The relevant embeddings are:


This should be read as follows: In the base layer, Miss Christine and the male character are not married. However, in the first conditional scenario (if…), they are imagined to be, and from this hypothetical situation (would…), it is judged that "he would not be beyond her control."

This imaginary situation is then equipped with an exceptional condition (unless…), namely that "she were extremely and foolishly indulgent." This condition is taken to imply to lack of control in a future scenario (not lexically expressed). (Cf. pp. 369–70.)

Pancake Upon Pancake

Another way to visualize the same thing is by keeping track of when different assumptions are introduced, inherited, and negated in the embedded structure:
Not marriage (assumed)
(No assumption about control)
(No assumption about indulgence)
Hypothetical scenario:
        Marriage (overridden)
        (No assumption about control)
        (No assumption about indulgence)
        Future scenario:
                Marriage (inherited)
                Control (assumed)
                (No assumption about indulgence)
                Exceptional scenario:
                        Marriage (inherited)
                        Control (inherited)
                        Indulgence (assumed)
                        Future scenario:
                                Indulgent (inherited)
                                No control (overridden)
                                Marriage (inherited)
The important grammatical fact is here which of its ancestor layers the exceptional condition (indulgence) is inconsistent with. The ancestor layers are here the base layer, the hypothetical scenario, and the future scenario.

Since the indulgence is neither assumed nor denied in any of these layers, the exceptional assumption stands in a relationship of contrasts with all of them. Both unless and except if are thus grammatical in this context.

Note that this conclusion crucially depends on the fact that the base layer assumed to be undecided about control. If we suppose that Miss Christine has no control over the male character in the base layer, except if should be ungrammatical

Tailless Escapes Captivity

This analysis should be compared to the example I should be on the table unless I had pulled my tail off (cf. p. 371):
Pulled tail (assumed)
Not on table (assumed)
Hypothetical scenario:
        (Assumptions about the tail deleted)
        On table (overridden)
        Exceptional scenario:
                Pulled tail (assumed)
                On table (inherited)
                Future consequence:
                        Pulled tail (inherited)
                        Not on table (overridden)
In this case, the exceptional condition (I pulled my tail off) is in contrast with the immediately preceding scenario, which has no assumptions about the tail. The exception is, however, not in contrast with the base layer. Consequently, except if is ungrammatical, and unless is grammatical.

But note again the assumptions going into this conclusion: Without the assumption of "forgetfulness" in the hypothetical scenario, the conclusion would not follow. It is thus crucial that the hypothetical If I were still on the table… "deletes" a particular one of our assumptions in a seemingly rather arbitrarily fashion.


And All the Other Cases…?

Let's look at the case I wouldn't have finished, which I assumed above was ungrammatical with both unless and except if. This sentence should presumably be analyzed as follows:
I have finished (assumed)
You helped me (assumed)
Hypothetical scenario:
        I have not finished (overridden)
        You helped me (inherited)
        Exceptional scenario:
                I have not finished (inherited)
                You did not help me (overridden)
                Future consequence:
                        I have not finished (inherited)
                        You did not help me (inherited?)
In this case, the conditional exception (you did not help me) contrasts with both of the two layers above. So why not a grammatical use of unless and except if?

I think questions like these point to the fact that the cognitive assumptions behind Dancygier's theory are pretty vague and very, very speculative. It is by no means clear when assumptions are "forgotten," or exactly which or how many assumption that have to contrast with the intensional context.

In fact, it is not even quite clear to me whether except if requires contrast to all ancestor levels, and whether it requires contrast on a particular, high-focus parameter, or just any single parameter. Without an answer to these questions, it will be very difficult to assess the theory to any interesting degree.

Thursday, April 12, 2012

Beate Hampe: "When down is not bad, and up is not good enough" (2005)

This thoughtful paper by Beate Hampe (see also From Perception to Meaning, 2005) is a nice empirical counterweight to some of the wildly speculative claims that are thrown around in cognitive linguistics. In this particular case, the issue is whether there is an inherent and global good/bad valence to the dichotomies up/down, in/out, on/off, front/back, etc.

The Theory
This claim has apparently been made most clearly by the Polish linguistic Tomasz P. Krzeszowski. Krzeszowski himself sees the claim as a specific version of the "Invariance Principle," since it claims that up/down metaphors inherit the good/bad valencies that standing and lying down have in our preconceptual lives.

This is a neat little fairytale about the genesis of meaning, but as usual, the empirical data spoils everything. The first, obvious sign comes from looking at bad things going up:
  • Unemployment is up. (bad)
  • Employment is up. (good)
Faced with such examples, one would have to say something like this: up and down have an inherent emotional value, but this emotional value is not very strong itself; a strongly laden context can thus pull the words in an "unnatural" direction. However, on average or in neutral contexts, up will have a weak tendency to lean towards positive emotional valence, and down a tendency towards negative.

The Counterevidence
However, this doesn't seem to be the case. Hampe has done a medium-sized corpus study of the constructions finish off, finish up, slow down and (the rare) slow up. For each occurrence, she looked for valence clues in the immediate context and categorized the example according to how positive it was. The result was, surprisingly, that up was much more likely to be used for negative purposes than down or off.

Now, it is of course a little unfortunate that she only looked at two verbs, and that one of her particle pairs (down/off) were not antonyms. A safer strategy would be to pick two antonym verbs and two antonym particles and then combine them in a table like the following:


in out
give 28.0 kk 27.3 kk
take 53.5 kk 118.0 kk

The numbers in this table are the number of occurrences (in millions), estimated by Google searches for the exact phrases. This excludes, for instance, split uses like take me out (or in general verb + NP + particle). However, Hampe's study seems to have the same weakness.

Similar tables could be made for the following:
  • come/go in/out
  • come/leave in/out
  • push/pull on/off
  • break/make up/down
  • break/fix up/down
  • etc.
Other Sources of Evidence
Several empirical hypotheses could be tested for such data.

For instance, one test whether there is a statistical tendency for some positive particle (e.g., in) to attach to the positive set of verbs (come, make, stand, keep) compared to its negative counterpart (out). This could be done with respect to grammaticality or with respect to empirical counts. It would essentially amount to collapsing all of the tables into a single 2 × 2 table.

Another question would be whether the positive and negative contexts where distributed evenly across the two columns of any such table. In order to do so, one would have to develop a valence assessment method like Hampe's, preferably an automatic one. After having trained such a model, one could use Fisher's exact test on the contingency table consisting of positive/negative valence × positive/negative column.

The automatic valence assessment might perhaps be achieved through semi-supervised learning. We can imagine starting from a valence function v0 defined on a small set of good and bad words:
  • v0(w) = 1 for w in some finite set G = {good, pleasant, improve, victory, truth, ...};
  • v0(w) = –1 for any w that is an antonym to a word in G;
  •  v0(w) = 0 for all other words w.
Assuming that good words co-occur with good words, this could be used to train a more fine-grained function, say, v1000.There are a number of problems with this assumption (it ignores rhetorical contrast effects as well as negation), but experiments would show whether it worked or not.

Monday, November 7, 2011

Sardinha: "Metaphor probabilities in corpora" (2008)

In his contribution to Confronting Metaphor in Use, Tony Sardinha argues that metaphor researchers should care more about the probabilities that a given word will be used metaphorically in a given corpus genre.

To illustrate his ideas, he reports a large number of metaphor probabilities taken from a highly specialized corpus of Brazilian Portuguese.

The article cites three interesting sources that are quite alien to metaphor theory proper:

Friday, October 7, 2011

Monday, September 26, 2011

Caroline Gevaert: "The ANGER IS HEAT question" (2005)

In defense of Geeraerts and Gondelaers' (1995) hypothesis about the historical origin of the ANGER IS HEAT metaphor, Caroline Gevaert examines some Old English corpus material and concludes that the metaphor was indeed in all likelihood inherited from Latin scholarship.

As a part of her argument, she acknowledges that test subjects' body heat increases slightly (about .1 degree celsius) when they make an angry face, apparently supporting a physiological basis for ANGER IS HEAT. However, she then comments that this temperature increase also occurs when subjects mimic other emotions, suggesting that I was boiling with sadness should be equally natural (p. 197).

She also notes that the evidence from Chinese actually seems to contradict the theory, referring only to chilies (and, presumably, their red color) and not to heat or flames (p. 196). The Native American language Chikasaw similarly does not seem to support the hypothesis.

Saturday, September 24, 2011

Mike Thelwall: "Fk yea I swear" (2008)

This is a corpus-based study on swearing on UK Myspace profiles. From my perspective, the article is mostly interesting because it contains some valuable statistics on the uses of swear words like fuck, cunt, twat, and shit.

It was published in the journal Corpora, but a preprint is available Mike Thelwall's website.

Metaphors with taboo source domains
As one part of the study, Thelwall and a helper cateogized 427 swear words from their custom-tailored corpus of Myspace comments and profiles.

They used a category scheme borrowed from a similar study on the British National Corpus. It includes categories like "Predicative negative adjective," "Emphatic adjective," etc. This taxonomy was proposed by Tom McEnery and Richard Xiao in "Swearing in modern British English" (2004).

Thelwall and his helper found, out of the 427 cases, not a single case of metaphorical use of a swear word. Thus, fuck was never used in a sense that extended its sexual meaning, as in, I suppose, I'm going to take this delicious cake back to my room and fuck it. It's even hard to force such a metaphorical reading on this sentence.

Literal use of taboo terms
He did find some literal (sexual, religious, etc.) uses of some swear words, but they only constituted 3% of all cases.

Using my definition of "literal," I would probably have to categorize the emphatic use of curse words as literal, since that was their most frequent use in the corpus. This boils down to saying that when you process a phrase like fucking tired, you do not retrieve or need to retrieve the sexual meaning of fuck.

A word like bloody (used emphatically) might have a slightly higher tendency to evoke the "blood-stained" meaning, since it is less common as an emphatic adjective relative to its "literal" meaning.

Tuesday, September 13, 2011

Steen et al.: "Metaphor in Usage" (2010)

This is a report of a project done at the VU here in Amsterdam under the supervision of Gerard J. Steen. It's an effort to tag the tokens in the British National Corpus as metaphorical or not metaphorical.

Section 1.1 of the paper includes some very handy references to various work critical of cognitive metaphor theory from the perspective of psychology (p. 766), comparative linguistics (p. 767), and linguistics in general (p. 767).

The methodology behind the tagging procedure requires that the human annotators (six PhD students) compare the meaning of a disambiguated word to any "more basic contemporary meaning" of that word (p. 769).

They write that basic meanings "tend to be" characterized by being
– more concrete; what they evoke is easier to imagine, see, hear, feel,
smell, and taste.
– related to bodily action.
– more precise (as opposed to vague).
– historically older.  (p. 769)
This seems to involve a certain amount confusion of criteria. They even state immediately after:
Basic meanings are not necessarily the most frequent meanings of the lexical unit. (p. 769)
This claim, however, comes without quantitative evidence (or footnote).