Showing posts with label basic level categories. Show all posts
Showing posts with label basic level categories. Show all posts

Thursday, June 14, 2012

Rosch: "Principles of Categorization" (1978)

Elanor Rosch's contribution to Cognition and Categorization really consists of two independent parts: An overview over her experiments investigating basic-level categories (pp. 30-35), and an overview of her experiments with prototype effects (pp. 35-41). In will only deal with the first part now.

Cue Validity

The most central concept in Rosch's discussion of basic-level categories is the notion of the cue validity. This is defined for a category such as "bird," which is more or less reliably identified by cues such as "wings." She explains:
The cue validity of an entire category may be defined as the summation of the cue validities for that category of each of the attributes of the category. (pp. 30-31)
This immediately raises two questions:
  1. Do all cues count in the summation with equal weight? There are infinitely many possible cues and only a few highly valid ones. This suggests that more explicit assumptions about "salience" are needed.
  2. With what weight do the various members of a category contribute to the average? Equally? Weighted by the frequency of the linguistic label? Weighted by the frequency of the thing?
While these questions may seem like technical remarks, they do in fact relate to some deeper issues that I will mention below.

The Ambiguity of "Basic"

There are two competing characterizations of "basic" in Rosch's work, an ostensive and a perceptual. It's not always clear which one she is taking as definitive, and this sometimes introduces problems.

Both definitions apply to concept trees and are meant to pick out a particular depth in such a tree. They do so by locating the level of abstraction at which either
  1. the categories "car," "chair," "tomato," and "hammer" are found; or
  2. average cue validity is maximized.
My worry is that her cross-cultural, developmental, and evolutionary claims may turn out to be tautologies when we look closer at the ups and downs of her theory.

For instance, if the "basic" means "maximal cue validity," then of course children learn names from this level first. On the other hand, if Rosch gets to pick what counts as "basic" in each branch of the English category system ("chair," "car," "tomato," ...), then she can obviously just pick the level that fulfills the second definition.

Learned Perception

The fact that she might unknowingly be making the tautological point that "normal things are normal" is hinted at when she comments that English-speakers tend to be less able to distinguish between plants than the ostensive definition suggests.

This is observation was echoed more recently my Jerome Feldman:
For many city dwellers, tree is a basic category—we interact the same way with all trees. But for the professional gardener, tree is definitely a superordinate category (Feldman 2006: 186)
With those kinds of qualifications, basic level categories will certainly guaranteed to have all of the properties that Rosch claims. But any claim about their universal centrality will also become an empty verbalism.

Note how this also ties in with the sticky issue of trained perception:
One influence on how attributes will be defined by humans is clearly the category system already existent in the culture at a given time. This our segmentation of a bird's body such that there is an attribute called "wings" may be influenced not by perceptual factors [...] but also by the fact that at present we already have a cultural and linguistics category called "birds." (p. 29)
She is apparently aware of this problem, but not willing to face the implication that complex cues like plumage are themselves categories that are open-ended and ambiguous.

Mutual Dependence and Iterated Learning

She does note, however, that attributes might be extracted from categories just as well as categories might be based on attributes. However:
Unfortunately, to state the matter in such a way is to provide no clear place at which we can enter the system as analytical scientists. What is the unit with which to start our analysis? (p. 42)
To me, this suggests a game-theoretical analysis. A category system is invented by people, but also has to be transmitted; fixed points in such an iterated learning process will be the systems that trade off difficulty of acquisition for pragmatic necessity, I guess.

This process could probably be modeled relatively easily in a multi-agent system with a set of Bayesian learners. However, such a model will probably be highly sensitive to the assumptions made about the environment of learning (e.g., the frequency of birds and the frequency of winged-ness).

Tuesday, June 12, 2012

Berlin: "Ethnobiological Classification" (1978)

Brent Berlin considers data from two "prescientific" cultures and concludes that their category systems are based on appearance above the basic level, and based on utility below.

Since appearances "cry out to be named" (p. 11) not all plant names will reflect practical concerns:
This finding would seem to controvert the view that preliterate man names and classifies only those organisms in the environment that have some immediate functional significance for survival. More than one-third of the named plants in both Tzeltal and Aguaruna, for example, lack any cultural utility, and these are not pestiferous plants that must be avoided due to poisonous properties or the like. (p. 11)
A couple of times in the paper (e.g., p. 20), he raises the issue of simple versus compound names for categories. He seems to think that the basic level should, normally and on average, be the lowest level that has simple names (tree, pine, etc.), but he doesn't discuss the topic specifically and in detail, perhaps he didn't have enough quantitative data for a meaningful claim.

His conclusion is that
folk biological classification is based on a recognition of natural discontinuities in the biological world that are considered to be similar or different because of gross, readily perceivable characteristics of form and behavior. (p. 24)

Tenenbaum and Xu: "Word Learning as Bayesian Inference" (2007)

Joshua Tenenbaum and Fei Xu report some experimental findings with concept learning and simulate them in a computational model based on Bayesian inference. It refers to a 1999 paper by Tenenbaum for mathematical background.

Elements of the Model

The idea behind the model is that the idealized learner picks a hypothesis (a concept extension, a set of objects) based on a finite set of examples. In their experiments, the training sets always consist of either one or three examples. There are 45 objects in the "world" in which the learner lives: some vegetables, some cars, and some dogs.

As far as I understand, the prior probabilities fed into the computational model were based on human similarity judgments. This is quite problematic, as similarity can reasonably be seen as a dual of categories (with being-similar corresponding to being-in-the-same-category). So if I've gotten this right, then the answer is to some extent already built into the question.

Variations

A number of tweaks are further applied to the model:
  • The priors of the "basic-level" concepts (dog, car, and vegetable) can be manually increased to introduce a bias towards this level. This increases the fit immensely.
  • The priors of groups with high internal similarity (relative to the nearest neighbor) can be increased to introduce a bias towards coherent and separated categories. Tenenbaum and Xu call this the "size principle."
  • Applying the learned posteriors, the learner can either use a weighted average of probabilities, using the model posteriors as weights, or simply pick the most likely model and forget about the rest. The latter corresponds to crisp rule-learning, and it gives suboptimal results in the one-example cases.
I still have some methodological problems with the idea of a "basic level" in out conceptual system. Here as elsewhere, I find it question-begging to assume a bias towards this level of categorization.

Questions

I wonder how the model could be changed so as to
  • not have concept learning rely on preexisting similarity judgments;
  • take into account that similarity judgments vary with context.
Imagine a model that picked the dimensions of difference that were most likely to matter given a finite set of examples. Dimension of difference are hierarchically ordered (e.g., European > Western European > Scandinavian), so it seems likely that something like the size principle could govern this learning method.

Tuesday, October 4, 2011

From Molecule to Metaphor (2006), chapter 10

Here's a strange quote from Feldman's book:
Even as adults, the experience we associate with a word and thus its meaning differs depending on our age, gender, profession, and so on. People who only watch a sport event or artistic performance cannot fully understand participants' conversation about the activity. (p. 130)
It's difficult to disagree with the first sentence here, and we should probably applaud that a cognitive linguist acknowledges this fact.

But it's almost equally difficult to agree with the second sentence or see any logical relation between them. Where does this extreme solipsism come from?

Perhaps this kind of thinking is the very core of the problem in cognitive metaphor theory. If everything become a matter of private experience, then our evident ability to understand each other seems like a miracle. Feldman needs a shot of Wittgenstein.

Basic Level Nonsense
Speaking of Wittgenstein, I'm still shocked at the sheer amount of nonsense associated with the term "basic level category." Elanor Rosch (in all due respect) seems to have assumed without the slightest shed of argument that all concepts come in triples like vehicle > car > truck.

That's a bizarrely unfounded claim, conceptually and empirically. What would, for instance, be the "basic level" in the following strings of inclusions?
entity  >  ( . . . )  >  artifact  >  toy  >  doll  >  puppet  >  hand puppet
As far as I can see, the only way to pin these levels onto an absolute scale with a "ground floor" would be to randomly pick one by intuition. This would essentially fold the data into the definition and empty the theory completely of any meaning.