Kisseberth & Zonneveld on exceptionality

[This post is part of a series on theories of lexical exceptionality.]

In a paper entitled simply “The treatment of exceptions”, Kisseberth (1970) proposes an interesting revision to the theory of exceptionality. Many readers may be familiar with the summary of this work given by Kenstowicz & Kisseberth 1977:§2.3 (henceforth K&K). Others may know it from the critique by Zonneveld (1978: ch. 3) or Zonneveld’s (1979) review of K&K’s book. I will discuss all of these in this post.

Kisseberth (1970)

A quick sidebar: Kisseberth’s paper is a fascinating scholarly artifact in that it probably could not be published in its current form today. (To be fair it was published in an otherwise-obscure journal, Papers in Linguistics.) For one, all the data is drawn from Matteson’s (1965) grammar of Piro;¹ the only other referenced work is SPE. Kisseberth (henceforth K) gives no page, section, or example numbers for the forms he cites. I have tried to track down some of the examples in Matteson’s book, and it is extremely difficult to find them. K gives no derivations, only a few URs, and the entire study hinges around a single rule. But it’s provocative stuff all the same.

K observes that Piro has a rule which syncopates certain stem-final vowels. He gives the following formulation:

(1) Vowel Drop:  V -> ∅ / VC __ + CV 

For example, [kama] ‘to make, form’ has a nominalization [kamlu] ‘handicraft’ with nominalizing suffix /-lu/, and [xipalu] ‘sweet potato’ has a possessed form /n-xipa-lu-ne/ [nxipalne] ‘my sweet potato’.² One might think that (1) is intended to be applied simultaneously, as this is the convention for rule application in SPE, but this would predict *[nxiplne], with a medial triconsonantal cluster. Left-to-right application gives *[nxiplune]; the only way to get the observed [nxipalne] is via right-to-left (RTL) application, which I’ll assume henceforth. As far as I know, the directionality issue has not been noticed in prior work.

Of course, there are exceptions of several types. (I am drawing additional data from the unpublished paper by CUNY graduate student Héctor González, henceforth G. I will not make any attempt to make González’s transcriptions or glosses comparable to those used by K, but doing so should be straightforward.)

One type is exemplified by /nama/ ‘mouth of’, which does not undergo Vowel Drop, as in /hi-nama-ya/ [hinamaya] ‘3sgmpssr-mouth.of-Obl.’ (G 5a); under RTL application we would expect *[hinmaya]. This is handled easily in the SPE exceptionality theory I reviewed a few weeks ago by marking /nama/ as [-Vowel Drop].

However, other apparent instances of exceptionality are not so easily handled. Consider two forms involving the verb root /nika/ ‘eat’. In /n-nika-nanɨ-m-ta/ [hnikananɨmta] ‘1sg-eat-Extns-Nondur-Vcl’ (G 5b) both vowels of the root satisfy (1) but do not undergo syncope. One might be tempted to mark this root as [-Vowel Drop], but it does undergo deletion in other derivations, such as in /nika-ya-pi/ [nikyapi] ‘eat-Appl-Instr.Nom’ (G 4d). Rather, it seems to be that the following /-nanɨ/ fails to trigger deletion. This is not easily handled in the SPE approach. K gives a number of similar examples involving the verbal theme suffixes /-ta/ and /-wa/, which also do not trigger syncope. If a morphemes vary in whether or not they undergo and whether or not they trigger Vowel Drop, one can imagine that these properties might cross-classify:

Mutable, catalytic The nominalizing suffix /-lu/, discussed above, is both mutable (i.e., undergoes syncope), and catalytic (triggers syncope) in /n-xipa-lu-ne/ [nxipalne].
Inalterable, catalytic I have not found any relevant examples in Piro; Kenstowicz & Kisseberth 1977 (118f.) present a hastily-described example from Slovak.
Mutable, quiescent /meyi-wa-lu/ [meyiwlu] ‘celebration’ shows that intranstive verb theme suffix /-wa/ is mutable but quiescent (does not trigger syncope; *[meywalu]).
Inalternable, quiescent /yimaka-le-ta-ni-wa-yi/ shows that the imperfective suffix /-wa/ (not to be confused with the homophonous intransitive /-wa/) is inalterable; /r-hina-wa/ [rɨnawa] ‘3-come-Impfv’ (G 6c) shows that it is quiescent (*[rɨnwa]).

According to G, there is also one additional category that does not fit into the above taxonomy: the elative suffix /-pa/ triggers deletion of the penultimate (rather than preceding) vowel, as in /r-hitaka-pa-nɨ-lo/ [rɨtkapanro] ‘3-put-Evl-Antic-3sgf’ (G 7a). Furthermore, /-pa/ appears to lose its catalytic property when it undergoes syncope, as in /cinanɨ-pa-yi/ [cinanɨpyi] ‘full-Elv-2sg’ (G 7c). Given the rather unexpected set of behaviors here, apparently confined to a single suffix, I wonder if this is the full story.

Having reviewed this data, I don’t have an abundance of confidence in it, particularly given K’s hasty presentation. However, K has identified something not obviously anticipated by the SPE theory. K’s proposal is a simple extension of the SPE theory; in addition to rule features for the target, we also need rule features for the context. For instance, inalterable, quiescent imperfective marker /-wa/,⁴ which neither undergoes nor triggers Vowel Drop, would be underlyingly [-rule Vowel Drop, –env. Vowel Drop]. Then, the rule interpretative procedure applies a rule R when its structural description is met, when the target is [+rule R], and when the all morphemes in the context are [+env. R].

Zonneveld (1978, 1979)

I have already gone on pretty long, but I should briefly discuss what subsequent writers have had to say about this proposal. Kenstowicz & Kisseberth (1977, henceforth K&K), perhaps unsurprisingly endorse the proposal, and provide some very hasty examples of how one might use this new mechanism. Zonneveld (henceforth Z), in turn, is quite critical of K’s theory. These criticisms are laid out in chapter 3 of Zonneveld 1978 (a published version of his doctoral dissertation), which reviews quite a bit of contemporary work dealing with this issue. The 1978 book chapter (about 120 typewritten pages in all) is a really good review; it is well organized and written, and full of useful quotations from the sources it reviews, and while it is somewhat dense it is hard to imagine how it could be made less so. Z reprises the criticisms of K’s theory briefly, and near verbatim, in his uncommonly-detailed review of K&K’s book (Zonneveld 1979). Z has several major criticisms of rule environment theory.

First, he draws attention to an example on where the conventions proposed by K will fail; I will spell this out in a bit more detail than Z does. The key example is /w-čokoruha-ha-nu-lu/ [wčokoruhahanru] ‘let’s harpoon it’. The anticipatory /-nu/ is mutable (but quiescent) and it is in the phonological context for syncope. To its left is the ‘sinister hortatory’ /-ha/, and this is known to be quiescent because it does not trigger deletion of the final vowel in /čokoruha/; cf. /čokoruha-kaka/ [čokoruhkaka] ‘to cause to harpoon’, which shows that the substring /…čokoruha-ha…/ does not undergo deletion because /-ha/ is quiescent rather than because /čokoruha/ is inalterable. To its right is the catalytic /-lu/. By K’s conventions, syncope should not apply to the /u/ in the anticipatory morpheme because /-ha/, in the left context, is [-env. Vowel Drop], but in fact it does. Z anticipates that one might want to introduce separate left and right context environment features: maybe /-ha/ is [+left env. Vowel Drop, -right env. Vowel Drop]. The following additional issues suggest the very idea is on the wrong track, though.

Seccondly, Z shows that rule environment features cause additional issues if one adopts the SPE conventions. The /-ta/ in /yona-ta-nawa/ [yonatnawa] ‘to paint oneself’ is presumably quiescent because it fails to trigger syncope in /yona/.⁵ Thus we would expect it to be lexically [-env. Vowel Drop], and for this specification to percolate to the segments /t/ and /a/. (I referred to this as Convention 2 in my previous post, and K adopts this convention.) However, it is a problem for this specification to be present on /t/, since that /t/ is itself in the left context for Vowel Drop, and this would counterfactually block its application to the second /a/! This is schematized below.

(2) Structural description matching for /yona-ta-nawa/

   VCVCV
yonatanawa

As a related point, Z points out that there many cases where under K’s proposal it is arbitrary whether one uses rule or environmental exception features. For instance, in the famous example obesity, the root-final s is part of the structural context so the root could be marked [-rule Trisyllabic Shortening], which would percolate to the focus e, or it could be marked [-env. Trisyllabic Shortening], which would percolate to the right-context s, or both; all three options would derive non-application. This is also schematized below.

(3) Structural description matching for obesity:

  VC-VCV
obes-ity

Z continues to argue that a theory that distinguishes between leftward and rightward contextual exceptionality also will not go through. Sadly, he does not provide a full analysis of the Piro facts in his preferred theory.

Z has a much more to say about the (then-)contemporary literature on rule exceptionality. For example, he discusses an idea, originally proposed by Harms (1968:119f.) and also exemplified by Kenstowicz (1970), that there are exceptions such that a rule applies to morphemes that do not meet its (phonologically defined) structural description. While he does seem to accept this, possible examples for such rules is quite thin on the ground, and the very idea seems to reflect the mania for minimizing rule descriptions and counting features that—and this is not just my opinion—polluted early generative phonology. If one rejects this frame, it is obvious that the effect desired can be simulated with two rules, applied in any order. The first will be a phonologically general one (with or without negative exceptions); the second will be the same change but targeting certain morphemes using whatever theory of exceptionality one prefers. Indeed, most examples of rules applying where their structural description is not met are already disjunctive, and I doubt whether such rules are really a single rule in the first place.

The ultimate theory Z settles on is one quite similar to that proposed by SPE. First, readjustment rules introduce rule features like [+R] and these handle simple exceptions of the obesity type. Z proposes further that such readjustment rules must be context-free, which clearly rules out using this mechanism for phonologically defined classes of negative exceptions; cf. (4-5) in my previous post. Secondly, Z proposes that so-called morphological features like Lightner’s [±Russian] will be used for deriving what we might now call “stratal” effects: morphemes that are exceptions to multiple rules. For instance, if we have three rules A, B, C that all [-Russian] morphemes are exceptions to, then context-free redundancy rules will introduce the following rule features.

(4)
[-Russian] -> {-A}
[-Russian] -> {-B}
[-Russian] -> {-C}

Z replays several arguments from Lightner about why morphological of this sort should be distinguished from rule features; I won’t repeat again them here. Finally, Z derives minor rules via readjustment rules triggered by so-called “alphabet” features. For instance, let us again consider umlauting English plurals like goose–geese. Z supposes, adding some detail to a sketchier portion of the SPE proposal, that morphemes targeted by umlaut are marked [+G] (where G is some arbitrary feature). There are two ways one could imagine doing this.

First, either the underlying form, /guːs/ perhaps, could be underlyingly [+G]. Then, let us assume that umlauting is simply fronting in the context of a Plural morphosyntactic feature, and that subsequent phonological adjustments (like the diphthongization in mouse–mice) are handled by later rules. Then it is possible to write this as follows:

(5) Umlaut (variant 1): [+Back, +G] -> {-Back} / __ [+Plural]

This rule is phonologically “context-free”, but its application is conditioned by the presence of the alphabet feature specification in the focus and the morphosyntactic feature in the context. I will take up the question of whether such rules are always phonologically context-free in a (much) later post.

I suspect that the analysis in (5) is the one Z has in mind, and it is also seems to be the orthodoxy in Distributed Morphology (henceforth DM); see, e.g., Embick & Marantz 2008 and particularly their (4) for a conceptually similar analysis of the English past tense. Applying their approach strictly would lead us to miss the generalization (if it is in fact a linguistically meaningful generalization) that umlauting plurals all have a null plural suffix. Umlauting plurals have an underlying feature [+G] (there is no “list” per se; it is just), but their rules of exponence also need to “list” these umlauting morphemes as exceptionally selecting the null plural rather than the regular /-z/. It seems to me this is not necessary because the rules of exponence for the plural maybe could be sensitive to the presence or absence of [+G]. This would greatly reduce the amount of “listing” necesssary. (I do not have an analysis of—and thus put aside—the other class of zero plurals in English, mass nouns like corn.)

(6) Rules of exponence for English noun plural (variant 1):

a. [+Plural] <=> ∅ / __ [+G]
b. <=> -ɹɘn / __ {√CHILD, …}
c. <=> -ɘn / __ {√OX, …}
d. <=> -z / __

Secondly and more elaborately, one could imagine that [+G] is inserted by—i.e., and perhaps, is the expression of—plurality for umlauting morphemes. In piece-based realizational theories like DM, affixes are said to expone (and thus delete) syntactic uninterpretable features. One possibility (which brings this closer to amorphous theories without completely discarding the idea of morphs) is to treat insertion of [+G] as an exponent of plurality.

(7) Rules of exponence for English noun plural (variant 2):

a. [+Plural] <=> {+G} / __ {√GOOSE, √FOOT, √MOUSE, …}
b. <=> -ɹɘn / __ {√CHILD}
c. <=> -ɘn / __ {√OX}
d. <=> -z / __

(7a) and (7b-d) implicate different types of computations—the former inserts an alphabet feature, the latter inserts vocabulary items—but I am supposing here that they can be put into competition. Under this alternative analysis, umlaut no longer requires a morphosyntactic context:

(8) Umlaut (variant 2): [+Back, +G] -> {-Back}

Beyond precedent, I do not see any reason to prefer analysis (5-6) over (7-8). Either can clearly derive what Lakoff called minor rules, though they differ in how exceptionality information is stored/propagated, and thus may have interesting consequences for how we relate the major/minor class distinction to theories of productivity. I have written enough for now, however, and I’ll have to return to that question and others another day.

Endnotes

I too will refer to this language as Piro, as do Matteson and Kisseberth. It should not be confused with unrelated language known as Piro Pueblo. Some subsequent work on this phenomenon refer to the language as Yine (and say it “was previously known as Piro”), though I also found another source that says that Yine is simply a major variety of Piro. I have been unable to figure out whether there’s a preferred endonyms.
I am not prepared to rule out the possibility that /xipa/ is itself an exception (“inalterable”), but all evidence is consistent with RTL application.
In his endnote 2, K says the rule is even narrower than stated above, since it does not apply to monosyllabic roots. However, he might have failed to note that this condition is implicit in his rule, if we interpret (11) strictly as holding that the left context should be tautomorphemic. Piro requires syllables to be consonant-initial, so the minimal bisyllabic roots is CV.CV. Combining this observation with (1), we see that the shortest root which can undergo vowel deletion is also bisyllabic, since concatenating the left context and target gives us a bisyllabic VCV substring. In fact, things are more complicated because monosyllabic suffixes do undergo syncope; many examples are provided above. Clearly, the deleting vowel need not be tautomorphemic with the preceding vowel, contrary to what a strict reading of the “+” in (1) would seem to imply. According to González, syncope imposes no constraints on the morphological structure of its context except that it only applies in derived environments—CVCVCV trisyllables like /kanawa/ ‘canoe’ surfaces faithfully as [kanawa], not *[kanwa]—and is subject to lexical exceptionality discussed here.
K glosses this as ‘still, yet’.
As was the case with /xipa/ in endnote 2, we’d like to confirm that /yona/ is mutable rather than inalterable, but one does not simply walk into Matteson 1965.

References

Embick, D. and Marantz, A. 2008. Architecture and blocking. Linguistic Inquiry 39(1): 1-53.
González, H. 2023. An evolutionary account of vowel syncope in Yine. Ms., CUNY Graduate Center.
Harms, R. T. 1968. Introduction to Phonological Theory. Prentice-Hall.
Kenstowicz, M. 1970. Lithuanian third person future. In J. R. Sadock and A. L. Vanek (ed.), Studies Presented to Robert B. Lees by His Students, pages 95-108. Linguistic Research.
Kenstowicz, M. and Kisseberth, C. W. 1977. Topics in Phonological Theory. Academic Press.
Kisseberth, C. W. 1970. The treatment of exceptions. Papers in Linguistics 2: 44-58.
Matteson, E. 1965. The Piro (Arawakan) Language. University of California Press.
Zonneveld, W. 1978. A Formal Theory of Exceptions in Generative Phonology. Peter de Ridder.
Zonneveld, W. 1979. On the failure of hasty phonology: A review of Michael Kenstowicz and Charles Kisseberth, Topics in Phonological Theory. Lingua 47: 209-255.

Pied piping and style

I find pied-piping in English a bit stilted, even if it is sometimes the prescribed option. Consider the following contrast:

(1) I’m not someone to fuck with.
(2) I’m not someone with whom to fuck.

In (1) the preposition with is stranded; in (2) it is raises along with the wh-element. What are your impressions of a speaker who says (2)? For me, they sound a bit like a nerd, or perhaps a cartoonish villain. I thought about this the other day because I was watching Alien Resurrection (1997)—it’s okay but not one of my favorite entries in the Weyland-Yutani cinematic universe—and one of the first bits of characterization we get for mercenary “Ron Johner”, played by badass Ron Perlman, is the following bit of dialogue (here taken directly from Joss Whedon’s screenplay):

This would work if Johner was a sort of evil genius, or if it was some kind of callback to something earlier, but I think this is probably just unanalyzed language pedantry ruining the vibe a little.

Generativism and anti-linguistics

I strongly identify with the generativist program. I recognize and accept that there are other ways to study language; some of these (e.g., any reasonably careful documentary work) contribute to generativist discourse and many of those that don’t are still prosocial. I for one would love to see the humanist aspect of documentation get more recognition. (Why don’t humanities programs hire linguists engaged in documentation and translation efforts?) But I’m most interested in the scientific aspects of language and think that generativism basically encompasses the big questions in this area, and some of the questions it doesn’t encompass just aren’t very important.

I don’t think it’s really ideal to brand generativism as Chomskyanism, which is the term anti-generativists tend to use. Certainly Chomsky is the plurality contributor to the program, but I think it gives undue credit to a single individual when there are so many others worth recognizing. I suspect the reason anti-generativists prefer it is they tend to see generativism as a cult of personality and perhaps want to trade on the repute of Chomsky’s (admittedly, extremely idiosyncratic but conceptually unrelated) political commitments. In evolutionary biology, it is common to refer to the modern theory of natural as the neo-Darwinian synthesis or modern synthesis. This makes sense because in 2024 there are no “strict Darwinists”, since subsequent work has integrated his monumental contributions with Mendelian and molecular genetics. Similarly, linguistics has no “strict Chomskyans”, even though we linguists eagerly awake our Mendel and our Crick & Watson.

The thing that sticks with me about the anti-generativist contingent is how disunited and disorganized they are. Anti-generativists are mostly a sincere lot (generativists too), but their attitudes are greatly shaped by negative polarization and as such, they have strange bedfellows. On the anti-generativist internet, you’ll see Adorno-disciple social constructivists talking at cross-purposes with construction grammarians, self-identified leftist/radical sociolinguists palling around with neocon consent-manufacturing journalists, experimental psycholinguists who reserve all their respect for exactly one out-of-practice fieldworker, tensorbros who don’t read books, and a few really mad, really old Boomers who never managed to build a movement around their heresy. By all accounts these people ought to hate each other. (And maybe, deep down, they do.)

In the worst case these conservations tend to veer away from constructive critique to a kind of anti-linguistics which devalues any form of language analysis that isn’t legible either as social activism or white-coat-wearing lab science. I for one can’t take your opinion about the science of language seriously if you can’t do the “armchair linguistics” that forms the descriptive-empirical base of the field. There are anti-generativists who clear this low bar, but not many. You don’t have to be a genius to do linguistics, but you do sort of have to be a linguist.

In my opinion, generativism has never been hegemonic beyond the level of individual departments, and claims otherwise are simply scurrilous. (Even MIT is a hotbed of anti-generativist reaction, after all.) But I think it would be a shame for college students to get a liberal arts education without learning about these very interesting ideas about human nature (in addition to standard consciousness-raising about prescriptivism and language ideology, which is important too).

A thought about academic jobs

I try not to pontificate about the academic job market. I recognize that I incredibly fortunate to have the job I have. I recognize that it is hard to get such a job, that it in some sense it comes down to luck, that there are more PhDs than faculty jobs, and finally that my job is not my friend. That said…

A colleague of mine had a PhD advisee who was offered a more or less ideal tenure-track job, at an excellent state school specializing in the advisee’s subarea, in a very pleasant town. The student, believe it or not, turned it down, and is now starting more or less from scratch on the alt-ac path. I genuinely don’t understand this. Earning a PhD in your field is the one always-necessary condition for getting an faculty job, even if the skills transfer to other pursuits. The demands of a graduate program expects from you are, to a great degree, necessary to get a faculty job. There are of course extra steps—that qualifying paper has to be sent off to a journal, and so on—but in terms of effort they are nothing compared to the work needed to get your degree. If you are doing well in your PhD program and if you are enjoying your studies, why not, for as long as you are able, consider applying for faculty positions? If you are not meeting your program’s expectations, your pessimism about the academic job market is besides the point, and if you are meeting or exceeding those expectations, you really might want to consider it.

Underspecification in Barrow Inupiaq

Dresher (2009:§7.2.1) discusses an interesting morphophonological puzzle from the Inuit (Canadian) and Inupiaq (Alaskan) dialects of Eskimo-Aleut. These dialects descend from a four-phoneme vowel system *i, *u, *ə, *a, but in most dialects *ə has merged into *i, yielding three surface vowels: [i, a, u]. However, in some dialects (including Barrow Inupiaq), there appears to be a covert contrast between two “flavors” of i: “strong i” triggers palatalization of a following coronal consonant whereas “weak i” does not.

(1) Barrow Inupiaq (Kaplan 1981:§3.22, his 27-29):

a. iglu ‘house’, iglulu ‘and a house’, iglunik ‘houses’
b. ini ‘place’, inilu ‘and a place’, ininik ‘places’
c. iki ‘wound’, ikiʎu ‘and a wound’, ikiɲik ‘wounds’

Presumably, the stem-final i in (1b) is weak and the one in (1c) is strong.

Following some prior work, Dresher supposes that there is an underlying contrast between weak and strong i. He posits the following featural specification:¹

(2) Features for Barrow Inupiaq vowels (to be revised):

strong i: [Coronal, -Low]
weak i: [-Low]
/u/: [Labial, -Low]
/a/: [+Low]

This analysis has a close relationship to the theory of underspecification used in Logical Substance-Free Phonology (henceforth, LP); I assume familiarity with the assumptions and operations of that theory, which have been discussed at length by Reiss and colleagues (Bale et al. 2020, Reiss 2021), including in an introductory textbook (Bale & Reiss 2018). Just a few modifications are needed, however.

For Dresher, who hypothesizes that non-contrastive features (computed using an algorithm he describes in detail; op. cit.:16) are the only ones which are phonologically active, it is not clear why strong i palatalizes coronal segments: clearly it is not that they are spreading the privative [Coronal], since that certainly would not trigger palatalization! One could, of course, adopt an analysis in which palatalization is not assimilatory. Alternatively, we could identify another feature specification which is characteristic of i and which might trigger palatalization of coronal consonants. Let us suppose this is in fact just [Palatal].² This gives us the following minimally-modified feature specification:

(2) Features for Barrow Inupiaq vowels (revised):

strong i: [Palatal, -Low]
weak i: [-Low]
/u/: [Labial, -Low]
/a/: [+Low]

According to Kaplan (§1.2), there are both plain and palatal coronal phonemes, so this seems to be a feature-changing process. Following the assumptions of LP that feature-changing processes derive from a deletion rule followed by an insertion rule, two rules are needed here; we give these below.

(3) [+Consonantal] {Coronal} / [Palatal] __
(4) [+Consonantal] ⊔ {Palatal} / [Palatal] __

Crucially, vowels other than strong i lack the [Palatal] specification to trigger (3-4).

Endnotes

Dresher’s analysis assumes privative features, but he notes elsewhere in the book that he usually adopts the features of his sources unless there is some relevant reason to dispute them.
If preferred, it is easy to translate the proposed analysis into one in which palatals are [-Back] and plain coronals are [+Back], à la Padgett (2003).

References

Bale, A., Papillon, M., and Reiss, C. 2014. Targeting underspecified segments: a formal analysis of feature-changing and feature-filling rules. Lingua 148: 240-253.
Bale, A., and Reiss, C. 2018. Phonology: a Formal Introduction. MIT Press.
Dresher, B. E. 2009. The Contrastive Hierarchy in Phonology. Cambridge University Press.
Kaplan, L. D. 1981. Phonological issues in North Alaskan Inupiaq. Alaska Native Language Center.
Padgett, J. 2003. Contrast and post-velar fronting in Russian. Natural Language and Linguistic Theory 21: 39-87.
Reiss, C. 2021. Towards a complete Logical Phonology model of intrasegmental changes. Glossa 6:107.

The phenomenology of assimilation

[This is adapted from part of paper I’m working on with Charles Reiss.]

Assimilation is a key notion for many phonological theories, and there are even intimations of it in the Prague School. Hyman provides an early formalization: he defines it as the insertion of a feature specification αF on a segment immediately adjacent to another segment specified αF.

(1) Assimilation schemata (after Hyman 1975: 159):
X → {αF} / __ [αF]
X → {αF} / [αF] __

In autosegmental phonology, assimilation is instead conceptualized as the sharing (rather than the “copying”) of feature specification via the insertion of association lines and phonological tiers provide a more general notion of adjacent, but the basic notion remains the same.

Substance-free phonology (SFP) also makes use of these “Greek letter” coefficients to express segmental identity (or non-identity) between features on various segments. What SFP denies is that there is any need to recognize or formalize notions like assimilation (or dissimilation) in the first place, because SFP rejects the notion of formal markedness. Yes, there are rules that cause an obstruent to agree in voicing with an obstruent to its immediate right, or which delete a glide between identical vowels, or which raise mid vowels before high vowels, and SFP can easily express such rules. However, there are also rules which raise mid vowels before low vowels, before nasals, or before a word boundary. The following principle expresses this position in general terms:

(2) Substance-freeness of structural change: featural specifications changed by rule application need not be present in the rule’s structural environment.

This principle is a claim that proposed phonological rules need not “make sense” in featural terms. It holds that the whatness of a rule, what feature is being added to a segment, is logically independent of the whereness, the triggering environment. This in turn echos Chomsky & Halle’s (1968:428) claim that “the phonological component requires wide latitude in the freedom to change features.” Note that principle (2) is not itself an axiom of SFP, but rather something which is not part of the theory.

References

Chomsky, N. and Halle, M. 1968. The Sound Pattern of English. Harper & Row.
Hyman, L. M. 1975. Phonology: Theory and Analysis. Holt, Rinehart and
Winston.

High school as signaling behavior

When you meet an adult for the first time in Cincinnati—where I grew up—it is customary to ask them where they went to high school. Even though I have had basically nothing to do with Cincinnati since I reached the age of majority, I can learn so much about someone by learning they went to St. Ursula, or Walnut Hills, or Elder, Summit Country Day, or Wyoming. (This is helped along by the fact that Cincinnati is, for historical reasons, rather Catholic.) It’s one of the first things I ask born and raised New Yorkers too, and it tends to yield a lot of information. I know half a dozen graduates of Bronx Science (including the president of my college); I believe David Pesetsky is one of several well-known linguists who attend Horace Mann; Hunter High is also a very promising sign, as is Stuyvesant. I even know about some of the elite high schools of Illinois at this point.

While virtually all the focus on “elite institutions” is directed at undergraduate colleges, I think this is something of a misdirection. While this may seem self-serving, I think high school choice might be a stronger signal than college choice, at least in parts of the country where it is common for one (with the help and possibly financial support of one’s parents, of course) to more or less pick a high school, with many magnet and private options.

My personal experience bears this out. I went to a very good suburban public school system (Lakota) until I was 14 and the strongest students at 14 who continued on to high school in that system are not living particularly impressive lives. In contrast, my class at my very good Catholic high school (St. Xavier) includes, among other impressive individuals, two centimillionaires (though one of those two is a phony and a scoundrel). I for one did not gain much personal ambition from St. Xavier, but I did acquire a love of learning (as someone once described it to me, “a pseudo-erotic attachment to knowledge”). Also, without any particular intentionality, I attended a good (but not selective) “R1” public college, and I feel like high school left me particularly well-positioned to take advantage of it. I didn’t even seriously consider elite colleges; I grew up in a solidly middle class family where there was no particular knowledge of elite institutions, to the point that I didn’t even find out what the Ivy League was until after I’d been accepted to Penn for my PhD. Had I been drawn from a slightly higher class stratum, I might have applied to Ivys, or at least one of those pricy private liberal arts schools on the East Coast like Vassar, and had I done so, I would have taken on an onerous load of personal debt in the process. And for what? It wouldn’t have made me any better a scholar.

Lees on underspecification

I know very little about the life of American linguist Robert Lees but he dropped two bangers in the early 1960s: his 1960 book on English nominalizations is heavily cited, and his 1961 phonology of Turkish has a lot of great ideas. In this passage, he seems to presage (though not formally) the idea that underspecified segments do not form singleton natural classes, and he correctly note that that’s a feature, not a bug.

The rest of the details of vowel- and consonant-harmony we shall discuss later; but there is one unavoidable theoretical issue to be settled in connection with this contrast between borrowed and native lexicon. The solution we have proposed for Turkish vowels, namely that they be written “archiphonemically” in all contexts where gravity is predictable by the usual progressive assimilation rules of harmony but be split into grave/acute pairs of “phonemes” elsewhere, will have the following important theoretical consequence. The phonetic rule which ensures the insertion of a “plus” or a “minus” sense for the gravity feature under harmonic assimilation must distinguish between occurrences of columns of features in which gravity is unspecified, as in the case of the “morphophoneme” /E/, from occurrences of the otherwise identical columns of features in the utterance being generated in which gravity has already been specified in a lexical rule, as in the corresponding cases of the “phonemes” /e/ and /a/, for now both the morphophoneme /E/ and the phonemes /e/ and /a/ occur simultaneously in the transcriptions. The gravity rule is intended to apply to /E/, not to /a/ or /e/. But this is tantamount to entering, as it were, a “zero” into the feature table of /E/ to distinguish it from the columns for /a/ and /e/, in which the feature of gravity has already been determined, or perhaps from other columns in which there is simply no relevant indication of this feature.

Thus the system of phonetic decisions will have been rendered trinary rather than the customary binary. The objection to this result is not based on a predilection for binary features, though there are good reasons to prefer a binary system. Rather, it arises because in a system of phonological decisions in which rules may distinguish between columns of binary features differing solely in the presence of absence of a zero for some feature one may also ipso facto always introduce vacuous reductions or simplifications without any empirical knowledge of the phonetic facts.

As a brief illustration of such an empty simplification, we might note that if a rule be permitted in English phonology which distinguishes between the features of /p/ and /b/ on one hand and on the other, the set of these same features with the exception that voice is unspecified (a set which we shall designate by means of the “archiphonemic” symbol /B/), then we could easily eliminate from English phonology, without knowing anything about English pronunciation, the otherwise relevant feature of Voice from all occurrences of either /p/ or /b/, or in fact from all occurrences of any voiced stop. Clearly, if a rule could distinguish /B/ from /p/ by the presence of zero in the voice-feature position, then that feature can be restored to occurrences of /b/ automatically and is thus rendered redundant. The same could then be done for inumerable [sic] other features with no empirical justification required.

Thus, we must assume that any rule which applies to a column of features like /B/ also at the same time applies to every other type of column which contains that same combination of features, such as /p/ and /b/. This is tantamount to imposing the constraint on phonological features that they never be required to identify unspecified, or zero, features. To the best of our present knowledge, there seems to be no other reasonable way to prevent the awkward consequences mentioned above.

To return to Turkish this decision means that the grammar is incapable of distinguishing native vowel-harmonic morphemes from borrowed non-vowel-harmonic morphemes simply be the presence of the archiphoneme /E/ in the former versus /e/ or /a/ in the latter. (Lees 1961: 12-14)

References

Lees, R. B. 1961. The Phonology of Modern Standard Turkish. Indiana University Press.

Stop capitalizing so much

One of the absolute scourges of student writing is the tendency to capitalize just about every multi-word noun phrase. The rule in English is pretty simple: you only capitalize proper names, and these are, roughly, the names of people, locations, or organizations. Technical concepts do not qualify. It doesn’t matter if it’s part of an acronym: we capitalize the acronym but not necessarily the full phrase. Natural language processing is not a proper name; cognitive science isn’t either; logistic regression certainly is not a proper name nor is conditional random fields or hidden Markov model or support vector machine or…

SPE & Lakoff on exceptionality

Recently I have attempted to review and synthesize different theories of what we might call lexical (or morpholexical or morpheme-specific) exceptionality. I am deliberately ignoring accounts that take this to be a property of segments via underspecification (or in a few cases, pre-specification, usually of prosodic-metrical elements like timing slots or moras), since I have my own take on that sort of thing under review now. Some takeaways from my reading thus far:

This is an understudied and undertheorized topic.
At the same time, it seems at least possible that some of these theories are basically equivalent.
Exceptionality and its theories play only a minor role in adjudicating between competing theories of phonological or morphological representation, despite their obvious relevance.
Also despite their obvious relevance, theories of exceptionality make little contact with theories of productivity and defectivity.

Since most of the readings are quite old, I will include PDF links when I have a digital copy available.

Today, I’m going to start off with Chomsky & Halle’s (1968) Sound Pattern of English (SPE), which has two passages dealing with exceptionality: §4.2.2 and §8.7. While I attempt to summarize these two passages as if they are one, they are not fully consistent with one another and I suspect they may have been written at different times or by different authors. Furthermore, it seemed natural for me to address, in this same post, some minor revisions proposed by Lakoff (1970: ch. 2). Lakoff’s book is largely about syntactic exceptionality, but the second chapter, in just six pages, provides important revisions to the SPE system. I myself have also taken some liberties filling in missing details.

Chomsky & Halle give a few examples of what they have in mind when they mention exceptionality. There is in English a rule which laxes vowels before consonant clusters, as in convene/conven+tion or (more fancifully) wide/wid+th. However, this generally does not occur when the consonant cluster is split by a “#” boundary, as in restrain#t.¹ The second, and more famous, example involves the trisyllabic shortening of the sort triggered by the -ity suffix. Here laxing also occurs (e.g., serene–seren+ity, obscene–obscen+ity) though not in the backformation obese-obesity.² As Lakoff (loc. cit.:13) writes of this example, “[n]o other fact about obese is correlated to the fact that it does not undergo this rule. It is simply an isolated fact.” Note that both of these examples involve underapplication, and the latter passage gives more obesity-like examples from Lightner’s phonology of Russian, where one rule applies only to “Russian” roots and another only to “Church Slavonic” roots.

SPE supposes that by default, that there is a feature associated with each rule. So, for instance, if there is a rule R there exists a feature [±R] as well. A later passage likens these to features for syntactic category (e.g., [+Noun]), intrinsic morpho-semantic properties like animacy, declension or conjugation class features, and the lexical strata features introduced by Lees or Lightner in their grammars of Turkish and Russian. SPE imagine that URs may bear values for [R]. The conventions are then:

(1) Convention 1: If a UR is not specified [-R], introduce [+R] via redundancy rule.
(2) Convention 2: If a UR is [αR], propagate feature specification [αR] to each of its segments via redundancy rule.
(3) Convention 3: A rule R does not apply to segments which are [-R].

Returning to our two examples above, SPE proposes that obese is underlylingly [−Trisyllabic Shortening], which accounts for the lack of shortening in obesity. They also propose rules which insert these minus-rule features in the course of the derivation; for instance, it seems they imagine that the absence of laxing in restraint is the result of a rule like V → {−Laxing} / _ C#C, with a phonetic-morphological context.

Subsequent work in the theory of exceptionality has mostly considered cases like obesity, the rule features are present underlyingly but with one exception, discussed below, the restraint-type analysis, in which rule features are introduced during the derivation, do not seem to have been further studied. It seems to me that the possibility of introducing minus-rule features to a certain phonetic context could be used to derive a rule that applies to unnatural classes. For example, imagine an English rule (call it Tensing) which tenses a vowel in the context of anterior nasals {m, n} and the voiceless fricatives {f, θ, s, ʃ} but not voiced fricatives like {v, ð}.³ Under any conventional feature system, there is no natural class which includes {m, n, f, θ, s, ʃ} but not also {ŋ, v}, etc. However, one could derive the desired disjunctive effect by introducing a -Tensing specification when the vowel is followed by a dorsal, or by a voiced fricative. This might look something like this:

(4) No Tensing 1: [+Vocalic] → {−Tensing} / _ [+Dorsal]
(5) No Tensing 2: [+Vocalic] → {−Tensing} / _ [-Voice, +Obstruent, +Continuant]

This could continue for a while. For instance, I implied that Tensing does not apply before a stop so we could insert a -Tensing specification when the following segment is [+Obstruent, -Continuant], or we could do something similar with a following oral sonorant, and so on. Then, the actual Tensing rule would need little (or even no) phonetic conditioning.

To put it in other words, these rules allow the rule to apply to a set of segments which cannot be formed conjunctively from features, but can be formed via set difference.⁴ Is this undesirable? Is it logically distinct from the desirable “occluding” effect of bleeding in regular plural and past tense allomorphy in English (see Volonec & Reiss 2020:28f.)? I don’t know. The latter SPE passage seems to suggest this is undesirable: “…we have not found any convincing example to demonstrate the need for such rules [like my (4-5)–KBG]. Therefore we propose, tentatively, that rules such as [(4-5)], with the great increase in descriptive power that they provide, not be permitted in the phonology.” (loc cit.:375). They propose instead that only readjustment rules should be permitted to introduce rule features; otherwise rule feature specifications must be underlyingly present or introduced via redundancy rule.

As far as I can see, SPE does not give any detailed examples in which rule feature specifications are introduced via rule. Lakoff however does argue for this device. There are rules which seem to apply to only a subset of possible contexts; one example given are the umlaut-type plurals in English like foot–feet or goose–geese. Later in the book (loc. cit./, 126, fn. 59) the rules which generate such pairs are referred to these as minor rules. Let us call the English umlauting rule simply Umlaut. Lakoff notes that if one simply applies the above conventions naïvely, it will be necessary to mark a huge number of nouns—at the very least, all nouns which have a [+Back] monophthong in the final syllable and which form a non-umlauting plural—as [-Umlaut]. This, as Lakoff notes, would wreck havoc on the feature counting evaluation metric (see §8.1), and would treat what we intuitively recognize as exceptionality (forming an umlauting plural in English) as “more valued” than non-exceptionality. Even if one does not necessarily subscribe to the SPE evaluation metric, one may still feel that this has failed to truly encode the productivity distinction between minor rules and major rules that have exceptions. To address this, Lakoff proposes there is another rule which introduces [–Umlaut], and that this rule (call it No Umlaut) applies immediately before Umlaut. Morphemes which actually undergo Umlaut are underlying -No Umlaut. Thus the UR of an noun with an umlauting plural, like foot, will be specified [–No Umlaut], and this will not undergo a rule like the following:

(6) No Umlaut: [ ] → {–Umlaut}

However, a noun with a regular plural, like juice, will undergo this rule and thus the umlauting rule U will not apply to it because it was marked [-U] by (6).

One critique is in order here. It is not clear to me why SPE introduces (what I have called) Convention 2; Lakoff simply ignores it and proposes an alternative version of Convention 3 where target morphemes, rather than segments, must be [+R] to undergo rule R. Of his proposal, he writes: “This system makes the claim that exceptions to phonological rules are morphemic in nature, rather than segmental.” (loc. cit., 18) This claim, while not necessarily its 1970-era implementation, is very much in vogue today. There are some reasons to think that Convention 2 introduces unnecessary complexities, which I’ll discuss in a subsequent post. One example (SPE:374) makes it clear that for Chomsky & Halle, Convention 3 requires that that for rule R the target be [+R], but later on, they briefly consider what if anything happens if any segments in the environment (i.e., structural change) are [-R].⁵ They claim (loc. cit., 375) there are problems with allowing [-R] specifications in the environment to block application of R, but give no examples. To me, this seems like an issue created by Convention 2, when one could simply reject it and keep the rule features at the morpheme level.

I have since discovered that McCawley (1974:63) gives more or less the critique of this convention in his review of SPE.

A correction: after rereading Zonneveld, I think Lakoff misrepresents the SPE theory slightly, and I repeated his mispresentation. Lakoff writes that the SPE theory could have phonological rules that introduce minus-rule features. In fact C&H say (374-5) that they have found no compelling examples of such rules and that they “propose, tentatively” that such rules “not be permitted in the phonology”; any such rules must be readjustment rules, which are assumed to precede all phonological rules. This means that (4-5) are probably ruled out. Lakoff’s mistake may reflect the fact that the 1970 book is a lightly-adapted version of his 1965 dissertation, for which he drew on a pre-publication version of SPE.

[This post, then, is the first in a series on theories of lexical exceptionality.]

Endnotes

The modern linguist would probably not regard words like restraint as subject to this rule at all. Rather, they would probably assign #t to the “word” stratum (equivalent to the earlier “Level 2”) and place the shortening rule in the “stem” stratum (roughly equivalent to “Level 1”). Arguably, C&H have stated this rule more broadly than strictly necessary to make the point.
It is said that the exceptionality of this pair reflects its etymology: obese was backformed from the earlier obesity. I don’t really see how this explains anything synchronically, though.
This is roughly the context in which Philadelphia short-a is tense, though the following consonant must be tautosyllabic and tautomorphemic with the vowel. Philadelphia short-a is, however, not a great example since it’s not at all clear to me that short-a tensing is a synchronic process.
Formally, the set in question is something like [−Dorsal] ∖ [+Voice, +Consonantal, +Continuant, −Nasal].
This issue is taken up in more detail by Kisseberth (1970); I’ll review his proposal in a subsequent post.

References

Chomsky, N. and Halle, M. 1968. The Sound Pattern of English. Harper & Row.
Kisseberth, C. W. 1970. The treatment of exceptions. Papers in Linguistics 2: 44-58.
Lakoff, G. 1970. Irregularity in Syntax. Holt, Rinehart and Winston.
McCawley, J. D. 1974. Review of Chomsky & Halle (1968), The Sound Pattern of English. International Journal of American Linguistics 40: 50-88.