Research

Dissertation

Register and Language Change in Ancient Egyptian
Humboldt-Universität zu Berlin · CRC 1412, project B03

The problem

Texts of the New Kingdom written in classical Middle Egyptian sometimes admit forms of the emerging Late Egyptian norm. In a paper given in 1985, Schenkel singled such texts out as following no binding norm but vacillating between several.1 He called the alternation unmotiviert — so unmotivated that he took it to indicate a certain deficit of linguistic competence. Whether that was in fact the reason he expressly declined to decide, asking instead for what was then unavailable: quantitative estimates, if not statistics.

The thesis

The history of Egyptian is conventionally told as a succession of stages, one written norm giving way to the next. In such an account a text that mixes two of them is an anomaly — a transitional specimen, or the work of someone who has not fully mastered the older norm.

This study argues instead that a period does not have one norm per kind of text, but a repertoire of conventions available at the same time, each tied to the situations it serves. Registers, on this account, are not the place where a change merely becomes visible; they are the condition under which it takes place, in that they decide whether an innovation is admitted at a given point or not. Change then looks less like one stage giving way to another than like a language differentiating functionally — a form is admitted where the situation tolerates it long before it is admitted where the situation does not, and the scribe who alternates is not failing at one norm but moving within several.

The frame is the variationist study of language change, which treats a language as carrying coexistent layers that a community has available at once and asks how a change is embedded in them — in the structure of the language, and in the social structure of those who use it. Variationist work usually pursues the second: a variable is correlated with class, age, or network. For Ancient Egypt that is mostly out of reach, since a scribe’s standing is rarely known and his circle almost never. What the texts do show is the occasion they were written for, so the question here is how the change is embedded in the functional layers rather than in a social stratification — a shift of evidence, not a claim that the social does not matter, since registers are socially constituted and socially valued themselves. The tools variationists built for living speakers do not come with the programme either: there are no recordings across situations, no judgements of what sounds careful, no vernacular to observe. What does carry over is the central discipline of counting a form against the occasions on which it could have occurred, which is why this has to be a corpus study. The corpus for it has existed since the Thesaurus Linguae Aegyptiae went online in 2004; the debate has gone on being conducted from examples all the same.

The thesis takes a testable shape from an observation of Kroch’s (1989) about syntactic change:2 as a new form spreads, it tends to advance at the same rate in every context in which it is possible, while the contexts differ in how far it has already got. The change itself belongs to the grammar; the contexts differ in what they admit. If registers are niches of admission in this sense rather than engines of change in their own right, that is exactly the pattern they should produce — one shared rate, different points of entry. So the prediction is not that innovation moves faster in some registers, but that it starts from a different level in each.

What makes the thesis testable on a language with no speakers is that it asks nothing of them. It does not require that any scribe meant to admit or to block a form: the pattern is what a great many situated choices add up to, no one of which need have been aimed at the outcome. And it does not require knowing how a scribe understood his own situation, which is irrecoverable in any case — the niches are found in the texts themselves and characterised by the occasions those texts were written for.

That register conditions the distribution is not a new claim in Egyptology. Jansen-Winkeln (1995) tested the rival model — innovation climbing step by step through a hierarchy of text classes — against texts that do not fit it, and put in its place a division by domains that comes close to the thesis held here.3 What is missing is not the claim but a way of deciding it.

The rival explanations

Schenkel’s verdict is the null hypothesis, and it is testable in a precise sense: the variation is unmotivated only if it covaries with no situational-functional parameter of the texts. But a register effect is not the only reading the same distribution admits, and the study is built to separate three.

  • Register. The Late Egyptian forms reflect deliberate choices conditioned by the communicative context a text was written for.
  • Interference. They seep from the writer’s spoken variety into classical written production where stylistic monitoring is reduced, the effect indexed to the communicative situation and to how far a given feature was open to conscious control. This is close to Schenkel’s competence-based reading, but mechanistically specified. Its sociolectal variant — innovation rising through the social hierarchy — is not excluded by definition but carried as a weakly downweighted sub-hypothesis that the data can override.
  • Transmission. They come from the compositional history of the text: copied passages preserve the language stage of their exemplar, freely composed ones carry the chronolect of the time of writing. This reading, argued in Egyptology against register-based explanations, is tested on the subcorpus of multiply attested texts.

Only separating the three carries the inference. Restating the claim as a quantity with an interval around it is the contribution, and so is setting it against its rivals on the same data rather than arguing between them.

Corpus and method

The corpus is drawn from the JSON export of the Thesaurus Linguae Aegyptiae, lemmatised and morphologically tagged. Two nested windows are analysed: from the Eleventh Dynasty to the end of the Eighteenth (1,678 texts) and on to the end of the New Kingdom (2,243 texts, 413,250 word tokens).

Every text is annotated for the situation it was written for, by a scheme of five parameter levels — use situation, recording purpose, communicative constellation, content, mode of production — plus a metadata block, each level with a defined role in the model: the externally determinable ones predict, the text-internally reconstructed one validates and segments, the production level moderates and filters, the metadata control confounds. Its Egyptological base is Jansen-Winkeln’s (1994) systematics of use situation, recording purpose and text carrier;4 the treatment of the participant relationship — who stands to whom in what role, at what social distance — follows Butt (2003),5 and the architecture of the parameters follows the annotation scheme developed in the CRC by Lehmann (2024),6 which is what keeps the Egyptian material comparable with the other projects there. The conditions by which Koch and Oesterreicher (1985) place an utterance between immediacy and distance — whether the partners are present to one another, how familiar they are, how much is planned in advance — are distributed across these levels rather than collapsed into a single scale.7 Two things in it are answers to problems the general frameworks do not pose: a recorded Egyptian text has two situations, the action it stems from and the act of recording it, and both are annotated; and because much of the communicative situation of a mortuary text can only be inferred, every value carries a flag saying what evidence licenses it — material, conventional, or textual — with form-based evidence barred from predictor roles, so that the model cannot be fed by the very language it is meant to explain. The unit of annotation is the part-text: where a parameter changes within a text, the annotation follows it. Field, tenor and mode — the three dimensions along which systemic functional linguistics describes a situation: what is going on, who is involved, and what part the medium plays — are laid over the scheme as one of three mapping layers, alongside Biber’s situational framework and Lehmann’s, which keeps it comparable with other annotation projects without organising it.

Where the writing system offered a classical and a Late Egyptian realisation of the same content, the feature is measured within its variation envelope: the contexts in which the two variants compete for the same functional slot, with the observation being which of them was chosen. A rate defined this way does not depend on how often a construction happens to be called for. Features with no classical counterpart are recorded as present or absent instead. The feature list comes from the Egyptological literature on Late Egyptianisms, Kroeber’s collection (1970) and its successors,8 and currently holds 26 features measured as presence and 17 as envelopes.

The evidence that results is thin, unevenly preserved and hierarchically structured: many texts contribute a handful of observations, and no text is an independent draw. Multilevel models with partial pooling are built for that — they let a sparsely attested text borrow strength from the rest instead of being either discarded or over-read. A hierarchical binomial model with an observation-level random effect estimates the effect of the situational parameters on the choice within the envelope, with varying intercepts across texts and across the levels of a feature’s accessibility to conscious control. Mean and dispersion are modelled separately, which is what makes the first two explanations distinguishable: register governs how much innovation a situation admits, monitoring how reliably the admitted norm is maintained.

The estimation is Bayesian, because what the study wants to state is how strongly the surviving evidence favours one reading over another. A Bayesian model says that directly: for each quantity of interest it returns the whole range of values compatible with the data, so that every result carries its own uncertainty. Between models the criterion is prediction — whether one accounts for texts it was not fitted on better than another does.

Findings so far, and their reach

On the prediction that register governs levels rather than speed, the evidence is so far consistent: on the specialisation axis the incoming features rise at one and the same rate across strata that lie far apart in how much innovation they admit — the strata differ in where they start, not in how fast they move. These figures are provisional in a strong sense. One decision on corpus membership is still open, and settling it would mean recomputing them; they are estimates whose sampling frame is still being fixed, not estimates awaiting more data.

Two limits bound what such results can mean. They are, in the first instance, facts about a system of written varieties: apparent movement through time has to be located within that system before it can be read as change at all, since stylistically layered variation can remain stable for centuries. And the claim they license is about the scribal culture, not about individuals — the record preserves too few writers with enough range to ask whether one and the same hand commands the whole repertoire and moves through it by situation. A repertoire available to a writing culture is what this corpus can show; a repertoire commanded by a person is not.

Further work

Multi-dimensional analysis for Egyptian. Biber’s multi-dimensional analysis derives dimensions of register variation from the covariation of linguistic features instead of from assumed categories (Biber 1988; an overview in Biber and Conrad 2009);9 it had not been applied to Egyptian. I have built it for Old, Middle and Late Egyptian on data from the Thesaurus Linguae Aegyptiae and tested how far it holds — split-half reliability of the dimensions, their dependence on the length of the unit analysed, and the effect of the uneven transmission of text types. It proves feasible under stated reservations, with one to three reproducible dimensions per language stage. This strand runs alongside the dissertation rather than forming part of its argument, and is the subject of three talks.

Lexical diversity in the TLA. A question that grew out of that probe: how much text does a measure need? Taking Hintze’s lexicostatistical work (1975, 1976) as a starting point,10 an evaluation of vocabulary-richness measures — Hintze’s S*, Maas, MATTR, MTLD and HD-D — on the TLA corpus: how they agree with one another, how far they depend on text length, and how much text an Egyptian composition must contain before any of them mean anything. With Eliese-Sophia Lincke; submitted, see Publications.

Register in Egyptology. With Svenja Damm, a conceptual history and critical review of the term. We compiled an annotated bibliography of the field, tagged on five axes, and derived a citation network from it, which lets one see who was read, when, and with what grasp of the concept. The argument is that register is firmly established in Egyptology yet rarely defined and inconsistently applied; that the citation record does not penalise this, so the case for conceptual care must be made on methodological grounds; and that the models proposed so far should be turned into testable hypotheses. Bibliography, network, figures and scripts are published openly, the network as a self-contained interactive page.

Reconstructed vocalisation. Egyptian is written without vowels; they are reconstructed from Coptic, from cuneiform and Greek transcriptions, and from the regularities of Egyptian word formation. A prototype tool of mine takes a lemma with its Coptic reflexes and returns the possible vocalised forms, each with its formation class and a confidence rating. A second strand passes such forms to an Amharic speech synthesiser, whose inventory comes closest to what is needed. The vowels are reconstructed and the prosody is Amharic, so the audio illustrates rather than shows how Egyptian sounded.

Footnotes

  1. Schenkel, Wolfgang. 1989. “Sprachforschung und Textquellen. Integrierte Datenverarbeitung als konkrete Utopie.” In Akten des Vierten Internationalen Ägyptologenkongresses München 1985, vol. 3, edited by Sylvia Schoske, 1–27. Studien zur Altägyptischen Kultur, Beihefte 3. Hamburg.↩︎

  2. Kroch, Anthony S. 1989. “Reflexes of Grammar in Patterns of Language Change.” Language Variation and Change 1 (3): 199–244.↩︎

  3. Jansen-Winkeln, Karl. 1995. “Diglossie und Zweisprachigkeit im Alten Ägypten.” Wiener Zeitschrift für die Kunde des Morgenlandes 85: 85–115.↩︎

  4. Jansen-Winkeln, Karl. 1994. Text und Sprache in der 3. Zwischenzeit: Vorarbeiten zu einer spätmittelägyptischen Grammatik. Ägypten und Altes Testament. Wiesbaden: Harrassowitz.↩︎

  5. Butt, David G. 2003. Parameters of Context: On Establishing the Similarities and Differences between Contexts. Centre for Language in Social Life, Macquarie University.↩︎

  6. Lehmann, Nico. 2024. Classifying Communicative Situations and Assessing Formality. Unpublished.↩︎

  7. Koch, Peter, and Wulf Oesterreicher. 1985. “Sprache der Nähe – Sprache der Distanz. Mündlichkeit und Schriftlichkeit im Spannungsfeld von Sprachtheorie und Sprachgeschichte.” Romanistisches Jahrbuch 36: 15–43.↩︎

  8. Kroeber, Burkhart. 1970. Die Neuägyptizismen vor der Amarnazeit. Studien zur Entwicklung der ägyptischen Sprache vom Mittleren zum Neuen Reich. PhD diss., Universität Tübingen.↩︎

  9. Biber, Douglas. 1988. Variation across Speech and Writing. Cambridge: Cambridge University Press. — Biber, Douglas, and Susan Conrad. 2009. Register, Genre, and Style. Cambridge: Cambridge University Press.↩︎

  10. Hintze, Fritz. 1975. “Die statistische Struktur des Wortschatzes ägyptischer Literaturwerke 1: der Reichtum des Vokabulars.” Zeitschrift für ägyptische Sprache und Altertumskunde 102: 100–122. — Id. 1976. “Die statistische Struktur des Wortschatzes ägyptischer Literaturwerke. Fortsetzung 2: die Verteilung der Häufigkeiten innerhalb des Vokabulars.” Ibid. 103: 22–30.↩︎