Five dimensions of register variation in Egyptian
Egyptian texts differ from one another in ways every reader notices: the letters of Heqanakht do not read like the Pyramid Texts, Papyrus Ebers not like the Victory Stela of Piye. Egyptology has mostly described those differences in one of two ways — as stages of the language, or as genres. This post is about a third way of describing them, through the linguistic concept of register, and about an attempt to measure register directly from the co-occurrence of linguistic features across a corpus.
The method is Douglas Biber’s multi-dimensional analysis, applied here to 244 texts from the Thesaurus Linguae Aegyptiae. It yields five dimensions of variation, three of them stable, and they separate not only groups of texts but also the parts of a single text.
Register
In the definition used by our Collaborative Research Centre, register is “intra-individual variation in language use, that is, recurrent and conventionalized linguistic patterns that speakers employ depending on the situational-functional context”.1 The same speaker uses language differently in different situations. That variation can affect every linguistic level, from lexicon and morphosyntax to discourse structure, and it is shaped by purpose, participants and setting. Register is therefore usage-based rather than user-based: unlike a dialect or a sociolect, it is not a stable property of a group of speakers.
The idea did not have to reach Egyptology from linguistics. The discipline saw functional variation from its beginnings, even if it mostly cast it as written against spoken language and as a succession of stages. In 1837 Richard Lepsius distinguished a “dialecte sacré”, used for religious and scholarly purposes, from a “dialecte populaire” that served legal and private affairs. In 1889 Adolf Erman set the archaising idiom of literary texts apart from the prose of contemporary scientific compendia and rejected the idea that Papyrus Westcar reflected the vernacular of its time. In 1925 Kurt Sethe described Demotic and Coptic not simply as stages but as functionally distinct varieties, one written and tied to official and religious discourse, the other spoken. And in 1933 Erman distinguished an ordinary way of speaking from an elevated style used in poetry and in references to the king and the gods, noted that educated scribes could move between the two, and pointed out that in Medinet Habu minor figures speak more simply than the main narrative.2
Register studies have three main traditions, and they differ above all in which direction they work.
Systemic Functional Linguistics works from the top down. Its line runs from the anthropologist Bronisław Malinowski through J. R. Firth and the London School of Linguistics to Michael Halliday and his students, who built a theory in which context determines linguistic choice.3 Field (the activity or subject matter), tenor (the roles and relationship of the participants) and mode (the channel and organisation of the text) define a register before any linguistic analysis begins. It is a strong tool for interpretation, but not in itself empirical.
Close to the beginning of this line, though outside it, stands Alan Gardiner, author of the seminal Egyptian Grammar. Encouraged by Bertrand Russell, he developed a theory of language that put communicative purpose and speaker intention at the centre, an idea he had already worked out in his correspondence with Malinowski from 1917 onwards.4 His Theory of Speech and Language (1932) treats utterances not primarily as representations of thought but as actions directed at influencing a listener. It distinguishes grammatical sentence form from sentence function, and it introduces a “special sentence quality” for the communicative aim behind an utterance — what speech act theory would later call illocutionary force. Gardiner speaks of an “act of speech” a year before Karl Bühler’s Sprechakt, which is usually credited as the first.5 He himself thought this work “of far greater importance” than his grammar,6 but it found little response. Malinowski was institutionalised through his student Firth and the London School of Linguistics; Gardiner, financially independent and without teaching duties, worked outside academic structures, and Firth, while mentioning him in connection with the context of situation, found his writing hard to read. Yet the classification of utterance types, the interrelation of morphosyntactic form and pragmatic function, and the modelling of language as context-bound social action all anticipate core tenets of later register theory.7
As a theory of register, this is also the tradition Egyptology adopted first. Pascal Vernus used “register” in the linguistic sense as early as 1979, recasting Fritz Hintze’s distinction between discourse and narration as one between registers, and in the same volume introduced égyptien(s) de tradition as a spectrum of mimetic registers rather than a single stage of the language.8 Karl Jansen-Winkeln’s habilitation, finished in 1990, classifies text types in a way that closely mirrors field, tenor and mode.9 Around the same time Orly Goldwasser met British register theory during her doctoral studies; her reading of the coexistence of Middle and Late Egyptian forms in Papyrus Anastasi I as register variation remains the most cited discussion of register in Egyptology.10
Variationist sociolinguistics works from the bottom up. William Labov showed, on Martha’s Vineyard and in the New York department stores, that language varies with class, ethnicity and situation, and that such patterns are found in the data rather than imposed on it.11 Its principle is that the frequencies of linguistic features reflect contextual pressures.
Multi-dimensional analysis carries that principle over to written and spoken corpora. Douglas Biber takes linguistic features that pattern together as the trace of an underlying communicative function, and treats register not as a category but as a position in a space of several dimensions at once.12
In Egyptology the call for this kind of analysis is explicit. Todd Gillen wrote in 2014 that registers should be defined “according to observable clusters of regularities”, and that this “requires the methods of machine-based corpus linguistics for statistical analysis and multidimensional comparison”.13 Stéphane Polis adopted every theoretical component of Biber’s framework — registers as continuous constructs, distinguished by relative frequencies, with co-occurrence as the decisive relation — and concluded that “the use of quantitative methods is necessary”.14 His study of the scribe Amennakhte in the same volume calls its core “a multidimensional analysis of Amennakhte’s linguistic registers”, but takes over the name rather than the procedure: there is no factor analysis, no modelling of feature covariation, no latent dimensions.15 As far as I know, the analysis below is the first one for Egyptian that actually runs it.
How a multi-dimensional analysis works
A linguistic feature, in this sense, is anything in a text that can be counted: a part of speech such as the noun, a grammatical form such as the imperative or the passive, a class of words such as the suffix pronouns of the second person. Its rate is how often it occurs per word of text. The input to the analysis is a table: one row per text, one column per feature, and in each cell the rate of that feature in that text. The analysis runs in four steps.
- Correlations. For every pair of features the analysis measures how far they rise and fall together across texts. This correlation runs from +1 (always together) through 0 (unrelated) to −1 (the one rises as the other falls).
- Latent variables. Factor analysis then looks for a small number of hidden variables, the factors, that would explain this web of correlations. It first extracts the combination of features that accounts for most of what the features have in common — their shared variance — then the next combination independent of the first, and so on.
- Rotation. The first solution is mathematically optimal but hard to read, because most features are moderately tied to most factors. Rotation reorients the solution, without changing how much it explains, until each feature is tied strongly to as few factors as possible.
- Output. Loadings, between −1 and +1, say how strongly each feature is tied to each factor; a negative loading means the feature goes with the opposite end of the scale. Scores say how much of each factor’s pattern is present in each text.
Each factor is a dimension: a continuous scale, not a yes-or-no category, on which every text has a position. The interpretation rests on a single assumption — features that keep occurring together probably serve a shared communicative function. If questions, first-person pronouns and short clauses cluster, and the texts in which they cluster read as personal and interactive, the dimension is read as something like “involved versus informational”. The label follows the pattern. What factor analysis cannot do is name the factors or say what they mean; that remains the researcher’s job.
Short texts
A multi-dimensional analysis presupposes a corpus that is representative, sufficiently large and balanced. No corpus of an ancient language meets that. The record is fragmentary, its distributions are skewed, and above all its texts are short. Earlier Egyptian before Amarna may serve as an example: in the TLA it comprises 12,975 texts. The longest has 4,088 words, the mean is 40.7, the median 9.
A rate computed over nine words is either zero or very high: one occurrence of a feature makes it frequent, none makes it absent. Long texts are rare: before Amarna, 1,128 texts reach 100 words, 263 reach 300, and only 142 reach 500.
How long does a text have to be for its rates to mean something? One way to find out is to take long texts, cut them to a given length, and then cut them in half. If the two halves give the same picture, the length suffices; if they do not, what one sees is largely chance. How closely the halves agree is the split-half reliability: 1 means they agree completely, 0 not at all, and from about 0.8 upwards a measurement counts as reliable. The figure shows this for 89 long texts from the Eleventh to the end of the Eighteenth Dynasty, cut to lengths from 40 to 640 words.
The line reaches the reliable zone at around 160 words. The two dotted lines mark what this means in practice. At 51 words, the typical length of a dated text from this period, it lies far below that zone. At 500 words, the minimum length of the texts used in this study, it is well inside it.
More texts do not make up for short ones. This can be tested by simulating the whole analysis with texts of realistic length and comparing the dimensions that come out with the true ones. The measure for that comparison is the congruence of two dimensions: 1 means they are identical, values from about 0.85 upwards count as similar, lower values as different. With texts of realistic length the congruence stayed at about 0.5, whether the simulation used 89 texts or 800. With 500 words per text, 89 texts already reached 0.81.
For Earlier Egyptian alone, there are too few long texts. A factor analysis needs about five texts for every feature it measures. The analysis presented below works with 43 features and so needs at least some 215 texts; before Amarna, only 142 texts reach 500 words.
There is a way around this in principle: pooling short texts into longer samples. A simulation that cuts long texts into fragments of realistic length and recombines them shows that samples of some 200 to 300 words reproduce the dimensions almost exactly, with a congruence of about 0.97, provided the fragments pooled together come from the same genre and period. Pooled at random, the samples still yield dimensions, but the differences between registers disappear: at 500 words about 60 per cent of the differences between samples are lost. Pooling within genre and period therefore needs reliable labels for both. Egyptology has no agreed classification of its texts by genre, and the notion of genre is itself contested.16 The analysis below therefore does not pool, but uses texts that are long enough on their own.
A workaround corpus of 244 texts
The corpus used here therefore draws on all pre-Demotic stages of the language: 244 texts of at least 500 words each, selected by hand and arranged in six groups. Of texts attested in several copies only one was kept, the longest, and funerary texts were thinned out — the Book of the Dead, for instance, enters as fifteen spells, three of each type — so that they would not swamp rarer groups such as the letters. It is a diachronic corpus as a deliberate workaround, and in size it is comparable to the corpora of multi-dimensional analyses of living languages — Biber and Hared’s analysis of Somali, for instance, worked with 158 written and 121 spoken texts.17 What mixing the stages does to the dimensions is taken up further below.
| Group | Texts | Language stage |
|---|---|---|
| Religion (REL) — Pyramid Texts, Book of the Dead, liturgies, rituals, hymns | 52 | Old and Middle Egyptian |
| Decorum (DEC)18 — royal narratives, biographies, decrees, building inscriptions, stelae | 62 | Middle and Late Egyptian |
| Magic (MAG) — apotropaic spells, ritual texts, magico-medical texts, oracular decrees | 47 | Middle and Late Egyptian |
| Literature (LIT) — teachings, narratives, dialogues, love songs, panegyric | 40 | Middle and Late Egyptian |
| Medicine (MED) — diagnostic treatises, recipe collections | 29 | Middle Egyptian |
| Letters (LET) — private and administrative correspondence | 14 | Middle and Late Egyptian |
The groups are a heuristic for sampling and description. The factor analysis never sees them: it works on the feature table alone, and relabelling or dissolving the groups changes no loading.
Why 500 words? Since reliability keeps rising with length, the threshold should be as high as the factor analysis allows. With 43 features at roughly five texts per feature, at least some 215 texts are needed. At 600 words only 191 remain, at 800 only 136. Five hundred is the strictest cut the analysis can sustain. Whether the threshold itself shapes the dimensions is taken up below, under the checks.
Features
Biber’s original inventory for English comprised 67 features in 16 categories, identified through part-of-speech patterns, dependency relations and lexical matching. The Egyptian inventory was built functionally: starting from what multi-dimensional analyses of ten living languages measure, and asking which Egyptian forms carry those functions. The result is 70 features. Not all of them enter the factor analysis. A feature that shares little of its variation with the others — its communality, the share it has in common with the factors, is below 0.20 — contributes noise rather than pattern and is dropped. That leaves 43. Three features are kept on theoretical grounds although they fall below the threshold: lexical diversity, the stative, and stance verbs. The table lists all 70 features; the 43 that enter the factor analysis are set in bold.
| Category | Features |
|---|---|
| Pronouns and reference | personal pronouns of the three persons, independent pronouns, demonstratives (total), distal demonstratives (pf/tf/nf), interrogatives, relative pronouns, possessive suffixes, the quantifier nb, emphatic reflexive ḏs |
| Noun phrase | proper nouns, common nouns, status constructus, indirect genitive with nï, nisbe adjectives, articles (pꜣ/tꜣ/nꜣ), numerals, phrasal coordination with ḥnꜥ |
| Verbal morphology | sḏm.n=f, stative, first-person stative, consecutive (sḏm.ꞽn=f), emphatic forms, imperatives (positive and negative), negated subjunctive with ꞽm, active participles, infinitives, relative forms, causative stems |
| Periphrasis and future | progressive ḥr/m + infinitive, future r + infinitive |
| Passive | tw-passive, passive participles, posterior passive, impersonal tw |
| Lexical verb classes | public verbs, private verbs, suasive verbs, stance verbs; verbs of self-propelled motion; ḫpr; markers of reported speech; wnn; the copula pw |
| Particles and discourse | discourse particles, clause markers, presentative m=k, desiderative particles, consecutive ꞽḫ/kꜣ/ḫr, question particle ꞽn-ꞽw, narrative auxiliaries ꜥḥꜥ.n/wn.ꞽn |
| Clause and information structure | conditional ꞽr, topicalising ꞽr, cleft sentences with ꞽn, temporal subordinators (m-ḫt, ḏr, ḫft, m-ḏr + verb), purpose clauses, attributive adjectives, predicative adjectives |
| Deixis | temporal deictics, local deictics (ꞽm, ꜥꜣ, dy) |
| Lexicon and measures | Late Egyptianisms, loanwords, amplifiers, adverbs, interjections, prepositions, mean word length, lexical diversity (a composite of MATTR, Maas and HD-D) |
One absence stands out. Multi-dimensional analyses usually draw heavily on clause-level syntax: types of subordinate clause, coordination, dependency hierarchies. All of that is out of reach here. The TLA is lemmatised and annotated for morphology, but it supplies no syntactic trees or dependency relations, and extracting syntactic features would require either manual annotation on a scale that is not affordable or a syntactic parser for Egyptian, which does not yet exist. What the inventory captures is the morphological and lexical surface of functional variation, not its syntactic depth.
Five dimensions
How many dimensions should the analysis extract? Factor analysis can produce as many factors as there are features, but each further factor explains less of what the features have in common. How much a factor explains is its eigenvalue. Plotted in order, the eigenvalues form a scree plot, named after the rubble at the foot of a slope: where the steep descent turns into a flat run, further factors add little. A second check, parallel analysis, repeats the calculation on random data of the same size. A factor is worth keeping only if it explains more than the corresponding factor of random data.
Here the curve bends twice: sharply after the second factor, and once more after the fifth, after which it runs flat. By the scree plot alone, three factors would be defensible, and so would five. Parallel analysis sets the upper limit: after the seventh factor the curve falls to the level of random data. The two checks therefore mark a range, from three to seven, not a number. Within that range, Biber’s procedure is to compare the solutions and to choose the one whose factors can all be interpreted plausibly, erring towards more factors rather than fewer;19 here that is five. That the fourth and fifth factors lie between the two bends shows again in the next test: they are the two dimensions that do not hold up under every check.
Whether a dimension holds up is tested by running the whole analysis again with other settings and on other corpora, and comparing each dimension with its counterpart, using the congruence introduced above (known as Tucker’s φ). By a common convention, φ ≥ 0.95 counts as practically identical and φ ≥ 0.85 as similar. The four alternatives are a lower threshold of 0.15 for dropping features, no dropping of features at all, only texts of at least 600 words, and 110 additional texts of 400 to 499 words.
| Dimension | communality 0.15 | no features dropped | only texts ≥ 600 words | plus 110 texts of 400–499 words | verdict |
|---|---|---|---|---|---|
| 1 | 0.98 | 0.99 | 0.95 | 0.97 | stable |
| 2 | 0.95 | 0.95 | 0.85 | 0.89 | stable |
| 3 | 1.00 | 0.98 | 0.98 | 0.96 | stable |
| 4 | 0.56 | 0.59 | 0.32 | 0.68 | tentative |
| 5 | 0.38 | 0.93 | 0.18 | 0.91 | tentative |
Dimensions 1 to 3 come out virtually identical every time. Dimensions 4 and 5 change shape whenever the corpus changes and are therefore labelled tentative. The two right-hand columns also answer the question whether the 500-word threshold makes the dimensions: raising it (191 texts) and lowering it (354 texts) both reproduce the stable core.
Each of the five dimensions is shown below in two panels. The upper panel lists the features that load on the dimension, with their loadings: bars to the right belong to one end of the scale, bars to the left to the other. The lower panel shows where the texts of each group fall on the scale, as a box plot. The box covers the middle half of the texts of a group, the line inside it marks the median, the whiskers reach to the remaining texts, and dots mark single texts far outside the rest.
The names of the dimensions are interpretations, not results of the calculation. Before a dimension was named, it was checked whether something other than a shared communicative function could explain why its features occur together — for instance the subject matter of the texts, or two features that are counted from partly the same forms.
Dimension 1: Vernacular-analytic versus formulaic density
On the positive pole Late Egyptianisms, articles and infinitives reflect the analytic morphosyntax of Late Egyptian. Consecutive particles, the progressive and purpose clauses add the grammar of someone telling what is happening, what follows and what for. Relative pronouns belong here too: in Late Egyptian the analytic relative converter increasingly replaces the synthetic relative forms.
The negative pole is not Classical Egyptian as such but nominal density — common nouns, chains in the status constructus, possessive suffixes, the copula pw. These are the features of texts that describe and classify rather than narrate or instruct: epithet sequences, ingredient lists, titularies. Things are identified and named; nothing unfolds.
This dimension carries a confound that the corpus cannot resolve. Its positive pole is above all Late Egyptian: of the twenty highest-scoring texts, seventeen are Late Egyptian letters and literary texts, while the two Middle Egyptian letters of Heqanakht stay close to zero. Formal royal texts that take up Late Egyptian forms land in the upper middle: the treaty with the Hittites at +10.7, the Kadesh Bulletin at +7.8, the Amarna boundary stelae between +3.7 and +7.5, the Kadesh Poem only at +0.7. The more Late Egyptian a text admits, the higher it scores; but in this corpus the texts that admit most are also the least formal. Whether the pole measures a stage of the language or an informal register therefore cannot be decided — and the dimension shows precisely why the diachronic and the register reading are entangled in the Egyptian record.
Dimension 2: Elaborated reference versus procedural lists
The positive pole belongs to texts about particular people and their relations. Third-person pronouns and possessive suffixes track named individuals and what belongs to them, prepositions encode the relations between them, sḏm.n=f gives narrative sequence, active participles describe the properties of actors, private and stance verbs express internal states and evaluations, and lexical diversity reflects the varied vocabulary of narrative as opposed to the repetition of formulae. The negative pole is dominated by three features: numerals give quantities, predicative adjectives state conditions, the tw-passive gives impersonal instruction — “one does X”. Together they are the grammar of recipes and prescriptions.
Dimension 3: Directive address versus descriptive classification
The positive pole is the grammar of directing someone: speaking to an addressee and telling them what to do or to attend to. The negative pole brings together attributive adjectives, nisbes, heavy nominal structures and nb “every” — the language of epithets, titles and descriptive prose. There is no addressee, only characterisation.
Dimension 4: Generic versus specific reference (tentative)
The negative pole is anchored almost entirely by proper nouns: texts about specific named persons, places or gods. The positive pole, where the medical texts sit, has impersonal passives, predicative adjectives and conditionals — the grammar of texts that treat their subject as a type. “If the wound is X, do Y”: no name, no particular patient, a generic case. Dimension 2 captures how such texts are structured, Dimension 4 what they are about. The dimension rests essentially on a single feature, though, and when the corpus is extended it loses its proper-noun anchor and restructures entirely. Hence the label is tentative.
Dimension 5: Case-based classification (tentative)
This dimension does not contrast two styles; it measures the presence of one construction cluster. The cluster is coherent: “it is X” with pw, “if …” and “as for X …” with ꞽr, emphatic forms and demonstratives are the grammar of reasoning case by case, protasis and apodosis with a copular classification: “if the situation is X, it is Y”. That is why it separates the groups so poorly: it picks out not a register but a rhetorical mode that cuts across registers. It too is sensitive to the composition of the corpus and dissolves under some extensions; its label is tentative as well.
Where texts and parts of texts fall
A first check is whether the groups fall where situational knowledge predicts. They do. Of all groups, the letters have the highest average on Dimension 1; they are also positive on Directive Address, the pragmatic core of letter writing, and negative on Elaborated Reference — letters act rather than describe. Medical texts are their mirror image: the lowest values on Dimension 2 in the corpus, recipe lists instead of connected prose; high on Generic Reference, with unnamed patients and counted substances; negative on Dimension 1, nominal and classificatory rather than verbal and sequential.
Group means, however, also hide structure. The medical group is positive on Dimension 5 on average, but that average conceals a split.
The diagnostic treatises, with their if-then structure of case and verdict, reach up to +22; the pure recipe collections sit at the negative end. “Medical text” as a category covers two quite different ways of organising discourse that happen to share a subject. The same holds across categories: among the texts with the highest scores on Dimension 5 are, besides medical diagnostics, spells of the Book of the Dead, the Dream Book, one of the Heqanakht letters and the Teaching of Ptahhotep. Casuistic framing is a resource available to several domains. Directive address behaves the same way — ritual invocations and letters share it. Linguistic configurations track communicative tasks, not domains.
Within a single text: Ptahhotep
The same thing happens inside a single composition. The Teaching of Ptahhotep was annotated for goal types, the basic functional labels Systemic Functional Linguistics uses for what a stretch of text is doing: argumentation, exposition, instruction and narration.
Instruction peaks on Directive Address, at about +12 — these are the maxims themselves, full of imperatives, jussives and second-person address. Argumentation and exposition carry the profile of Elaborated Reference, argumentation at about +12, exposition at about +8. Exposition is also by far the most formulaic segment on Dimension 1: dense nominal structure combined with elaborated reference, the profile of informationally dense presentation. A genre label such as “teaching” averages over four distinct linguistic configurations. Register variation is a property of stretches of text, not of whole texts.
Narration and dialogue
Splitting narrative texts into narration and dialogue tests the dimensions the same way.
On Directive Address all three texts show the expected contrast: the dialogue scores higher than the narration. In Westcar and the Shipwrecked Sailor the narration sits near zero; in Wenamun (+2.2 against +5.1) it is positive as well. The magnitude differs, though. The dialogue in the Shipwrecked Sailor, above all the serpent’s speeches, reaches +6.9 — among the highest values in the corpus — against +2.7 for Westcar. Westcar’s narrator, on the other hand, builds an elaborated, reference-rich world (+4.1 on Dimension 2) that disappears when the characters speak (−0.1); in the Shipwrecked Sailor both parts stay positive (+3.5 and +2.5). On Dimension 1 the Shipwrecked Sailor is formulaic throughout, and its dialogue more so than its narration (−3.4 against −1.9). One would expect dialogue to pull towards the vernacular pole; here it pushes further into formulaic territory, as if the dialogue itself were highly conventional. Westcar and the Shipwrecked Sailor are both Middle Egyptian narrative literature, and the dimensions still separate them sharply.
The Report of Wenamun behaves differently. On Dimension 1 its narration scores +17.3, its dialogue +17.6 — practically no difference. The whole text is saturated with Late Egyptian syntax; the narrator’s voice uses it as fully as the characters do. In Wenamun this is a property of the text, not of the speech situation.
Separately by language stage
For a comparison by language stage, the texts were not selected by hand but drawn from the TLA automatically: Old and Middle Egyptian by date (Third to Eighth Dynasty; Eleventh Dynasty to before Amarna), Late Egyptian by the TLA’s language tag. That yields 328, 261 and 117 documents of at least 150 words each, many of them shorter than the 500 words of the main corpus: the median length is about 220 words for Old, 300 for Middle and 400 for Late Egyptian. Run separately for each stage, and in a joint analysis of all three, the analysis gives a more sober picture. The check here is a different one: the documents of a stage are divided at random into two halves, the analysis is run on each half, and its dimensions are compared with those of the whole stage. Repeated many times over, only one to three dimensions per stage keep their shape; the others change from one division to the next. The three stages share no common set of dimensions: between stages the congruence reaches at most 0.74, well below the level at which dimensions count as similar. Only 22 features vary at all in all three stages, because much of the new grammar — articles, the progressive, loanwords — only exists from Late Egyptian onwards. What recurs is a basic contrast in each stage between nominal, monumental or enumerative texts and involved, verbal ones.
As a diachronic signal, one narrow strand survives every check: the shift from synthetic to analytic, with the status constructus declining and the indirect genitive, the article, demonstratives and possessive articles rising. It holds within a single group of texts as well — the decorum texts, the only group usable across all three stages. Beyond that, the separation of language stage and register has a principled limit, not only a statistical one: change runs through registers, the register categories themselves are not stable over fifteen hundred years, and what survives from each stage depends on its registers, because stone outlasts papyrus.
What is gained
The obvious objection is that all this was known. Every philologist knows that Papyrus Ebers contains both diagnostics and recipes. In that generality, yes. What was not known is which forms covary, that the diagnostic sections share their configuration with wisdom literature rather than with the recipes of the same papyrus, and that this emerges from a model that never saw the categories. Where the method reproduces what we knew, that is validation, and it is what makes the deviations worth taking seriously.
Knowing that texts differ is also not knowing on how many independent axes, by how much and with what internal spread. A text scored on the dimensions can be placed, whatever its genre label, including fragments and unprovenanced pieces. And Egyptian becomes comparable with the other languages analysed this way, to which it adds a time depth none of them has.
A second objection is that comparing the dimensions against situational expectations brings back the subjectivity that statistics claims to remove. It does. The claim is not that the interpretation is objective, but that the pattern is: anyone may read the same loadings differently, but they have to read the same loadings. Quantification does not remove subjectivity; it makes it inspectable, documentable and revisable. The relevant contrast is not objective against subjective, but checkable against uncheckable subjectivity.
Tagging, and the missing syntax
With a fully lemmatised corpus, Egyptian is comparatively well placed for this kind of analysis. Feature tagging in multi-dimensional analysis was always at least partly automatic, and Biber himself considered a coverage of “90 per cent or better” — after manual editing — good enough.20 For English it is now fully automatic, and fast. One estimate takes Common Crawl, a freely available archive of text collected from the web, at around 1.1 trillion tokens (roughly: words). A tagger based on a neural network, run on eight graphics processors in parallel, would need about 13.6 days to tag it; the established Multidimensional Analysis Tagger, a conventional program running on ordinary processors with the same degree of parallelism, about 756 days, more than two years.21
For Egyptian, neural models for lemmatisation and part-of-speech tagging exist too.22 The table gives the share of words, in per cent, for which their answer is correct, separately by language stage and by the form in which the text is fed in: as transliteration or as a sequence of hieroglyphs. Detailed part of speech (XPOS) is a simplified version of the TLA’s own word classes, which, for instance, tell divine, royal and personal names apart and separate titles and epithets from other nouns. Basic part of speech (UPOS) is a coarse class such as noun or verb from a scheme shared across languages. Lemma is the dictionary entry a word belongs to, down to which of several homonymous entries it is.
| Correct, in per cent | Demotic | Earlier Egyptian, transliteration | Earlier Egyptian, hieroglyphs | Late Egyptian, transliteration | Late Egyptian, hieroglyphs |
|---|---|---|---|---|---|
| Detailed part of speech (XPOS) | 97.1 | 96.2 | 93.0 | 94.0 | 93.1 |
| Basic part of speech (UPOS) | 97.5 | 96.6 | 93.6 | 94.5 | 93.5 |
| Lemma | 92.2 | 87.6 | 80.2 | 80.0 | 76.6 |
What is missing is syntax. Lemmas and word classes can now be assigned by machine with the accuracy shown above; for syntax no such model for Egyptian exists yet, and one can only be trained on texts whose syntax has already been annotated by hand. Wolfgang Schenkel set out to create such material early on. In 1968 he demonstrated M.A.A.T. (Maschinelle Analyse Altägyptischer Texte), a system he had built to analyse spells of the Coffin Texts lexically, morphologically and syntactically, and called its principle “integrated” capture: one machine-readable corpus, annotated once on every level, for every kind of study.23 In 1985 he added what such a corpus was for: the linguistic character of a text was to be judged not by collecting individual Late Egyptianisms and classicisms but by “quantitative estimates, if not statistics”.24 A multi-dimensional analysis is a statistic of this kind. His Coffin Text data, around 400,000 fully syntactically annotated tokens, are now being integrated into the TLA. Two more recent projects work towards the same end on a smaller scale: Carlos Gracia Zamacona’s database of the Coffin Texts, with more than 27,000 records, most of them a simple sentence each, with fields for syntactic analysis,25 and the Egyptian-UJaen Treebank of Roberto A. Díaz Hernández, in its current release some 34,000 tokens from the Pyramid Texts.26 Material of this kind could serve to train models for automatic syntactic annotation — and with that, a complete multi-dimensional analysis would become possible.
The limits that the corpus imposes — few text types, little overlap between the stages, short units — cannot be calculated away. The limit set by the missing syntax can be lifted.
244 texts from the TLA raw data, at least 500 words each, all pre-Demotic stages. 70 features extracted in Python, 43 in the factor analysis; five factors. Robustness checked by alternative feature selection, alternative corpora and split-half comparison. Code and feature definitions are not published; I am glad to share them on request.
Footnotes
Pescuma, Valentina N., Dina Serova, Julia Lukassek et al. 2023. Situating language register across the ages, languages, modalities, and cultural aspects: Evidence from complementary methods. Frontiers in Psychology 13, 1–31.↩︎
Lepsius, Richard. 1837. Sur l’alphabet hiéroglyphique: lettre à monsieur le Prof. Hippolyte Rosellini. Annali dell’Instituto di Corrispondenza Archeologica 9 (1), 5–100, here 70–73. — Erman, Adolf. 1889. Die Sprache des Papyrus Westcar. Abhandlungen der Königlichen Gesellschaft der Wissenschaften zu Göttingen 36. Göttingen. Pp. 4–10. — Sethe, Kurt. 1925. Das Verhältnis zwischen Demotisch und Koptisch und seine Lehren für die Geschichte der ägyptischen Sprache. Zeitschrift der Deutschen Morgenländischen Gesellschaft 79, 290–316, here 301–314. — Erman, Adolf. 1933. Neuaegyptische Grammatik. 2nd edition. Leipzig. — See Damm, Svenja Kristina and Tobias B. Paul. Forthcoming. Register studies in Egyptology: A conceptual history and critical review. Sections 3.1–3.2.↩︎
Halliday, Michael A. K. 1978. Language as Social Semiotic: The Social Interpretation of Language and Meaning. Baltimore.↩︎
Nerlich, Brigitte. 1996. Anthropology, Egyptology, and linguistics: Malinowski and Gardiner on the functions of language. In Vivien Law and Werner Hüllen (eds.), Linguists and Their Diversions: A Festschrift for R. H. Robins on His 75th Birthday, 361–394. Münster.↩︎
Gardiner, Alan H. 1932. The Theory of Speech and Language. Oxford. Pp. 50–51 and § 26. — Bühler, Karl. 1933. Die Axiomatik der Sprachwissenschaften. Kant 38 (1–2), 19–90, here p. 48.↩︎
Gardiner, Alan H. 1962. My Working Years. London. P. 43.↩︎
Damm and Paul, forthcoming, section 3.3.↩︎
Vernus, Pascal. 1982. Diachronie et synchronie dans la langue égyptienne. In L’Égyptologie en 1979: Axes prioritaires de recherches 1, 17–18. Paris. — Vernus, Pascal. 1982. Deux particularités de l’égyptien de tradition: nty ꞽw + présent I; wnn.f ḥr sḏm narratif. Ibid., 81–89.↩︎
Jansen-Winkeln, Karl. 1994. Text und Sprache in der 3. Zwischenzeit: Vorarbeiten zu einer spätmittelägyptischen Grammatik. Ägypten und Altes Testament 26. Wiesbaden.↩︎
Goldwasser, Orly. 1990. On the choice of registers: Studies on the grammar of Papyrus Anastasi I. In Sarah Israelit-Groll (ed.), Studies in Egyptology Presented to Miriam Lichtheim, 200–240. Jerusalem.↩︎
Labov, William. 1963. The social motivation of a sound change. Word 19, 273–309. — Labov, William. 1966. The Social Stratification of English in New York City. Washington.↩︎
Biber, Douglas. 1985. Investigating macroscopic textual variation through multifeature/multidimensional analyses. Linguistics 23 (2), 337–360. — Biber, Douglas. 1988. Variation across Speech and Writing. Cambridge.↩︎
Gillen, Todd J. 2014. Ramesside registers of égyptien de tradition: The Medinet Habu inscriptions. In Eitan Grossman, Stéphane Polis, Andréas Stauder and Jean Winand (eds.), On Forms and Functions: Studies in Ancient Egyptian Grammar. Hamburg. Quotations pp. 44 and 73.↩︎
Polis, Stéphane. 2018. Linguistic variation in Ancient Egyptian: An introduction to the state of the art (with special attention to the community of Deir el-Medina). In Jennifer Cromwell and Eitan Grossman (eds.), Scribal Repertoires in Egypt from the New Kingdom to the Early Islamic Period, 60–88. Oxford. Quotation p. 76.↩︎
Polis, Stéphane. 2018. The scribal repertoire of Amennakhte son of Ipuy: Describing variation across Late Egyptian registers. Ibid., 89–126.↩︎
Damm and Paul, forthcoming, section 4.6.↩︎
Biber, Douglas and Mohamed Hared. 1992. Dimensions of register variation in Somali. Language Variation and Change 4, 41–75.↩︎
“Decorum” designates texts that publicly represent social order, norms and hierarchy; the term follows John Baines’s concept of decorum. Baines, John. 1990. Restricted knowledge, hierarchy, and decorum: Modern perceptions and ancient institutions. Journal of the American Research Center in Egypt 27, 1–23, here pp. 20–21. — Baines, John. 2007. Visual and Written Culture in Ancient Egypt. Oxford. P. 15.↩︎
Biber, Douglas. 1988. Variation across Speech and Writing. Cambridge. Pp. 82–84. — Biber, Douglas. 1995. Dimensions of Register Variation: A Cross-Linguistic Comparison. Cambridge. Pp. 120–121.↩︎
Biber 1988, 217.↩︎
Alkiek, Kenan, Anna Wegmann, Jian Zhu and David Jurgens. 2025. Neurobiber: Fast and interpretable stylistic feature extraction. arXiv.↩︎
Sahala, Aleksi and Eliese-Sophia Lincke. 2025. Neural models for lemmatization and POS-tagging of Earlier and Late Egyptian (supporting hieroglyphic input) and Demotic. In Adam Anderson, Shai Gordin, Bin Li et al. (eds.), Proceedings of the Second Workshop on Ancient Language Processing, 77–82. Stroudsburg, PA. Here Table 3, whole dataset.↩︎
Schenkel, Wolfgang. 1969. Der Computer als Hilfsmittel für die lexikalische und grammatische Beschreibung des Altägyptischen: Möglichkeiten und Grenzen. In Wolfgang Voigt (ed.), XVII. Deutscher Orientalistentag vom 21. bis 27. Juli 1968 in Würzburg: Vorträge 1, 97–105. Wiesbaden.↩︎
Schenkel, Wolfgang. 1989. Sprachforschung und Textquellen: Integrierte Datenverarbeitung als konkrete Utopie. In Sylvia Schoske (ed.), Akten des Vierten Internationalen Ägyptologenkongresses München 1985 3, 1–27. Studien zur Altägyptischen Kultur, Beihefte 3. Hamburg. Here 16; quotation translated.↩︎
Gracia Zamacona, Carlos. 2013. A database for the Coffin Texts. In Stéphane Polis and Jean Winand (eds.), Texts, Languages & Information Technology in Egyptology, 139–155. Aegyptiaca Leodiensia 9. Liège.↩︎
Díaz Hernández, Roberto Antonio and Marco Carlo Passarotti. 2024. Developing the Egyptian-UJaen treebank. In Proceedings of the 22nd Workshop on Treebanks and Linguistic Theories (TLT 2024), 1–10. Hamburg. The size given is that of the current release in Universal Dependencies, UD_Egyptian-PC: 3,089 sentences, 34,234 tokens.↩︎