Tobias B. Paul

Egyptology and ancient Sudan, corpus linguistics and digital humanities

I am a doctoral researcher at Humboldt-Universität zu Berlin, working in project B03 Register Variation and Asymmetric Communication in Ancient Egypt within the Collaborative Research Centre 1412 Register.

My dissertation asks what governs the mixing of classical Middle Egyptian and Late Egyptian forms in texts of the New Kingdom — whether the situation a text was written for determines which form a scribe chose. I test this quantitatively on some two thousand texts from the Thesaurus Linguae Aegyptiae, annotated for their communicative setting and modelled with Bayesian methods. More about the project →

Beyond it I work on what quantitative measures can and cannot tell us about short ancient texts, and on the conceptual history of register in Egyptology. A second line of work belongs to Sudan archaeology: I have argued that the Meroitic war elephant is a modern projection rather than a historical fact. I am also executive secretary of the Sudan Archaeological Society in Berlin.

Background

I took a BA in Cultural Studies at Universität Leipzig (2009) and an MA in Archaeology and Cultural History of Northeast Africa at Humboldt-Universität zu Berlin (2024), having studied Historical Linguistics alongside that subject in Berlin beforehand. I worked as a student assistant at the Institute of Archaeology there, and have been a research associate in CRC 1412 since 2024.

That path is why the dissertation sits where it does: historical linguistics supplies the question, Egyptology the material. Sudan archaeology is my second field, and cultural studies the perspective I brought to both — how forms, images and practices travelled in ancient Northeast Africa, and what became of them on arrival, is a question I keep returning to, most recently in the case of the elephant, whose Hellenistic imagery Meroe took up and turned to ends of its own.

Languages and material — Egyptian across its language stages, from Old Egyptian through to Coptic; Ancient Greek; the archaeology, history and written culture of ancient Egypt and Sudan; the Thesaurus Linguae Aegyptiae and its data model.

Methods and tools — corpus linguistics and quantitative text analysis, annotation design, reproducible pipelines under version control. I prepare and analyse corpus data in Python and R, and am currently working my way into Bayesian modelling with Stan.

Digital humanities — most of my work begins by making a resource queryable: modelling philological data for databases, retrieving and parsing it, and building the tools that turn an edition or a corpus into something one can ask questions of. I also contribute to open-source software for working with hieroglyphic text.

Glossing and encoding — I gloss fluently and as a matter of course. Interlinear glossing is a basic discipline of linguistic work rather than an ornament: it commits the analysis to paper, where it can be checked. For Earlier Egyptian I follow the glossing conventions developed at HU Berlin. Beside the editor I keep a glossing bench that applies them: it lines the morphs up from the segmentation, sets the small caps by rule, checks the glosses against the attested paradigms, and copies the result into Word, LaTeX or a plain file.

I set hieroglyphs in Unicode rather than as images or transliteration substitutes, and would encourage anyone working with Egyptian to do likewise: encoded text can be searched, quoted and reused; a picture of a text cannot. To make that easier I keep a fork of Mark-Jan Nederhof’s hieroglyphic editor that runs in the browser, with keyboard entry, sign variants and rearrangement by mouse — no installation needed. For Meroitic there is Merotype, which turns transliteration into encoded text and back. All three are alpha versions, so expect rough edges: bug reports and suggestions for features are very welcome.

Public communication — before returning to research I spent years in newspaper production for Axel Springer, latterly leading teams in editorial production, reader services and social media. Writing and editing for readers outside the field is familiar ground, and so is typography, which comes from the same years.

Technical fluency — I looked after the department’s technical administration as a student assistant, and colleagues still tend to arrive at my desk when something has stopped working. I build and run websites, from the design to the server they sit on. For my own research I have built an environment fitted to the work rather than the other way round, down to the operating system: where a program does not do what the work requires, I rewrite it, and where the system gets in the way, I reconfigure it. That is not tinkering for its own sake: it raises what one person can get done, to the point where a corpus study on the scale of my dissertation remains a one-person undertaking instead of needing a team.

Agentic coding — I work with coding agents daily across the whole of my research: building and auditing corpus pipelines, keeping data, code and manuscript in step, and writing the guardrails that make such collaboration dependable rather than merely fast. The useful part is often not the code but the questioning: a model that has the chapter and the data in view can point to a gap in an argument, ask what a category is meant to do, or make me say plainly what I had left vague. It is also easily led, which is why the discipline around it matters more than the prompting — what an agent may touch, what has to be verified before it counts, and how mistakes are caught rather than committed. I am glad to advise colleagues who want to bring these tools into their own work.