Attestation density in the TLA
A way of plotting when a lemma is attested, given that Egyptian texts are dated to spans rather than to years.
Egyptian texts are almost never dated to a year, but to a reign, a dynasty, a period — to a span from ten to several hundred years wide.
The curves here spread each token evenly across the dating span of its text: a token from a text dated to a hundred years contributes one hundredth to each of those years, a token from a text dated to ten years contributes one tenth. Summed over the corpus this gives a weight for each year, and those weights are the curve.
Reading the plot
The y-axis is the sum of the weights, in tokens per year. The pale line is the raw sum, the bold line the same data smoothed with a Gaussian kernel of 25 years; the distance between them shows how much of the shape comes from the data and how much from the smoothing. The bands in the background are the conventional periods.
In absolute mode the curve shows the lemma’s own weight. Since closely dated texts give off more weight per year than loosely dated ones, and the corpus is unevenly dense across the centuries, it rises and falls with the corpus as much as with the word. In relative mode the same weight is divided by the weight of all lemmatised tokens of the same year.
What is counted
Every token carrying a lemma ID is counted; broken and unidentified words contribute nothing. Alternative readings of a passage are stored as separate files in the TLA and would be counted twice, so only the canonical reading is used. Where a text carries several dating statements, its weight is divided among them.
Two exclusions: texts whose dating carries no year interval in the thesaurus drop out, and lemmata with fewer than three tokens in the corpus are not exported at all.
Comparing two forms
I built this not for the single curve but for the comparison of two. Where a Middle Egyptian form and its Late Egyptian counterpart compete, plotting both in one frame shows whether the one recedes as the other advances, whether they stand side by side for centuries, and where they cross — if they do.
This needs the relative mode (Figure 2). In absolute terms both curves rise and fall with the corpus, and two forms attested in the same texts will look alike however different their histories; only as shares of the same year’s evidence can they be read against each other.
The tool
Attestation density in the TLA takes one or more lemma numbers, separated by commas, and draws them in one frame. Demotic lemmata carry a leading d; since the TLA keeps two lemma lists, linked entries can be merged to follow a word across the change of script.
Built from the TLA raw data, corpus issue v20. The underlying index holds, per lemma, the dating intervals of the texts it occurs in with a token count, aggregated and without text identifiers.