A large language model named Apollo is now filling in the gaps left by torn, faded and burned ancient Greek writing, giving papyrologists a tool that can propose missing words in texts up to 2,000 years old. Built by the Austrian Academy of Sciences with the AI company Mistral AI and the consultancy Sail Reply, Apollo was trained on roughly 600 million words of historical Greek and is free for anyone to use online. The academy calls it the first large language model built specifically for Ancient Greek, and its backers are already pitching it as a way to make a backlog of undeciphered antiquity move faster than any team of scholars could manage alone.
The project’s public face is Anna Dolganov, a papyrologist at the Austrian Archaeological Institute, who has spent years reading damaged inscriptions the slow way, letter by letter, guessing at what erosion or fire took away. Papyrology has traditionally worked that way for more than a century: a specialist compares a torn fragment against a database of known formulas and idioms, then proposes a restoration that other scholars can accept, dispute or refine in print, sometimes decades later.
Apollo’s training set folds in more than the raw word count suggests. Beyond the 600 million words of running Greek text, the academy says the model absorbed tens of thousands of published inscriptions and papyri, the same primary material specialists already cite in critical editions, giving its suggestions a grounding in real epigraphic and papyrological convention rather than generic modern Greek.
An 80 percent hit rate on deliberately hidden text
To test Apollo before release, the team masked passages in known texts and asked the model to guess what belonged in the gaps. According to HeritageDaily’s account of the launch, Apollo achieved a hit rate of around 80 percent on those deliberately concealed passages. In a separate round of expert review, human papyrologists judged Apollo’s proposed reconstructions to be at least as good as human-produced versions in 77 percent of cases — and in a meaningful share of those, rated the machine’s version better.
That is a strange thing to sit with for a field built on the assumption that only trained specialists can responsibly guess at a dead language’s missing syllables.
Dolganov calls it inspiration, not a shortcut
Heinz Faßmann, president of the Austrian Academy of Sciences, framed the project as a deliberate pairing of old and new: “Ancient languages and artificial intelligence are not a contradiction.” Dolganov, who oversaw the model’s historical accuracy, put the payoff in more personal terms, saying “this LLM marks the beginning of an exciting journey in the study of antiquity” and that Apollo “doesn’t just accelerate work, it provides genuine inspiration” by surfacing readings a human eye might skip past.
Guillaume Lample, co-founder and chief scientist of Mistral AI, supplied the underlying model architecture, while Sail Reply, a division of the Reply Group, built the surrounding infrastructure. All three partners describe the setup as running on secure, European-based servers, a detail the academy has repeated in announcing the collaboration to European researchers wary of sending fragile source material to outside servers.
Roughly a million papyri are still waiting to be read
The scale of the backlog is what makes the 600-million-word training run matter. HeritageDaily’s report puts the number of Greek papyri worldwide still awaiting decipherment at 92 percent of everything known to survive, a mountain of scorched Herculaneum scrolls, tax receipts and private letters that no papyrology department has the staff to work through by hand. Apollo does not claim to read any of it outright; it proposes candidate reconstructions that a human still has to check against grammar, context and, where one survives, a parallel text.
The model cost roughly €400,000 to build, a modest sum against the scale of the collection it is meant to help unlock. It is now hosted at apollo.vbc.ac.at, where the academy has made it freely available to researchers and the public rather than licensing it out. Beyond filling gaps, the tool is also built to run semantic searches across its Greek corpus, letting a researcher look for a phrase’s meaning rather than its exact spelling, and to help with handwritten manuscript passages that standard optical character recognition tends to mangle.
A Latin model is next, at twenty times the data
The team is already planning a successor built for Latin, which according to Reply’s announcement of the partnership would train on approximately 12 billion words, roughly twenty times the corpus behind Apollo. That gap reflects how much more Latin material survives compared with Greek, from legal texts to graffiti, and how much of it has never been digitized in a form a language model could learn from.
For now, Apollo’s real test is ordinary use rather than another benchmark. The Austrian National Library holds tens of thousands of papyri in its own collection, one of the largest of its kind anywhere, and it is papyrologists working through that specific archive who will decide, edition by edition, whether Apollo’s proposed readings earn a place in print.
This article was produced with the assistance of AI and reviewed by Morning Overview editors prior to publication.
More from Morning Overview
- Amazon’s Prime refunds are rising to $200 as millions more customers become eligible
- A handful of car engines are so tough mechanics say they almost never wear out
- The NSA is again telling phone owners to switch off one location setting
- Supplements now rank as the fifth-leading cause of death from liver disease.