Researchers have read an entire rolled Herculaneum papyrus scroll end to end using AI, recovering Greek text from layers of papyrus carbonized by the eruption of Mount Vesuvius in 79 CE. The scroll, catalogued as PHerc. 1667, was digitally unrolled and decoded without any physical contact, according to a preprint detailing the work. The result builds on years of incremental progress, from early X-ray imaging experiments to a global competition launched by Brent Seales at the University of Kentucky, and it raises a pointed question: what kinds of lost ancient writing will emerge as these tools scale to hundreds of unopened scrolls still stored in Naples?
From single letters to full columns of Greek text
The path to reading a complete scroll without opening it stretches back at least a decade. A 2016 peer-reviewed study demonstrated that X-ray phase-contrast tomography and virtual unrolling could reveal short Greek sequences inside rolled Herculaneum papyri. Those early results proved the concept but yielded only fragments, not sustained readable passages, and the ink often appeared as faint contrasts that were difficult to distinguish from the papyrus structure itself.
The breakthrough that made longer recovery possible came from a training method described in the EduceLab-Scrolls dataset. Researchers created labels for their machine-learning models by aligning spectral photography of opened fragments with CT scans of the same pieces. That alignment taught the system to distinguish carbon-based ink from carbonized papyrus in CT data alone, a difference invisible to the naked eye and barely detectable even with conventional imaging. Once trained, the pipeline could be pointed at sealed scrolls where no spectral photography was possible, because the text remained buried inside dense, warped layers.
Seales, a computer scientist at the University of Kentucky, then opened the work to the world. He launched a global competition tied to demonstrated AI letter extraction from X-ray images, drawing participants from machine learning, classics, and imaging science. According to a Nature news report, contest participants obtained the first clearly readable letters from inside an unopened Herculaneum scroll. That contest accelerated development, produced new segmentation and ink-detection algorithms, and attracted talent that might never have encountered papyrology otherwise.
The latest step, described in a preprint on arXiv, claims that PHerc. 1667 is the first scroll fully digitally unrolled and read for extended study without physical opening. The authors report reconstructing continuous columns of Greek prose across the length of the roll, mapping the warped internal layers into a flat representation and then using trained models to highlight ink regions. That claim carries a caveat: both the earlier contest results and the 2016 X-ray work also described reading text from sealed scrolls, though at far smaller scales. The distinction rests on scope. Earlier efforts recovered isolated words or short sequences. The new work describes enough continuous text to treat PHerc. 1667 as a coherent document rather than a puzzle of scattered letters.
Technically, the pipeline combines several steps. High-resolution CT scans capture the internal geometry of the scroll. Segmentation algorithms identify individual layers, despite tears, folds, and compressions caused by the eruption and subsequent handling. A “virtual unrolling” stage then flattens those layers into two-dimensional sheets while trying to preserve the relative positions of fibers and ink traces. Finally, the trained ink-detection network produces probability maps that human readers and philologists can inspect, tracing letters and words much as they would on a fragile, partially burned manuscript under a microscope.
Why philosophy may dominate what the scrolls reveal
The Herculaneum collection is not a random cross-section of ancient literature. The scrolls come from a single villa, widely identified by scholars as belonging to the family of Lucius Calpurnius Piso Caesoninus, Julius Caesar’s father-in-law. The villa’s library, based on scrolls already opened over the past two centuries, skews heavily toward Epicurean philosophy, especially works by Philodemus of Gadara. That pattern matters for predicting what AI will find next, because the unopened rolls almost certainly reflect the same collecting habits as the ones already studied.
If the same CT-to-ink mapping pipeline is applied to the remaining unopened rolls held in Naples, the rate of new philosophical fragments recovered will likely exceed that of poetic or epic ones by a wide margin. The existing opened scrolls from this collection contain philosophical prose at a ratio that dwarfs poetry, drama, or epic. Nothing in the digital reading process changes the content of the scrolls themselves. The AI reads what is there, and what is there, based on every physical opening so far, is overwhelmingly Epicurean treatises on ethics, theology, and the nature of the cosmos.
That does not rule out surprises. Scholars have long speculated that a second, lower level of the villa may contain Latin literary works, including lost texts by major Roman authors. No excavation has reached that level, and the scrolls now under study all come from the already explored upper structures. Even within the known library, there remains room for novelty: previously unknown works by Philodemus, variant versions of familiar treatises, or texts by other Hellenistic philosophers whose writings survive only in fragments. But among the scrolls already recovered and not yet opened, the statistical expectation points toward more philosophy, not lost epics. The headline promise of recovering unknown Homer or Sappho is real in principle but unlikely to be fulfilled by this particular collection.
For historians of philosophy, that skew is a feature, not a bug. Each newly legible column can clarify how Epicurean thinkers argued about pleasure, friendship, political life, and the gods. For papyrologists and textual critics, the digital workflow also offers a new window into scribal practices, corrections, and marginal notes, which can be preserved in situ rather than scraped away or lost during physical unrolling.
Gaps in verification and what to watch next
Several significant questions remain open. No full transcription of the recovered text from PHerc. 1667 has been published in a peer-reviewed journal. The preprint describes the achievement, outlines the imaging and machine-learning methods, and presents sample passages, but it does not release the complete Greek text or identify the author of the work with certainty. Without that information, scholars cannot independently assess how much new literary content the scroll actually contains or whether it duplicates texts already known from other sources.
There is also no public dataset release or peer-reviewed follow-up confirming the extent of readable text described in initial news coverage. The ink-detection accuracy metrics that would let outside researchers evaluate the pipeline’s reliability remain limited to what appears in preprint summaries rather than full methodological disclosures. Key details, such as false-positive and false-negative rates on independent test fragments, will matter when classicists decide how confidently to reconstruct damaged words and lines.
The competing “first” claims add another layer of ambiguity. The contest work, the earlier X-ray experiments, and the new PHerc. 1667 study all describe reading text from sealed scrolls, but they emphasize different thresholds: first letters, first connected words, first substantial passage, first complete scroll. Until a consensus emerges on how to define “read,” public communication will likely continue to feature overlapping superlatives that can obscure the continuity between these efforts.
Over the next few years, the key developments to watch will be less about single headlines and more about infrastructure. Peer-reviewed editions of digitally read scrolls, with side-by-side images, ink probability maps, and transcriptions, will allow specialists to test readings and propose alternatives. Open or at least accessible imaging datasets will let independent teams try rival algorithms on the same material, clarifying which methods are robust and which are brittle. And as more scrolls are processed, patterns in handwriting, dialect, and library organization may emerge that were invisible when only a handful of damaged texts were available.
For now, PHerc. 1667 stands as a proof of concept at scale: a demonstration that AI-assisted imaging can move beyond isolated letters to sustained ancient argument. The technical details and philological interpretations still require scrutiny, and the balance between excitement and caution remains delicate. But each additional column of legible Greek pulled from a lump of carbonized papyrus shifts the boundary between what seemed irretrievably lost and what can once again be read, studied, and debated.
More from Morning Overview
*This article was researched with the help of AI, with human editors creating the final content.