💻 AI

It Cost $1.8 Million to Read One Burned Scroll. There Are 600 More in Storage and Possibly Thousands Underground.

AI and synchrotron X-rays just produced the first-ever complete non-invasive reading of a scroll carbonized by Mount Vesuvius in 79 AD. Twenty-two columns of lost Stoic philosophy, recovered for roughly $900 per character. Scanning throughput, not algorithmic capability, now determines how fast we recover what may be the largest cache of lost ancient literature on Earth.

A carbonized Herculaneum scroll next to its AI-recovered text, with synchrotron light beams illuminating the ancient papyrus layers

On June 25, 2026, researchers at a press conference in Naples announced something that had been considered impossible three years earlier: a complete, non-invasive reading of a sealed Herculaneum scroll. PHerc. 1667, a carbonized lump of papyrus assigned a readability score of zero after a failed 1980s opening attempt, yielded 22 columns of ancient Greek text across nearly five feet of papyrus. Its content appears to be a Stoic philosophical treatise, possibly written by Chrysippus, one of antiquity's most prolific thinkers, whose 705 works were thought entirely lost.

Since launching in March 2023, a Silicon Valley-funded competition called the Vesuvius Challenge has awarded $1.8 million in prizes to teams that used AI and synchrotron X-ray imaging to read text no human eye had seen for 1,947 years. But prize money tells only part of the economic story. What matters more is the per-character cost curve, how fast remaining scrolls can be processed, and what fraction of the largest surviving ancient library might still sit under a modern Italian town.

From One Word to One Scroll

Results have accelerated dramatically, following a pattern familiar to anyone who has watched an ML pipeline mature.

Date Milestone Characters recovered Cumulative investment
March 2023 Challenge launched; CT scans released to public 0 ~$1M (initial prize pool + scanning)
October 2023 Luke Farritor deciphers first word ("porphyras") ~9 ~$1.04M
February 2024 Grand Prize: 15 columns, 4 passages of 140+ chars ~2,000 ~$1.74M
June 2026 Complete scroll: 22 columns, ~1.5m continuous text ~7,000 (estimated) ~$1.8M (prizes) + scanning costs

Notice the shape. Seven months produced a single word, then four months produced roughly 2,000 characters across four passages, and 28 months after that a complete scroll of approximately 7,000 characters emerged from a lump of carbonized papyrus that a 1980s team had given up on entirely. Not exponential growth. What it looks like instead is a fundamental research problem getting solved, after which the remaining work becomes engineering: refining ML models, improving virtual unwrapping geometry, building papyrological validation pipelines that let classicists confirm what the algorithms surface.

Proving feasibility was the hardest part. Luke Farritor, a 21-year-old University of Nebraska student, earned $40,000 for detecting "porphyras" by training an ML model on a texture pattern that another contestant, Casey Handmer, had identified in CT scans, a pattern so faint that researchers had stared at similar data for years without recognizing it as ink. That texture, dubbed "crackle," turned out to be the signature of carbon ink on carbon papyrus. No imaging technique had detected it in decades of prior attempts.

A Cost-Per-Character Calculation Nobody Has Published

Here is an analysis that, as far as we can find, does not appear in any existing coverage. It matters because it determines whether reading the rest of the library takes decades or years.

Vesuvius Challenge prizes to date: $1.8 million (Reuters, June 2026).

Synchrotron scanning costs: Beam time at facilities like Diamond Light Source (Didcot, UK) and the European Synchrotron Radiation Facility (Grenoble) runs roughly $2,000 to $5,000 per hour for commercial users. Academic rates are lower, often subsidized. A full high-resolution CT scan of one scroll takes 24 to 48 hours of beam time, putting scanning costs at $50,000 to $250,000 per scroll. Forty-five scrolls and fragments scanned since the project began suggest total scanning costs of $2.25 million to $11.25 million.

Conservative total investment through June 2026: $4 million to $13 million, covering prizes, scanning, computational infrastructure, and personnel not captured in prize payouts.

Total characters recovered across all scrolls: Approximately 9,000+, combining the Grand Prize's ~2,000 characters, PHerc. 1667's estimated ~7,000 characters, and smaller recoveries from other scanned fragments.

Divide through and you get a cost-per-recovered-character range of roughly $440 to $1,440, with a midpoint around $900, which is either absurdly expensive if you think of it as printing costs or absurdly cheap if you think of it as resurrecting the intellectual output of a civilization that was buried under sixty feet of volcanic rock and written off for two millennia.

For comparison: conserving and digitizing the Dead Sea Scrolls from 2009 to 2016 cost approximately $10 million to image roughly 25,000 fragments containing an estimated 2.5 million characters, or about $4 per character, but those scrolls were physically accessible and could be read with magnification and infrared imaging even in their damaged, fragile state. Herculaneum scrolls cannot be opened. Touch them wrong and they crumble.

Nine hundred dollars per character buys you something no amount of money could buy before 2023: text from a scroll that was considered permanently, irrecoverably lost, where every character recovered is a character that every papyrologist alive had written off after the 1980s attempt at physical opening yielded literally nothing.

Marginal cost is collapsing fast. Grand Prize models and virtual unwrapping algorithms are now open-sourced, meaning that once a scroll is scanned, running the ink-detection pipeline costs GPU-hours rather than millions of dollars, a shift from capital expenditure to commodity compute that mirrors how every successful AI application scales after the initial research investment pays off. Reading the first scroll cost $1.8 million in prizes. A new $1 million prize now targets any second complete scroll. By the tenth, per-scroll cost will be dominated almost entirely by scanning.

Why Scanning, Not AI, Is Now the Bottleneck

Forty-five scrolls and fragments have been scanned in roughly three years. Six hundred or more remain in storage at the Biblioteca Nazionale di Napoli. At about 15 scrolls per year, clearing the backlog would take 40 years.

This is not a fundamental constraint but a logistics problem, because approximately 70 synchrotron radiation facilities operate worldwide and Diamond Light Source and ESRF are only the two that have participated so far. A coordinated scanning campaign across even five facilities, running dedicated beam-time allocations, could process 50+ scrolls per year. Twelve years. Done.

But photon throughput is not actually the binding limitation; institutional coordination is, because Italian cultural heritage authorities enforce strict lending and handling protocols for artifacts this fragile, and moving carbonized papyrus internationally requires climate-controlled transport, conservation oversight, bilateral agreements, and the kind of insurance policy that makes underwriters nervous. Each scroll is a one-of-a-kind artifact. Drop one, and that conversation is over forever.

An alternative: bring a portable CT scanner to the scrolls. Micro-CT systems capable of ~7-micron resolution exist but run slower than synchrotron beamlines. What they sacrifice in speed they recover in deployability. A dedicated scanning station installed at the Biblioteca Nazionale could run year-round without moving a single scroll across a border.

Lost Works: Running the Numbers on What Chrysippus Left Behind

Here the math gets staggering.

Herculaneum's Villa of the Papyri contained the only intact library to survive from the ancient Greco-Roman world. Roughly 1,800 scrolls and fragments have been catalogued from 18th-century Bourbon-era excavations. Most were Epicurean philosophical texts by Philodemus of Gadara, a first-century BC philosopher who lived in the villa. But discovering a Stoic treatise in PHerc. 1667, possibly by Chrysippus, reveals the collection was more diverse than previously assumed.

Consider the stakes, because they are unlike anything else in contemporary scholarship. Chrysippus of Soli (c. 279-206 BC) co-founded Stoic logic, wrote 705 works according to Diogenes Laertius, was reportedly the third most-quoted author in antiquity after Homer and Euripides, and yet before the Vesuvius Challenge exactly zero of his complete works survived; everything scholars knew of Chrysippus came from secondhand quotations by later authors, many of whom disagreed with him and had every incentive to misrepresent his arguments.

If even 5% of 1,800 catalogued scrolls contain Stoic philosophy rather than Epicurean texts, that is roughly 90 scrolls. At an average of 7,000 characters per scroll (based on PHerc. 1667), that yields 630,000 characters of potentially lost Stoic philosophy. For scale, Plato's complete surviving works total roughly 600,000 words. One ancient library, largely unread, could contain a comparable volume of material from a tradition we currently know almost entirely through hostile summaries.

What Remains Underground

No one has fully excavated the Villa of the Papyri, a fact that bears repeating because it means the library already catalogued may represent only a fraction of what exists. Bourbon-era tunnels reached the library level but left vast areas unexplored, and the site, sitting beneath modern Ercolano, faces repeated postponement of deeper excavation due to cost, political complexity, and structural risk to buildings above.

Archaeological surveys suggest the villa extended significantly beyond what those tunnels reached, and some scholars believe the known scroll collection came from a single room while additional storage areas remain sealed under volcanic debris, a hypothesis supported by the sheer scale of the structure whose peristyle, at roughly 100 meters long, ranks among the largest ever found in the Roman world.

If additional rooms with scroll storage exist, the total library could substantially exceed 1,800 catalogued items. Estimates range from speculative to disciplined, but the possibility of thousands of additional scrolls is taken seriously by archaeologists familiar with the site. Italy allocated funds for a limited excavation campaign beginning in 2024, though progress has been slow.

Limitations

Several caveats apply to an optimistic reading of these numbers, and the most important is that different scrolls exhibit very different preservation states, meaning PHerc. 1667 may have been among the most readable while some scrolls are so severely carbonized that even synchrotron imaging cannot distinguish layers from each other, and the ML models trained on one scroll's ink signature may fail entirely on scrolls with different ink compositions since some ancient inks contained metallic compounds detectable by X-ray fluorescence while others were pure carbon, virtually indistinguishable from the carbonized papyrus substrate surrounding them.

Second, the Chrysippus attribution is preliminary, requiring papyrological identification that involves reading and contextualizing recovered text, comparing it against known fragments from secondary sources, and achieving scholarly consensus across multiple institutions, all based on a preprint not yet peer-reviewed.

Third, character counts are estimates. No exact figures for the complete reading have been published, and our estimate of ~7,000 characters for 22 columns derives from typical column density for ancient Greek philosophical texts on papyrus rolls, which average 30 to 40 lines per column with 15 to 20 characters per line.

Against Optimism: Selection Bias

One strong counterargument deserves full-strength presentation: selection bias may mean the easiest scrolls got scanned first, since challenge organizers chose scrolls with relatively intact external geometry where internal layers were more likely separated enough for virtual unwrapping algorithms to trace, while remaining scrolls include many that are crushed, fragmentary, or geometrically distorted beyond what current algorithms handle.

If tractable scrolls represent only 10% of the collection, the project's impact on recovering lost ancient literature could plateau well before all 1,800 are processed, and an exponential improvement narrative would give way to a long tail of diminishing returns where each successive scroll demands new algorithmic approaches for increasingly degraded material.

Real concern. But it mirrors every ML application in early deployment. Stopping is not the answer. Every degraded scroll that partially yields text becomes training data for next-generation models, and progress compounds even when individual scroll quality declines, a dynamic that has played out identically in medical imaging, satellite reconnaissance, and autonomous driving.

What You Can Do

If you are a machine learning researcher: all Vesuvius Challenge data, code, and models are open-sourced at scrollprize.org. A new $1 million prize awaits the first team to produce a complete reading of any other scroll. Segment layers in 3D CT data, unwrap them into 2D surfaces, detect ink signatures with trained classifiers. All you need is a good GPU and an internet connection.

If you work at a synchrotron facility: contact EduceLab at the University of Kentucky about dedicated beam-time allocations. Scanning throughput is the binding constraint, not algorithmic capability. Every additional facility that participates compresses the timeline by years.

If you work in cultural heritage policy: excavating the Villa of the Papyri may be the single most consequential archaeological decision available to any government on Earth right now. Sealed beneath volcanic debris could be the only surviving copies of works by Aristotle, Chrysippus, Epicurus, and other foundational thinkers known today only through fragments and hostile summaries. Excavation costs sit in the tens of millions of euros. What may be recovered has no price.

Bottom Line

For $1.8 million in prize money and perhaps $10 million in total investment, a competition launched by two Silicon Valley investors and a University of Kentucky computer scientist recovered the first complete text from a library sealed by a volcanic eruption in 79 AD, achieving a per-character cost around $900 that is falling rapidly as tools are open-sourced and models improve. AI capability is no longer the bottleneck. Scanning throughput is. Six hundred scrolls sit in a Naples library, possibly thousands more under an Italian hillside, and the technology to read them exists right now, meaning the difference between a 12-year timeline and a 40-year one comes down to how many beam-time hours the world decides to allocate to the only surviving ancient library on the planet.