TL;DR
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Software engineer Jesse Waites says an AI-assisted search across digitized archives surfaced a forgotten meteorite report, records of three rhinos and possible unrecorded volcanic eruptions. He says candidate findings were checked against original document scans, but their historical significance and novelty still require specialist review.
Software engineer Jesse Waites says an AI-assisted search of digitized historical archives surfaced a forgotten meteorite report, records of three rhinoceroses and possible unrecorded volcanic eruptions. The project shows how automated screening can help researchers find candidates in collections too large to read page by page, though the claims depend on examining original records and comparing them with specialist catalogues.
Waites describes the work in a report inspired by historian Benjamin Breen, who had used AI to locate a 1615 Dutch East India Company ship’s journal describing sailors catching tortoises and dodos on Mauritius. Waites sought to extend that approach by searching for several historical questions at once, including reports of animals, meteorites, eruptions, earthquakes and missing ships.
The search covered digitized sources including 4.35 million pages of Dutch East India Company records, Dutch national library newspapers, two centuries of American newspapers and some ship logbooks. Waites says his system represented 5.7 million passages from the company archive as semantic search fingerprints, intended to find relevant material despite spelling variation and errors in text recognition.
Rather than sending every result to a large model, the workflow used a smaller model called Jev to screen passages, followed by Claude Haiku for closer reading, translation and extraction of dates and places. Waites says he then checked candidate passages against scans of the original documents and consulted relevant catalogues before describing findings as new. The source report provides the headline results, but the supplied account does not give full archival citations for each one.
Searching Archives at Scale
The project’s practical significance is its screening method: AI narrows a large volume of imperfectly transcribed material to a smaller set that a researcher can inspect. Waites estimates that reading the Dutch East India Company archive manually at two minutes per page would take about 70 years under a stated work schedule; he says his system processed the archive in a 12-hour overnight run. Those figures are his comparison, not an independently assessed measure of research time saved.
Finding a passage is not the same as establishing a historical discovery. Handwriting recognition can misread words, language models can misinterpret context, and an old record may already be known under a different spelling or catalogue description. Waites’s account describes a process designed to address these risks by checking original scans and existing catalogues. The method may help direct scholarly attention, but the reported results need documentary references and expert assessment before their novelty and interpretation can be independently judged.
The approach also illustrates a division of labor between people and models. Automated tools handled broad retrieval and early triage, while Waites chose questions, adjusted the workflow and examined stronger candidates. That distinction matters: the AI did not independently validate the historical claims, and the report presents the work as human-led research aided by automation.
As an affiliate, we earn on qualifying purchases.
From Dodo Record to Wider Search
Waites says the project began after reading Breen’s account of a newly located 1615 eyewitness description of dodos in Dutch East India Company records. Those documents span the company’s operations from the 17th century into the 1790s. A separate digitization initiative, GLOBALISE, has made millions of handwritten pages available as searchable transcriptions, opening them to searches that would be difficult to conduct manually.
Waites’s report describes thirteen possible research questions selected with help from an AI research assistant. The criteria included the completeness and accessibility of online records, whether a finding could be checked against an original page, and whether the question appeared to have been addressed already. The list included a large eruption in 1808 whose source had not been identified, as well as animal records and historical earthquake reports.
Because historic spellings and machine transcriptions are inconsistent, Waites used semantic search rather than relying only on exact keywords. He says an initial search could return tens of thousands of passages, so the smaller model screened results before a more capable model translated and extracted details from a limited set. His account says the pipeline was built with Claude Code and involved multiple AI tools, but the source material does not provide an independent technical evaluation of the system.
“Adding even one new data point to a couple of niche fields felt like a small but worthwhile contribution to make.”
— Jesse Waites, describing his research goal
historical record digitization software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Finds Are Verified
The supplied report does not include the full document references, transcriptions or images for the meteorite and rhinoceros records, nor enough detail to assess the possible eruption reports independently. It is therefore not possible from this material alone to confirm the exact contents of each record, whether specialists already know of them, or whether the descriptions support Waites’s interpretation.
Waites says he checked original scans and specialist catalogues, but the account provided here does not name those catalogues or document the outcome of each comparison. It also does not describe independent review by historians or subject specialists. The claims should be treated as reported candidate findings pending access to the underlying references and expert scrutiny.
The project’s processing figures and cost estimate are also reported by Waites. No independent benchmark is provided for the accuracy of its transcription search, the models’ screening decisions or the amount of researcher time saved. The report does not establish that every promising result from the wider search has been fully checked.
AI research assistant for archives
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Primary Records and Review
The next step for readers and researchers seeking to assess the claims is access to the archive references and original scans for each reported finding. Those materials would allow specialists to check the reading, date and location, assess the surrounding record, and compare it with prior catalogues and scholarship.
Waites’s account does not specify a publication schedule for fuller citations, a formal peer-review process or planned follow-up work. Until those details are available, the meteorite, rhino and eruption results remain claims reported by the project’s author rather than independently confirmed discoveries. Further reporting should establish which records can be verified and whether they change existing historical accounts.
historical document translation tool
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did the AI-assisted archive search find?
Jesse Waites reports locating a forgotten meteorite report, records involving three rhinoceroses and possible unrecorded volcanic eruptions. The supplied account does not include full citations for each record, so the findings cannot be independently checked from this material alone.
Which historical records did the project search?
Waites says the search included 4.35 million pages of Dutch East India Company records, Dutch national library newspapers, two centuries of American newspapers and selected ship logbooks.
How did the AI system screen the records?
Waites describes a staged process: semantic search identified passages by meaning, a smaller model screened candidates, and Claude Haiku reviewed selected passages for translation and details. He says he then checked candidates against scans of original documents.
Are the reported findings independently confirmed?
The source account says Waites compared candidates with original scans and specialist catalogues, but the supplied material does not provide full archival citations or describe independent specialist review. The findings should be treated as reported results awaiting verification.
Why use AI to search historical archives?
Large archives contain more material than a researcher can read manually, while old spellings and transcription errors make exact keyword searches unreliable. AI-assisted retrieval can narrow the material to passages for human review, but it does not replace checking the original records.
Source: hn
Columbus Day / Indigenous Peoples' Day Picks
long weekend sales
As an affiliate, we earn on qualifying purchases.
