Working with Historical Documents


Historical document research depends on organizing large amounts of handwritten records, archived files, and disconnected information.

How do historians and researchers actually work through decades of documents to find what matters?

In this episode of The AiCR Exchange, Joe Furlong sits down with Dr. Jonathan DeCoster, Kathleen Miller, and Miles Gabrielli-Burke from the University of New England to talk about historical document research, community-based learning, and the challenges of working through archived materials. The conversation breaks down how researchers approach historical records, how students participate in real-world projects, and why organizing information is often one of the hardest parts of the process.

  •  How historians approach document research and archives
  •  The challenges of working with handwritten and historical records
  •  Community-based research projects and student involvement
  •  What historical document workflows actually look like day to day

The AiCR Exchange features conversations with leaders across mortgage and financial services on document workflows and technology decision-making.

Getting text out of a historical document sounds like it should be solved. It is not. Making digitized documents searchable, structured, and usable for research is one of the harder problems in archival work today. 

In Episode 8 of The AiCR Exchange, Joe Furlong sits down with Dr. Jonathan DeCoster, professor of history at the University of New England, Kathleen Miller, librarian and archivist, and Miles Gabrielli-Burke, a history student, to talk about how historical document research actually works, what community-engaged learning looks like in practice, and what happens when you test document extraction tools on 18th-century handwritten court records. 

What kinds of historical documents do researchers actually work with? 

The documents that show up in community historical research are varied and often unexpected. Miles Gabrielli-Burke worked on a digitization project involving hotel menus from 1939, collected by a local historian in Biddeford, Maine, whose family had deep roots in the area. Another project involved photographs of artifacts from the Abbey Museum with minimal existing documentation. The challenge in all of these cases is the same: taking a physical document or image and turning it into something searchable, structured, and usable for research purposes. That gap between a document existing and a document being usable is where most of the work actually happens. 

Why is digitizing historical documents so difficult? 

Taking a picture of a document is straightforward. Making the text extractable, searchable, and usable is a different problem entirely. Jonathan DeCoster describes extracting text from handwritten historical records as one of the hardest things he has ever done. The challenge compounds when you are working with handwriting from hundreds of years ago, inconsistent formatting, abbreviations that require historical knowledge to interpret, and documents that were never designed to be read by a machine. A researcher who wants to identify trends across hundreds of court cases cannot do that by reading each one individually. The text has to be extracted, structured, and organized at scale. 

How do document extraction tools perform on historical records? 

Jonathan DeCoster ran a live comparison on a 1770 handwritten court register using four tools: Apple’s built-in OCR, Adobe Acrobat, Google Drive’s conversion tool, and AiCR. Apple’s tool produced poor results. Adobe Acrobat could not even identify that text was present. Google Drive performed somewhat better but was still limited. AiCR produced substantially better results, successfully extracting names, identifying plaintiff and defendant roles even when they appeared only as abbreviations, and recognizing town names. It was not perfect, and Jonathan noted the corrections he had to make by comparing against the original image. But the difference in output quality was significant, and the ability to process hundreds of records rather than one at a time is where the research value becomes real. A separate document, a textile mill store register that was eight gigabytes and 360 pages of dense handwritten entries, proved too challenging for any tool to handle cleanly. 

How does AiCR handle hallucinations in document Q&A? 

AiCR’s document Q&A feature allows researchers to ask questions directly about the content of their documents rather than reviewing everything manually. The system is designed so that if it does not know the answer based on what is in the document, it says so rather than generating an answer from outside sources. The knowledge base is constrained to the extracted document content. Confidence scoring gives users transparency about how certain the system is when it does return a value. Joe Furlong describes the design principle as preferring a false negative, where the system says it does not know, over a hallucination where the system invents an answer. For research applications where accuracy and source integrity matter, that distinction is critical. 

What is the librarian and archivist perspective on AI tools? 

Kathleen Miller, librarian and archivist at the University of New England, describes her current position as skeptical but open. Her concerns center on how AI tools are being deployed broadly, questions about intellectual property, and the reliability of outputs that still require significant human review and correction. She has used AI tools to help transcribe oral history recordings and found them marginally useful, getting her closer to a usable transcript but requiring careful fact-checking throughout. When she ran the hotel menus from the digitization project through AiCR, she found the results genuinely impressive for typed text, with only minor formatting adjustments needed. Her view is that document-specific AI tools with constrained knowledge bases and transparent confidence scoring are meaningfully different from broad large language model applications, and represent a more defensible approach for archival and research work. 

How are professors addressing AI use in academic work? 

Jonathan DeCoster addresses AI in academic work by designing projects that require students to do things AI cannot do. Community-based research that involves physical document handling, interviews with community members, and database construction shaped by real research decisions is inherently resistant to AI shortcuts. The real learning in historical research is not in transcription. It is in reading, thinking, talking to people, understanding who the audience is, and deciding what the output needs to accomplish. AI tools that speed up the mechanical parts of that work free students to focus on the parts that actually require judgment. 

Frequently Asked Questions About Historical Document Research 

What is the difference between digitization and making documents searchable? 

Digitization is the process of converting a physical document into a digital image. Making a document searchable requires extracting the text from that image so it can be queried, analyzed, and used in research databases. Many historical documents exist only as digital images without any extracted text, which means researchers still have to read them manually. Extracting text at scale is what enables trend analysis and large-scale historical research. 

Why is extracting text from handwritten historical documents so hard? 

Historical handwriting does not follow the consistent letterforms that OCR tools are trained on. Older documents use different abbreviations, formatting conventions, and vocabulary. The physical condition of the document, ink fading, paper degradation, and scanning quality all affect how clearly text can be read. Most standard OCR tools fail entirely on 18th-century handwritten records because they were not designed for that document type. 

How does AiCR handle document Q&A without hallucinating? 

AiCR’s document Q&A feature constrains the system’s knowledge base to the extracted content of the documents being processed. It does not draw on outside sources to generate answers. When the system is uncertain or the information is not in the document, it returns a low confidence score or indicates it does not know rather than generating an answer from general knowledge. This design makes the Q&A output traceable back to the source document. 

About The AiCR Exchange

The AiCR Exchange is a live conversation series hosted by Joe Furlong. New episodes air live on LinkedIn on the second and fourth Tuesday of each month at 12pm ET. Follow AiCR on LinkedIn to catch episodes as they air and join the conversation.

About Dr. Jonathan DeCoster 

Dr. Jonathan DeCoster is a professor of history at the University of New England, where he teaches and researches using a community-engaged model that connects students with real archival projects and community partners. He can be connected with on LinkedIn

About Kathleen Miller 

Kathleen Miller is a librarian and archivist at the University of New England. She can be connected with on LinkedIn