Historical Document and Media Processing
Summary
This course introduces historical document processing, focusing on concepts and methods that enable the transformation of digitised materials into structured and searchable data. Grounded in machine learning and document processing, it also covers data curation and copyright considerations.
Content
Large-scale digitisation has created extensive collections of digitised historical documents. Beyond their obvious value for preservation and access, these digitised sources open new possibilities for the automatic processing and analysis of their textual and visual content. How to extract and link the complex multimodal information enclosed in digitized historical documents - especially historical newspapers?
The course follows the main stages of historical document and media processing, from digitisation and text acquisition to information access and computational analysis. Topics include text acquisition from digitised documents through optical layout recognition (OLR) and optical character recognition (OCR); OCR quality assessment, improvement, and data standards; text representation and information retrieval; computational methods for exploring and structuring collections, including document clustering and topic modelling; information extraction, with a focus on historical named entity processing; and data annotation and dataset creation as foundations for training and evaluating machine learning systems.
These approaches are grounded in the core concepts of machine learning and information retrieval that underpin them. Modern AI is addressed throughout the course as part of this evolving toolbox. In addition to technical methods, the course introduces standards and best practices for data preparation, as well as copyright considerations.
Finally, beyond technical aspects, the course situates historical information extraction within the broader contexts of digital scholarship and the cultural heritage ecosystem.
By the end of the course, students will have acquired a solid understanding about how a historical collection can be prepared, processed, explored, and made accessible for a specific research or public-use need.
Outline (tentative)
- Part 1- Collections, Research Needs, and Processing Pipelines (two classes): historical media and cultural heritage collections; digitisation; the historical document processing pipeline; links between research and access needs, information retrieval and information extraction tasks, and machine-learning approaches.
- Part 2 - Text Acquisition, Representation, and Information Retrieval (three classes): OLR and OCR; quality assessment and improvement; OCR tools, formats, and standards; text representation and information retrieval.
- Part 3 - Collection Analysis, Information Extraction, and Data Annotation (three classes): document clustering and topic modelling; historical named entity processing and contextualised representations; data annotation, dataset creation, and evaluation.
- Part 4 - Modern AI, Copyright, and Course Synthesis (three classes): LLMs, RAG, and multimodal discovery; copyright and the reuse of digitised collections; review of the pipeline, methods, trade-offs, and connections across the course.
Keywords
historical document, natural language processing, machine learning, information extraction, information retrieval, cultural heritage data, generative AI, digital humanities
Learning Prerequisites
Recommended courses
- DH-405 Foundations of Digital Humanities (recommended alongside this class)
Important concepts to start the course
Basic knowledge of Machine Learning is recommended.
For those who wish to go deeper into ML, CS-433 Machine learning is recommended (to be followed in parallel or later).
Learning Outcomes
By the end of the course, the student must be able to:
- Characterize the main stages, actors, and challenges of historical document and media processing.
- Analyze a historical collection and formulate an appropriate processing workflow for a defined research or access need.
- Apply and explain core methods for OCR/OLR, information retrieval, information extraction, collection exploration, and data annotation.
- Assemble a high-quality dataset for machine learning training purposes.
- Contextualise the ecosystem surrounding historical text processing, including challenges, approaches, resources, actors, recent developments, and interdisciplinary aspects.
- State the main questions surrounding copyrights and use of digitised archive collections.
Teaching methods
- Lectures in the class room
- Exercice / lab sessions: Hands-on practical work on historical document datasets in Jupyter notebooks, combining computational exercises with short reflection activities.
- Guided discussion of research literature and real-world cultural-heritage infrastructures.
Expected student activities
- Attend and engage with lectures, exercise sessions, and readings.
- Complete practical notebook exercises and reflect on methods and results.
- Develop a semester-long synthesis assignment on a selected course topic.
- Prepare for written assessments through course materials, exercises, and assigned research papers.
Assessment methods
- Mind-map assignment (15%): a semester-long individual synthesis on a freely selected topic related to the course, including a formative oral work-in-progress presentation and a final mind map with a short accompanying note.
- Reading exam (15%): a 1h15 written assessment of three assigned research papers, based on analytical short-essay questions.
- Graded notebook assignments (20%): two practical Jupyter-notebook assignments applying and reflecting on course methods.
- Final written assessment (50%): a 3-hour written exam covering concepts, methods, and practice-oriented applications from lectures, readings, and exercises.
The respective weightings of the assessment components are provisional and will be confirmed at the beginning of the semester.
Resources
Bibliography
The course bibliography will be distributed during the first class.
Moodle Link
In the programs
- Semester: Fall
- Exam form: During the semester (winter session)
- Subject examined: Historical Document and Media Processing
- Courses: 2 Hour(s) per week x 14 weeks
- Exercises: 2 Hour(s) per week x 14 weeks
- Type: mandatory
- Semester: Fall
- Exam form: During the semester (winter session)
- Subject examined: Historical Document and Media Processing
- Courses: 2 Hour(s) per week x 14 weeks
- Exercises: 2 Hour(s) per week x 14 weeks
- Type: mandatory
- Semester: Fall
- Exam form: During the semester (winter session)
- Subject examined: Historical Document and Media Processing
- Courses: 2 Hour(s) per week x 14 weeks
- Exercises: 2 Hour(s) per week x 14 weeks
- Type: optional