- Have an interface for all my records that I've collected throughout the years
- My vision is worsening rapidly, I need to have transcriptions listed with images to keep working on projects (accessibility!)
- How to organise many hundreds of records, which results in 70k+ pages of scans
- How to identify each record, they come from many places
The KNAW Humanities Cluster / Huygen's Institute's setup for large scale digitisation projects; toolings like Loghi, TextRepo, AnnoRepo, TextAnnoViz.
- IIIF3 compatible server: Cantaloupe
- TextRepo
- Loghi
- TIFY
- Creating the viewer application: Python 3.12 with Flask/iiif_prezi3/textrepo python client/elasticsearch7 python client
- Describe records
- Have scans located in an organised structure that makes sense for accessing/usage, which is added to the record description; base on originating archive's organisation where possible.
- Transcribe records, using Transkribus or Loghi. Store the PageXML exports with the original records in a
pagesubfolder. - Load records into TextRepo
- Create IIIF3 Presentation manifests
- Spreadsheet to describe what records there are and where they're stored
- what is the record
- what area is the contents about
- what kind of record is this
- unique urn for the record, with prefix for the library project, including where they come from. Eg.
urn:<library-prefix>:matricula:deutschland:muenster:rees:st-mariae-himmelfahrt:kb001:68for https://data.matricula-online.eu/en/deutschland/muenster/rees-st-mariae-himmelfahrt/KB001/?pg=68, orurn:<prefix>:ecal:3019:762:69when describing the baptism of Theodora Johanna Roes (ECAL DTB (collection 3019), Gendringen RK baptismal book 1733-1761 (inventory 762), scan 69). - where is the record located digitally (relative path to library root folder)
- number of pages
- a couple flags of what has been done to the record so far:
- are the scans present?
- has it been transcribed?
- has it been loaded into TextRepo yet
- has an IIIF3 presentation manifest been generated for it yet?
- [in case of problems]:
- are there problems with the scans, such as them being incomplete or files damaged?
- are there problems with the transcriptions, such as them being incomplete or files damaged