Skip to content

Instantly share code, notes, and snippets.

@boisei0
Created February 22, 2025 11:14
Show Gist options
  • Select an option

  • Save boisei0/350c6f2b2dfcfbb90174f4b4fc233e66 to your computer and use it in GitHub Desktop.

Select an option

Save boisei0/350c6f2b2dfcfbb90174f4b4fc233e66 to your computer and use it in GitHub Desktop.
Digital library setup

Digital library setup

Goals

  • Have an interface for all my records that I've collected throughout the years
  • My vision is worsening rapidly, I need to have transcriptions listed with images to keep working on projects (accessibility!)

Challenges

  • How to organise many hundreds of records, which results in 70k+ pages of scans
  • How to identify each record, they come from many places

Inspiration

The KNAW Humanities Cluster / Huygen's Institute's setup for large scale digitisation projects; toolings like Loghi, TextRepo, AnnoRepo, TextAnnoViz.

Plan

Tools used

  • IIIF3 compatible server: Cantaloupe
  • TextRepo
  • Loghi
  • TIFY
  • Creating the viewer application: Python 3.12 with Flask/iiif_prezi3/textrepo python client/elasticsearch7 python client

Set up a pipeline to prepare any records for viewing

  1. Describe records
  2. Have scans located in an organised structure that makes sense for accessing/usage, which is added to the record description; base on originating archive's organisation where possible.
  3. Transcribe records, using Transkribus or Loghi. Store the PageXML exports with the original records in a page subfolder.
  4. Load records into TextRepo
  5. Create IIIF3 Presentation manifests

Pipeline

  • Spreadsheet to describe what records there are and where they're stored
    • what is the record
    • what area is the contents about
    • what kind of record is this
    • unique urn for the record, with prefix for the library project, including where they come from. Eg. urn:<library-prefix>:matricula:deutschland:muenster:rees:st-mariae-himmelfahrt:kb001:68 for https://data.matricula-online.eu/en/deutschland/muenster/rees-st-mariae-himmelfahrt/KB001/?pg=68, or urn:<prefix>:ecal:3019:762:69 when describing the baptism of Theodora Johanna Roes (ECAL DTB (collection 3019), Gendringen RK baptismal book 1733-1761 (inventory 762), scan 69).
    • where is the record located digitally (relative path to library root folder)
    • number of pages
    • a couple flags of what has been done to the record so far:
      • are the scans present?
      • has it been transcribed?
      • has it been loaded into TextRepo yet
      • has an IIIF3 presentation manifest been generated for it yet?
      • [in case of problems]:
        • are there problems with the scans, such as them being incomplete or files damaged?
        • are there problems with the transcriptions, such as them being incomplete or files damaged
import pathlib
import iiif_prezi3
import json
LIBRARY_DRIVE = pathlib.Path("REPLACE_THIS")
cfg = {
"cantaloupe_base_dir": LIBRARY_DRIVE / "REPLACE_THIS",
"cantaloupe_base_url": "REPLACE_THIS"
}
license_mapper = {
"CC BY-SA 4.0": "https://creativecommons.org/licenses/by-sa/4.0/"
}
def create_manifest(urn, is_genealogie_domein=False):
with open('REPLACE_THIS_WITH_INVENTORY.csv', newline='') as csvfile:
invreader = csv.DictReader(csvfile, delimiter=',', quotechar='"')
# print(invreader.fieldnames)
for row in invreader:
if row["URN"] == urn:
print('Record found')
manifest_file_name = '_'.join(urn.split(':')[2:]) + '.json' # cut off the prefix, replace `:` by `_` to make it url-safe
manifest = iiif_prezi3.Manifest(
id="http://example.com/tbd-url-to-this-manifest-file.json", # FIXME
label={"en": [row["What"]]},
)
manifest.add_metadata(
label={"en": ["URN"]},
value={"en": [urn]}
)
# TODO: add permalink, dates, other information;
# TODO: add this information to the inventory sheet?
# Add image files in inventory dir as individual canvases with IIIF linked.
# row["Relative location on disk"]
# LIBRARY_DRIVE / row["Relative location on disk"].replace(' ', '')
img_dir = LIBRARY_DRIVE / row["Relative location on disk"].replace(' ', '')
for img in sorted(img_dir.glob("*.jpg")):
# img : PosixPath
# filename -> `img.name`
iiif_url = cfg["cantaloupe_base_url"] + img.relative_to(cfg["cantaloupe_base_dir"]).as_posix().replace("/", "%2F")
if is_genealogie_domein:
page_number = int(img.stem.split('-')[0])
else:
page_number = int(img.stem.split('-')[-1])
# canvas = manifest.make_canvas_from_iiif(
# url -> iiif url for service
# id -> items.items.target
# anno_id -> items.items.id
# anno_page_id -> items.id
#)
canvas_id = manifest.id + "/canvas/p" + str(page_number)
canvas = manifest.make_canvas_from_iiif(
url=iiif_url,
id=canvas_id,
anno_id=iiif_url + "/full/max/0/default.jpg",
anno_page_id=canvas_id + "/1"
)
with open(manifest_file_name, 'w') as manifest_file:
manifest_file.write(manifest.json(indent=4))
# Finished?!
break
else:
raise Exception("Record not found in inventory. Make sure the URN is correct!")
What Where Type URN Relative location on disk #Pages Scans Transcribed TextRepo IIIF3 Manifest Scans incomplete or damaged Transcription incomplete or damaged
Nederduits Doopboek 1665-1732 Aalten DTB urn:<prefix>:ecal:3019:1 Records / ECAL / 3019_Retroacta-DTB / Aalten / Nederduits-gereformeerde-gemeente / 0001_Doopboek_1665-1732 239 TRUE TRUE FALSE FALSE FALSE FALSE
Nederduits Doopboek 1733-1811 Aalten DTB urn:<prefix>:ecal:3019:2 Records / ECAL / 3019_Retroacta-DTB / Aalten / Nederduits-gereformeerde-gemeente / 0002_Doopboek_1733-1811 305 TRUE TRUE FALSE FALSE FALSE FALSE
Nederduits Trouwboek 1665-1732 Aalten DTB urn:<prefix>:ecal:3019:3 Records / ECAL / 3019_Retroacta-DTB / Aalten / Nederduits-gereformeerde-gemeente / 0003_Trouwboek_1665-1732 315 TRUE TRUE FALSE FALSE FALSE FALSE
Nederduits Trouwboek 1733-1811 Aalten DTB urn:<prefix>:ecal:3019:4 Records / ECAL / 3019_Retroacta-DTB / Aalten / Nederduits-gereformeerde-gemeente / 0004_Trouwboek_1733-1811 160 TRUE TRUE FALSE FALSE FALSE FALSE
import pathlib
import glob
import csv
from io import StringIO
from textrepo.client import TextRepoClient, DocumentIdentifier
from pagexml.parser import parse_pagexml_file
from pagexml.helper.pagexml_helper import pretty_print_textregion
from icecream import ic
# === CONSTANTS ===
TR = TextRepoClient(base_uri='[REPLACE_THIS]')
library_drive = pathlib.PurePath("[REPLACE_THIS]")
cantaloupe_base_dir = library_drive / "REPLACE_THIS"
# === PAGE FINDERS ===
ecal_dtb_page_find = lambda x: x.rsplit('-')[-1].split('.')[0]
ga_page_find = lambda p: p.split('-')[-1].split('.')[0]
genealogie_domein_page_find = lambda p: p.split('-')[0]
class TextRepoFolderUploadFailure(Exception):
def __init__(self, urn, base_dir, func_page_extract, step):
pass
def upload_folder_to_tr(urn, base_dir, func_page_extract, dry_run=False):
"""Uploads the transcriptions of a single folder to TextRepo.
Keyword arguments:
urn -- The urn of the folder to upload. Each document in the folder will get a `:<page>` suffix conform the library standard to form the individual document URNs
base_dir -- The base directory of the folder, specifically, the `page` subfolder created by Loghi.
func_page_extract -- Function with one argument which will extract the page number from any given pagexml file name
dry_run -- Debugging step; Print steps rather than uploading files to TR
"""
tr_docs = []
pagexmls = sorted(glob.glob("*.xml", root_dir=base_dir))
for pxml in pagexmls:
# Step 1: extract the page from the filename
try:
page = int(func_page_extract(pxml))
except Exception as e:
# jeez, don't do this, but what other way to catch anything thrown by a lambda
raise TextRepoFolderUploadFailure(urn, base_dir, func_page_extract, "page_extract")
urn_external_id = f'{urn}:{page}'
# Step 2: create document in TR:
if not dry_run:
tr_doc = TR.create_document(urn_external_id)
TR.find_document_metadata(urn_external_id)
# Step 3: upload PageXML to TR
with open(base_dir / pxml, 'r') as fp:
if dry_run:
print(f'`TR.import_version(external_id={urn_external_id}, type_name="pagexml", contents=..., allow_new_document=True, as_latest_version=True)` with file `{base_dir / pxml}`')
else:
pxml_version_info = TR.import_version(external_id=urn_external_id, type_name="pagexml", contents=fp, allow_new_document=True, as_latest_version=True)
# Step 4: parse plaintext
plaintext = ''
doc = parse_pagexml_file(base_dir / pxml)
for region in doc.text_regions:
for line in region.lines:
if line.text is not None:
plaintext += line.text
plaintext += "\n"
# Step 5: upload plaintext to TR
if dry_run:
print(f'`TR.import_version(external_id={urn_external_id}, type_name="text", contents=..., allow_new_document=True, as_latest_version=True)` with file contents `{plaintext[0:25]} [...] {plaintext[-25:0]}`')
else:
text_version_info = TR.import_version(external_id=urn_external_id, type_name="text", contents=plaintext, allow_new_document=True, as_latest_version=True)
# finishing off
if not dry_run:
tr_docs.append({
"external_id": urn_external_id,
"txt_version": text_version_info,
"pxml_version": pxml_version_info
})
return tr_docs
def upload_records(dry_run=False):
records = []
with open('REPLACE_THIS', newline='') as csvfile:
invreader = csv.DictReader(csvfile, delimiter=',', quotechar='"')
# print(invreader.fieldnames)
for row in invreader:
if row["Type"] == "Geestelijke goederen" and row["Scans"] == "TRUE" and row["Transcribed"] == "TRUE" and row["TextRepo"] == "FALSE" and row["Scans incomplete or damaged"] == "FALSE" and row["Transcription incomplete or damaged"] == "FALSE":
records.append((row["URN"], row["Relative location on disk"], row["#Pages"])). # REPLACE THIS BY THE SUBSET OF RECORDS YOU'RE IMPORTING
print(f'Total records to upload: {len(records)}.')
print('This will take a while...')
for record in records:
print(f'Uploading: {record[0]}. Total pages: {record[2]}')
upload_folder_to_tr(record[0], library_drive / record[1].replace(' ', '') / "page", ga_page_find, dry_run=dry_run)
print(f'Finished uploading {record[0]}')
print('')
if __name__ == '__main__':
upload_records(dry_run=False)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment