Skip to content

Transcription

The transcription workspace at /library/transcribe/<itemKey> is for turning facsimile images into one editable text surface beside the evidence.

It suits photographs from the archive, uploaded images and linked IIIF facsimiles. See Manuscript facsimiles for linking a shelfmark to its digitised copy.

On a manuscript’s Transcribe tab, choose Open editor; in its facsimile copy on Read, the leaf view has Transcribe. The desk opens in a dedicated editor with a link back to the library item. A catalogued manuscript offers Transcribe beside its photographs, and the link then returns to the catalogue.

Query parameters can pin the facsimile source (?source=) or open a saved snapshot (?snapshot=).

Layout switches between horizontal split, vertical split, image only and text only. The choice is remembered on this device.

Edit presents the transcription as one continuous text area. Physical line numbers remain in a narrow gutter, and selecting a passage reveals its text type and markup controls. Clicking the image highlights its corresponding text; text selections highlight the image. Line actions offers explicit split/join operations that retain those links.

The live draft autosaves. Save a version stores an immutable snapshot, including its text-type definitions, which you can reopen read-only or restore. If saving fails, your edits remain visible and Retry save lets you try again.

Read shows the current leaf or the whole manuscript. Manuscript line breaks are kept unless you enable Flowing text. Page/folio references are always visible. What I see controls other text types, such as margin notes, catchwords and types defined for this manuscript. Hiding text does not delete it or change what is classified as Main text. Untranscribed leaves are marked as gaps.

Copy with reference, and keyboard copying inside the transcription, copy the selected text without inline page numbers and append its manuscript reference. Consecutive page locators become pp. 1–3; foliation becomes fos 1–3, preserving recto/verso when recorded. Separate selections retain gaps, such as pp. 1–3, 5. With no selection, the Copy button copies the visible reading.

To read a finished transcription without the desk around it, or several in order, open the reading room.

All unclassified text starts as Main text. Select words or lines to change their type. Text types & AI guidance lets you add or edit a name and description: these guide future AI readings without changing existing classifications. AI can suggest new types; you can override any suggestion. User corrections survive later reads.

The Markup menu records additions, deletions, uncertainty, illegible passages and editorially supplied text. These marks are independent of text type. When a mark cannot be aligned safely between diplomatic and expanded readings, Review alignment lets you identify the corresponding passage. It never silently reuses an offset from a reading of a different length.

Transcribe with AI reads the current leaf and checks its draft against the image. Uncertain alternative readings appear when you select the relevant line. Choosing one is an ordinary, undoable edit.

The reader does better when it knows the hand. A manuscript with no catalogue entry is read with the language and date on its library record. A catalogued manuscript carries a Language and a Period beside its upload drop, the same fields as its ISAD(G) description, and the first AI reading of a manuscript with neither filled asks for them, optionally. Lines you have corrected on the same manuscript are shown to the model as examples of the hand on every later reading.

An early printed book is transcribed on the same desk. An ESTC edition in your library has copies on Read and a Transcribe tab like a manuscript: add its page images, or a PDF, with Add a copy on Read, or link a library’s IIIF copy there. Pangur does not fetch images from EEBO or any other subscription service, so download the pages from the library that gives you access and upload them.

The reader is told it is looking at type, and what that changes. It keeps the long s and the printed u, v, i and j, writes a ligature as its letters, reads black letter, roman and italic as one text, and files catchwords, signatures and running titles under their own text types so they stay out of the main text. The language and date on the item’s record tell it the rest. A book with no catalogue entry needs nothing more than that record.

Where the edition has an EEBO-TCP text, import it from ESTC and read it in the reading room, or match its pages to your images from Open the TEI transcription desk.

Use Use a supplied transcription to align existing text to detected physical lines. Use Markup → Illegible for unreadable passages and Supplied for editorially supplied letters. Status messages report save progress and errors.

Export, in the editor’s toolbar, saves the transcription as a file. When a named version is open, it saves that version.

  • TEI P5: the whole transcription in one document, each page a <pb/> and each line an <lb/> pointing at the line’s place on the page (its polygon and baseline, in a <facsimile>). Lines are grouped by text type. Markup becomes TEI’s own: <add>, <del>, <unclear>, <gap>, <supplied>, and <seg type> for a passage of another text type. Both readings keeps the diplomatic and expanded readings side by side in <choice>; or choose one of them.
  • PAGE XML: one file per page in a ZIP, with each line’s polygon, baseline and readings. Text types are regions named as eScriptorium and Transkribus name them (SegmOnto), so either tool can take the transcription on, for training a model or for further work.
  • ALTO: one ALTO 4 file per page in a ZIP, laid out as eScriptorium writes it. Each text type is a block tagged with its SegmOnto name; each line keeps its polygon and baseline. The diplomatic reading is the line’s text and the expanded reading its alternative, or choose one reading on the command line.
  • Plain text: the main text page by page, in either reading.

Each file names its image (<shelfmark>_<page>_<label>.jpg) the same way, and TEI points at the IIIF manifest when there is one. With page images puts the images in the same ZIP, at full size: PAGE XML in a page/ folder beside them, which Transkribus imports as it is, ALTO in alto/, or one TEI file. For eScriptorium, upload the images to a document, then import the page/ or alto/ files. Fetching the images from a library’s server takes a while; a long transcription comes as several ZIPs of up to sixty pages each. If a page’s image cannot be fetched, its transcription is still there and the ZIP’s README.txt names the page.

TEI, PAGE XML and ALTO are checked against their official schemas. A project export holds every transcription as TEI, PAGE XML and plain text, without the images.

Connected assistants can open facsimiles, transcribe a leaf, read the saved transcription and keep named versions with the tools under Assistant tools. zcp manuscript opens facsimiles, transcribes a leaf, prints the saved transcription and exports it (zcp manuscript export). Both say when an item is a printed book, read as type. To read a finished transcription as continuous text, use read_text or zcp reading. The web workspace is the place to correct what the model got wrong.