Corpus Building in DraCor#

Warning

This chapter is a draft. It has not yet been proofread or formally reviewed. Content, terminology, and examples may change.

Chapter metadata

Authors: Daniil Skorinkin; Julia Jennifer Beine
Version: 0.2 (2026-09-02)
Review status: in progress
Planned reviewers: Antonio Rojas Castro; Frank Fischer

1. Overview#

In this chapter, we describe how the existing corpora in DraCor have been built. In most cases, the corpus building has started by inventorying available online resources while applying certain selection criteria. We elaborate on these selection criteria, illustrated by the resulting DraCor corpora.

2. Requirements and Competences#

  • Web browser and internet access required.

3. Learning Outcomes#

After completing this chapter, learners will be able to:

  1. Describe the principles and focuses of corpus building in DraCor.

  2. Describe the types of sources of the existing DraCor corpora.

4. Theoretical Background#

4.1. Distributed Corpus Building#

DraCor is a collaborative, community-driven project. While a core team at several universities maintains the shared platform – the Application Programming Interface (API), the front-end, the text encoding schema – the corpora themselves come from a wide variety of contributors. Individual scholars, research groups, and digital humanities projects across many countries build, maintain, and contribute corpora. In building the corpora, scholars may pursue their own research questions (also see Schöch [Schoch17], p. 224). This distributed model of corpus building [GST+23] is one of the crucial characteristics of DraCor and, at the same time, an important factor shaping the diversity and heterogeneity of the DraCor corpora. The common ground is that all DraCor corpora include dramatic texts. In the following chapters, we will look at the DraCor corpora from different perspectives, considering their selection criteria and focuses as well as their sources.

4.2. Sources of the Corpora#

In DraCor, most corpus builders follow an opportunistic approach, inventorying available online resources while applying certain selection criteria (on the different types of corpus building, see Schöch [Schoch17], pp. 225–226). Overall, the DraCor corpora may be divided into two broad types of corpora, based on how they were built: imported and in-house corpora.

The first type is collections that already existed as standalone digital projects and were subsequently converted to the DraCor format and integrated into the platform. The DraCor format is files in eXtensible Markup Language (XML) that follow the encoding guidelines by the Text Encoding Initiative (TEI) (see chapter 2 “TEI Encoding: Preparing Texts for Programmable Corpora”) and also adhere to the DraCor TEI schema. When importing existing text collections, the selection of dramatic texts was not made by the DraCor team; it was determined by the scope and criteria of the original project. The role of the DraCor team was primarily one of format conversion and homogenisation. A prominent example is the French Drama Corpus (FreDraCor, Milling et al. [MFGobelnt]), which is derived entirely from Paul Fièvre’s “Théâtre Classique” project [fient], a long-running, meticulously curated collection of dramatic texts from French classicism. Similarly, the Shakespeare Drama Corpus (ShakeDraCor, DraCor [DraCornta]) inherits its texts from the digital editions of the Folger Shakespeare Library [fol]; the Swedish Drama Corpus (SweDraCor, DraCor [DraCorntb]) takes its texts from the “eDrama” project [for]; the English Drama Corpus (EngDraCor, Giovannini and DraCor team [GDraCorteamnt]) is almost entirely based on TEI/XML files from the “EarlyPrint” project [mue]; and the Greek Drama Corpus (GreekDraCor Beine and Fischer [BFnt]) and Roman Drama Corpus (RomDraCor, Beine et al. [BFI24]) both draw on the “Perseus Digital Library” [crat]] to provide the extant ancient plays. Note that GreekDraCor additionally features a play from Wikisource.

The second type is in-house corpora – collections that were built specifically for DraCor by its maintainers, drawing on multiple digital sources and sometimes even newly digitised materials. Here, the DraCor corpus maintainers themselves decide which plays to include, based on criteria of their own choosing. The corpus grows organically over time as new texts are found, encoded, and added. The German Drama Corpus (GerDraCor) is the most prominent example: it began with plays from the TextGrid Repository [tex, FGobela, FGobelb, FT], but has since expanded to include texts from over a dozen different digital sources, such as Wikisource [wiknt], Projekt Gutenberg-DE [latnt], Google Books [goot]], Deutsches Textarchiv [deut]], the Internet Archive [intt]], and various academic libraries. The Russian Drama Corpus (RusDraCor, Fischer and Skorinkin [FSnt]) was similarly assembled from a range of Russian online libraries, such as lib.ru [Mos], ilibrary.ru [Kom], rvb.ru [a], feb-web.ru [b], and Wikisource [wiknt]. The Dutch Drama Corpus (DutchDraCor, van der Deijl [vdDnt]) derives most of its plays from the “Digitale Bibliotheek voor de Nederlandse Letteren” (DBNL, Taalunie et al. [taat]]), plus a handful of texts added from the “Ceneton” project [Har07, Hart]]. The Ukrainian Drama Corpus (UDraCor, Tokarskyi and Skorinkin [TSnt]) draws on UkrLib online library [ukr], dedicated editions of individual authors, and several other Ukrainian digital sources, such as litopys.org.ua [lit] and myslenedrevo.com.ua [mys]. The Polish Drama Corpus (PolDraCor, Pastuch et al. [PMWkasinskant]) combines materials mainly from Polona (the Polish national digital library, Polona [pol]), with a few plays from other Polish library platforms.

In practice, the boundary between imported and domestic corpora is not always sharp. Some corpora started as imports from a single project but were later supplemented with plays from other sources, making them hybrids. For instance, the Italian Drama Corpus (ItaDraCor, Fischer et al. [FMGnt]) draws roughly 95% of its texts from the “Biblioteca Italiana” [sapt]], with occasional additions from Google Books, the Internet Archive, and Wikisource. Still, the imported-versus-domestic distinction captures a crucial and consequential difference: it determines whether the DraCor curators could freely choose how to build their corpus or were constrained by the selection decisions of a source project.

4.3. Scope of the Corpora#

On a more fine-grained level, DraCor corpora differ in their selection criteria. All DraCor corpora share a genre-specific selection criterion, as they only include dramatic texts. Further selection criteria may be a certain time period, a certain region of production, a certain language, or a certain author (also see Schöch [Schoch17], p. 224).

Language is by far the most common selection criterion. The majority of DraCor corpora provide plays originally written in a particular language, regardless of the authors’ nationality, the creation period, or the place of publication. GerDraCor, for instance, includes German-language dramas by authors from what is now Germany, Austria, and Switzerland, spanning from the 1530s to the 1940s. RusDraCor contains dramas written in Russian by citizens of the Russian Empire and the Soviet Union (whose ethnic identity might or might not be Russian in a narrow sense). The Romanian Drama Corpus (RoDraCor, Terian et al. [TOU+nt]) collects Romanian-language dramas; the Hungarian Drama Corpus (HunDraCor, Department of Digital Humanities at Eötvös Loránd University and National Laboratory for Digital Heritage [DepartmentoDHaEotvosLorandUniversityNationalLfDHeritagent]) Hungarian ones; the Polish Drama Corpus (PolDraCor) Polish ones; and so on. Language-specific corpora form the largest corpora group in DraCor.

An adjacent case is a corpus whose scope is a dialect or a regional variety of a language. The Alsatian Drama Corpus (AlsDraCor, Ruiz Fabo [RFnt]) comprises plays written in Alsatian, which is variously classified as a dialect of Alemannic German or as a language in its own right – a useful reminder that the line between “language” and “dialect” is often more political and cultural than strictly linguistic. The Argentinian Drama Corpus (ArDraCor, RosDH, University of Rostock and HD LAB-CONICET, Argentina [RosDHUoRostockHDLCONICETArgentinant]) presents a related but different kind of boundary question: it collects plays written in Spanish, but specifically by Argentinian authors, reflecting a decision by its creators to treat Argentinian dramatic literature as a distinct cultural tradition rather than folding it into a pan-Hispanic corpus [dRioR26]. Here, the selection criterion is better described as national culture or place of production rather than language alone.

Some corpora combine language and time period as joint selection criteria. GreekDraCor is not a corpus of all drama ever written in Greek; it is a corpus of ancient Greek drama – the tragedies and comedies of classical Athens. Similarly, RomDraCor focuses on ancient Roman drama (Plautus, Terence, Seneca), not on Latin drama in general. The Neo-Latin Drama Corpus (NeoLatDraCor) complements RomDraCor by collecting dramatic texts written in Latin starting from the early-modern period. The Spanish Drama Corpus (SpanDraCor) is also limited by both language (Spanish) and period (1868–1936), since it inherits the temporal boundaries of the “BETTE” corpus [CFernandez] from which it derives.

Finally, a few DraCor corpora are defined by authorship as they contain the dramatic works of a single playwright. ShakeDraCor is the most prominent example, bringing together Shakespeare’s 37 plays from the Folger Shakespeare Library editions. The Calderón Drama Corpus (CalDraCor, Ehrlicher et al. [ELKnt]) collects the comedies of Pedro Calderón de la Barca. The German Shakespeare Drama Corpus (GerShDraCor, Fischer [Fisnt]) is a special case: it contains Shakespeare’s plays in the canonical German translation by the translator group around August Wilhelm Schlegel and Ludwig Tieck, making it an author-focused corpus that simultaneously documents a major act of cultural transfer (also see Beine [Bei17]). The Ibsen Drama Corpus (IbsDraCor, Centre for Ibsen Studies [CentrefIStudiesnt]) collects all of Henrik Ibsen’s dramatic works from the scholarly edition “Henrik Ibsen’s Writings” maintained by the Centre for Ibsen Studies at the University of Oslo.

It is worth noting that these categories are not mutually exclusive. A language-based corpus like GerDraCor implicitly spans several centuries and many authors; an author-based corpus like ShakeDraCor is also implicitly a period corpus (roughly 1590–1613) and a language corpus (English). The previous explanations focused on the primary selection criterion – the principle that first determines what is included and what is not.

4.4. Source Types of the DraCor Corpora From the Technical Viewpoint#

TEI/XML encoding in DraCor requires that the dramatic text is available in digital form. This text can either come from a pre-existing digital resource – an online library, a text collection, a digital archive – or it can be digitised from a print source such as a book or a manuscript. In the vast majority of cases, DraCor corpora are the result of what we might call secondary digital processing: the plays have already undergone primary digitisation by someone else (a library, a scholarly project, a volunteer community), and the DraCor encoding work consists of converting those existing digital texts into the DraCor TEI/XML format. Only in a minority of cases are plays digitised from print sources specifically for inclusion in DraCor.

The type of digital source significantly influences the process of corpus building, because sources differ enormously in the extent of usable structural markup they already contain. The following overview captures the main source types, roughly ordered from the sources easiest to include in a DraCor corpus to those more difficult to incorporate.

TEI/XML sources. The most straightforward scenario is when the source texts are already encoded in TEI/XML and provide at least the core dramatic structure: acts, scenes, speakers, speech, and stage directions. This was the case for the TextGrid Repository texts that seeded GerDraCor, for the Théâtre Classique files behind FreDraCor, for the Folger Shakespeare editions in ShakeDraCor, for the eDrama files in SweDraCor, for the EarlyPrint files in EngDraCor, and for many texts in DutchDraCor (via DBNL). However, “TEI/XML” is a broad standard. Despite being nominally standardised, these sources differ considerably in the granularity of their markup, in how strictly they followed the TEI guidelines, which version of the TEI schema they used, and how they split texts across files. In TextGrid, for example, Goethe’s “Faust” was stored as five separate TEI documents, and some plays counted double due to co-authorship metadata. These inconsistencies require substantial homogenisation work using eXtensible Stylesheet Language Transformations (XSLT) or Python scripts, and manual editing. Most recently, Large Language Models (LLMs) have also been used to assist with certain encoding tasks.

Structured non-TEI markup. In many cases, source texts come not as TEI/XML but in some other format that encodes the structural properties of plays in a consistent, semi-machine-readable way. The Hypertext Markup Language (HTML) is the most common example. Online libraries typically display plays as HTML pages, and while HTML was designed for visual presentation rather than semantic encoding, it often distinguishes consistently between spoken text and stage directions, marks act and scene headings as different levels of header, and wraps speaker names in recognisable formatting. Exploiting this structural regularity, the DraCor corpus builders may write conversion scripts to produce a first TEI draft that requires only moderate manual correction. This has been the approach for many plays in RusDraCor (sourced from online libraries like ilibrary.ru and rvb.ru), for UDraCor (with plays from ukrlib.com.ua), and for plays sourced from various national digital libraries. DOCX/DOC files with consistent formatting represent a similar case: the formatting (styles, bold text, indentation) can carry structural information that aids conversion. NeoLatDraCor has developed a related From Word to TEI workflow [BMnt].

Weakly structured or plain text. Sometimes a play is available as an HTML document whose tags are inconsistent or carry no useful structural information, or as mere plain text. In the most fortunate plain-text cases, the text may still contain rudimentary implicit markup, such as speaker names in capital letters, stage directions in brackets, consistent blank lines between scenes, that may be exploited with regular expressions to automate part of the encoding. In less fortunate cases, the encoding must be done largely by hand. Some plays in DutchDraCor, for example, were sourced from the Ceneton project [Har07, Hart]], whose HTML was too inconsistent to be converted to TEI automatically; the text was stripped to raw form and re-encoded from scratch using a combination of regular expressions and manual work.

Digital images (OCR required). Some plays are only available as page images, for instance, as PDF files without a text layer, or as scanned pages in a digital library. In these cases, Optical Character Recognition (OCR) had to be performed first to extract the text, which then required correction – OCR output for older typefaces and historical orthography is rarely clean – before TEI encoding could begin.

Print sources (scanning and OCR). In the rarest cases, the DraCor corpus maintainers digitise plays from physical print sources: scanning the pages, performing OCR, correcting the output, and then encoding the text in TEI/XML. In very rare cases, corpus builders may also include manuscript sources from scratch, using Handwritten Text Recognition (HTR). This is the path of last resort, used only for dramatic texts that existed in no digital form at all.

It should be noted that a single corpus often draws on sources of several different types. GerDraCor is a good example: the majority of its plays (528 out of 768 as of early 2026) comes from the TextGrid Repository as TEI/XML, but the corpus also includes 115 plays sourced from Google Books (typically as scanned page images requiring OCR), 39 from Projekt Gutenberg-DE, 32 from Wikisource, and smaller numbers from the Berlin State Library [sta], Deutsches Textarchiv, the Internet Archive, and various academic libraries – each with its own format and its own conversion challenges.

5. Practical Examples#

5.1. Example A: The Building of GerDraCor[1]#

GerDraCor is the oldest and one of the most extensively documented corpora in DraCor. Its history illustrates many of the general principles discussed above – the role of inherited collections, the gradual expansion from a single source to many, and the nature of corpus curation as an ongoing process. Because its development is closely intertwined with the origins of the DraCor platform itself, telling the story of GerDraCor is also, in part, telling the story of how DraCor came to be.

The roots of GerDraCor lie in the DLINA project (Digital Literary Network Analysis, Fischer et al. [FTK+]), a research group whose members included Frank Fischer, Mathias Göbel, Dario Kampkaspar, Christopher Kittel, and Peer Trilcke. The goal of DLINA was to study the network structure of dramatic texts (who appears on stage together with whom), using computational methods. In order to do so, the group needed a large, consistently encoded corpus of German-language plays.

The starting point was the TextGrid Repository [tex], the largest freely available TEI-encoded collection of German literature that contains thousands of texts released under a CC BY licence. But extracting the dramatic texts from TextGrid turned out to be a challenge [FGobela]. The seemingly straightforward query “How many dramatic texts are in the TextGrid Repository?” led through a thicket of complications: multi-part works stored as separate TEI documents (Goethe’s “Faust” was split into five files, Wagner’s “Ring” into four); doublets caused by co-authorship metadata (a play by two authors was counted twice because each author’s file contained a reference to it); and inconsistencies in genre classification. After systematic cleaning using XML Query Language” (XQuery) against an eXist-db instance, the DLINA team arrived at 666 dramatic texts.

From these 666 texts, the team then selected 465 plays to form the DLINA Corpus 15.07 (Codename: Sydney), named after the DH2015 conference in Sydney, where the results would be presented. The selection criteria narrowed the corpus to a specific period and type [FT]:

  • The temporal scope began with the German Enlightenment, specifically with Gottsched’s “Der sterbende Cato” (printed 1732), a widely recognised turning point in the history of German drama. This meant discarding 147 pre-Gottsched texts.

  • Foreign-language originals and translations were excluded – the corpus was to contain only works originally written in German.

  • Pantomime plays lacking speech elements (<sp>) were discarded, since they could not be analysed for dramatic dialogue.

  • Fragments – texts clearly left unfinished by their authors – were removed.

A further 32 texts were sorted out during editing because of severely defective TEI markup, because they turned out to be fragments that had been overlooked, or because their dramatic structure was too complex for the initial project tools to handle [FT].

The DLINA project also had to deal with inconsistent metadata [FGobelb]. The date information in the TextGrid Repository was unreliable – sometimes the <creation> element was empty, sometimes it contained only the author’s birth and death years as a vague bracket. The team developed a pragmatic decision tree: first, look for an exact year in the <creation> element; if unavailable, take the earliest year mentioned in the <note> element (which often contained information about first print or premiere dates); and as a last resort, use the author’s year of death as a terminus ante quem. These approximate dates were encoded into filenames (e.g. 1772-Lessing_Gotthold_Ephraim-Emilia_Galotti-lina.xml) to enable chronological sorting – a practical workaround that the team openly described as an approximation, not a substitute for proper metadata curation.

For their network analysis, the DLINA team did not initially need the full texts; they worked with a custom intermediate format they called the “Zwischenformat”, which stored only metadata and detailed structural information about acts, scenes, and which characters appeared in which segment. This was sufficient for extracting co-presence networks. The full texts from TextGrid were preserved alongside the “Zwischenformat” files, but were not the primary object of analysis at this stage [KFT].

On 2 December 2016, Mathias Göbel – during a hackathon at the University of Potsdam – made the first commit to what would become the GerDraCor repository on GitHub [BornerT24], p. 13. This initial commit added 465 TEI/XML files (one per play) to a data folder. The commit message reads: “inital commit: converted text based on LINA and TextGrid.” The files combined the full text from the TextGrid Repository sources with metadata from the DLINA “Zwischenformat”, now encoded following the TEI Guidelines rather than the custom format of the project.

This moment marks the “birth” of GerDraCor as a distinct entity, though it would take until 2017/2018 for the corpus to become fully integrated into the emerging DraCor platform. In September 2017, two significant organisational commits reshaped the repository: the data folder was renamed from data to tei (establishing the convention used across all DraCor corpora), and all filenames were standardised to match the playname identifier format (e.g. lessing-emilia-galotti.xml instead of 1772-Lessing_Gotthold_Ephraim-Emilia_Galotti-lina.xml).

Throughout 2017, the corpus remained stable at 465 plays – no new texts were added, though the existing files underwent significant internal changes (markup homogenisation, formatting). The first sign of growth beyond the original TextGrid seed came on 6 January 2018, when a play not derived from TextGrid was added for the first time: “Die Überschwemmung” by Franz Philipp Adolph Schouwärt, converted from Wikisource. This small event – one play from one new source – marked the beginning of GerDraCor’s transformation from a static inherited collection into a living, actively curated corpus [BornerT24], p. 15.

The growth accelerated markedly from 2020 onwards. As documented in the CLS INFRA D7.3 report by Börner and Trilcke [BornerT24], p. 15, the number of plays added per year increased substantially:

Year

New plays

Cumulative total

2016

465

465

2017

0

465

2018

7

472

2019

7

479

2020

41

520

2021

32

552

2022

45

597

2023

66

663

2024

15

678

By March 2026, GerDraCor contains 768 plays from over a dozen different digital sources. The TextGrid Repository remains the largest single contributor (528 plays), but Google Books has become the second-largest source (115 plays), followed by Projekt Gutenberg-DE (39), Wikisource (32), the Berlin State Library (9), Project Gutenberg (9), a scholarly edition of Marie von Ebner-Eschenbach (9), Deutsches Textarchiv (4), and smaller numbers from the Internet Archive, the Göttingen State and University Library, the Herzogin-Anna-Amalia-Bibliothek, and several other academic libraries [2].

This diversification of sources also brought an expansion of the temporal scope of the corpus. The original DLINA selection had focused on the period from the German Enlightenment (1730s) through the early 20th century. As new plays were added from different sources, GerDraCor grew to cover German-language drama from the 1530s to the 1940s – a much broader span that now includes Renaissance and early Baroque texts as well as works from the Weimar Republic era [BornerT24], pp. 16–17.

It is important to understand that the evolution of a living corpus like GerDraCor is not only a matter of adding new plays. The existing files also change over time – sometimes substantially. The CLS INFRA D7.3 report identifies several categories of change [BornerT24], p. 20:

Batch edits affect all files in the corpus simultaneously. Major examples include the addition of Wikidata identifiers to all plays (September 2018), the replacement of DLINA-era identifiers with DraCor IDs (May 2019), the introduction of a RelaxNG schema reference (September 2019), the normalisation of genre terms and introduction of Wikidata-linked genre classification (December 2020), and the addition of a <standOff> element for contextual metadata (May 2022). Each of these batch edits represents a structural evolution of the encoding, often driven by new features being developed in the DraCor API [BornerT24], pp. 20–24.

Individual file edits correct errors, enrich metadata, or improve the encoding of specific plays. The D7.3 report traces the file for Lessing’s “Emilia Galotti” through 26 distinct versions, documenting changes that range from the addition of text formatting (line breaks and indentation) in the earliest commits to the gradual enrichment of the TEI header with Wikidata links, genre information, and character relation data [BornerT24], pp. 18–20.

Encoding of character relations – social and family relationships between dramatic characters – was introduced to GerDraCor between September and November 2019, in cooperation with the QuaDramA project [reit]]. This feature was added to 358 plays, not all at once but file by file, illustrating how new encoding features can enter a corpus incrementally rather than through a single batch operation.

This ongoing evolution is what makes GerDraCor a “living corpus” in the full sense of the term: it is not merely a collection that grows by accretion, but one whose individual components are themselves subject to continuous revision and enrichment. For researchers working with GerDraCor, this has practical implications for reproducibility – a topic addressed in detail in the CLS INFRA D7.3 report, which recommends citing specific Git commits as stable references to particular states of the corpus [BornerT24], p. 32.

6. Exercises#

Exercise 1: Imported and In-house Corpora#

Exercise 2: Selection Criteria in Corpus Building#

7. Teaching Notes#

Lecturers may select DraCor corpora not discussed in this chapter and ask their students to research how the corpora have been built, preferably in small groups of two. Note that a basic knowledge of the DraCor front-end may be useful, e.g. how to find the editorial information of a corpus in the front-end (or GitHub) or related publications in the DraCor research bibliography. Thus, lecturers may include Chapter 4, “Front-End: Navigating DraCor” in the learning session.

8. Further Reading and Resources#

Readers interested in general approaches to corpus building may read Schöch [Schoch17] as a beginner-friendly introduction. Those who would like to delve into “digital corpus archaeology”, i.e. investigating the history of a digital corpus, may consult Börner and Trilcke [BornerT24].

9. Glossary Entries#

Term

Definition[3]

Corpus

A corpus is a text collection selected based on specific criteria.

(to) encode / Encoding

In the context of TEI/XML, the verb “encode” refers to the process of adding information to an electronic text, e.g. in the form of XML tags. The noun “encoding” refers to the result of this procedure, e.g. the XML markup in a file.

HTML

An abbreviation for the “Hypertext Markup Language” commonly used on websites.

HTR

An abbreviation for “Handwritten Text Recognition” that refers to the process of generating machine-readable text from an image of a manuscript, e.g. from a scan.

LLM

An abbreviation for “Large Language Model”, a type of artificial intelligence that generates text. Large Language Models may serve as the basis of Chatbots.

(to) mark up / markup

In the context of TEI/XML, the noun “markup” refers to the information added to an electronic text in the form of XML tags. The verb “mark up” refers to the process of adding this information to the text.

OCR

An abbreviation for “Optical Character Recognition” that refers to the process of generating machine-readable text from an image of said text, e.g. from a scan.

Python

Python is a programming language.

Regular expression

Regular expressions consist of one or more characters or symbols through which a text may be searched for certain patterns. In find-and-replace actions or programming scripts, regular expressions may serve as placeholders to address these patterns to encode them in a certain way in TEI/XML.

TEI

An abbreviation for “Text Encoding Initiative” which may refer to that organisation, its encoding guidelines, or files that follow those guidelines.

XML

An abbreviation for “eXtensible Markup Language”, a method for marking up texts and encoding information.

XQuery

An abbreviation for “XML Query Language”, a programming language.

XSLT

An abbreviation for “EXtensible Stylesheet Language Transformations”, a programming language.

10. Next Steps#

Continue with chapter 2, “TEI Encoding: Preparing Texts for Programmable Corpora”, to learn how the dramatic texts are encoded in DraCor.

11. AI Use Declaration#

In chapters 4 and 5, Daniil Skorinkin used Claude Opus 4.6 in the process of literature review – literature search and systematisation, writing and editing – text generation, and writing and editing – formulation of conclusions. In the other chapters, no generative AI was used.

12. Author Contributions#

Daniil Skorinkin – investigation, writing – original draft Julia Jennifer Beine – conceptualisation, writing – original draft, writing – review & editing

13. References#

[sta]

Digitalisierte sammlungen der staatsbibliothek zu berlin. URL: https://digital.staatsbibliothek-berlin.de/ (visited on 2026-07-12).

[mue]

Earlyprint. URL: https://earlyprint.org/ (visited on 2026-07-09).

[fol]

Folger shakespeare library. URL: https://www.folger.edu/ (visited on 2026-07-09).

[pol]

Polona. URL: https://polona.pl/ (visited on 2026-07-12).

[tex] (1,2)

Textgrid repository (textgridrep). URL: https://textgridrep.org/ (visited on 2026-07-09).

[for]

Edrama (dramawebben). URL: http://www.dramawebben.se/sida/edrama (visited on 2026-07-12).

[lit]

Ізборник. історія україни ix–xviii ст. першоджерела та інтерпретації. URL: http://litopys.org.ua/ (visited on 2026-07-12).

[ukr]

Бібліотека української літератури укрліб. URL: https://www.ukrlib.com.ua (visited on 2026-07-12).

[mys]

Мислене древо — багатоцільовий український сайт. URL: https://myslenedrevo.com.ua/ (visited on 2026-07-12).

[latnt]

Projekt gutenberg. 1994–present. URL: https://projekt-gutenberg.org/ (visited on 2026-07-12).

[wiknt] (1,2)

Wikisource. 2003–present. URL: https://wikisource.org/ (visited on 2026-07-12).

[fient]

Théâtre classique. 2007–present. URL: https://www.theatre-classique.fr/ (visited on 2026-07-09).

[crat]]

Perseus digital library. [1994–present]. URL: https://www.perseus.tufts.edu/ (visited on 2026-07-12).

[intt]]

Internet archive. [1996–present]. URL: https://archive.org/ (visited on 2026-07-12).

[taat]]

Digitale bibliotheek voor de nederlandse letteren (dbnl). [1999–present]. URL: https://www.dbnl.org/ (visited on 2026-07-12).

[sapt]]

Biblioteca italiana. [2000–present]. URL: http://bibliotecaitaliana.it (visited on 2026-07-12).

[goot]]

Google books. [2004–present]. URL: https://books.google.com/ (visited on 2026-07-12).

[deut]]

Deutsches textarchiv (dta). [2007–present]. URL: https://www.deutschestextarchiv.de/ (visited on 2026-07-12).

[reit]]

Quadrama – quantitative drama analytics. [2017–present]. URL: https://quadrama.github.io/ (visited on 2026-07-12).

[Bei17]

Julia Jennifer Beine. Die re- bzw. dekonstruktion des "schlegel-tieck-shakespeare" anhand der kritischen edition des "hamlet". "forsch!". Studentisches Online-Journal der Universität Oldenburg, 1:75–86, 2017. URL: https://zenodo.org/doi/10.5281/zenodo.17564010 (visited on 2026-04-27), doi:10.5281/ZENODO.17564010.

[BFnt]

Julia Jennifer Beine and Frank Fischer. Greek drama corpus (greekdracor). 2019–present. URL: dracor-org/greekdracor (visited on 2026-07-12).

[BFI24]

Julia Jennifer Beine, Frank Fischer, and Viktor J. Illmer. Just the type: analysing character typology in roman comedy with romdracor. In Jajwalya Karajgikar, Andrew Janco, and Jessica Otis, editors, DH2024 Book of Abstracts. Zenodo, 2024. URL: https://zenodo.org/doi/10.5281/zenodo.13801481 (visited on 2026-06-06), doi:10.5281/zenodo.13801481.

[BMnt]

Julia Jennifer Beine and Carsten Milling. Neo-Latin Drama Corpus (NeoLatDraCor). 2024–present. URL: dracor-org/neolatdracor (visited on 2026-05-20).

[BornerT24] (1,2,3,4,5,6,7,8,9,10)

Ingo Börner and Peer Trilcke. CLS INFRA D7.3 On Versioning Living and Programmable Corpora: (Executable) Report and Prototypes for Reproducible Research. February 2024. URL: https://zenodo.org/doi/10.5281/zenodo.11081934 (visited on 2026-05-20), doi:10.5281/ZENODO.11081934.

[CFernandez]

José Calvo and Teresa Santa María Fernández. Ghedi/bette: versión 1.0 de bette (10.2017). doi:10.5281/zenodo.1010140.

[dRioR26]

María Gimena del Río Riande. Abordajes del teatro argentino (siglos xviii y xix) desde las humanidades digitales: codificación, publicación digital y análisis automatizado. Información, cultura y sociedad, 54:117–121, 2026. URL: https://dialnet.unirioja.es/servlet/articulo?codigo=10800420 (visited on 2026-08-17), doi:10.34096/ics.i54.18354.

[ELKnt]

Hanno Ehrlicher, Jörg Lehmann, and Simon Kroll. Calderón drama corpus (caldracor). 2019-present. URL: dracor-org/caldracor (visited on 2026-07-12).

[Fisnt]

Frank Fischer. German shakespeare drama corpus (gershdracor). 2021-present. URL: dracor-org/gershdracor (visited on 2026-07-12).

[FGobela] (1,2)

Frank Fischer and Mathias Göbel. A (not so) simple question and a somewhat diabolic answer. URL: https://dlina.github.io/A-Not-So-Simple-Question/ (visited on 2026-07-09).

[FGobelb] (1,2)

Frank Fischer and Mathias Göbel. Working with inconsistent metadata. URL: https://dlina.github.io/Working-With-Inconsistent-Metadata/ (visited on 2026-07-09).

[FMGnt]

Frank Fischer, Carsten Milling, and Luca Giovannini. Italian drama corpus (itadracor). 2020–present. URL: dracor-org/itadracor (visited on 2026-07-12).

[FSnt]

Frank Fischer and Daniil Skorinkin. Russian drama corpus (rusdracor). 2020–present. URL: dracor-org/rusdracor (visited on 2026-07-09).

[FT] (1,2,3)

Frank Fischer and Peer Trilcke. Introducing dlina corpus 15.07 (codename: sydney). URL: https://dlina.github.io/Introducing-DLINA-Corpus-15-07-Codename-Sydney/ (visited on 2026-07-09).

[FTK+]

Frank Fischer, Peer Trilcke, Dario Kampkaspar, Mathias Göbel, and Hanna-Lena Meiners. Dlina - digitally-driven literary network analyis (of dramatic texts). URL: https://dlina.github.io/ (visited on 2026-08-17).

[GST+23]

Luca Giovannini, Daniil Skorinkin, Peer Trilcke, Ingo Börner, Frank Fischer, Julia Dudar, Carsten Milling, and Petr Porízka. Distributed corpus building in literary studies: the dracor example. In Walter Scholger, Georg Vogeler, Toma Tasovac, Anne Baillot, and Patrick Helling, editors, DH2023 Book of Abstracts. Zenodo, 2023. doi:10.5281/zenodo.8107457.

[GDraCorteamnt]

Luca Giovannini and DraCor team. English Drama Corpus (EngDraCor). 2024–present. URL: dracor-org/engdracor (visited on 2026-04-12).

[Har07] (1,2)

A. J. E. Harmsen. Ceneton, of de bibliotheek en de schouwburg. Nieuw letterkundig magazijn, 25:34–37, 2007. URL: https://www.dbnl.org/tekst/_nie012200701_01/_nie012200701_01_0009.php.

[Hart]] (1,2)

A. J. E. Harmsen. Census nederlands toneel (ceneton). [1992–present]. URL: https://www.let.leidenuniv.nl/Dutch/Ceneton/index.html (visited on 2026-07-09).

[KFT]

Dario Kampkaspar, Frank Fischer, and Peer Trilcke. Introducing our zwischenformat. URL: https://dlina.github.io/Introducing-Our-Zwischenformat/ (visited on 2026-07-09).

[Kom]

Aleksey Komarov. Интернет-библиотека алексея комарова (ilibrary.ru). URL: https://ilibrary.ru/ (visited on 2026-07-12).

[MFGobelnt]

Carsten Milling, Frank Fischer, and Mathias Göbel. French Drama Corpus (FreDraCor): A TEI P5 Version of Paul Fièvre's "Théâtre Classique" Corpus. 2021–present. URL: dracor-org/fredracor (visited on 2026-04-12).

[Mos]

Maxim Moshkow. Библиотека максима мошкова (lib.ru). URL: http://lib.ru/ (visited on 2026-07-12).

[PMWkasinskant]

Magdalena Pastuch, Barbara Mitrenga, and Kinga W\k  asińska. Polish drama corpus (poldracor). 2023–present. URL: dracor-org/poldracor (visited on 2026-07-12).

[RFnt]

Pablo Ruiz Fabo. Alsatian drama corpus (alsdracor). 2019–present. URL: dracor-org/alsdracor (visited on 2026-07-12).

[Schoch17] (1,2,3,4)

Christof Schöch. Aufbau von datensammlungen. In Fotis Jannidis, Hubertus Kohle, and Malte Rehbein, editors, Digital Humanities: Eine Einführung, pages 223–233. J.B. Metzler, 2017. URL: http://link.springer.com/10.1007/978-3-476-05446-3_16 (visited on 2026-05-08), doi:10.1007/978-3-476-05446-3_16.

[TOU+nt]

Andrei Terian, Ovio Olaru, Aura Cristina Udrea, Victor Cobuz, Alexandra Oprescu, Andreea Popescu, Teodora Susarenco, and Cristina Cojocaru. Romanian drama corpus (rodracor). 2025–present. URL: dracor-org/rodracor (visited on 2026-07-12).

[TSnt]

Bohdan Tokarskyi and Daniil Skorinkin. Ukranian Drama Corpus (UDraCor). 2022–present. URL: dracor-org/udracor (visited on 2026-04-12).

[vdDnt]

Lucas van der Deijl. Dutch drama corpus (dutchdracor). 2023–present. URL: dracor-org/dutchdracor (visited on 2026-07-12).

[CentrefIStudiesnt]

Centre for Ibsen Studies. Ibsen drama corpus (ibsdracor). 2025-present. URL: dracor-org/ibsdracor (visited on 2026-07-12).

[DepartmentoDHaEotvosLorandUniversityNationalLfDHeritagent]

Department of Digital Humanities at Eötvös Loránd University and National Laboratory for Digital Heritage. Hungarian drama corpus (hundracor). 2021–present. URL: dracor-org/hundracor (visited on 2026-07-12).

[DraCornta]

DraCor. Shakespeare drama corpus (shakedracor). 2018–present. URL: dracor-org/shakedracor (visited on 2026-07-12).

[DraCorntb]

DraCor. Swedish drama corpus (swedracor). 2018–present. URL: dracor-org/swedracor (visited on 2026-07-09).

[OxfordEDictionary23]

Oxford English Dictionary. Mark-up, n., sense 2.c. December 2023. URL: https://doi.org/10.1093/OED/1115679047.

[OxfordEDictionary25a]

Oxford English Dictionary. Encode, v. September 2025. URL: https://doi.org/10.1093/OED/7215232060.

[OxfordEDictionary25b]

Oxford English Dictionary. Regular expression, n. June 2025. URL: https://doi.org/10.1093/OED/3705754966.

[RosDHUoRostockHDLCONICETArgentinant]

RosDH, University of Rostock and HD LAB-CONICET, Argentina. Argentinian drama corpus (ardracor). 2025–present. URL: dracor-org/ardracor (visited on 2026-07-12).

[a]

Русская виртуальная библиотека. Русская виртуальная библиотека (рвб / rvb). URL: https://rvb.ru/ (visited on 2026-07-12).

[b]

Фундаментальная электронная библиотека. Фундаментальная электронная библиотека «русская литература и фольклор» (фэб / feb). URL: http://feb-web.ru/ (visited on 2026-07-12).

14. Footnotes#