- 1 According to data provided by GWI for the Digital 2024 Global Overview Report, 71% of respondents b (...)
1When, at the end of the last century, the Deep Blue computer played a series of winning chess games against grandmaster Garry Kasparov, the world watched with carefree curiosity. Here was a machine created by humans, fed by the experience of many outstanding chess players who provided it with ready-made moves, and it had begun to surpass human intelligence. This event from the world of elite sports and equally inaccessible Silicon Valley laboratories – a world far removed from the perspective of the lesser mortals – was the harbinger of radical technological changes that have accumulated in recent years and, in the opinion of many, have begun to threaten the established order.1
2ChatGPT was launched as a free online tool at the end of November 2022. Netizens could not ignore this debut. Finally, there was a bot that could be talked to about any topic at any time and place, and could also be used for various unexciting tasks – in other words, a virtual companion, personal assistant and patient advisor. The brainchild of the OpenAI engineers was a huge success and surpassed the previous leaders such as TikTok or Instagram in terms of how quickly it gained new users.
3The impact of the new chatbot was unprecedented,2 but the initial enthusiasm quickly gave way to deep concern: what if a versatile and super-efficient artificial intelligence replaces us all? Not only chess players, but also doctors, lawyers, journalists, artists, programmers, office workers, economists, translators, editors and teachers? As long as expensive AI projects remained the domain of high-tech, the fear of losing one's job or a reduction in salary did not seem to exist, or at least was not widespread. When AI tools suddenly became widely available, many people began to fear for their future. Judging by the topic of the nineteenth conference of The European Society for Textual Scholarship in Budapest (Textual Scholarship, Artificial Intelligence, Corpora and Intelligent Editions),3 a similar curiosity tinged with uncertainty also accompanies editors. This is true even though it would seem that this particular profession has a long tradition of awareness of and participation in technological progress, dating back to the 1940s, when the Jesuit Roberto Busa authored the digital corpus of the works of St. Thomas Aquinas.4
- 5 A. Silberling, “Why AI Can’t Spell ‘Strawberry’”, https://techcrunch.com/2024/08/27/why-ai-cant-spe (...)
- 6 A. Tong, K. Paul, “Exclusive: OpenAI Working on New Reasoning Technology under Code Name ‘Strawberr (...)
- 7 Sora is another OpenAI engineering project, capable of generating realistic films that are very dif (...)
4What does the future hold for scholarly editing? How will we work on texts of works and will we even do it the same way as we used to? Or maybe artificial intelligence will first search through all available online library catalogues, access the texts stored in digital libraries, automatically read their text layer, collate the collected material, select the basis for the edition, and then – having at its disposal the knowledge of the entire Internet, including metadata in the form of articles on the practice of scholarly editorship – introduce conjectures, comment on questionable passages, explain difficult words, attempt an erudite foreword with an exhaustive characterization of the transmission of the text? Will it generate its own edition by using already published digital scientific editions – especially their open code? As long as we are making fun of the absurd clumsiness of the supposedly powerful chatbot, which cannot count the number of letters ‘r’ in the word strawberry,5 this vision seems distant; on the other hand, artificial intelligence engineers are already looking for solutions to create new models capable of conducting multi-stage research.6 The fact that the ‘artificial’ is becoming more and more ‘human’ and realistic is demonstrated by the following: most Internet users doubt their cognitive abilities when asked on social media to distinguish a bot from a human being or a deepfake generated by Sora7 from a real video recording.
- 8 During the DSE Communities conference at the Institute of Biology of the Polish Academy of Sciences (...)
5The ESTS conference program shows that the editing community is currently exchanging ideas on how to use AI to create scholarly editions. No longer digital scholarly editions, but intelligent editions, AI editions or at least AI-driven editions are becoming the strongest stimulus for the philological imagination.8 This approach does not seem to deviate from the general trend in other professions ‘threatened’ by competition with intelligent computers: instead of sticking to a losing position, it is better to start using AI for one's own purposes, to become its operator. Interest in artificial intelligence among humanists is growing, especially since it is supported by a strong current of quantitative literature research known as distant reading, which treats literary texts as large data sets to be processed (big data) – after all, generative artificial intelligence works in a similar way: it draws its ‘wisdom’ from enormous data resources. Digital source editing (born-digital) is also becoming increasingly popular, which should come as no surprise given that humanity has been creating various electronic documents on a large scale for over half a century and that the service known as the World Wide Web has been in existence for over thirty years. Countless digital works are stored on the web and on computer hard drives. Among these are works that are important for literary researchers and that require proper handling. However, it may be impossible to access these works without the help of an artificial intelligence assistant.
6This situation has led to a veritable flood of information at universities and research institutes. Tools and platforms are being developed independently in many places. They are often imperfect, ‘under construction’, ‘currently being transferred to cloud servers’, ‘only partially accessible’. These solutions are impossible to keep up with and generally not easy to implement in one's own editorial projects because they are not versatile or user-friendly enough. The Social Sciences and Humanities Open Marketplace catalogue9 gives an idea of how numerous and diverse these resources are.
7With so many tools at their disposal, editors are trying to reorganize their workshop and build work standards in a digital environment. This would involve carefully selecting existing services, applications or even useful scripts – for example, for automatic transcription, collation or annotation – which can then be accessed as needed. Such attempts result in case studies in which the authors describe the use of specific tools for creating editions.
- 10 Cf: E. Pierazzo, “What Future for Digital Scholarly Editions? From Haute Couture to Prêt-à-Porter”, (...)
- 11 R. Viglianti, G. del Rio Riande, “Against Infrastructure. Global Approaches to Digital Scholarly Ed (...)
8In view of this dynamic development, it is also important to consider the long-term viability of digital editing projects (future-proof editing), as there is a real risk of overloading the editing process with expensive, custom-made original solutions10 that no longer display correctly in browser windows after a few years. Some believe that the answer to these challenges is so-called minimal processing, which involves using the most standard programming and coding languages that are as independent as possible from changing internet standards;11 research data repositories are also being created where the source code of the edition can be deposited.
- 12 H. Hollender, “Czy świat czeka przyszłość średniowiecza?” [Is the world facing a medieval future?], (...)
9The absence of a single convenient method of digital editing, which discourages many researchers, is sometimes compared to the early days of printing – at that time there was also no catalogue of good practices for creating incunabula, or even rules for writing vernacular languages, but despite this, Gutenberg's invention gradually gained acceptance, changing the social relations and intellectual culture of subsequent generations of readers.12 As Peter L. Shillingsburg wrote:
- 13 P.L. Shillingsburg, From Gutenberg to Google. Electronic Representations of Literary Texts (Cambrid (...)
It is easy to get lost or discouraged in the field of electronic texts. Every new whoop-tee-doo in these areas soon becomes last week’s news in the face of even newer ones. We are tempted to wait out the turmoil, perhaps hoping to come in at the home stretch with the winners, like one who cheats in marathon races by joining for the last mile or two. The finish line, however, seems, like the horizon, to recede.13
10It seems that the emergence of a new factor in the form of artificial intelligence significantly changes the existing rules of the game, and that is precisely why the temptation to wait it out is something that should not be given in to too much.
- 14 The team coordinated by Magdalena Komorowska operates within the framework of the Digital Humanitie (...)
- 15 These are: revitalization of the Library of Old Polish and New Latin Literature “Neolatina” (https: (...)
11The team of the Digital Editing Laboratory (LabEdyt),14 founded at the beginning of 2023 at the Jagiellonian University, has set itself several tasks: experimenting with available tools to organize workflow systems, exchanging experiences as scientific digital editions are developed, constantly monitoring technological innovations, teaching students, and supporting scientists in the implementation of digital projects. LabEdyt is currently working on several pilot projects,15 one of which – devoted to Moralia of Wacław Potocki – explores the possibilities of using machine learning to automatically create transliterations, transcriptions and XML semantic tagging.
- 16 R. Grześkowiak, “Stary druk jako podstawa edycji krytycznej. Preliminaria” [Early printed books as (...)
- 17 Cf: “Potocki Wacław (1621–1696)”, in: Bibliografia literatury polskiej “Nowy Korbut” [The “New Korb (...)
12The 17th century in the Polish-Lithuanian Commonwealth was, according to researchers, the ‘age of manuscripts’. As Radosław Grześkowiak wrote: ‘the most interesting works of the era were entrusted to manuscripts and reproduced in an informal circulation’. This applied to the ‘works not only of such luminaries as Jan Andrzej Morsztyn, Wacław Potocki or Stanisław Herakliusz Lubomirski, but also of second-tier figures important for the history of our literature, such as Daniel Naborowski or Hieronim and Zbigniew Morsztyn’.16 Wacław Potocki, a nobleman from Łużna who was extremely prolific in his literary output, and in this respect has been compared to Józef Ignacy Kraszewski, left virtually all his works in manuscript form – including his most famous epic poem Transakcja wojny chocimskiej [The transaction of the Chocim war]. Only one major work was published at the end of his life: Poczet herbów szlachty Korony Polskiej i Wielkiego Księstwa Litewskiego [Coats of arms of the nobility of the Polish Crown and the Grand Duchy of Lithuania]. It was published in 1696 in the Cracow printing house of Mikołaj Aleksander Schedel. Other texts, including Moralia abo rzeczy do obyczajów nauk i przestróg w każdym stanie żywota ludzkiego z łacińskich i z polskich przypowieści ojczystym krótko napisane wierszem [Moralia, or things pertaining unto manners, lessons and admonitions for each estate of man’s life, briefly set down in native verse from Latin and Polish proverbs], remained unpublished until the following centuries. The abundant work of the Old Polish writer did not attract interest until the turn of the 20th century, and Aleksander Brückner made the greatest contribution to its popularization at that time.17
- 18 L. Kukulski, Prolegomena filologiczne do twórczości Wacława Potockiego [Philological Prelegomena to (...)
- 19 W. Potocki, Moralia, manuscript, ca. 1688–1696, The National Library, manuscript 3049 III, Polona.p (...)
- 20 E. Roterodamus, Adagiorvm Chiliades Des. Erasmi Roterodami Qvatvor Cvm Dimidia Ex Postrema Avtoris (...)
13From around 1688 until his death, Potocki worked on his most extensive work – Moralia – using a 1551 edition of Erasmus of Rotterdam's Adages printed by Froben in Basel,18 which is a collection of short poems composed around ancient maxims selected and commented on by Erasmus. The fair copy of Moralia, made by Potocki, stored in the National Library,19 has grown to an impressive size of 712 sheets, or 1424 pages, while a copy of Adages, which the writer used, as evidenced by his handwritten notes in the margins, was found among the duplicates of the Jagiellonian Library and subsequently transferred to the Warsaw Scientific Society (today the Library of the Institute of Literary Research of the Polish Academy of Sciences).20
- 21 A. Brückner, “Wacława Potockiego Moralia (1688), wyd. Tadeusz Grabowski, Jan Łoś, [Cracow] 1915–191 (...)
- 22 See, e.g.: L. Kukulski, Prolegomena filologiczne…, p. 14, footnote 28.
- 23 A. Brückner, “Wacława Potockiego Moralia…”, p. 161.
- 24 W. Potocki, Dzieła, vol. 3: Moralia i inne utwory z lat 1688–1696 [Works, vol. 3: Moralia and other (...)
14The work has only been published once in its entirety: it was edited by Tadeusz Grabowski and Jan Łoś and published in three volumes between 1915 and 1918 in the series Biblioteka Pisarzów Polskich [Library of Polish Writers]. This edition, which is now over a century old and therefore considered a historical document by contemporary readers, has been criticized from the outset for being inaccurate and too sparse in its explanations. Aleksander Brückner pointed out many shortcomings,21 and Leszek Kukulski, an expert on the work of the Sarmatian poet, added to the list of errors.22 Brückner, as if he had the gift of clairvoyance, pessimistically predicted that Potocki's work would not be published in a revised edition anytime soon: ‘so the excess of frugality has been achieved at the expense of comprehensibility, which is very regrettable, because Moralia will probably not see a better, more careful edition’23 – he wrote, and he was not wrong. Fragments of this collection appeared later only in selections, including the third volume of the extensive edition of Potocki's writings.24
- 25 R. Grześkowiak, Stary druk…, p. 12.
- 26 S. Grzeszczuk put it plainly: ‘Potocki has a hopeless advantage over an individual researcher, no m (...)
15Re-editing Moralia is a thankless task for many reasons. First of all, it is not an ‘unweeded garden’, to paraphrase the title of another collection by Potocki. Certainly, there is a greater temptation to break new ground and deal with unpublished texts. Moralia, having already had ‘some’ edition, lose the competition with works still awaiting publication. Besides, in order to edit them, one would have to refer to the manuscript, which editors, as diagnosed by Radosław Grześkowiak, are clearly not fond of.25 If we add to this the extraordinary size of the text (the question immediately arises: how to convince grant committees to finance the printing?), the necessity to examine the connections with Adages and the fact that most of the editorial work was done by Leszek Kukulski before his death (he was also the one who made the most interesting discoveries), it is easy to come to the pragmatic conclusion that in times of scholarly haste it would be difficult to devote so much time to studying one work.26
- 27 Cf: L. Kukulski, “Klucz do Moraliów” [The Key to Moralia], in: idem, Prolegomena filologiczne…, pp. (...)
- 28 “KorBa” electronic corpus of Polish texts from the 17th and 18th century, https://korba.edu.pl.
- 29 J.S. Gruchała, Wacław Potocki…, p. 19.
16However, digital editing tools may hold some hope for Moralia. They make it possible to create editions faster and more accurately, at a lower cost. Such an e-edition further expands the possibilities of printing, as it can include different versions of the text (transliteration, modernizing transcription, text with editorial commentary, facsimile of the manuscript, perhaps also images of the Froben's edition of the Adages), create an interactive ‘key to Moralia’,27 provide thematic indexes, an advanced search module or even a frequency list – all without having to worry about printing sheets. Tagging the text and integrating it with existing databases (such as WikiData) also links it to the Linked Open Data network, which is a collection of open, linked data on the Internet. Creating these links enables further research, especially with the help of AI tools. Chatbots that use large language models, such as the Polish Bielik AI, are also able to read the edition and support its users. Since tagged text fragments have a semantic surplus in the form of metadata, they can be processed by a computer in a more advanced way than in the case of plain text. The editors of language corpora such as KorBa28 could benefit from such a version of Moralia, for example. Even though the corpus does contain excerpts from Potocki's works, they are presented in a modernized form, as Janusz Gruchała stated: ‘popular scientific rather than anything else’.29 However, before we can start thinking about the benefits of digital editions, we first need to source the text.
- 30 P.L. Shillingsburg, From Gutenberg…, p. 27.
17Transliteration, also known as ‘diplomatic transcription’, is the basis for any editing of texts ‘born’ before the digital age – and will probably also be used in the future for many works that were written on a computer but not saved in digital form. Many have written about its indisputable significance. Peter Shillingsburg described transliteration as a form of ‘reincarnation’ – the subsequent embodiment of a certain intangible idea, i.e. a text. ‘Reincarnation’ is therefore the adoption of a new ‘body’ by the text (however puzzling it may sound in the context of digital space).30 In editing, according to the Platonic concept, this embodiment of the perfect idea becomes its corruption at the same time, because matter always remains imperfect. In the process of ‘reincarnation’, mistakes are inevitable – and everyone makes them: copyists, editors, proofreaders, typesetters and their modern counterparts – DTP graphic designers, printers, bookbinders and even the author, making unintentional slips of the pen.
- 31 A. Brückner, “Wacława Potockiego Moralia…”, pp. 159–161.
18In the case of Moralia, it is not difficult to make an error when rewriting. The enormity of the rather monotonous material is conducive to mistakes. This is all the more likely, the more people are involved in copying a single work – there exists a risk that not everyone will be able to adhere to the accepted arrangements despite their best intentions. In addition, Moralia is a text that is convoluted from the perspective of today's reader, full of archaisms, and therefore incomprehensible in places – a great deal of this is due to baroque poetics. This was already Aleksander Brückner’s opinion about the work a hundred years ago,31 so what can a 21st-century Polish speaker say? Failure to understand the text can lead to erroneous, hasty readings, and the form of the message – a manuscript – increases the difficulty of the task, although it must be emphasized that Potocki's handwriting is legible.
Fig. 1. Sample of handwriting from the first pages of the Moralia manuscript. Source: Polona (scan 4v–5r).
- 32 Cf: J. Łoś, “Wstęp” [Preface], in: Wacława Potockiego “Moralia” 1688, vol. 3, ed T. Grabowski, J. Ł (...)
Fig. 2. Sample of handwriting from the last pages of the Moralia manuscript – fragments written by a ‘boy copyist’ come from the Ogród fraszek32 [Garden of epigrams]. Source: Polona (scan 641v–642r).
19With these reservations in mind, the Digital Editing Laboratory team decided to perform the transliteration using an automatic handwriting recognition tool – the Transkribus application. This program (although in its current state of development it should probably be called a SaaS – Software as a Service – tool, as the so-called desktop client has already been discontinued) uses user-prepared samples of images and their associated transcriptions to train specialized text reading models. The ‘training’ itself consists of applying a basic artificial intelligence function, namely machine learning, to the provided training set. The trained model can be saved in a private library assigned to the account or shared with the Transkribus community.
20The application offers a number of ready-made models that recognize both handwritten and printed texts, but none of them were suitable for transliterating Moralia. When selecting a model, several parameters must be taken into account. Of course, the most important is the close visual similarity between the text we want to automatically transliterate and the text used to train the model. In practice, this means that a model for English cursive (Copperplate) will not work on a sample written in uncial – similarly, if the author's handwriting differs even slightly from that of the model, the results of automatic transliteration will not be satisfactory. Apart from this obvious issue, it is also worth remembering the CER (Character Error Rate), which determines how often the computer makes mistakes when reading characters; the higher the CER (above 5%), the less accurate the reading. A CER of 15% means about fifteen incorrectly recognized characters per hundred, or almost fifteen corrections per standard line of 12-point Times New Roman text in Microsoft Word – this is a lot, so correcting such a transcription can take more time than creating it from scratch.
- 33 This is not always desirable. For example, the Polish Schwabacher model, designed to generate trans (...)
- 34 As of December 2024.
21When considering a ready-made solution, one should also take into account the language of the text on which the AI was trained. A model trained on English-language documents will certainly not recognize Polish diacritical marks and will also make more mistakes because it has not learned the character combinations that are typical for Polish but absent in English. What is more, observations of transcriptions generated in Transkribus show that the model learns not only characters, but also the shape of entire words, which is why it is able to correct scanned text to a certain extent (!).33 Few public models in Transkribus are suitable for use by Polish editors. In the gallery of ready-made solutions containing over two hundred models – from Church Slavonic to Tibetan cursive – only three are intended for Polish.34 Two of these are large models that are constantly being developed and fed with large data sets – one for print, the other for manuscripts (multilingual Transkribus Print M1 – CER 2.2%, Transkribus Polish M2 – CER 4.1%). The third model, The Polish Schwabacher (CER 0.87%), was developed in early 2024 by editing students at the Jagiellonian University as part of a course on creating digital scholarly editions and is used to automatically transliterate Schwabacher.
22To achieve the best possible results, a new model specifically designed to read Potocki's manuscripts had to be trained. This involved preparing a sample of real data (ground truth) that the algorithm would treat as an ideal model to follow – in other words, even if we wanted to automate the work, we first had to do a considerable amount of it ourselves. The transliteration of the one first hundred pages of the manuscript was done by Lidia Nowak, a graduate of editing and currently a doctoral student at the Jagiellonian University Doctoral School in the Humanities. The manuscript had to be read as accurately and unambiguously as possible, because in the ground truth sample, each character in the manuscript should have a fixed equivalent in the transliteration – otherwise, the AI training would not be as effective. For a computer, every character is equally abstract and meaningless: if we consistently show it that the capital letter A on the scan is actually a lowercase g, it will begin to recognize it as such.
23In accordance with previously adopted rules, the transliteration of Moralia, among other things:
-
retained the layout and breaks in pages, lines, and marginalia,
-
retained punctuation marks appearing in the original (without retaining the inconsistent spacing preceding these marks),
-
decomposed ӕ, œ, & ligatures into ae, oe and et,
-
standardized the three variants of the grapheme z, ƶ, Ʒ to z (analogously to the sign ż – if the variants had a dot above),
-
rendered long s as ſ,
-
retained the author's decisions regarding the use of capital letters (including the inconsistent but clear distinction between capital letters I and J),
-
expanded abbreviations,
-
used < > brackets to mark damaged and illegible places.
Fig. 3. Transliteration made in Microsoft Word.
- 35 The activities described took place in the first half of 2023, when Transkribus was still operating (...)
24The transliteration was done in Word, so it had to be transferred to Transkribus and linked to the scans of the document. This procedure was carried out in several stages. First, scans of Moralia were downloaded from the Polona Digital National Library, then, using a batch function (a feature of Photoshop), the scans were reduced in size and divided into separate files with individual pages. The edited graphics were transferred to Transkribus35 servers, where the next step was to build and organize the page layout. The program offers an automatic layout recognition feature and detects columns and lines of text on its own, but the editor must check this process because automatic recognition is not perfect. Often lines of text are out of order, baselines are broken, or a single line is interpreted as two separate lines (for example, when the author used larger than usual spacing between words). It is particularly important to ensure that the baselines reach the end of the text lines, as characters not on the baseline will not be read later. At the layout stage, much depends on the quality of the scan. For example, if the digitized book did not open fully, the inner part of the column on the scan will ‘curl’ towards the spine. Such curling causes two problems for Transkribus: first when recognizing the layout, and later when reading the characters (because they are distorted and tilted – it is not without reason that CAPTCHA images displayed on some websites and in electronic forms as protection against bots contain ‘wavy’ letters that humans can recognize without any problems, but computers cannot... at least until recently).
25The finished transliteration of Moralia was assigned to previously mapped text fields and lines of text. Of course, Transkribus allows users to enter transliterations directly in the browser window, but preparing a sample in Word first has its advantages, primarily related to the more advanced text editing options available in this editor (searching, replacing, using regular expressions); its stability, which still is an issue in Transkribus, is also important.
Fig. 4. Transkribus browser application – preview of the facsimile and transliteration. On the pages of the manuscript, apart from the main text, text fields with pagination, marginalia and a running header have been highlighted.
26The creators of the service suggest in the documentation that a sample of five to fifteen thousand words is sufficient to train an effective model.36 Printed text recognition models usually perform well even on small samples, while manuscripts require much more material. Originally, according to the instructions, the LabEdyt team intended to use fifty pages of the manuscript (approximately fifteen thousand words), but the model obtained from this sample had a fairly high CER of 5.9%. It was only after adding an additional fifty pages of transliteration that the CER was reduced to a satisfactory level of 2.6%.
Fig. 5. Validation set for the Potocki-2 model. On the right, text read by a computer, no editor intervention.
- 37 Transkribus allows data to be exported directly to a TEI file, but this option is only available to (...)
27What conclusions can be drawn from this experiment? First and foremost, it is important to consider when it is really worth investing time in training an AI model to recognize text. Had the Moralia manuscript been 300 pages long, it might have been quicker, easier, and more accurate to transcribe it manually into a text file. After all, the use of artificial intelligence does not relieve the scholarly publisher of the obligation to carefully proofread the text several times; the preparation of the training sample itself is also time-consuming. Nevertheless, the procedure presented here has a significant advantage if we intend to create a digital scholarly edition in the long term. Transkribus' user-friendly graphic interface is equipped with useful tagging functions, thanks to which it is possible, already at the transcription stage (or transliteration, depending on the user's needs), to mark illegible, damaged, deleted, supplemented or corrected places, and also to indicate elements of the source text's structure, such as pagination, running headers, catchwords, signatures, marginalia, titles, and the like. Recently, it has become possible to use the <persName> and <placeName> tags, as well as to define one’s tags. All this can be achieved without entering a single line of code, which is an undoubted advantage for editors who are reluctant to use markup languages. The text obtained in this way is not actually plain text (although it can be downloaded from Transkribus, as the application allows export in many different formats), but has an additional layer of data that will later be included in the exported Page XML file converted to TEI XML.37 This greatly facilitates work on digital scholarly editions, as it allows one to obtain pre-tagged and mapped files in a semi-automatic way, in which lines of text or even individual words are assigned coordinates on the digital facsimile.
- 38 Dynamic changes in access to software pose a significant threat to projects financed under the gran (...)
- 39 The developers of TEI Publisher, a tool for publishing digital scholarly editions, have announced p (...)
28Currently, the correction of the transliteration of the entire Moralia is nearing completion. The next stage of work will begin soon – the semi-automatic creation of a transcription modernizing the spelling. To this end, the team intends to use a word frequency list obtained by processing a clean text file using a simple script written in Python. While work was underway on reading Potocki's manuscript (the Potocki-2 model was created on 16 June 2023), significant progress was made in this field. Transkribus itself has undergone a huge transformation, evolving from a simple, virtually free tool offering unrivalled functionality at the time to a commercial service operating in a freemium model.38 However, it is worth remembering that digital scholarly editing has exactly the same goal as traditional editing – the scholarly preparation of a text in accordance with accepted guidelines. Therefore, one should not overly fetishize digital humanities tools or become particularly attached to them. If a program that better fulfills its purpose appears today or in the near future, and there are already opinions39 that it could be made available under an open license such as eScriptorium, then nothing will prevent subsequent transcriptions of literary works from being created in it.
*
- 40 A. Frutiger, “OCR-B. Znormalizowane pismo o czytelności optycznej” [OCR-B: A standardized character (...)
- 41 Ibidem, p. 102.
- 42 Ibidem.
29In the 1960s, Swiss typographer Adrian Frutiger collaborated with the European Computer Manufacturers Association (ECMA) to develop a typeface that would be recognizable by optical readers. This resulted in the ‘standardized Latin alphabet’ OCR-B (Optical Character Recognition – font B), which was a compromise between the requirements of the digital machines of the time and the centuries-old tradition of typography.40 The designer tried to draw the individual characters of the numbers and alphabet so that none of them (for example, I and l or B and 8) were similar to each other, as ‘flawed’ computers could not cope with small differences in shape. Even then, Frutiger had no doubt that ‘one day, a reading machine will become so advanced that it will be able to read the characters and forms of any style of contemporary alphabets without error’.41 The typographer considered his work on the OCR-B typeface, which was not particularly attractive, to be a ‘success in the field of ethics’, because ‘it is not the machine that forces man to use a “mechanized” style of characters; it is man who tries to “teach” the machine to read the writing that is in common use, the writing that has developed over centuries from stone hieroglyphs, through the pen and parchment, the engraver's chisel, to the current methods of graphic designers, typesetters, and printers of our time’.42 Is that so?
30The artificial intelligence revolution we are witnessing today is the future Frutiger dreamed of. Its course raises many ethical and existential questions: about the future of existing professions, about ecology (supercomputers consume enormous amounts of energy), about data privacy. Scholarly editors are also asking themselves these questions. Automatic transcription or transliteration using AI seems like a fairly innocent procedure. It is obvious that in the end the text will be read and corrected by a ‘human instance’. But what about attempts to create commentaries with the use of chatbots? Who will be the author of such an edition, who will take responsibility for it – the editor writing the prompts or the computer processing the knowledge of all humanity? Is a text written at our request our text? Will the editor continue to be an editor – the one who knows the work best – or perhaps a specialized machine operator?