Skip to navigation – Site map

HomeIssuesIssue 183D Text Encoding and TEI

3D Text Encoding and TEI

Text, Editions, and Spatiality
Jun Ogawa, Kiyonori Nagasaki, Ikki Ohmukai and Asanobu Kitamoto

Abstract

This paper explores the integration of 3D technology in digital humanities, focusing on the increasing enthusiasm for 3D research practices in the field. It discusses the adoption of 3D methods in research and the challenge of advancing the field, particularly the representation of “text” in 3D models, such as inscriptions. The paper suggests integrating TEI (Text Encoding Initiative) markup with 3D space to represent rich text information in three-dimensional space, using inscriptions as a case study. It outlines methods for 3D text encoding, leveraging TEI, especially using the EpiDoc schema and extending it for 3D applications. The paper emphasizes the advantage of adopting the genetic edition–like markup using <sourceDoc>, to document spatial data, particularly 3D coordinates, to fully represent text in 3D space. The paper contributes to the discussion on encoding textual data in 3D environments, proposing adjustments to existing schemas to accommodate 3D information and highlighting the need for further development in this emerging field.

Top of page

Full text

1. Introduction

1Over the past thirty years, there has been a burgeoning discussion on the role of 3D technology within the realm of digital humanities. This conversation explores not only the creation of 3D models through scanning and modeling techniques, such as LiDAR, photogrammetry, and 3DCG software like Blender or SketchUp, but also their integration and use within the broader framework of humanities research and education (Münster 2022, 1).

2The field of archaeology, in particular, has shown significant enthusiasm for adopting 3D technologies. Numerous institutions, projects, and even individual archaeologists are actively embracing the challenge of advancing their research through the application of 3D methods. For instance, the Archaeology Data Service (ADS) has long been engaging in the application of 3D technologies to archaeological study,1 and more recently, the Extended Matrix advocates a methodology that integrates archaeological principles such as stratigraphy to describe and visualize 3D data (Demetrescu 2018, 102). This initiative now involves a number of partners who are committed to and actively developing this approach. Similarly, the PURE3D project proposes expanding the influential concept of “scholarly editions,” which is closely related to the activities conducted in the TEI, to 3D elements. It promotes the creation of 3D scholarly works and eventually academic discussion forums that harness the immersive advantages of 3D, presenting scholarly information in three-dimensional environments (Schreibman and Papadopoulos 2019).

3Research into 3D visualization–oriented academic communication naturally progresses alongside efforts to establish the foundation for constructing 3D data. This is because building an academically robust 3D environment and facilitating scholarly discourse within it requires a comprehensive data infrastructure that encompasses precise information. These data comprise various elements crucial for academic inquiry, including contextual details such as the equipment used for model creation and the scanning conditions, as well as parameters like polygon counts and textures of the generated models, along with annotations containing scholarly interpretations. To address this need, the ADS is developing data models to meticulously document the para- and metadata around 3D models during their creation process (Trognitz et al. 2016, table 3). Similarly, the above-mentioned initiatives like the Extended Matrix and PURE3D are presenting models for structuring information relevant to their respective endeavors. Additionally, there are efforts to represent such data as Linked Open Data (Kettler 2021). For instance, SCOTCH, proposed by V. Vitale, introduces an ontology for articulating the “multiplicity” of interpretations associated with 3D models and spaces as RDF knowledge graphs (Vitale 2017).

  • 2 The term “3D humanities” is well explained on the website of the National Endowment for the Humanit (...)

4While some studies have addressed methods for structuring and leveraging data associated with 3D materials, with careful consideration given to data structuring and management approaches, there has been insufficient exploration into how to represent the “text” in 3D models. This gap is particularly noticeable when considering materials such as inscriptions. Although methods for reconstructing and annotating them as 3D models have been explored, well-accepted techniques for handling text information in 3D space have yet to be proposed. By text information, we mean annotating models not merely with literal strings, but rather with “rich” text data, imbued with layers of textual interpretation and critique, as TEI has long done. We contend that representing such rich text information in 3D space is a crucial task and challenge for advancing the emerging field of “3D humanities.”2

Figure 1. Image of 3D encoded texts in TEI/XML displayed in 3D space.

Figure 1. Image of 3D encoded texts in TEI/XML displayed in 3D space.
  • 3 The inscription was published in Bernard Rémy et al. (2004), Inscriptions Latines de Narbonnaise (I (...)

5For example, figure 1 shows an image of 3D-encoded texts placed within a three-dimensional space (Ogawa et al. 2023). The inscription itself was discovered in Gallia Narbonensis,3 one of the provinces of the Roman Empire, and is positioned here within a model of the Forum of Augustus. Of course this particular combination of the inscription and the space would not actually have existed, but it is a historical fact that many inscriptions were installed in the forum and people encountered them in their daily lives, so the situation sufficiently reflects historical reality. Yet how, in practice, did people interact with and respond to such spatialized texts? To address this question, if textual information can be displayed in a 3D space at appropriate positions—presenting, in parallel, multiple possible readings and emendations—it becomes possible to explore textual interpretation within the three-dimensional environment not only on the basis of traditional textuality, but also by incorporating aspects that depend heavily on materiality and spatiality, such as the visibility of the text and the spatial relations among texts. Here, the advantage of integrating such “rich” textual information with materiality becomes evident.

6As mentioned earlier, as a long-standing international standard for digital text encoding in the humanities, TEI has long focused on digital representation of rich text information. It follows that to facilitate the 3D representation of such information, one might explore integrating TEI markup with 3D space. While this may appear to be a straightforward solution, there have been few full-swing attempts to describe such three-dimensional information in TEI thus far, as pointed out by Vitale (Vitale 2017, 30). Therefore, this paper will investigate methods for representing 3D information in TEI by extending the concept of <sourceDoc>, using epigraphic materials as a case study. To begin, the following section will outline why inscriptions were selected as a case and review previous research on data structuring associated with them.

2. Why Take Inscriptions As an Example?

2.1 Inscriptions As Physical 3D Objects

7Inscriptions, with their distinctive characteristics as historical sources, offer an excellent opportunity to explore methods for encoding 3D text in TEI. One notable aspect of inscriptions, differing from other forms of documents, is their inherent “materiality.” In antiquity, inscriptions adorned forums, sanctuaries, and various other public and private spaces, serving diverse functions ranging from tombstones to commemorative monuments and public decrees. They varied greatly in size, from small plaques of less than 30 centimeters to relatively huge monuments exceeding several meters. Their impact on viewers was paramount, suggesting meticulous consideration of size, shape, and placement. Consequently, when interpreting inscriptions, it is not only the textual content that matters but also the insights gleaned from their materiality. They hold significant value, as Cenati et al. indicate when they wonder, “How were these texts experienced and ‘felt’ by those who were trying to make sense of the inscribed environment that they inhabited?” (Cenati et al. 2022, 144–45). Additionally, physical details such as surface wear or nail marks often provide clues about their original installation locations.

  • 4 In an ancient Roman context, projects like Rome in 3D or Rome Reborn have already made some attempt (...)
  • 5 Such attempts to change the appearance of objects depending on the way light hits them and simulati (...)

8Because of this intrinsic material quality, there is substantial merit in representing an inscription as a 3D model. For instance, comprehending an inscription’s dimensions is easier when it is visualized in 3D, and even more so if it is experienced in an immersive environment like virtual or augmented reality (VR/AR),4 rather than merely viewing it as a 2D image. Moreover, in terms of physical placement, recreating the surrounding environment to position the inscription within it could enhance understanding of the spatial context. Furthermore, when examining the text itself, it remains crucial to consider both its materiality and its readability. For instance, if an inscription was mounted high on a temple wall, inquiries arise regarding the distance from which the text could have been legible. Additionally, readability may have varied depending on factors such as lighting conditions.5

2.2 TEI Markup of Epigraphic Texts

9Thus examining inscriptions, which functioned as both texts and objects, in 3D offers substantial advantages in terms of interpretive breadth. At the same time, however, traditional philological endeavors, such as reconstructing texts through the meticulous creation of a critical apparatus, remain indispensable in epigraphic studies. Particularly with inscriptions, where sections may be illegible due to poor preservation, and where unclear omissions can lead to significant variations in interpretation, accurately describing such nuances as data is crucial. Therefore the challenge lies in effectively integrating detailed textual information with 3D models or spaces. Achieving this integration would enable the investigation of unreadable or difficult-to-read sections of inscriptions within 3D space while referencing their apparatus, potentially revealing insights overlooked in flat 2D images.

10For epigraphic text editing, EpiDoc, a subset of TEI, offers a meticulously designed schema for annotating inscriptions, encompassing elements like text lacunae, wear, abbreviations, symbols, and even stone decorations, while leveraging the existing tagsets defined in TEI (EpiDoc 2023). Similarly, SigiDoc, which is a derivative of EpiDoc specifically designed for Byzantine seals, is proposed (SigiDoc 2023). For these schemas, the availability of publishing systems like EFES indicates the existence of a comprehensive ecosystem, spanning from text markup to publication.6 Their adoption is widespread across numerous databases so far developed, establishing itself as the de facto standard for producing digital scholarly editions of these materials.7 Considering this context, it seems fitting for this study to adhere as closely as possible to these schemas during the markup process. However, given that they have so far primarily catered to 2D text or image editing, they lack adequate provisions for representing inscriptions as 3D objects. Consequently, the following sections will explore the integration of text data structured according to TEI, specifically EpiDoc, with 3D space, exemplifying this through real instances of 3D text encoding.

3. Method for 3D Text Encoding

3.1 Concepts and Principles

11The primary approach advocated in this paper for 3D text encoding involves leveraging existing TEI schemas that are already designed to describe the shape or physical attributes of 2D materials by slightly extending them to suit our specific needs.

12The basic premise of this research aligns closely with Chapter 11 of the TEI P5 Guidelines, titled “Representation of primary sources,” which serves as a pivotal point of reference (TEI Consortium 2023, chapter 11). This chapter extensively addresses the description of the physical attributes of the source materials themselves. In Chapter 11, the <facsimile> element takes precedence, serving to represent digital images of original materials as structured data rather than transcribed text. By portraying the script and spatial layout of the original materials through images, the focus remains firmly on their physical characteristics. Additionally, TEI facilitates the description of <surface> as a subcomponent of <facsimile>, and <zone> as one of <surface>, with predefined attributes aimed at conveying the positional data of text within the materials using x-y coordinates. The guidelines’ examples illustrate the use of <facsimile> for digital images of entire works, <surface> for images of individual pages, and <zone> for delineating specific areas on a page. While these elements are primarily designed for 2D image materials, TEI’s interest in capturing the form and spatial arrangement of text remains evident. But while the use of <facsimile> and its associated elements aids in delineating the visual and spatial attributes of text and images within materials, it does not inherently link such spatial information with transcribed text data.

13The TEI Guidelines offer two principal approaches to integrating visual and spatial information with transcriptions. One method, known as parallel transcription, involves segregating the descriptions of <facsimile> and transcribed text data, while establishing a referential relation between them. Here, the relation is delineated primarily from the perspective of the text data, employing @facs. The other approach, termed embedded transcription, incorporates the transcription directly within elements used to describe visual and spatial information, thereby combining both aspects. When employing this method, elements like <surface> and <zone> are used within a parent element named <sourceDoc>, which contains transcription text alongside, or in place of, image data of the materials. Structuring text data using <sourceDoc>, which is defined in TEI guidelines to contain “a transcription or other representation of a single source document potentially forming part of a dossier génétique or collection of sources,” places greater emphasis on the appearance and layout of text in primary materials, as compared to using the general <text> for description.

14In pursuit of the goal of integrating 3D space and text data, either the parallel transcription or the embedded transcription method could be employed. This study, however, opts for the latter approach using <sourceDoc>, based on the conceptualization of 3D text within this research that emphasizes its materiality. To clarify this decision, it is beneficial to refer to the notion of the “genetic edition,” which has been a topic of discussion in TEI-related research.

15Genetic edition is an approach that aims to systematically organize the diverse editorial interventions conducted during the creation process of a text, including actions like cut-and-paste, symbol insertion, quotations, annotations, corrections, deletions, or even physical alterations. Prominent examples in the digital humanities include the Samuel Beckett Digital Manuscript Project,8 Elena Pierazzo’s project on Proust’s Cahier 46,9 and the very impressive practices of the Shelley-Godwin Archive.10 In constructing such genetic editions, the TEI Guidelines offer structuring methods that use <sourceDoc>, and more recently there has been a growing trend to employ this method in research focused on modern Japanese novel drafts (Shioi and Nagasaki 2023). Given that novel drafts often feature numerous traces of deletions, corrections, marginal notes, and occasionally irregular text layouts, the choice of a “document-oriented markup” using <sourceDoc> appears to facilitate the faithful reproduction of these drafts.

16We encounter similar complexities when considering the textual content of inscriptions. Inscriptions often exhibit textual transformations, such as inconsistencies in the inscription period or amendments and erasures on the stone surface. Prominent instances include additional decrees inscribed on Athenian decree stones or the Roman custom known as damnatio memoriae, which frequently involved the removal of condemned emperors’ names from stones, and sometimes even the obliteration of other inscriptions. Given the potential to address the dynamic process of text creation that evolves over time and changing environmental conditions, the markup of textual data in this study is guided by insights from the construction of genetic editions using <sourceDoc>. Consequently, textual information in this study is primarily delineated within <sourceDoc>, supplemented by 3D spatial data.

3.2 Markup Method

17As noted earlier, the foundation of 3D text encoding in this study rests on the markup of epigraphic texts using the vocabulary offered by EpiDoc. Consequently, we will leverage the TEI/XML resources provided by the Epigraphic Database Heidelberg (EDH) and expand upon them to illustrate a practical example of 3D text encoding.

Example 1. TEI XML-encoded text provided by EDH (https://​edh.​ub.​uni-heidelberg.​de/​edh/​inschrift/​HD019338).

<text>
  <body>
    <div type="edition" xml:space="preserve">
      <head>Text</head>
      <ab>
        <lb n="1"/><expan><abbr>D</abbr><ex>is</ex></expan><expan><abbr>M</abbr><ex>anibus</ex></expan>
        <lb n="2"/><gap reason="lost" extent="unknown" unit="character"/>MI
        <lb break="no" n="3"/><gap reason="lost" extent="unknown" unit="character"/>R
        <lb break="no" n="4"/><gap reason="lost" extent="unknown" unit="character"/>XXXXVII<expan><abbr>stip</abbr><ex>endiorum</ex></expan>
        <lb n="5"/>XXVIII<expan><abbr>Iul</abbr><ex>ius</ex></expan>Martialis
        <lb n="6"/><expan><abbr>dupl</abbr><ex>icarius</ex></expan>alae Britan<expan><ex>n</ex><abbr>ic</abbr><ex>a</ex><abbr>e</abbr></expan>
        <lb n="7"/>heres Ermius
        <lb n="8"/><expan><abbr>lib</abbr><ex>ertus</ex></expan>eius</ab></div>
    <div type="commentary">
      <p/>
    </div>	
    <div type="bibliography">
      <listBibl>
        <bibl>AE 1955, 0132.</bibl>
        <bibl/>
      </listBibl>
    </div>
  </body>
</text>

3.2.1 Use of <sourceDoc> and Overall Structure

18In the data provided by EDH, the text is encapsulated within the <text> element, as depicted in example 1, with details such as supplements for omissions, lacunae, and annotations of additions and deletions meticulously marked up. This study adopts a different approach, opting to represent the text itself within the <sourceDoc>. Consequently, segments within the existing <text> identified by @type="edition" will be newly transcribed into a <sourceDoc>. Although the current EpiDoc Guidelines do not explicitly accommodate storing text data within <sourceDoc>, this study will proceed with such markup as a case study. Given the discourse surrounding genetic editions, the authors advocate allowing the use of <sourceDoc> within the EpiDoc schema. It is important to note that elements with @type="commentary" and @type="bibliography", lacking a physical entity akin to inscribed characters or text on a monument, would be unsuitable to be included within <sourceDoc>.

19Moving forward, the composition of the <sourceDoc> will adhere primarily to the TEI Guidelines. Since the current TEI Guidelines do not account for the representation of 3D objects or spaces, however, it is necessary to clarify the role of <surface> and <zone> within <sourceDoc>. According to the TEI Guidelines, <surface> is defined as “a written surface as a two-dimensional coordinate space, optionally grouping one or more graphic representations of that space,” which implies the presence of a single surface containing text or symbols. But this definition presupposes a predominantly two-dimensional context, as suggested by the term “surface” and the guidelines’ explanation. Conversely, <zone> is defined as “any two-dimensional area within a surface element,” making it suitable for delineating specific segments within a single surface. For instance, a particular page of a manuscript could be designated as a <surface>, with each of potentially multiple illustrations on that page marked as an individual <zone>. Additionally, there exists an element called <surfaceGrp>, positioned hierarchically above <surface>, which clusters multiple surfaces based on the unit that the data creator considers a cohesive entity. This is particularly applicable in scenarios where a two-page spread is regarded as a single unit.

20When applying the <surface> and <zone> within the <sourceDoc> to inscriptions or any 3D objects, how should they be conceptualized? In this study, we consider that a <surface> is designated as the element corresponding to the entire 3D object. Directly beneath this <surface>, a <graphic> will be used to refer to the model file that represents the 3D object. Under the <surface>, all the areas, that is to say either semantic or textual divisions, are encoded with <zone>. For example, if the front face of an inscription encompasses both text and relief elements, each would be demarcated with its respective <zone>, as shown in figure 2. Naturally, this constitutes no more than a semantic division, as construed by the encoder, for it is impossible to establish exact areas of reliefs or texts, and any designation of such regions can only be regarded as provisional.

Figure 2. Front surface of an inscription containing both relief and text (https://​edh.​ub.​uni-heidelberg.​de/​edh/​foto/​F000013).

Figure 2. Front surface of an inscription containing both relief and text (https://​edh.​ub.​uni-heidelberg.​de/​edh/​foto/​F000013).

21With these redefinitions in mind, let us examine how the TEI/XML data from EDH would be restructured to fit our 3D text encoding (example 2).

Example 2. Proposed method for encoding 3D text information in <sourceDoc>.

<sourceDoc>
  <surface xml:id="inscription01" type="3Dobject">
    <figure>
      <graphic type="model" subtype="glb" url="inscription01.glb"/>
      <graphic type="texture" url="texture.png"/>
    </figure>
    <zone xml:id="inscription01_front">
      <zone xml:id="inscription01_front_relief"/>
      <zone xml:id="inscription01_front_text" points="0,1,0 1,0,0 -1,0,0 0,-1,0">
        <lb n="1"/><expan><abbr>D</abbr><ex>is</ex></expan><expan><abbr>M</abbr><ex>anibus</ex></expan>
        <zone xml:id="inscription01_ft_l2" points="0,0,0 0.5,0,0 -0.5,0,0 0,-0.5,0">
          <lb n="2"/><gap reason="lost" extent="unknown" unit="character"/><zone>M</zone><zone>I</zone>
        </zone>
        <zone xml:id="inscription01_ft_l3" points="0,-0.5,0 -0.5,0,0 0.5,0,0 0,0.5,0">
          <lb n="3" break="no"/><gap reason="lost" extent="unknown" unit="character"/><zone>R</zone>
        </zone>
        <lb n="4" break="no"/><gap reason="lost" extent="unknown" unit="character"/> XXXXVII<expan><abbr>stip</abbr><ex>endiorum</ex></expan>
        <lb n="5"/> XXVIII <expan><abbr>Iul</abbr><ex>ius</ex></expan> Martialis
        <lb n="6"/><expan><abbr>dupl</abbr><ex>icarius</ex></expan> alae Britan<expan><ex>n</ex><abbr>ic</abbr><ex>a</ex><abbr>e</abbr></expan>
        <lb n="7"/> heres Ermius 
        <lb n="8"/><expan><abbr>lib</abbr><ex>ertus</ex></expan> eius 
        </zone>
    </zone>
    <zone xml:id="inscription01_back"/>
  </surface>
</sourceDoc>

22It is apparent that while the introduction of <sourceDoc> has altered the overarching structure, the sections containing text information continue to use the original markup. By embedding text markup within <sourceDoc> that denotes the physical attributes of the original materials, text information can be seamlessly integrated with 3D space. But there is room for discussion regarding the markup of textual lines. According to the TEI Guidelines, <line> is preferred over <lb> for line markup within <sourceDoc>, aligning with the intention of <sourceDoc> to represent the physical layout and spatial relations of text. On the contrary, within the EpiDoc schema, line markup predominantly employs <lb>, and the use of <line> is not anticipated. The advantage of <lb> lies in its facilitation of markup across lines without obstruction and its ability to document word splits by line with @break. Technically, both methods are viable, but this study prioritizes the benefits of <lb> in facilitating line-crossing markup and will adhere to EpiDoc’s use of <lb> for line markup. It is worth noting, however, that marking up lines within <sourceDoc> is not mandatory, and for text segments that do not constitute lines, direct description within <zone> or structuring with <seg> is also an option.

23Furthermore, concerning the markup using <zone>, example 2 only denotes relatively broad divisions such as “relief zone” and “text zone.” But given that TEI permits nesting within <zone>, the areas represented by this tag can be described more flexibly. For instance, if there is a need to represent spatial information for each line, corresponding <zone> tags can be depicted as illustrated in example 3, and if spatial information for each character is required, corresponding tags for each character can be depicted similarly. This enables the adaptable depiction of 3D spatial information, ranging from relatively sizable areas containing a certain amount of text to exceedingly fine elements such as individual lines, words, and even each character.

Example 3. <zone> representing each line in the text.

<zone>
  <lb n="2"/><gap reason="lost" extent="unknown" unit="character"/><zone>M</zone><zone>I</zone>
</zone>
<zone>
  <lb n="3" break="no"/><gap reason="lost" extent="unknown" unit="character"/><zone>R</zone>
</zone>

3.2.2 Description of 3D Coordinates and Rotation

24When documenting essential 3D spatial information, it is imperative to include the z-axis alongside the traditional x- and y-axes. Additionally, a challenge arises concerning the rotation angle, which in two dimensions is insignificant. There are no doubt several possible solutions to achieve this, but here we examine two of them.

25The first option is based on 2D surface description. In this approach, we first define the central point of the surface in the <zone> element with the existing attributes @ulx and @uly for the x- and y-coordinates, while the nonexistent z-axis can be indicated using @ana="z:" for its value. While we are provisionally using the attributes @ulx and @uly, since they are formulated with two-dimensional coordinate representation in mind, it would be essential in the future to introduce new attributes to denote the x- and y-coordinates in a 3D context, along with the currently absent z-coordinate. In case these attributes are added to the schema, they might be part of an att.coordinated class, and the description of this class in the current guidelines should be modified to “att.coordinated provides attributes that can be used to position their parent element within a two- and three-dimensional coordinate system.”

26Regarding text rotation, the current TEI Guidelines already propose using CSS transforms within the @style attribute, for example <ab style="transform:rotate(-45deg)">. But while using CSS for such descriptions is effective for reusing style declarations within XML files intended for text display in browsers, this advantage does not extend to 3D visualization contexts. In fact, it could complicate subsequent processing, particularly if rotations occur across all x-, y-, and z-axes, necessitating the documentation of all values in the same style. Therefore, in this study aimed at implementing 3D text encoding from the outset, it is appropriate to establish unique attributes representing rotation in 3D space. As an interim measure, this paper introduces attributes not currently defined in TEI—@rotX, @rotY, @rotZ—to denote the degree of rotation around each of the x-, y, and z-axes respectively. The sample data encoded in this approach are shown in example 4.

Example 4. <zone> attributes when applying the “center-point” approach.

<zone xml:id="inscription1_front_text" ulx="-0.02" uly="5.51" ana="z:-0.27" rotX="0" rotY="0" rotZ="0">
  <lb n="2"/><gap reason="lost" extent="unknown" unit="character"/><zone>M</zone><zone>I</zone>
</zone>

27The second option is in some sense a more straightforward approach, using @points to draw a polygon. This can overcome the limitation of the first approach, in which only specific shapes based on a single central point (e.g., a square) could be represented, since a variety of shapes can now be described by specifying all their vertices. Moreover, because the plane described by @points already inherently contains information on rotation, there is no need to record rotation as a separate attribute. An example of data representation using this approach is already included in example 4, and zoomed in example 5. The slight modification we must propose in this approach is that we allow giving three values for each point described in the @points attribute, as it is currently only allowed to have two values, corresponding to the x- and y-axes, for each.11

Example 5. <zone> attributes when applying the “points” approach.

<zone xml:id="inscription01_ft_l2" points="0,0,0 0.5,0,0 -0.5,0,0 0,-0.5,0">
  <lb n="2"/><gap reason="lost" extent="unknown" unit="character"/><zone>M</zone><zone>I</zone>
</zone>

28Although there are two possible approaches to describing <zone> elements in a 3D context, the second approach should certainly be preferable, as it is more flexible and requires less effort to modify current TEI schema. Of course, as these approaches are not mutually exclusive, we can consider introducing both ways in the TEI guidelines to provide more possibilities for 3D encoding in TEI.

3.3 Other Considerations

29To this point we have proposed a concrete method for placing TEI-marked text data within a 3D space. Of course, when considering the description of 3D objects, parameters such as texture, color, and geometry could also be envisioned, but the method proposed here is specifically designed for positioning textual information in three-dimensional space, and assumes that the object models themselves will be handled externally. For this reason, the present approach does not address the description of these object-level parameters at this stage. Likewise, if one were to model the entire 3D scene, parameters such as lighting and camera settings would also become relevant, but as our current focus is limited to encoding textual data on a single object, these aspects are not treated either. That said, the representation of a 3D scene could be extended, for example, by using the <surfaceGrp> element to correspond to the spatial context, within which multiple objects represented by <surface> could be included. By defining new subelements and attributes to be contained within <surfaceGrp>, like newly defined <camera> and <light> elements for 3D scene parameter description, it would be possible to describe fully information such as lighting and camera settings.

4. Contributions and Future Considerations

30The foregoing discussion on 3D text encoding using <sourceDoc> reveals the adaptability of concepts traditionally developed for 2D document–oriented markup or genetic editions to textual representations in 3D spaces. Naturally, some adjustments to the specific entities represented within <sourceDoc> are warranted, as illustrated by the epigraphic materials outlined in this paper. Furthermore, there is an inherent need to document supplementary information such as 3D spatial coordinates and rotational angles, which are not essential in 2D settings. Nevertheless, the overarching structure can still adhere to traditional methodologies.

31As mentioned in section 1, before this study was initiated in 2022 (Ogawa et al. 2022), there were few established methods for encapsulating 3D spatial data in TEI, making this research a pioneer of comprehensive exploration in this domain. Consequently, it has been demonstrated that there is no need for a wholly new framework; rather, a minor extension of existing methodologies facilitates the linkage between TEI’s meticulous text-annotation and 3D spatial data.

32Whether it is accomplished through the existing 2D-like “surface” method or a more intricate 3D area specification, the challenge lies in acquiring and inputting such coordinate and rotational information. Tools such as the Image Map in the oXygen XML Editor, which automatically integrates the selected image area’s coordinates into the XML file, will enhance markup efficiency. A comparable feature tailored for 3D markup would be highly anticipated within the field.

33Finally, this study, which uses textual data marked up according to EpiDoc, proposes suggestions for extending this schema. We have observed that the current EpiDoc schema does not accomodate the representation of 3D information, resulting in a limited tag set available within the <zone> element. A notable example is the <expan> tag used to denote abbreviations in text. In EpiDoc, the use of <expan> deviates somewhat from the original TEI, where it is typically employed within <choice>. In EpiDoc, <expan> functions as a top-level tag containing both <abbr> and <ex>. Unfortunately, within the EpiDoc schema, <expan> is not permitted within <zone>. Consequently, if we adopt the approach proposed by this study to describe text information within <sourceDoc>, we are unable to document abbreviations according to the schema, rendering the markup examples depicted in example 4 technically incorrect according to the EpiDoc schema. Such challenges do not arise in TEI itself, as <choice> can be described within <zone>. These examples underscore the potential need for future schema expansions in both TEI and its subsets to accommodate 3D spatial representation, particularly if the goal is to describe textual information within 3D spaces within <sourceDoc>.

Top of page

Bibliography

Cenati, Chiara, Victoria González Berdús, and Peter Kruschwitz. 2022. “When Poetry Comes to its Senses: Inscribed Roman Verse and the Human Sensorium.” In Dynamic Epigraphy: New Approaches to Inscriptions, edited by Eleri H. Cousins. Chapter 7. Oxford. https://​books.​casematepublishing.​com/​Dynamic_Epigraphy.​pdf.

Demetrescu, Emanuel. 2018. “Virtual Reconstruction as a Scientific Tool: The Extended Matrix and Source-Based Modelling Approach.” In Digital Research and Education in Architectural Heritage (UHDL 2017/DECH 2017), edited by Sander Münster, Kristina Friedrichs, Florian Niebling, and Agnieszka Seidel-Grzesińska. Communications in Computer and Information Science 817: 102–16. https://​link.​springer.​com/​chapter/​10.​1007/​978-3-319-76992-9_7.

EpiDoc. 2023. EpiDoc Guidelines: Ancient Documents in TEI XML. Version 9.5. Last updated April 26. https://​epidoc.​stoa.​org/​gl/​latest/​index.​html.

Kettler, Hanna Scates. 2021. “Linked Open Data for 3D Models and Environments.” In Sarah E. Bond, Paul Dilley, and Ryan Horne, eds., Linked Open Data for the Ancient Mediterranean: Structures, Practices, Prospects (ISAW Papers 20). https://​dlib.​nyu.​edu/​awdl/​isaw/​isaw-papers/​20-5/.

Münster, Sander. 2022. “Digital 3D Technologies for Humanities and Research and Education: An Overview.” Appl. Sci. 12, 2426. https://​www.​mdpi.​com/​2076-3417/​12/​5/​2426.

Ogawa, Jun, Kiyonori Nagasaki, Ikki Ohmukai, Yusuke Nakamura and Asanobu Kitamoto. 2022. “Text as Object: Encoding the data for 3D annotation in TEI.” TEI 2022 Conference Book. 86–88. https://​zenodo.​org/​records/​7120027.

Ogawa, Jun, Kiyonori Nagasaki, and Asanobu Kitamoto. 2023. “3D Text Encoding and TEI: Text, Editions, and Spatiality.” TEI 2023 Book of Abstracts, 19–23. https://​zenodo.​org/​records/​10427826.

Schreibman, Susan, and Costas Papadopoulos. 2019. “Textuality in 3D: Three-Dimensional (Re)constructions as Digital Scholarly Editions.” International Journal of Digital Humanities 1: 221–33. https://​link.​springer.​com/​article/​10.​1007/​s42803-019-00024-6.

Shioi, Sachiko and Kiyonori Nagasaki. 2023. “A Preliminary Proposal for Digital Scholarly Editing that Uses Modern Japanese Autograph Manuscripts: How to Markup Autograph Manuscripts of Rampo Edogawa.” Poster presentation. Joint MEC and TEI Conference 2023. https://​teimec2023.​uni-paderborn.​de/​contributions/​167.​html.

SigiDoc. 2023. SigiDoc Guidelines: Byzantine Seals in TEI-XML. Version 1.1. https://​sigidoc.​huma-num.​fr/​Guidelines_XML/​xml-html/​SigiDoc_allguidelines.​html.

TEI Consortium. 2023. TEI P5: Guidelines for Electronic Text Encoding and Interchange. Version 4.7.0. Last updated November 16. https://​www.​tei-c.​org/​Vault/​P5/​4.​7.​0/​doc/​tei-p5-doc/​en/​html/​index.​html.

Trognitz, Martina, Kieron Niven, and Valentijn Gilissen. 2016. “Documentation and Metadata.” In Archaeology Data Service: Guides to Good Practice. https://​archaeologydataservice.​ac.​uk/​help-guidance/​guides-to-good-practice/​data-analysis-and-visualisation/​3d-models/​archiving-3d-data/​documentation-and-metadata/.

Vitale, Valeria. 2017. “Rethinking 3D Digital Visualisation: From Photorealistic Visual Aid to Multivocal Environment to Study and Communicate Cultural Heritage.” PhD thesis, King’s College London. https://​kclpure.​kcl.​ac.​uk/​ws/​portalfiles/​portal/​83196417/​2017_Vitale_Valeria_ethesis.​pdf.

Top of page

Attachment

Top of page

Notes

1 One of the important products from ADS 3D research is the ADS 3D viewer: https://​archaeologydataservice.​ac.​uk/​about/​projects/​ads-3d-viewer/.

2 The term “3D humanities” is well explained on the website of the National Endowment for the Humanities: https://​www.​neh.​gov/​blog/​3d-humanities-digital-visualizations-promote-new-research-and-discourse.

3 The inscription was published in Bernard Rémy et al. (2004), Inscriptions Latines de Narbonnaise (I.L.N.), vol. 2. Vienna, CNRS Editions, p. 268.

4 In an ancient Roman context, projects like Rome in 3D or Rome Reborn have already made some attempt in this direction. Cf. https://​www.​flyoverzone.​com/​rome-reborn-flight-over-rome/.

5 Such attempts to change the appearance of objects depending on the way light hits them and simulations based on such changes have been conducted for more than twenty years in connection with RTI (Reflectance Transformation Imaging). https://​mci.​si.​edu/​reflectance-transformation-imaging.

6 https://​github.​com/​EpiDoc/​EFES.

7 For example, the Epigraphic Database Heidelberg (https://​edh.​ub.​uni-heidelberg.​de/) and the Inscriptions of the Northern Black Sea (https://​iospe.​kcl.​ac.​uk/​index.​html) provide TEI/XML data encoded according to EpiDoc.

8 https://​www.​beckettarchive.​org/.

9 http://​peterstokes.​org/​elena/​proust_prototype/​about.​html.

10 The example edition of Mary Shelley’s text (Bodleian MS Abinger c.56) at the archive is accessible via http://​shelleygodwinarchive.​org/​sc/​oxford/​ms_abinger/​c56/​#/​p5/​mode/​rdg.

11 https://​github.​com/​TEIC/​TEI/​issues/​2816.

Top of page

List of illustrations

Title Figure 1. Image of 3D encoded texts in TEI/XML displayed in 3D space.
URL http://journals.openedition.org/jtei/docannexe/image/6727/img-1.jpg
File image/jpeg, 539k
Title Figure 2. Front surface of an inscription containing both relief and text (https://​edh.​ub.​uni-heidelberg.​de/​edh/​foto/​F000013).
URL http://journals.openedition.org/jtei/docannexe/image/6727/img-2.jpg
File image/jpeg, 463k
Top of page

References

Electronic reference

Jun Ogawa, Kiyonori Nagasaki, Ikki Ohmukai and Asanobu Kitamoto, “3D Text Encoding and TEI”Journal of the Text Encoding Initiative [Online], Issue 18 | 2024, Online since 16 April 2026, connection on 19 July 2026. URL: http://journals.openedition.org/jtei/6727; DOI: https://doi.org/10.4000/165rx

Top of page

About the authors

Jun Ogawa

Jun Ogawa obtained his master’s degree in history from the University of Tokyo in 2017. He is currently working as an assistant professor at the University of Tokyo, Graduate School for Humanities and Sociology. His main research interest is in the digital representation of historical sources and knowledge, especially knowledge organization using Linked Open Data. He participates in several cooperative research ventures in Japan and abroad, including as the chair of the Pelagios Network’s People activity. He is also active in promoting the use of digital methods in the field of (Greco-Roman) classics in Japan..

Kiyonori Nagasaki

Kiyonori Nagasaki, PhD, is a Professor of Library and Information Science in the Faculty of Letters at Keio University and a Senior Fellow at the International Institute for Digital Humanities. His main research interest is in the development of digital frameworks for collaboration in Buddhist studies. He is also engaging in investigation into the significance of digital methodology in the humanities and in promotion of DH activities in Japan. He has been participating in a number of digital humanities projects conducted at several institutions in Japan and abroad. He has also engaged with international standards such as ISO/IEC 10646 (Unicode), TEI Guidelines, and IIIF, so that East Asian DH will be viable globally.

By this author

Ikki Ohmukai

Ikki Ohmukai received his PhD degree in informatics from the Graduate University for Advanced Studies in 2005. He was an Associate Professor with National Institute of Informatics. He is currently an Associate Professor with the University of Tokyo. His research interests focus on digital humanities research methodologies and infrastructure, including digital archives and knowledge graphs.

Asanobu Kitamoto

Asanobu Kitamoto earned his PhD in electronic engineering from the University of Tokyo in 1997. He is now the Director of the Center for Open Data in the Humanities (CODH), Joint Support Center for Data Science Research (DS) in the Research Organization of Information and Systems (ROIS). He is also a Professor at the National Institute of Informatics and Sokendai (The Graduate University for Advanced Studies). He has developed various data-driven science approaches in fields such as the humanities, earth sciences, and disaster management.

Top of page

Copyright

The text only may be used under licence For this publication a Creative Commons Attribution 4.0 International license has been granted by the author(s) who retain full copyright. . All other elements (illustrations, imported files) may be subject to specific use terms.

Top of page
Search OpenEdition Search

You will be redirected to OpenEdition Search