Routely, Nicholas. 1985. “A Practical Guide to musica ficta.” Early Music 13: 59–72. doi: 10.1093/earlyj/13.1.59. https://doi.org/10.1093/earlyj/13.1.59.
Lessons from the Classroom: MEI for Data Scientists
Abstract
Many data science and computer science students today are familiar with JSON, and may even have worked with APIs to extract data from the web. Ask about XML,1 however, let alone TEI or MEI, and you are often met with quizzical looks. Yet XML files contain much information that can be productively analyzed with modern data science tools, so training students to leverage these materials is a worthwhile endeavor. The article shows some of the methods we use to help students understand XML as a hierarchical network of elements, how to traverse this network in search of relevant data, and how to harvest XML elements and attributes as tabular data for further analysis. It also reflects on some of the larger lessons learned through all of this work, as students were encouraged to consider the implications of representing the same knowledge in different ways, or what is gained or lost in the transformation of that knowledge from one representation to another.
Outline
Top of pageFull text
The development of the tools shared in this paper has been made possible by the generous support of the American Council of Learned Societies and Haverford College.
Project URL: https://github.com/RichardFreedman/Encoding_Music/tree/main/05_Beautiful_Soup_MEI_Paderborn.
- 2 For more about The Music Encoding Initiative (MEI), see https://music-encoding.org/resources/tu (...)
- 3 Pandas: https://www.w3schools.com/python/pandas/default.asp
- 4 Beautiful Soup: https://tedboy.github.io/bs4_doc/
- 5 XPath: https://www.w3schools.com/xml/xpath_intro.asp
- 6 Plotly Express: https://plotly.com/python/plotly-express/; NetworkX: https://networkx.org/do (...)
1Many data science and computer science students today are familiar with JSON, and may even have worked with APIs to extract data from the web. Ask about XML, however, let alone TEI or MEI, and you are often met with quizzical looks. Yet XML (eXtensible Markup Language) files contain much information that can be productively analyzed with modern data science tools, so training students to leverage these materials is a worthwhile endeavor. The authors’ work on Citations: The Renaissance Imitation Mass (CRIM; https://crimproject.org/) has demonstrated the value of both using MEI as the basis for representing a digital score (with the help of Verovio) and performing sophisticated queries over large corpora.2 Encodings can thus be created with these dual uses in mind—editorial and analytic alike. This report summarizes the authors’ experience teaching XML concepts in the data science classroom with—for largely practical reasons—Python tools already familiar to students, including Pandas,3 Beautiful Soup4 and XPath,5 and related libraries for plots and graphs, such as Plotly and networkX.6 It shows some of the methods used to help students understand XML as a hierarchical network of elements, how to traverse this network in search of relevant data, and how to harvest XML elements and attributes as tabular data for further analysis. It also reflects on some of the larger lessons learned through all of this work, as students were encouraged to consider the implications of representing the same knowledge in different ways, or what is gained or lost in the transformation of that knowledge from one representation to another.
- 7 Readers are welcome to contact the authors of this paper for additional resources. But all the code (...)
2The ideas discussed here were used in the context of an interdisciplinary course at an American liberal arts college, with a mix of undergraduate students, many of whom were new to code or musical notation, or even both. Indeed, the hybrid nature of this course (Encoding Music, taught by Freedman with Russo-Batterham’s help) was a virtue, signaling to students the need to collaborate across disciplinary ways of making knowledge.7 But elements of the approach could readily be adapted to more advanced or specialized settings with equal success, instilling a sense of discovery and exploration of a field not typically approached from this vantage point.
1. A First Example: MEI Up Close and at a Distance
3In introducing the subject of musical encodings to students, it was found to be helpful to begin with a clear, if minimal, example—in this case a single note from a Latin motet by one of the great composers of the sixteenth century, Giovanni Pierluigi da Palestrina (about 1525–1594), who was active for decades in important Roman churches. His music represents the pinnacle of the contrapuntal vocal music of the age, and continues to be heard throughout the world today. By examining variant encodings side by side, students can see that the same musical information can be represented in very different ways: first as a single voice part from a sixteenth-century printed source, then in a modern edition (with all four voice parts aligned in what musicians call a “score”), and finally in a small fragment of an XML (MEI) encoding (see [bad link to item: ]). The remarkable efficiency of European graphical and symbolic musical notation is one story to be told here, compressing into a very small space much of what the musician needs to know in order to sing a part—pitches, durations, and (often in very precise ways) how to align words and music, which was a key part of expression during this period.
Figure 1. Three views of one measure in Palestrina’s “Veni sponsa Christi,” showing the original source, a modern edition, and the corresponding portion of the MEI encoding. The penultimate note is marked in red. For the complete modern edition and MEI file, see https://crimproject.org/pieces/CRIM_Model_0019.
4But as we focus on the single note highlighted in figure 1, we also notice something else about this graphical notation: its lack of explicitness. Renaissance singers were used to figuring out many things for themselves: not only questions of pacing and dynamics (which were never indicated in these sources), but even some accidentals. In certain situations (and often at the ends of phrases like the one shown here), musicians had to supply unwritten sharps and flats in order to make the music sound as it should.
5Exactly why Renaissance composers (and editors) refrained from writing down something that in fact was required in performance is a story for another time. But the entire tradition had a name that hints at the reason: musica ficta, meaning a kind of imaginary music that was sung (and heard) but never written (Routely 1985). And so when modern editors approach this kind of situation they take care to indicate the accidentals above the staff (as seen in the middle portion of figure 1), making explicit both the intended musical result and the fact that the modern editor was the one responsible for it. The MEI encoding is the most explicit of all of these, since it records the data in logical terms that make it perfectly clear that the <accid> tag has been <supplied> and the reason for it, too. On the other hand, the MEI encoding is remarkably verbose—what we see in figure 1 represents just one measure of one voice part in this piece. It takes hundreds of lines of code to represent what can be seen in a few pages of a score.
6Students might reasonably wonder what such logical encodings are good for, or how to navigate something so sprawling. The first part of the answer must be that these are not for human eyes, at least not immediately. But it is worth explaining to students that like other XML, it has other virtues, most notably its hierarchical structure, which makes it possible to understand relations between elements and to query files intelligently in search of larger patterns.
2. XML Network as Graph Visualization
- 8 For an explanation of how these are created, see https://github.com/RichardFreedman/Encoding_Mu (...)
7One novel approach to explaining the hierarchical structure of XML was through colorful and interactive network visualizations.8 Teaching students the details of network theory and practice is well beyond the scope of the course. Instead, they are shown how to understand a hierarchical structure that is difficult to see in any other way, and is certainly opaque to any human reader of the XML file itself. What is more, these graphs helped students grasp the sheer scale and complexity of an encoded text, as well as the diversity of the elements used and where they appear, with some elements frequently “leaf” nodes and others structuring much broader sections.
Figure 2. A network graph representation of the MEI shown in figure 1. For the full MEI file, see https://crimproject.org/mei/CRIM_Model_0019.mei.
8In figure 2, for instance, we return to the same measure shown previously (the XML context is shown at the right). But now the information is displayed as a directed graph: the measure is the parent node (the large circle at the center), which then yields to four staves (the four voice parts of the complete score), and in turn the individual notes in each. The penultimate note of the superius part (the original focus of our figure 1 representations) has a <supplied> child and finally an accidental at the very tip of the chain (see the box highlighting the relevant portion of the graph and encoding). From the original encoding we could even provide further detail, displaying for each node the relevant attributes of pitch, duration, and other musical details (see figure 3).
9Thus prepared to understand how a local string of musical events can be represented as a hierarchical graph, students can begin to understand how staggeringly verbose but also impressively structured these encodings are when regarded from the top level down.
10And once confronted with a vision of the whole, they can also begin to understand in an intuitive way the relations between different high-level elements, such as the <body> of the MEI (with all the notes, and thus the central part of the dense network shown above), and the <head>, where all the metadata are contained (top right in figure 4).
Figure 4. The <head> of the MEI file from the previous examples. The red highlight shows the <respStmt>, with composer and three editors.
3. From Exploration to Investigation
11These network visualizations have been found to be a powerful way of helping students to conceptualize XML encodings of musical sources. But like any data representation, there are some underlying limitations and tradeoffs. Network visualizations help students understand the hierarchy of measures, staves, and notes, but they do little to highlight the order in which elements at the same level of the hierarchy (like notes within a particular staff in a particular bar) appear. The visualization also abstracts away the XML syntax, which students will eventually be required to interact with more directly. While they are important considerations, these caveats were not found to be problematic in a pedagogical context. Rather than shield students from having to engage with raw XML (or significantly delay them from doing so), these visual aids instead hastened the learning process by teaching core concepts without the immediate encumbrance of XML syntax.
- 9 Citations: The Renaissance Imitation Mass (CRIM), directed by Richard Freedman. URL: http://crimpr (...)
12What is more, now prepared to see the big picture, students were able to begin thinking about the kinds of questions they might ask of such encodings. And since they had already focused on the supplied accidentals, they quickly seized on the idea of investigating editorial practice across the corpus of files. Freedman’s current research on musical borrowing and similarity in the Renaissance (The CRIM Project)9 provided the perfect test bed for such work, featuring more than two hundred MEI encodings of works by dozens of different composers, which had been assembled by a team of editors. Would editorial habits reveal themselves across different composers? How could we reason over the notes (in the body of the MEI files) and the names of editors (in the head) in a systematic way in pursuit of the answers? Put more succinctly, how could we see the forest that emerges from the details of the XML trees?
4. Climbing the XML Tree with Beautiful Soup
13Beautiful Soup is a Python library that makes it relatively simple to climb XML trees. In their most basic form, Beautiful Soup queries allow students to find parents, children, and siblings of various tags. We can return these elements individually, or as lists, or as lists of elements that meet certain conditions. Perhaps more importantly, working with these files in a Python environment lets students transform their results into tabular data, and in turn apply other analytic methods to them. The data science ecosystem is rich with tools designed to work effectively with tabular data. Along the way, students are taught to reflect on the kinds of information they might need to find, how best to go about finding it, how to revise their approach when things do not go quite as expected, and what next steps they might try based on what they have learned. Consider, for instance, two approaches that were tried with students: a statistical approach to measuring compositional style, and an attempt to track editorial habits.
14One way to get students thinking was to ask them to translate a verbal query into a Pythonic one (and through it something that could be used to navigate the XML). How, for instance, would you go about finding the pitch of the last note of the first staff in our piece? This is a simple enough process with a score in hand, but what would this look like expressed in a series of encoded steps for a machine? Thinking at first in pseudocode, the student might:
-
find all the measures (as a list)
-
find the last measure in that list
-
find the first staff of that last measure
-
find all the notes in that staff (as a list)
-
find the last note in that list
-
find the pitch of that note
15These steps can be done individually in Beautiful Soup, or strung together as shown in figure 5, in a series of “find_all” steps (producing a list matching a particular element) and “find” steps, which return only the first matching element of a given type. Fluency with the code is crucial here, but no less important is the intuitive understanding of the XML tree depicted in the network graphs above. Although this stepwise approach is more verbose than using compact modes of querying XML trees such as XPath expressions, it helps students focus on one type of element at a time, learning the key features of each in turn. Beautiful Soup, moreover, is already familiar to many data science students, since it is frequently used to scrape data from the web (note that the notebook that accompanies this article also provides some XPath alternatives, for those familiar with that approach).
16What would a picture of all the notes look like? We have already seen this as a network graph in figure 4 above. But now prepared with an understanding of how to find a single pitch, students reasoned that it would also be possible to find all the pitches, tabulate the distribution of values, and then use various familiar charting libraries to see how they stack up across pieces. The thinking here was a bit more involved than in the previous steps, but the basic principles were the same: pseudocode to plan the query, implementation in Beautiful Soup, and finally plotting with various Python libraries (in this case we favor Plotly for its elegance and adaptability). In this case the pseudocode thinking was:
-
With Beautiful Soup, find all the notes, and find the pitches associated with them, as a list.
-
With Python, count the occurrences of each pitch class (in this case using the “counter” class, which makes rapid work of such categoricals).
-
Also with Python, keep track of the total number of notes in the piece, since it will be helpful to “scale” our results to allow comparison across pieces.
-
Finally, transform the results as tabular data with Pandas, and in turn use Plotly to show bar charts, radar plots, or other renderings.
- 10 See the complete piece on CRIM at: https://crimproject.org/pieces/CRIM_Model_0019.
17As we see in figure 6, the results give analysts a clear and useful way of understanding the pitch space of any piece. In this case we see the relative strength of G, D, and C in Palestrina’s motet, which fits nicely with our musical intuitions for this piece, at least based on the shape of the opening melody and the tones in the final sonority.10
Figure 6. Distribution of pitch classes in “Veni sponsa Christi,” showing code, tabular data, bar chart, and radar plot.
5. Looking at Editorial Practice
18Now prepared to find and count notes, how did students go about investigating editorial practice and “explicitness” in our corpus? First, they reviewed two aspects of our MEI encodings in order to understand where to find the relevant data. The editorial accidentals themselves are contained in <supplied> elements within the <body> of the encoding, as we noted above in figure 1. Meanwhile information about who was responsible for the edition is to be found in the <head> of the MEI file—specifically in the <titleStmt>. And so here we look for all <persName> whose @role attribute is "editor" (see figure 7).
Example 1. Code for finding the editors with Beautiful Soup.
editors = soup.titleStmt.find_all('persName', {'role' : 'editor'})
for editor in editors:
print(editor.prettify())
# this is the XML output from the previous code
<persName role="editor">
Marco Gurrieri
</persName>
<persName role="editor">
Vincent Besson
</persName>
<persName role="editor">
Richard Freedman
</persName>
# In `XPath` we can use this square bracket syntax to find the people with the "editor" role:
print(xpath_query(root, "//mei:persName[@role='editor']"))
# this is the XML output from the previous code
<persName xmlns="http://www.music-encoding.org/ns/mei" role="editor">
Marco Gurrieri
</persName>
<persName xmlns="http://www.music-encoding.org/ns/mei" role="editor">
Vincent Besson
</persName>
<persName xmlns="http://www.music-encoding.org/ns/mei" role="editor">
Richard Freedman
</persName>
19Reporting the overall distribution of such interventions is a bit more complicated than for the pitch-class distributions above. But the key is knowing where to find the relevant data in the MEI file and breaking things down into discrete steps:
-
Find all the <supplied> elements.
-
Find all of the <accid> elements within those (since we only want these sorts of accidentals and not all of them), and return them as a list.
-
Find the measure number in which each <accid> element appears (since it might be interesting to plot these interventions over the course of a piece).
-
Find the editors, and select the first (main) editor.
-
Find the title and composer of the work.
-
Iterate through the list of accidentals, creating a temporary Python “dictionary” (in a single “row” of a table) that would contain all the information noted above.
-
Append each of these temporary dictionaries to a list and make a dataframe from those, which in turn could be used to create various kinds of reports and charts.
20In figure 8 we see this process implemented in Beautiful Soup. Here students come to understand how to translate the logical steps into Pythonic actions. The result is a synthetic view of the data found in one MEI file.
Figure 8. A tabular summary of all the supplied accidentals in a single MEI file, using Beautiful Soup, Python, and Pandas.
Example 2. Code for generating a tabular summary of all the supplied accidentals in a single MEI file, using Beautiful Soup, Python, and Pandas.
# with Beautiful Soup:
list_supplied_accid_data = []
for supplied in soup.find_all("supplied"):
accidentals = supplied.find_all('accid')
for accidental in accidentals:
temp_dict = {
'title' : soup.find('title').text.strip(),
'composer' : soup.find('persName', {'role' : 'composer'}).text.strip(),
'editor' : soup.find('persName', {'role' : 'editor'}).text.strip(),
'parent_measure_number' : accidental.find_parent('measure').get('n'),
'pitch' : accidental.find_parent('note').get('pname'),
'accid_value' : accidental.get('accid')
}
# append each dictionary to a list of all the results
list_supplied_accid_data.append(temp_dict)
supplied_accids = pd.DataFrame(list_supplied_accid_data)
supplied_accids["parent_measure_number"] = supplied_accids["parent_measure_number"].apply(int)
supplied_accids
# tabular results
title composer editor parent_measure_number pitch accid_value
0 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 6 f s
1 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 6 f s
2 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 11 f s
3 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 17 f s
4 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 19 f s
5 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 41 g s
6 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 43 g s
7 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 45 c s
8 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 66 f s
# Fetching all this data looks a little different with `XPath`:
list_supplied_accid_data = []
title = xpath_query(root, '(//mei:title)[1]/text()')
composer = xpath_query(root, "(//mei:persName[@role='composer'])[1]/text()")
editor = xpath_query(root, '(//mei:persName[@role="editor"])[1]/text()')
# We don't use our xpath_query function here since it returns strings that can't themselves be queried
accidentals = root.xpath('//mei:supplied//mei:accid', namespaces=namespaces)
for accidental in accidentals:
temp_dict = {
'title': title,
'composer': composer,
'editor': editor,
'parent_measure_number': xpath_query(accidental, 'ancestor::mei:measure/@n'),
'pitch': xpath_query(accidental, 'ancestor::mei:note[1]/@pname'),
'accid_value': accidental.get('accid')
}
list_supplied_accid_data.append(temp_dict)
supplied_accids = pd.DataFrame(list_supplied_accid_data)
supplied_accids['parent_measure_number'] = supplied_accids['parent_measure_number'].astype(int)
supplied_accids
# tabular results
title composer editor parent_measure_number pitch accid_value
0 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 6 f s
1 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 6 f s
2 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 11 f s
3 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 17 f s
4 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 19 f s
5 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 41 g s
6 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 43 g s
7 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 45 c s
8 Veni sponsa Christi Giovanni Pierluigi da Palestrina Marco Gurrieri 66 f s
21The next thought problem was to imagine how we could look across an entire corpus of encoded files in order to understand how these interventions compare across composers, genres, and editors. The idea here was to think about what we might call “comparative explicitness” throughout the many layers of the CRIM Project. Here the exact count of supplied accidentals was less important than the relative proportion. And so the pseudocode steps looked like this:
-
Find title and editor for each piece.
-
Determine a “musica ficta factor” for each piece, which would be the count of <supplied> accidentals divided by the count of all notes; this would be a number between 0 and 1.
-
Apportion those musica ficta factors into “bins” (e.g., "high", "medium", and "low".)
-
Look for correlations between composers, editors, and musica ficta.
-
Chart the results in a way that might allow participants to see the big picture easily.
22Implementing these steps in Python is a bit more complex than what we have seen in previous examples. But the principles are the same: using Beautiful Soup to climb the XML Tree, then assemble the results as a list of temporary dictionaries, and finally aggregate those dictionaries as tabular data for visualization (see example 3).
Example 3. Finding all supplied accidentals and responsible editors across several MEI files with Beautiful Soup and Pandas.
list_piece_data = []
list_ficta_factors = []
# get each piece as BS object
for piece in corpus_list:
print(piece)
xml_document = getXML(piece)
soup = BeautifulSoup(xml_document, 'xml')
# find all the accidentals that are editorial
# divide the number of such editorial accidentals by the total number of notes in the piece ('ficta factor')
ficta_factor = len(soup.find_all('accid', {'func' : 'edit'})) / len(soup.find_all('note'))
# round those ratios to no more than three decimal points
list_ficta_factors.append(round(ficta_factor, 3))
# build a temporary dictionary of the data, with file_name, editor, composer, and the ficta factors
temp_dict = {
'file_name' : os.path.basename(piece),
'composer' : soup.find('persName', {'role' : 'composer'}).text.strip(),
'editor' : soup.find('persName', {'role' : 'editor'}).text.strip(),
'ficta_factor' : ficta_factor}
# append each dictionary to a list of all the results
list_piece_data.append(temp_dict)
# make a dataframe out of that list of results
df = pd.DataFrame(list_piece_data)
# create 'bins' of any number, based on the ranges of the ficta factors
df['ficta_bins'] = pd.cut(df['ficta_factor'], bins=3, labels=['low', 'medium', 'high'])
df
23And finally we can translate this tabular information into some fairly compelling graphical representations of the data. In figure 9, for instance, each dot represents the proportion of supplied accidentals in a list of dozens of pieces from the CRIM Project corpus (pieces with dots that are large and red contain a comparatively high proportion of such elements; those that are small and blue do not). Meanwhile we plot supplied accidentals for the editor of each piece (the vertical axis) against the composer of each piece (the horizontal axis). Some striking outliers emerge. Are they the result of editorial or compositional preference? On the one hand Marco Gurrieri (who edited a large number of pieces in our project) supplied a comparatively large number of accidentals across several very different composers, so perhaps he has something of a heavy editorial pen. On the other hand, María Elena Cuenca Rodríguez’s work with a piece by the Spanish composer Cristóbal de Morales (he was a contemporary of Palestrina’s) provokes still other questions. Is the relatively high proportion of musica ficta a reflection of María Elena’s habits (we note that another piece, by Nicholas Gombert, also has a high proportion of musica ficta)? Or is it something peculiar to Morales? Musicologists have often noted that Spanish composers in particular tend to require many accidentals in their polyphony.
Figure 9. Plotly visualization of the data collected with the code in example 3, showing the relative proportion of supplied accidentals per note in each piece and by each editor.
6. Student Reflections and Next Steps
24The classroom program focused on equipping students with the practical skills needed to understand, navigate, and analyze MEI and other forms of XML. Just as important, however, are the critical insights that students developed with respect to their own work and the work of others. This project on supplied accidentals and other features of MEI-encoded scores prompted students to reflect on the biases and values implicit in all editorial work, and the technologies for representing the perishable art of music.
25While these insights were as diverse as the students’ backgrounds, several key themes emerged. Students recognized the value of being explicit about their own assumptions and approach to formulating and evaluating what they found, carefully weighing biases in their motivations, intuitions, and methods. They learned that no encoding is wholly neutral with respect to what it describes: score, parts, MEI, and tabular data all both illuminate and obscure certain kinds of patterns. Part of data science is knowing how best to represent the data for a given type of analysis and which tools can be applied to transform it from one form to another. Through manipulating these encoded representations, students both with and without musical training developed a more nuanced understanding of how counterpoint actually works; not everything is “melody and chords.”
26Perhaps most crucially, students learned to operate in a multidisciplinary space, drawing on the strengths of peers and applying modes of analysis that come from other disciplines. Equally, they learned to teach others from different backgrounds in a setting that promoted peer-to-peer learning. These lessons from the data science classroom could be adapted to a range of academic contexts. While computational musicology continues to grow as a field, more could be done to leverage musical encodings for analysis. Working with music in this way teaches us to move fluidly from quantitative to qualitative, and from high-level to intimate readings of the musical text, interleaving insights from divergent methods and techniques. No doubt our students will apply their creativity to the tasks at hand in novel ways, inaugurating modes of inquiry and making new kinds of knowledge wherever they take themselves in the academic, cultural, or business world. We look forward to learning with them in the years to come.
Notes
1 XML: https://www.w3schools.com/xml/
2 For more about The Music Encoding Initiative (MEI), see https://music-encoding.org/resources/tutorials.html; on the rendering system Verovio, see https://www.verovio.org/index.xhtml.
3 Pandas: https://www.w3schools.com/python/pandas/default.asp
4 Beautiful Soup: https://tedboy.github.io/bs4_doc/
5 XPath: https://www.w3schools.com/xml/xpath_intro.asp
6 Plotly Express: https://plotly.com/python/plotly-express/; NetworkX: https://networkx.org/documentation/latest/tutorial.html; along with Pyvis: https://pyvis.readthedocs.io/en/latest/tutorial.html
7 Readers are welcome to contact the authors of this paper for additional resources. But all the code mentioned here can be found at https://github.com/RichardFreedman/Encoding_Music/blob/main/05_Beautiful_Soup_MEI_Paderborn/Beautiful_Soup_MEI.md.
8 For an explanation of how these are created, see https://github.com/RichardFreedman/Encoding_Music/blob/main/05_Beautiful_Soup_MEI_Paderborn/Beautiful_Soup_MEI.md#visualizing-xml-as-network-graphs.
9 Citations: The Renaissance Imitation Mass (CRIM), directed by Richard Freedman. URL: http://crimproject.org/
10 See the complete piece on CRIM at: https://crimproject.org/pieces/CRIM_Model_0019.
Top of pageList of illustrations
![]() |
|
|---|---|
| Title | Figure 1. Three views of one measure in Palestrina’s “Veni sponsa Christi,” showing the original source, a modern edition, and the corresponding portion of the MEI encoding. The penultimate note is marked in red. For the complete modern edition and MEI file, see https://crimproject.org/pieces/CRIM_Model_0019. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-1.png |
| File | image/png, 578k |
![]() |
|
| Title | Figure 2. A network graph representation of the MEI shown in figure 1. For the full MEI file, see https://crimproject.org/mei/CRIM_Model_0019.mei. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-2.png |
| File | image/png, 112k |
![]() |
|
| Title | Figure 3. The same network shown in figure 2, now including attributes for each element (node). |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-3.png |
| File | image/png, 182k |
![]() |
|
| Title | Figure 4. The <head> of the MEI file from the previous examples. The red highlight shows the <respStmt>, with composer and three editors. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-4.png |
| File | image/png, 120k |
![]() |
|
| Title | Figure 5. Beautiful Soup code to find the last note of the first staff in the given file. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-5.png |
| File | image/png, 149k |
![]() |
|
| Title | Figure 6. Distribution of pitch classes in “Veni sponsa Christi,” showing code, tabular data, bar chart, and radar plot. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-6.png |
| File | image/png, 140k |
![]() |
|
| Title | Figure 7. Finding the editors with Beautiful Soup. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-7.png |
| File | image/png, 102k |
![]() |
|
| Title | Figure 8. A tabular summary of all the supplied accidentals in a single MEI file, using Beautiful Soup, Python, and Pandas. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-8.png |
| File | image/png, 217k |
![]() |
|
| Title | Figure 9. Plotly visualization of the data collected with the code in example 3, showing the relative proportion of supplied accidentals per note in each piece and by each editor. |
| URL | http://journals.openedition.org/jtei/docannexe/image/5545/img-9.jpg |
| File | image/jpeg, 152k |
References
Electronic reference
Richard Freedman and Daniel Russo-Batterham, “Lessons from the Classroom: MEI for Data Scientists”, Journal of the Text Encoding Initiative [Online], Issue 18 | 2024, Online since 21 February 2025, connection on 19 July 2026. URL: http://journals.openedition.org/jtei/5545; DOI: https://doi.org/10.4000/13e5b
Top of pageCopyright
The text only may be used under licence For this publication a Creative Commons Attribution 4.0 International license has been granted by the author(s) who retain full copyright. . All other elements (illustrations, imported files) may be subject to specific use terms.
Top of page













