Mishnah – Digital Mishnah https://www.digitalmishnah.org Developing a Digital Edition of the Mishnah Wed, 01 Jan 2014 12:58:57 +0000 en-US hourly 1 https://wordpress.org/?v=4.9.6 Housekeeping https://www.digitalmishnah.org/uncategorized/housekeeping/?utm_source=rss&utm_medium=rss&utm_campaign=housekeeping https://www.digitalmishnah.org/uncategorized/housekeeping/#comments Thu, 26 Apr 2012 17:25:09 +0000 http://www.digitalmishnah.org/?p=119 Read More +]]> This Site

I’ve now updated the “Examples of Work” page to include viewable samples. Thanks to Kirsten Keister for setting up the light box format to view the samples. The examples include two samples of work that processes more than one text (collation, synopsis) and a number of examples of manuscripts.

The Project

I’ve been working on two issues. One is pointing. I now have a complete set of pointers from the reference file (ref.xml) to the witness files for locating spans of damaged text and page and fragment beginnings and ends for fragmentary texts. Of course, because nothing is simple, the direction of all of these will have to be reversed, so that the individual witnesses point into the reference text.
In addition, I’ve improved the tokenization process, so that I can process “rich” tokens, retaining data about the word in question (e.g., that it is an abbreviation, or deleted ….; hold a regularized spelling as well as the original) as well as simple tokens, and re-join a collation based on simple tokens with the complex tokens.

Text Geek Heaven

Along the way, I’ve discovered some joining Genizah fragments. The coolest by far on a technical, jigsaw-puzzle level is the four-way join between TS AS 78.69, TS AS 78.162, TS AS 78.235 and TS NS 329.286 (Cambridge). The four fragments adjoin yet another, TS E2.71. This will be featured as a Fragment of the Month of the Taylor-Schechter Genizah Research Unit. Look for it there!
Cool in that that they join material from multiple cities are:

  • TS E1.99 (Camb), MS heb. 8-11 (Oxf),  and TS F6.3, joining fragments from Cambridge and Oxford and:
  • TS AS 85.270 (Camb), TS F6.2, and MS R2339, fol. 1 (JTS), joining fragments from Cambridge and New York

 

]]>
https://www.digitalmishnah.org/uncategorized/housekeeping/feed/ 4
Progress, real but in small steps. https://www.digitalmishnah.org/uncategorized/progress-real-but-in-small-steps/?utm_source=rss&utm_medium=rss&utm_campaign=progress-real-but-in-small-steps https://www.digitalmishnah.org/uncategorized/progress-real-but-in-small-steps/#respond Thu, 15 Mar 2012 12:41:33 +0000 http://www.digitalmishnah.org/?p=73 Read More +]]> [Originally published on March 11, 2012 at http://blog.umd.edu/digitalmishnah]

I had been holding out for my next post for a new Digital Mishnah website, courtesy of MITH, and a new collation demo hosted on it, but, that will be for my next post, deo volente.

Since my last confession, I have:

  • Submitted a paper that details methods and progress to date. It’s for a Festschrift, and I’ve been asked not to state the venue openly, but can share a draft.
  • Thought a lot about (and only partly understand) multivariate statistics.
  • Completed the first round of markup for all the Genizah fragments for my sample chapter. A second round of markup linking the fragments to the reference text needs to be done (next bullet). Formatted versions of these texts will be viewable
  • Started rethinking how to handle the encoding of highly fragmentary texts. In particular, I’ve found four pieces of a single sheet of text in two different locations in the Taylor-Schechter collection (TS AS 78.69 + TS AS 78.162 + TS AS 78.235  + TS NS 329.286; the sheet adjoins another single sheet from a third box, TS E2.71). For the present, we are encoding each fragment as a document, and recording the extent of the lacunae at the edges of the fragment as fitting within the smallest properly oriented rectangle that encloses the fragment. What needs doing is a pointing scheme that will point into the reference text.
  • Identified the next fragments to work on to expand the work to Tractate Neziqin (aka the Bavot), and started to recruit people to work on it.

Next up, completing fragmentary texts; encoding the remaining Mishnah texts in the Babylonian Talmud mss., and learning some Java.

]]>
https://www.digitalmishnah.org/uncategorized/progress-real-but-in-small-steps/feed/ 0
Thinking about the end product https://www.digitalmishnah.org/uncategorized/thinking-about-the-end-product/?utm_source=rss&utm_medium=rss&utm_campaign=thinking-about-the-end-product https://www.digitalmishnah.org/uncategorized/thinking-about-the-end-product/#comments Wed, 25 Jan 2012 20:44:48 +0000 http://blog.umd.edu/digitalmishnah/?p=28 Read More +]]> Since my last post, I have been working on a grant application. This has afforded the opportunity of some stock taking. I’ve also had some very helpful conversations with scholars in the field: Juan Garcés and Matt Munson in Hebrew Biblical Studies, Tim Finney in New Testament and Desmond Schmidt in textual computing and classics.

1. Collation. Based on very simple normalization and tokenization and a few samples, CollateX will remain error prone, unless the algorithm changes significantly. Examples: (1) In a Mishnah section with repeated words, slight differences in spelling resulted in pushing a whole clause off to the second match. (2) In another passage, CollateX failed to diagnose a missing clause in the text and aligned non matching tokens. My estimate is that currently the error rate is above 10% (for one passage it was about 15%). Better normalization will improve this result. This raises the question of whether the normalization (or, which may amount to the same thing, having CollateX ignore certain characters in comparison) can be carried out automatically, and what this would look like, or whether, as Desmond Schmidt assures me, the whole enterprise is wrongheaded.

2. Statistical measures, now done by hand, but ideally automated. I have now invested in a license for SPSS. This, and my old friend Excel have allowed me to run some preliminary analyses. First: run collations on every Mishnah section in my sample chapter using a few representative witnesses. Transfer the output to Excel; manually fix the alignment (remember, high error rate). Then start flagging variations. I have opted for a method that is akin to what Schmidt and Tim Finney have used: effectively to create a master document with all possible readings, and use a binary encoding (1, 0) for each witness for whether the reading appears in a given witness. (Since the text is already tokenized, I used individual tokens, aka words, not characters, for estimating distance.) Use SPSS to generate a distance matrix, multi-dimensional scaling (MDS), and clustering. I have also experimented with sites providing a graphic interface to Bioinformatic software (FastME and Phylip) to produce phylogenetic trees.

The results were interesting enough that I wanted to see the results with more careful identification of variance (I’m doing these by hand, after all) and more witnesses. I used the sections with the fullest representation among witnesses (Chapter 2, Mishnah 1-2), choosing a total of 10 witnesses. The results I got were consistent with the larger text sample and fewer witnesses, but neither represented the accepted wisdom on the relationship between manuscripts. I therefore divided the cases between no-variation, substantive (different word, different gender, change in grammatical form), and orthographic (initial waw, matres lectiones, spacing between preposition and word). As an example, the Greek word emporia generated no fewer than six variant spellings, but all represented a recognizable version of the word: orthographic, not substantive variation.

MDS for Orthographic Differences, 10 Witnesses

MDS for Substantive Differences

MDS for Substantive Differences, 10 Witnesses

Now, there were some interesting results: the manuscripts thought to be of the “Palestinian type” clustered closely on substantive differences, considerably less so (and differently) on orthographic differences.

The lesson: Orthographic and substantive variations do not coincide, probably due to scribal decision-making (and inconsistency). Substantive differences  seem to be better for groupings of text families. (This may be easier to identify automatically as well: normalizing orthography to improve collation erases orthographic difference (by definition), while retaining non-orthographic difference.) But lingusitic and orthographic differences are of research significance too. We may need a way for the user to flag readings to be compared.

Rooted Tree (Phylip) for Substantive Differences, 10 Witnesses

Unrooted Phylogenetic Tree, 10 Witnesses

As for visualization, we are not yet ready for phylogenetic stemmata, certainly not of the rooted type. The underlying assumptions about a steady evolutionary clock, and the absence of the assumption of contamination make the results interesting from a heuristic point of view, but unreliable in fact. We might think of an unrooted tree as a way of imagining the MDS space with links showing connections. The phylogenetic links in my examples are identical in the rooted and unrooted trees, although from the point of view of grouping families the unrooted tree makes more intuitive sense of the data (closer MSS appear closer) but the trees make the various close relations (the so-called “Palestinian tradition”) into the ancestors or early descendants of distinct traditions. This would require more work to establish, but in more generally, a phylogenetic scheme will require a model better suited to the data.

]]>
https://www.digitalmishnah.org/uncategorized/thinking-about-the-end-product/feed/ 5
New Output https://www.digitalmishnah.org/uncategorized/new-output/?utm_source=rss&utm_medium=rss&utm_campaign=new-output https://www.digitalmishnah.org/uncategorized/new-output/#comments Mon, 02 Jan 2012 20:22:32 +0000 http://blog.umd.edu/digitalmishnah/?p=26 Read More +]]> Only spammers seem to be noticing this blog, but for web-trolling software that might be interested in digital humanities and philology I thought I might add that I have updated the sample output from Collatex.

collatex-table-apparatus.html shows output from user-specified witnesses in the form of (1) an alignment table based on user-specified order, (2) an extracted text of a base text (taking the first specified witness is the base text), (3) generating an apparatus.

CollateX is not perfect. Some of the output problems are the result of tokenizing (the samples used were tokenized very coarsly) and can be fixed. Abbreviations and the phenomenon of connected or unconnected prepositions (של, also words such as כיצד) can also be fixed. But some errors have to do with how CollateX deals with with edit distance. Not sure how we are going to handle this.

]]>
https://www.digitalmishnah.org/uncategorized/new-output/feed/ 1
Starting Out https://www.digitalmishnah.org/uncategorized/starting-out/?utm_source=rss&utm_medium=rss&utm_campaign=starting-out https://www.digitalmishnah.org/uncategorized/starting-out/#comments Tue, 18 Oct 2011 15:52:16 +0000 http://blog.umd.edu/digitalmishnah/?p=6 Read More +]]> This blog describes my progress on an born-digital critical edition of the Mishnah. For the various audiences who might read this, let me break out the terms and discuss them further.

  • Born-digital: An edition that uses or develops technology to record, store, present, search, analyze, and study textual material, rather than a static presentation of my research.
  • Critical edition: An edition that attempts to deal seriously with the state of a text, typically based on comparison of manuscripts. There is substantial debate about what text the edition is trying to recover (original? at some moment?), and whether this is even possible, and scholars of rabbinic texts have generally chosen one print or manuscript witness as a base “copy text” rather than a reconstructed text.
  • Mishnah: A legal text produced about 200 AD/CE. Still studied today in Jewish religious circles of all kinds, the text is the basis of the Talmuds (there is both a Palestinian and a Babylonian Talmud), and is of enormous historical significance as well.

The text is preserved in manuscripts the very earliest perhaps from the ninth or tenth centuries but most later, and was first printed in 1492 in Naples. There are several substantial manuscripts including the whole text or large blocks of it, but there are also fragmentary manuscripts of one or a few leaves from the Cairo Genizah, a cache of documents from Egypt, mostly from the tenth to the thirteenth centuries. These are among the earliest manuscripts. The textual transmission is also preserved in citations from the Talmud, and in citations of medieval scholars in their legal or commentary works.

I have recently begun as a faculty fellow at MITH, the Maryland Center for Technology and the Humanities, to develop a pilot edition using a single chapter, Bava Metsi’a Chapter 2, dealing with lost objects. To date, I have (with the assistance of students!) encoded several witnesses using the TEI encoding protocol, and developed some XSLT style sheets to display the results in HTML. I am now working with the good people at MITH to develop an ODD, and work on a pilot webservice to test the project, while continuing to work on transcriptions. Some sample texts are viewable at https://sites.google.com/site/digitalmishnah/files.

]]>
https://www.digitalmishnah.org/uncategorized/starting-out/feed/ 1