Foreword
The publication of this paper is meant as a tribute to the late C. Michael Sperberg-McQueen by the coauthors. In their mind, the paper represents the best possible "graceful" conclusion they managed to produce of the work on transcription they carried out with Michael over a span of more than two decades (a large portion of which the third author was absent from the team).
While we by no means consider ourselves in a position to do justice to the numerous enticing ideas and practical developments contributed by Michael on the topic, we have felt it important to at least try our best to make known some of the insights, hopes, and the enthusiasm he had regarding the possibilities of formalization for furthering the understanding of theoretical and practical aspects of transcription.
Whatever good ideas may be found in this account stem in one way or another from Michael. Only we, however, must bear the responsibility for the many shortcomings that may remain. Their number would have been much smaller had we been able, as we repeatedly found ourselves wishing, to rely on the wit and benevolence of our friend to keep us from going astray.
Our first formal publication on this topic came out in 2008 [Huitfeldt / Sperberg-McQueen 2008], followed by a Balisage paper in 2010 [Huitfeldt / Marcoux / Sperberg-McQueen 2010]. The notion of
transcriptional implicatureis partly based on (and what we saw as a low-hanging fruit from) ideas from these earlier publications. What is presented here was more or less finished some time around 2015, and last touched in 2018.[1] However, it was put on the back burner and has never been formally published until now.[2]At the time, we regarded transcriptional implicature as only one of the parts of a more ambitious and comprehensive "formal account", or a theory of "the logic", of transcription. Regrettably, we did not manage to accomplish this quite demanding task before Michael's untimely death.
Materials from that later period, going beyond what we present here, will hopefully be made available as part of the [Sperberg-McQueen Nachlass project].
We would like to thank participants of a number of conferences through the years, such as DH2009, Balisage2010, and DH2014, for their interest, questions, and suggestions.
In particular, we want to thank Michael's wife, Marian Sperberg-McQueen, for encouraging the posthumous publication of this paper.
— Claus Huitfeldt & Yves Marcoux
Introduction
The notion of transcriptional implicature
There is hardly any universal transcription practice: for every generalization we find exceptions. Is everything in the exemplar[3] transcribed? Not when deletions and irrelevant material are excluded. Does everything in the transcript reproduce some word or character in the exemplar? Not when line breaks are marked explicitly with vertical bars, or notes are added.[4] Many scholarly editions account for variations like these in an explicit statement of transcription practice. Such statements typically describe deviations from the usual practice, but rarely the ways in which it exemplifies usual practice. Common practice may sometimes be felt to be so obvious that it needs no mention or explanation.
By transcriptional implicature for a
given community, we mean informally the things, suggested
or entailed by the rules of transcription, that members of that
community may find it unnecessary to mention explicitly
.
Different communities of transcription practice have different sets of tacit assumptions and thus different rules of transcriptional implicature. Is there a common core of transcriptional practice shared by all communities? There may be; it is an empirical question. A definitive answer would require more detailed studies of a wider variety of communities of practice than we can currently manage. Our hypothesis, however, is that there is such a common core, at least in the sense that the transcriptional implicature of any community of practice can be described with reference to some common set of rules.
If (as we conjecture) there is such a thing as the
transcriptional implicature of a given community, then the
transcriptional practice of any project of that community could be
described by listing the ways in which it
deviates
from the transcriptional implicature.
Such deviations
can be addition, modification, or
withdrawal of rules. Thus, the rules of the transcriptional
implicature are defeasible for particular projects. They apply,
except if explicitly excluded.
If (as we conjecture) the transcriptional implicature of a community can in turn be described as a set of deviations from some default transcriptional implicature, then it follows that any project's transcription practice could be described with reference to the default transcriptional implicature, by merging the list of ways in which the community’s implicature deviates from the default implicature with the list of ways in which the project’s practice deviates from the default community practice.
The (conjectured) default transcriptional implicature is a set of rules which apply by default, but which may be overridden in particular cases, analogous to the rules of conversational implicature proposed by H. P. Grice as a way of explicating the logic of everyday conversation [Grice 1975].[5]
Earlier work
Since most digital transcriptions are represented today using markup (e.g., TEI), providing a formal model of transcription might prove useful for work on the semantics of markup. Indeed, prevalent approaches to markup semantics are based on formal languages like first-order predicate logic [Sperberg-McQueen / Huitfeldt / Renear 2001], which would a priori blend well with a formal model of transcription, yielding a framework in which meaning can be assigned to markup constructs whose colloquial definitions involve transcription, for example, p elements defined as transcription of a text block in a manuscript.
Some have proposed to explicate the
meaning of markup by specifying, for each
construct in a markup vocabulary, a sentence schema in a natural
language, with blanks to be filled in with data from the document
[Marcoux 2006]; others make a similar proposal but
allow sentence schemata in formal languages like first-order
predicate logic as well [Sperberg-McQueen / Huitfeldt / Renear 2001]. This appears
straightforward, although far from trivial, for metadata [Wickett / Renear 2009] and perhaps even for born-digital texts, but
how shall the meaning of a p element be formalized in
a markup language which defines it as containing a
transcription of a text block in a
manuscript? What does it mean for
a document to be a transcription of another document? Laying a
formal ground on which such questions can be tackled, may be seen
as contributions to the semantics of markup, as applied in
transcription.
Early work [Huitfeldt / Sperberg-McQueen 2008] has explored the nature of the similarity between transcripts and their exemplars. If documents are defined as sequences of characters, then perhaps the similarity consists simply in the exemplar and the transcript containing the same sequence of characters? This can be formalized but proves disappointing, partly because the definition of documents as sequences of characters omits text structures like division into paragraphs and partly because the model offers no way of describing disagreements among transcribers about how to read the exemplar, or about which character distinctions (e.g. i/j, u/v, s/ſ) to retain and which to level. It is also wrong for the reason that few transcripts have exactly the same character sequence as their exemplar.
Later work has extended the model by introducing boolean (disjunctive and conjunctive) types, which allowed explicit representation of ambiguity, as well as readings, which allowed modeling transcriber agreement and disagreement explicitly [Sperberg-McQueen / Huitfeldt / Marcoux 2009], and lifting the analyses from characters to higher-level textual structures [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
In this paper, we propose to assess the notion of transcriptional implicature using a formal approach, in the line of our earlier work on transcription [Huitfeldt / Sperberg-McQueen 2008] and [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
Outline of the formal model
Formalization seems interesting to us for two reasons: First, and most importantly, formalization encourages a level of explicitness and rigor which may help us better understand what transcription is, and thereby also inform our discussions and judgments in concrete cases. Second, now that transcription is mostly done digitally, a formal model may help us build better applications and better understand their possibilities and limitations.
Our goal is simply to describe formally the relationship that exists between two documents when the one is said to be a transcription of the other. It is not our goal to establish normative criteria which may serve to distinguish good from bad or right from wrong in transcription, nor to construct a theory of the psychological or physical processes involved in transcription.
The formal model we use is that of [Huitfeldt / Marcoux / Sperberg-McQueen 2010], but we present just enough technical details to be able to flesh out our discussions of transcriptional implicature. For further details, see [Huitfeldt / Sperberg-McQueen 2008] and [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
What is formalized
The term transcription may be used variously to refer to the act or process of transcribing a document, to the physical product of that act (that is, another document), or to the relation between the two documents.[6]
Our formal approach focuses on the relation that must obtain between two documents for the one to be a transcription of the other. Where such a relation obtains, we call one of the documents a transcript, and the other an exemplar.
We take as basic facts about this relation that:
-
any exemplar antedates its transcript,
-
the transcript has been made by someone for certain reasons and with certain intentions,
-
it has been made on the basis of access to the exemplar (or facsimile of the exemplar), and
-
the text of the transcript is the same as, or at least similar in some way to, the text of the exemplar.[7]
Of all those aspects of the relation between exemplar and transcript, our formalization addresses only the last one, that is, the similarity relation that is assumed to obtain between the texts of a transcript and of its exemplar.
The kinds, levels or degrees of text similarity required for a document to be counted as a transcription of another obviously vary from context to context, and it is not hard to find examples of disagreement about the matter. The formal model must be flexible enough to allow for such variations.
The basic notions
The main idea in our approach to formalization is that documents are textual objects[8] and, as such, can be analyzed in terms of a theory of types and tokens.[9] In turn, the similarity relationship that must obtain between two documents for one to be a transcript of the other (we will call it T-similarity) is defined in terms of type-token analyses of the documents.
Here are the basic notions on which the current account is based.
-
document: A document (in particular, an exemplar or a transcript) is a physical phenomenon containing or exhibiting marks. (For example, the copy of Moby Dick on the third author's nightstand, the Wittgenstein notebook cataloged as MS 108 in the Wren library.)
-
mark: A mark is a perceptually discernible arrangement of physical reality in a document; some (and perhaps, but not necessarily, all) marks are tokens (for example, words written in pen or pencil in a manuscript, or the printed words, fly specks and pencil marks of an old book).[10]
-
token: A token is a mark determined to instantiate a particular type while reading the document. (For example, a letter, punctuation mark, a word or a sentence on a piece of paper or parchment.) Tokens can be basic (atomic) or compound. A compound token is one that is read as the composition of smaller tokens, called its subtokens. A whole document is viewed as a single "highest-level" compound token, i.e., one that is not part of a larger compound token.
-
type: A type is a character, letter, word, sentence, chapter, or any other object instantiated by one or more tokens in one or more documents. (For example, the character 'A' of the Latin alphabet, or the sentence 'I know Verona'.) Like tokens, types can be basic or compound. A type entering in the composition of a compound type is said to be a subtype of the latter.
For example, the written word "cat" might be read as a compound token instantiating the compound word type 'cat', and composed of the three basic subtokens "c", "a", and "t", instantiating respectively the basic letter types 'c', 'a', and 't', which are subtypes of the word type 'cat'.[11]
Documents (including exemplars and transcripts), which are highest-level compound tokens, each instantiate a (highest-level) compound type.
Types can also be conjunctive or disjunctive, corresponding to ambiguous tokens. For example, in a manuscript where it is not possible to determine whether some token is of type 'i' or 'j', yet clear that it is not of any other type, the token can be said to be of the disjunctive type 'i/j'. In some documents, as for example so-called ambigrams, certain tokens are meant to be read in more than one way. Those tokens can then be regarded as conjunctive types.[12]
-
type system: A type system is a set of type repertoires, which in turn are sets of types. For example, a type system may contain:
All tokens instantiate at least one type in at least one repertoire. The same token cannot instantiate more than one type in any given repertoire. Strictly speaking, a token may represent different types in different repertoires. For example, the token "I" in English may instantiate both a character, a word, and a sentence, but not more than one of each. However, for the sake of convenience, we will usually speak of "the type" (singular) of any given token.-
a repertoire of characters used in 17th-century English manuscripts,
-
a repertoire of words used in the English of that period,
-
a repertoire of sentences formed using those words, and
-
repertoires of paragraphs, sections, and other textual structures.
-
The model allows for constraints to be defined on which types can be instantiated by a compound token, in terms of the subtypes instantiated by its subtokens. For example, it could be that any compound token must be assigned a compound type such that the subtokens of the compound tokens instantiate the subtypes of the compound type (as was the case in the 'cat' example above).
Identification of tokens and types
The process of reading of a document, in particular identifying the tokens in the document and the type that each of them instantiates, is outside the scope of our model. Whenever we refer to a document, we will allow ourselves to refer also to its tokens and to the type of each of these tokens, without further explanation.
Like most formal models, our work assumes that names can be assigned to individuals in the universe of discourse, in particular tokens and types. The practical issues involved in reaching agreement on the identity of individuals and the names to be used for them are outside the scope of the formalization. When individuals are as numerous as the tokens in a novel, of course, agreement on names for them all is likely to present practical difficulties.
The absence from our account of any description of a method for assigning names to the tokens of a document reflects not the belief that such an operation is easy, only the fact that it has no bearing on the logical structure of transcription, or of transcriptional implicature. Like the process of reading a document in the first place, the process of identifying the tokens in a document sufficiently well to enable productive disagreement about how to read them is of great practical importance but outside the focus of our work.
T-similarity
The model defines a relation between documents called T-similarity.[13] Two documents are T-similar if and only if, as high-level tokens, they are either of the same (high-level) type, or of similar types, where similarity among types is defined in a ad hoc manner, in the context of some given transcription project or set of projects.
T-similarity is the relation that captures transcription. Its formal definition reflects the rules of transcription.[14]
In the simplest case, T-similarity is based on type equality; in other words, exemplar and transcript must be of identical types, not just similar types. However, the definition can be relaxed, and this is how we can hope to capture reasonably realistic aspects of the similarity between exemplar and transcript in actual projects. Allowing for such loosening of the similarity conditions is an essential aspect of our model. We will later further relax the similarity conditions with the introduction of special tokens.
The intent behind T-similarity (with suitable versions of similarity among types) is that it corresponds to the kind of similarity expected to obtain in practice between an exemplar and a corresponding transcript, as discussed above. We do not expect, nor claim, that T-similarity can capture all possible views of exemplar-transcript similarity, but we suggest it represents a minimal necessary condition. Also, as already mentioned, similarity of text is but one aspect of the relationship between exemplar and transcript in real life: normally a transcript must follow temporally after its exemplar, be made in the presence of the exemplar, and be the result of an intentional effort.
As a mathematical relation, T-similarity is reflexive, transitive, and symmetric. The irreflexive, intransitive, and asymmetric nature of the factual relation observed between transcripts and their exemplars appears to us to be a consequence of contingent facts and not a property of the similarity relation holding between the documents.
We have informally defined transcriptional implicature for a
given community as the things, suggested or entailed by the
rules of transcription, that members of that community may find
it unnecessary to mention explicitly
. From the point of
view of our formal model, which focuses on the relation between
exemplar and transcript, this amounts to the things
about the T-similarity relation that
go without saying
.
Document pairs
We propose to characterize T-similarity as a set of formal
statements expressed in first order logic (FOL),
with variables ranging over the types and tokens of the related
documents. "Things" about the T-similarity relation (including
those that purportedly go without saying
) can then
be formally expressed as such statements.
Such a representation of T-similarity can be put to work in various ways. For instance, given a pair of documents, we can seek to formally establish whether they stand in a T-similarity relationship or not. This will be done for a simple pair of documents in [appendices].
Or, given one known document (say, a transcript) and assuming
that it is T-similar to some other document (an exemplar), we can
ask ourselves what can be inferred about that other document on
the basis of the known document. Exploratory work in this
direction have been described in Claus Huitfeldt and C. M.
Sperberg-McQueen, transcriptional Implicature: Using a
Transcript to Reason about an Exemplar
in [Digital Humanities 2017 Conference Abstracts], p. 266,
([http://blackmesatech.com/2017/08/MLCD/).
To do those things, we need to be able to represent pairs of documents in FOL. The goal of this section is to introduce FOL elements – predicates and constants – that can be used to "say things" about pairs of documents, one exemplar and one transcript, in terms of their types and tokens.
In the first-order formulae below, we use the following notation: The variables X, Y, Z, etc. range over all individuals in the universe of discourse (i.e., the types and tokens of the related documents). Each individual carries one (and only one) of the predicates:
-
type(X): X is a type
-
e_token(X): X is a token in the exemplar
-
t_token(X): X is a token in the transcript
We define the following binary relations (predicates):
-
typeof(X,Y) ⇔ the type of X is Y.
-
transcript(X,Y) ⇔ token Y in the transcript corresponds to token X in the exemplar.
-
exemplar(X,Y) ⇔ token Y in the exemplar corresponds to token X in the transcript.
As already mentioned, we are silent on how tokens are identified and assigned types. We are also silent on how the correspondences between exemplar and transcript tokens represented by the transcript() and exemplar() predicates are established.[15]
Finally, we introduce the following two individual constants:
-
e: the exemplar (as a single, usually compound, token)
-
t: the transcript (as a single, usually compound, token)
Note that we have:
-
e_token(e), and
-
t_token(t).
It should be clear that facts or hypotheses (to be proved or disproved) about pairs of documents can be formulated using the above apparatus. In particular, an exemplar-transcript pair can be entirely represented, in terms of its types and tokens, as a series of statements using the above predicates together with additional constants standing for the involved types and tokens. An example is given in [Appendices].
As previously mentioned, we do not prescribe nor describe any method for naming the constants corresponding to types and tokens, we simply take for granted the use of some adequate naming scheme.
Axiomatization
The above definitions can be expressed as the following set of axioms that document pairs must satisfy.
Axioms
-
transcript domain: For every X and Y, if X is the transcript of Y, then X is a t_token and Y is an e_token.
∀(X,Y)(transcript(X,Y) → (e_token(X) ∧ t_token(Y))) -
exemplar domain: For every X and Y, if X is the exemplar of Y, then X is an e_token and Y is a t_token.
∀(X,Y)(exemplar(X,Y) → (t_token(X) ∧ e_token(Y))) -
typeof domain: For every X and Y, if the type of X is Y then X is a t_token or an e_token, and Y is a type.
∀(X,Y)(typeof(X,Y) → ((t_token(X) ∨ e_token(X)) ∧ type(Y))) -
no class overlap: Every individual X is either an e_token or a t_token or a type.
∀(X)( (e_token(X) → ( ¬t_token(X) ∧ ¬type(X))) ∧(t_token(X) → ( ¬e_token(X) ∧ ¬type(X))) ∧(type(X) → ( ¬e_token(X) ∧ ¬t_token(X)))) -
Exemplar and transcript inverse: For every X and Y, if X is the exemplar of Y, then Y is the transcript of X.
∀(X,Y)(exemplar(X,Y) ⇔ transcript(Y,X)) -
At most one type per token: For every X and Y, if the type of X is Y then no other individual is the type of X.
∀(X,Y,Z)((typeof(X,Y) ∧ typeof(X,Z)) → Y=Z)[16]
Initial formulation of the default transcriptional implicature
The initial formulation of the conjectured
default transcriptional implicature which we
present here could be described informally as a transcript
contains the same text as its exemplar
. Formally, this
means that the exemplar and the transcript are tokens instantiating
the same type: the transcript re-instantiates the type of the
exemplar. Thus, only documents of equal types
are considered T-similar; this corresponds to the simplest case of
T-similarity presented earlier.
T-similarity could then be formalized in a single rule: ∃(X) (typeof(e,X) ∧ typeof(t,X)).
In interesting cases, however, the type X (of which e and t are both tokens) will be a compound type consisting of some structure of smaller compound types (which in turn consist of smaller ones still), instantiated by a compound token which similarly consists of smaller tokens.
It is thus more appropriate to define T-similarity as the conjunction of four conditions, as follows:
Default transcriptional implicature definition of T-similarity
-
Reciprocity: The transcript() and exemplar() predicates map tokens in a one-to-one fashion.
reciprocity ⇔ ∀(X,Y,Z)( (((transcript(X,Y) ∧ transcript(X,Z)) → Y=Z) ∧((exemplar(X,Y) ∧ exemplar(X,Z)) → Y=Z))) -
Completeness: There is nothing in the exemplar that is left out from the transcript.
In other words, for each and every token X in the exemplar there is at least one corresponding token Y in the transcript.
completeness ⇔ ∀(X)(e_token(X) → ∃(Y)(exemplar(Y,X))) -
Purity: There is nothing in the transcript except what comes from the exemplar.
In other words, for each and every token X in the transcript there is at least one corresponding token Y in the exemplar.
purity ⇔ ∀(X)((t_token(X) → ∃(Y)(transcript(Y,X)))) -
Type identity: The exemplar and the transcript are type-identical through and through (not just the tokens e and t).
In other words, for every pair of corresponding tokens X in the exemplar and Y in the transcript, X and Y are of the same type.
type_identity ⇔ ∀(X,Y)((transcript(X,Y) → ∃(Z)(typeof(X,Z) ∧ typeof(Y,Z))))
We can now define T-similarity as follows:
-
T-similarity: The relation between exemplar and transcript is one of reciprocity, completeness, purity, and type identity.
t_similarity ⇔ (reciprocity ∧ completeness ∧ purity ∧ type_identity)
With the axioms given earlier and a FOL representation of any
pair of documents, we can verify whether T-similarity,
i.e., the conjunction of the four rules of reciprocity,
completeness, purity, and type identity, comes out as a theorem or
not. This is described in more detail below.
Discussion
It may be worth stressing that the rules just given do not constitute a claim that every transcript is reciprocal, pure, complete, or thoroughly type similar to its exemplar. They amount to a claim that these things are true unless otherwise stated or unless the standards of a given community of practice dictate otherwise. That is, they express general assumptions about transcripts which are defeasible in particular cases.
Consider different transcripts of the same exemplar. They may vary for several reasons. As a first application of our model, let us use it to describe some of those reasons.
Transcripts may disagree about which of the marks in the exemplar instantiate types and are thus tokens (one transcript may read a mark as a decorative pen stroke, the other as a letter in a word). They may agree on the set of tokens found in the exemplar but disagree on which types they instantiate. Or they may use different type systems. One, for example, may distinguish the allographs i/j, u/v, and ſ/s, while the other treats the allographs as instantiations of the same type (as they are instances of the same grapheme).
The use of different type systems can lead to the same kinds of difference between transcripts as different understandings of the exemplar; failure to understand the nature of the disagreement (different reading of the exemplar? or different choice of type system?) can lead to confusion and acrimony. Analysis of conflicting transcripts through the common lens of a formal model might reduce such confusion and acrimony.
We observe that a number of common variations in transcription practice can be classified according to which rule of the initial formulation of the default transcriptional implicature they override.
Some transcripts omit deleted material, extraneous material, illegible material, or material in specific writing systems (mathematics, Greek, …). In other words, transcripts are not always entirely complete.
Some transcripts mark lines with bars, sometimes also adding line numbers. In some cases omissions may be marked by symbols or standard phrases ("[Illegible]", etc.) In other words, transcripts are not always entirely pure.
Some transcripts preserve allographic variations, while others level those distinctions. For example, some transcripts preserve the distinctions between long and short s, or vocalic and consonantal i and u, others do not. This might be seen as a limitation of the rule of thorough type similarity. We believe it is more natural, however, to assume in such cases that the transcripts employ different type systems, one in which the allographs are considered tokens of the same type, and another one in which they are not. In other words, different transcripts do not necessarily employ identical type systems.
Some transcripts silently expand abbreviations, normalize spelling and correct slips of the pen. At character level, the rules of completeness and purity seem to be broken. Even so, such transcripts may observe thorough type similarity at higher level tokens, such as words. In other words, transcripts are not always entirely type similar through and through.
Expansion of abbreviations in brackets or italics can preserve the rules of default transcriptional implicature on word and higher levels, but introduces characters in transcripts which lack corresponding characters in the exemplar. Again, transcripts are not always entirely pure and do not always preserve entirely thorough type similarity.
In cases of doubt, as in the case of parts of manuscripts which are hard to read because of wear, damage, or difficulties in handwriting, transcribers tend to interpret words as correctly spelled and sentences as grammatically well-formed, at least unless there is evidence to the contrary. This is often referred to as the principle of charity.[17] This tendency may be accounted for in at least two ways: Either we can regard the token in e as instantiating a disjunctive type and the token in t as instantiating one of the unobjectionable disjuncts, or else we can regard the transcriber as choosing to read the token in e as an instantiation of an unobjectionable type. In the former case, type similarity must be defined to allow pairs of the form "x/y" in e and either 'x' or 'y' in t (so the types are similar but not identical). In the latter case, type identity is preserved but the reading of e simply ignores the fact that the passage in e might plausibly be interpreted differently.
Reformulation of the default transcriptional implicature to permit exceptions
As we have just seen, in almost all actual transcription projects there are at least a few exceptions to each of the four rules a-d: material in the exemplar not transcribed (perhaps because irrelevant to the purpose of the transcript), material in the transcript not present in the original (if only page numbers and footnotes), and transcription conventions which transcribe selected tokens with tokens of similar, not identical type.
The examples which follow show a variety of such exceptions; the formalizations given show one way in which the defeasibility of the simple rules just given can be represented in logical formulae.
As will become clear below, concrete transcription practices can be summarized as a set of exceptions to the rules of purity, completeness, and type similarity. Since the rules given in the preceding section do not countenance exceptions, it will prove necessary, for reasoning about any actual transcript and its exemplar, to reformulate those rules to capture the qualification that they each apply except in special cases.
We therefore extend the document-pair model with the following predicates and relations:
-
se_token(X): X is a special exemplar token -
st_token(X): X is a special transcript token -
typesimilar(X,Y): types X and Y are considered similar
In order to account for the inclusion of special exemplar and
special transcript tokens we modify the typeof domain
and the no class overlap axioms as
follows:
-
typeof domain:
∀(X,Y)(typeof(X,Y) →((t_token(X) ∨ e_token(X) ∨ st_token(X) ∨ se_token(X)) ∧ type(Y))) -
no class overlap:
∀(X) ((e_token(X) ⇔ (¬t_token(X) ∧ ¬type(X) ∧ ¬se_token(X) ∧ ¬st_token(X))) ∧(t_token(X) ⇔ (¬e_token(X) ∧ ¬type(X) ∧ ¬se_token(X) ∧ ¬st_token(X))) ∧(type(X); ⇔ (¬e_token(X) ∧ ~t_t_token(X)) ∧ ¬se_token(X) ∧ ¬st_token(X)) ∧(se_token(X) ⇔ (¬t_token(X) ∧ ¬type(X) ∧ ¬e_token(X) ∧ ¬st_token(X))) ∧(st_token(X) ⇔ (¬t_token(X) ∧ ¬type(X) ∧ ¬se_token(X) ∧ ¬e_token(X)))
We add the following axioms:
-
typesimilar domain:
∀(X,Y)(type similar(X,Y) → (type(X) ∧ type(Y))) -
typesimilar reflexive:
∀(X)(type(X) → typesimilar(X,X)) -
typesimilar symmetric:
∀(X,Y)(typesimilar(X,Y) → typesimilar(Y,X)) -
typesimilar transitive:
∀(X,Y,Z)((typesimilar(X,Y) ∧ typesimilar(Y,Z)) → typesimilar(X,Z))
In the definition of T-similarity, we replace the type-identity rule with the following:
-
Type similarity: The exemplar and the transcript are type-similar through and through.
In other words, for every pair of corresponding tokens X in the exemplar and Y in the transcript, X and Y are of similar types.
type_similarity ⇔ ∀(X,Y)((transcript(X,Y) →( ∃(Z,V)(typeof(X,Z) ∧ typeof(Y,V) ∧ typesimilar(Z,V)))))
Then, naturally, we change the definition of
T-similarity to require type_similarity
instead of type_identity:
-
T-similarity:
t_similarity ⇔ (reciprocity ∧ completeness ∧ purity ∧ type_similarity)
Note that in the absence of special tokens and of unequal types that are typesimilar, the reformulation changes nothing to T-similarity. Thus, the reformulation does not per se change the default transcriptional implicature.
However, the reformulation allows a concise description of any transcription practice in terms of its deviations from the rules of the default transcriptional implicature. For any given project's transcription practice, we can define the extension of the predicates se_token, st_token, and typesimilar, and combine them with the rules just given to provide a basis for inference.
The skeptical reader may object to the explanatory value of this approach: In general, anything can be described as a deviation from any set of rules. More specifically, the skeptical reader will have observed that taken strictly, this reformulation amounts to saying that the properties of completeness, purity, and type-identity will apply in all cases, except when they do not; the formalization just given has no way to express the expectation that completeness, purity, and type-identity are the normal, expected, or usual case, and the cases covered by the predicates se_token, st_token, and typesimilar (when it does not coincide with type identity) are special, unusual, and less frequent cases.
A logic designed to formalize defeasible reasoning would perhaps capture that distinction better; testing the utility of such formalisms for reasoning about transcription remains a desideratum for the future. In the meantime, however, we hope that this formulation in terms of standard first-order logic will suffice for the purposes of our argument.
Readers may have been struck by some similarity (pun intended) between our initial formulation of the default transcriptional implicature and certain methods or criteria for identifying or measuring document similarity, such as, for example, the so-called Levensthein edit distance between strings of characters.[18] The edit distance between two such strings is measured in terms of the minimal number of deletions, insertions or substitutions of individual characters in the one string that is required in order to make it identical to the other.
Analogously, we might think of special character tokens as deletions, special transcript tokens as insertions, and type similar tokens as substitutions. The difference, however, is that while Levensthein similarity is well suited for operations on character strings, it is less well suited for work on documents with a non-trivial structure. The additions and modifications we made in the reformulation of the default transcriptional implicature to permit exceptions may be seen as an attempt to take care of that difference.
Can our formulation of the default transcriptional implicature be put to empirical test? In principle, yes. At least we can see whether it makes good sense when applied to the actual transcriptional practices of a wide range of real projects. In practice, however, we lack the resources required to establish any really broad empirical basis. In this paper, we have limited ourselves to looking at two examples of descriptions of transcription practice. First we discuss an example drawn from a U.S.-based historical documentary edition (the papers of William Penn), then the transcription in a literary edition of some manuscript notes by Hermann Melville.
Example 1: The papers of William Penn
Statement of practice
In the chapter Editorial Method
in [The Papers of William Penn], the editors provide what we regard as a
statement of practice
, stating that:[19]
...In this edition, we aim to print a completely faithful transcript of each original text, including blemishes and errors. ... In general, our editorial interpolations [enclosed in square brackets] within the text are minimal...
We observe that the first sentence seems at least implicitly to confirm our rules of completeness and cleanliness: Every token of e is transcribed by exactly one token of the same type in t, and every token in t is exemplified by exactly one token of the same type in e. The second sentence introduces an exception: transcripts also contain additional material in the form of editorial interpolations (that is, transcripts are not entirely clean). However, such added material is explicitly marked by square brackets. The fact that the editors find it worth pointing this out may be taken to suggest that it does not go without saying in the relevant community that blemishes and errors are retained, or that additions are always marked explicitly. The editors continue:
Our editorial rules may be summarized as follows:
1. Each document selected for publication in The Papers of William Penn is printed in full. ...
At first sight, this looks like a straightforward confirmation of the rule of completeness. So why does it not go without saying? Perhaps because it is not altogether common practice of this community to print every document in full. Or perhaps because the edition is a selection (that is, not complete in the sense of containing all the papers of William Penn) the editors found it important to make clear that although their transcript is not complete with respect to the entire body of material, it is complete with respect to each individual document.
2. Each document is numbered, for convenient cross-reference, and is supplied with a short title.
From this we can infer that transcription breaks the implicit rule of purity by adding document numbers and titles, which do not occur in the exemplar. It seems unlikely that the editors find it necessary to make this explicit statement on the assumption that otherwise readers would believe that William Penn himself numbered his letters and provided them with titles. It is more likely that the community in question does expect to find some indication of where one document ends and the next begins as well as some means of referring to individual documents, but has not agreed on a uniform way of doing so.
3. The format of each document (including the salutation and complimentary closure in letters) is rendered as in the original or copy...[20]
It is not entirely clear from this, or from the rest of the
declaration of practice, what the editors mean by the
format
of a document. Comparison of a transcript
to a facsimile of its exemplar,[21] however, suggests that what they mean is page layout.
At least the transcripts seem to preserve such features as blocks
or lines of text flushed to the right or left or indented, blank
lines, and so on. (We also note that, elsewhere in the edition,
poems are printed as lines of verse.[22])
It is possible that the editors' practice is based on an explicit or implicit distinction between phenomena such as paragraphs, salutations, signatures, date lines, poems, verse lines, etc. In either case, they seem to be referring to what we would call compound types, and to try to preserve T-similarity between transcript and exemplar also in this respect.
It is therefore perhaps striking to observe that the transcripts contain no indication of line or page breaks[23] in the exemplar, and that the editors do not mention this fact. Should we consider this omission a violation of the rule of completeness? (That is, are line and page breaks not considered tokens of some type?) If so, is the reason why the editors do not mention this omission that it is part of the transcriptional implicature of the relevant community of practice? Or do they simply rely on the appearance of the transcript on the printed page with regular, running lines etc. to make it too obvious to deserve mention that line and page breaks of the exemplar must be different?
The statement numbered 4 in [The Papers of William Penn] p. 17 deals with the rendering of datings according to Julian, Gregorian, and Quaker calendars, a complicated issue the details of which we do not go into here.
5. The text of each document is rendered as follows:
a. Spelling is retained as written. Misspelled words are not marked with an editorial [sic]. If the sense of a word is obscured through misspelling, its meaning is clarified in a footnote.
b. Capitalization is retained as written. In seventeenth-century manuscripts, the capitalization of such letters as "c," "k," "p," "s," and "w" is often a matter of judgment, and we cannot claim that our readings are definitive. Whenever it is clear to us that the initial letter in a sentence has not been capitalized, it is left lower case.
c. Punctuation and paragraphing are retained as written. When a sentence is not closed with a period, we have inserted an extra space.
d. Words or phrases inserted into the text are placed {within braces}.
e. Words or phrases deleted from the text are crossed through.
f. Slips of the pen are retained as written, and are not marked by [sic].
g. Contractions, abbreviations, superscript letters, and ampersands are retained as written. When a contraction is marked by a tilde, it is expanded.
5.a suggests that this transcript, by retaining original spelling, does not deviate from general transcriptional implicature, but also that it is normal practice in the relevant community to do so, i.e., by silently normalizing spelling or marking misspelling with [sic].[24] We take 5.f to mean that the editors make a distinction between slips of the pen (as in 5.a) from misspellings, but that they treat them the same way.
The insertion of footnotes to clarify the meaning of misspelled words mentioned in 5.a, however, does represent an exception to the default rule of purity, as does the expansion of contractions marked by tilde mentioned in 5.g. Other deviations from the rule of purity are identified in 5.c for the insertion of an extra space when a sentence is not closed with a period, and 5.d for the use of braces to surround editorial insertions into the text.[25]
5.e simply seems to confirm the default rules of transcriptional implicature in stating that deleted words or phrases in the exemplar are crossed through in the transcript. That the editors are explicit about this may suggest that normal practice in the community is different. (Perhaps inserting special markers for deleted text (breaking the rule of purity) or leaving deleted text out (breaking the rule of completeness)).[26]
Statement 5.b seems to be a paradigmatic application of what we referred to above as the principle of charity. We take it to be saying that deciding whether a given character in the manuscript is uppercase or lowercase requires judgment on the part of the transcriber.[27]
Thus, to a large extent the statements above explicitly confirm rules which form part of the default transcriptional implicature; that they are stated explicitly may suggest that in the relevant community of practice (or among the expected readers of the edition), it might be common to deviate from the default rules in these cases.
h. The thorn is rendered as "th," and superscript contractions attached to the thorn are brought down to the line and expanded: as "the," "them," or "that." Our justification for this procedure is that we no longer have a thorn, and modern readers mistake it for "y." Likewise, since modern readers do not recognize that "u" and "v" were used interchangeably in the seventeenth century, we have rendered "u" as "v," or "v" as "u," whenever appropriate.
i. The £ sign in superscript is rendered as "l."[28]
j. The tailed "p" is expanded into "per," "pro," or "pre," as indicated by the rest of the word.
k. The long "s" is presented as a short "s." The double "ff" is presented as a capital "F."
The statement contained in the first clause of the first sentence of 5.h, i.e., rendering the letter "Þ" as two characters, "t" and "h", may seem to deviate from all the four rules of our default transcriptional implicature: 1) it breaks the one-to-one correspondence between the tokens of the exemplar and the transcript (no reciprocity), 2) there are tokens (thorns) in the exemplar which do not appear in the transcript (incompleteness), 3) there are tokens ("th"-sequences) in the transcript which do not occur in the exemplar (impurity), and 4) none of the characters in question are of the same type (no type similarity). We observe that the statement may suggest that the normal practice of this community would be to represent "thorn" with tokens of a distinct type, which would not deviate from default transcriptional practice at all.
However, if we interpret the statement to the effect that the project employs a type system in which "Þ" and "th" are tokens of the same type, the practice is entirely in accordance with the rules of our default transcriptional implicature. On the one hand, therefore, such an interpretation seems attractive. On the other hand, it may seem objectionable: In all other contexts than where they occur in sequence and stand for a thorn, the tokens "t" and "h" will still be taken as tokens of two different types.
To this objection we may answer that it is not at all unusual
for the type of tokens, or even the decision about where to draw
boundaries between tokens, to be context dependent.[29] Also, we do have here a principled justification for
the context-dependent treatment of the thorn: A convention already
exists for representing Þ
as
th
in English spelling; "Þ" and "th" are
traditionally pronounced identically by modern readers; and, as
the statement itself argues, modern readers tend to mistake the
thorn for a "y". (The last argument seems to appeal to a principle
of charity with the reader.)
If we are not convinced by these arguments and thus reluctant to accept that "Þ" and "th" may be regarded as tokens of the same type, one alternative might be to interpret the first clause of 5.h to the effect that the thorn is regarded as a contraction which is expanded to "th". In that case, the practice constitutes a deviation from default rules on a par with other contractions — see the discussion of tilde in 5.g above. In any case, the same goes for the statement made in the second clause of the first sentence concerning expansion of superscript contractions attached to the thorn.
Now to the second sentence of 5.h, which observes that "u" and "v" were used interchangeably in the seventeenth century, and therefore the one is rendered as the other (or vice versa) "whenever appropriate" — presumably this means according to the expectations of modern readers (and thus this practice may be seen as another instance of "charity to the reader"). This may be a case for assuming that the project employs a type system according to which "u" and "v" are tokens of the same type, and in that case no deviation from default transcriptional implicature is implied.
What may seem awkward about this analysis is that even if "u"
and "v" are in free distribution in the seventeenth century texts
(alternate realizations of the same type, just as allographs are
alternate realizations of the same grapheme), they are certainly
distinct types (and indeed distinct graphemes) in modern English.
An alternative analysis, therefore, may be that the type system of
the exemplar subsumes that of the transcript: every type in the
transcript corresponds to exactly one type in the exemplar, but
more than one type in the transcript ( here, u
and
v
) may correspond to the same type in the exemplar
(here, the type instantiated by the tokens we read as
u
and v
in the exemplar). Either
way, no deviance from the default transcriptional implicature is
implied.
Statements 5.i and 5.k also describe the type system used: superscript £ is treated as a token of type "l",[30] and long and short "s" are regarded as tokens of the same type. This will produce less ambiguity in the transcript than was the case for "Þ" and "th" or for "v" and "u", since long "s" will not occur at all in the transcript. Regarding "ff" as a token of capital "F", however, could in principle lead to ambiguity.
Statement 5.j may seem to break the rule of purity for the expansion of the abbreviations mentioned. However, if tailed "p"s are regarded as tokens (depending on context) of the types "per," "pro," or "pre," the rule is preserved.
The statements 5.h-k may or may not be understood as representing deviations from the rules of default transcriptional implicature, deviations which are either not generally shared, or dealt with in other ways, by the community of practice. In either case, giving a formal account of all details involved here will lead to a certain amount of complication. As a first simplifying approximation, we may suggest that the statements imply that the project at hand employs a type system in which the following tokens are regarded as tokens of the same type:
-
"Þ" and "th"
-
"Þ" with superscripts and "the", "them", or "that" (depending on context and/or the nature of the superscript, — the statement is silent on this point)
-
"u" and "v"
-
"U" and "V"
-
Superscript "£" and "l"
-
tailed "p" and "per," "pro," or "pre," as indicated by the rest of the word.
-
long and short s
-
"ff" and "F"
l. Words underlined in manuscript are italicized.
m. Blanks in the manuscript, missing words, and illegible words are rendered as [blank] or [missing word] or [illegible word] or [illegible deletion]. If a missing word can be supplied, it is inserted [within square brackets]. If the supplied word is conjectural, it is followed by a question mark.
Statement 5.l, that words underlined in the exemplar are rendered in italics in the transcript, may be seen in different ways within our framework of rules of transcriptional implicature. (The following remarks are relevant also to statement 5.e above, where deleted words or phrases in the exemplar are rendered as crossed through in the transcript.) We might simply assume that the editors regard underlined and italicized tokens as typesimilar (like allographs of the same grapheme). If so, however, why do they preserve the distinction in the transcript? After all, the distinction between other allographs of the same grapheme, like long and short s, are not preserved.
It may seem more natural to assume that the editors understand underlining in the exemplar as signaling a higher-level feature, let us call it emphasis, which pertains not to the atomic letter tokens individually, but to the entire underlined word, and that they have chosen to signal this feature by other means, i.e., italics, in the transcript. On this assumption, an underlined word in the transcript is treated as a compound token of a type different from the same word without underlining. Also on this account, however, we might regard this as a way of preserving type similarity, and thus in accordance with the rules of transcriptional implicature.
Statement 5.m, that information about blanks, missing words, illegible words, and conjectural readings are represented by words or phrases within square brackets, signals a break with the rules of both completeness and purity. Again, we may interpret the editors' statement as a confirmation that these rules are otherwise followed. The reason why they see a need to make this explicit may be either that these rules are normally followed in such cases, or that the phenomena in question are normally signaled in other ways.
Interpretation in terms of T-similarity
Given the reformulation of the default transcriptional implicature offered above, we can interpret the statement of practice in terms of the three predicates se_token (special exemplar token), se_token (special transcript token), and typesimilar (similar types).
As far as we can tell, the only tokens in e that are not present in t are the tildes marking contraction. However, these contractions are expanded in t, so one might argue that this is a case of type similarity rather than special exemplar tokens.
The discussion above indicates that the following tokens in t are special transcript tokens, i.e., that they correspond to nothing in e:
-
footnotes (and their markers)
-
extra space at end of sentence lacking final punctuation
-
braces marking inserted material
-
expansions of contractions (a) marked by tilde, (b) using thorn + superscripts as the, them, that, (c) using tailed "p" as per, pro, pre.
-
constituent "t" and "h" of "th" = thorn.
-
constituent "l" and "." of "l." = superscript £.
We have concluded that the type system of t may be assumed to be the same as that of e, but also noted that the transcript seems to assume several cases of type similarity. Types t1 and t2 are similar iff:
-
t1 = t2
-
or t1 = thorn, t2 = 〈 "t", "h" 〉
-
or t1 = superscript £, t2 = "l"
-
or t1 = "ff", t2 = "F"
-
or (disjunctive(t1) ∧ t2 ∈ disjuncts(t1))
We conclude that together, the statement of practice and our study of the relation between the exemplar and the transcript, may confirm our hypothesis that the underlying idea of transcription in this edition is that of T-similarity.
Example 2: Melville's notes in an edition of Shakespeare
Statement of practice
On pages 955 to 970 of [Hayford et al. 1988], the editors
discuss notes made by the American author Herman Melville in an
edition of Shakespeare. On pages 967-970 they transcribe the notes
and provide facsimile images of the pages. (Facsimiles of
Melville's notes and the editor's transcript, taken from [Hayford et al. 1988], can be found in [Appendix B].) On page
967 they provide a guide to Symbols used
, which we
quote in full:[31]
[1] [...] revision or insertion enclosed in square brackets was made later than initial inscription of leaf
[2] <...> letters or words enclosed in diamond brackets were canceled by lining out
[3] <...>word letter(s) or word(s) written over are enclosed in diamond brackets closed up to the following word or letter that was superimposed
[4] ?word prefixed by a question mark indicates conjectural reading
[5] xxxx undeciphered letters (number of x's approximates numbers of letters involved)
[6] all words in roman are Melville's
[7] all words in italics outside brackets are words Melville underlined
[8] all words in italics inside brackets are editorial
There is no explicit statement, in this short text, that the exemplar has been transcribed in full or that nothing appears in the transcript that is not transcribing something in the exemplar. On the other hand, statement [6], that all words in roman are Melville's, and statements [4] and [8], which explain how editorial conjectures and additions are marked, indicate that great care has been taken to let the reader know which parts of the document are added by the editors. One way of making sense of this is to infer that for these transcribers it goes without saying that unless otherwise indicated, the transcript is complete and pure in our sense.[32]
Statements [1], [2], [3], and [7] indicate that the transcribers re-instantiate all insertions, cancellations, overwritings, and underlined text in the original, using the notations indicated. We model this by taking insertion, cancellation, overwriting, and underlining as compound types, which can like other types be recognized in the exemplar and re-instantiated in the transcript.[33] These statements have a further implication for the type system to be used in reading the transcript: the charitable reader will infer that the various forms of brackets used to signal insertions, cancellations, and overwriting do not appear in the exemplar, since if they did, any occurrence of them in the transcript would become ambiguous.[34]
Statement [5] describes an exception to the usual rule of type-identity between tokens in e and tokens in t, and also to the default 1:1 mapping between tokens in the two documents. Undeciphered tokens are transcribed by an approximate and not an exact number of x's, because tokens cannot be counted reliably until they have been identified as tokens of specific types. (To be a token is to be a token of a specific type.) And yet, if the occurrences of x in the transcript are tokens, then they must surely be tokens of some type.
This would call for treating a sequence of n occurrences of x as a single token of the type undeciphered sequence of letters about n characters wide. In the exemplar, the undeciphered sequence would be an atomic (or basic) token, not a compound one; in the transcript, it would be a compound token composed of an appropriate number of occurrences of x. The undeciphered token in the exemplar would then illustrate the principle that different readers may plausibly read an exemplar in different ways: a reader who manages to decipher a word will assign it and its characters to the appropriate types, while a reader who finds the word illegible will assign it to an undeciphered letter-sequence type of appropriate length.
(In the formalizations of the exemplar and the transcript
below, however, we rely on a source of information according to
which the sequence of letters in question, which are transcribed
as ?almxxxx
has been deciphered as
almanacks
. In this situation, the sequence in the
exemplar must be considered as a special exemplar token, while the
sequence in the transcript is a special transcript token.)
In contrast to the Penn Papers, page and line breaks are preserved in this edition. We observe that on this point none of the editions make any explicit statement on their choice. On our account, this suggests that the transcriptional implicatures of the two communities of practice are different: In this case it apparently goes without saying that line breaks are preserved, in the other it apparently goes without saying that they are not.
It is interesting to observe that the statement of practice in the case of Melville's notes is formulated in terms of what the reader will see in t and what it means, whereas the Penn Paper's statement of practice starts from phenomena in e and describes the rules the transcribers have followed in creating t. The statement of practice for Melville's notes, that is, specifies more or less directly what inferences the reader can make about e, given what is in t, whereas the Penn Papers rules specify more or less directly what the transcriber must do in t, given what is in e. To use the Penn transcripts to make inferences about e requires that the rules be applied backwards, so to speak.
Interpretation in terms of T-similarity
As before, we can interpret the statement of practice in terms of the three predicates se_token (special exemplar token), se_token (special transcript token), and typesimilar (similar types).
Among tokens present in e and not present in t are marks (arrows, circles, and carets to mark the desired point of insertion) which indicate where inserted material belongs. (Note that these are not mentioned explicitly in the list of symbols, perhaps because they do not appear as symbols in the transcript; their existence can be inferred only from the editorial remarks in the transcript, at line 18 of page 969.) None of these types are in fact transcribed in t. So, they can all be considered special exemplar tokens in e, having no corresponding tokens in t.
The phrase Conversation upon Gabriel, Micheal &
Raphel – gentlemanly &c
occurs in line 18 of t, but
only below in e. The occurrence in e may be considered a
special exemplar token.
Similarly, if one believes (as one probably should[35]) that the string represented as
?almxxxx
in line 23 of t can be read as
almanacks
, then almanacks
(or the
substring anacks
) of e should be considered a
special exemplar token.
The List of Symbols makes clear that italicized words within angle brackets are editorial, not authorial, i.e., they are special exemplar tokens.
Similarly, a question mark prefixed to a word whose reading is uncertain is a special transcript token.
Square brackets, angle brackets, angle brackets with immediately following text, question marks prefixed to words, and sequences of the character x are all identified by the List of Symbols as having special meaning to which no token in the exemplar directly corresponds, i.e., they are special transcript tokens.
We may take as a given that in normal cases it will be clear to a competent reader on inspection whether a token in t is or is not an instance of one of these types. However, one may also easily imagine cases in which it is not clear; these are not different in kind from legibility issues for the reader of t, both consisting in uncertainty as to the type instantiated by a token (or uncertainty as to whether a given mark is a token or not). As already pointed out, the process of reading lies outside the scope of our work.
There are no clear examples of type similarity between tokens in e and t. However, see our remarks on a normalized transcript further below.
Proof strategy
We have now discussed the Melville edition's statement of practice in terms of the rules of transcriptional implicature. By doing so we have formulated an account in more or less concise prose of which parts of e and t should be assumed special exemplar or transcript tokens, and which types should be assumed type-similar in order for e and t to be considered T-similar. How can we establish a formal proof to show that, given these assumptions, t and e are T-similar?
After all, e and t are concrete, visual objects. They cannot directly be subjected to the formal procedures required for such proofs, and definitely not to the digital tools we will use to test such proofs.
What we can do, however, is to let digital representations of the documents stand in for e and t (and all the tokens contained therein), and then perform the formal proof on these representations.
There are many ways this can be done. The way we have chosen, is to create TEI-XML documents representing t and e. (As a sanity test, we also produce HTML visual replicas of e and t from the XML.) From these XML documents we create one FOL representation of both documents and of the relations between them. By feeding this representation alongside our transcriptional implicature axioms to a theorem prover we check whether T-similarity between the two documents can be proven.
The theorem prover we have used is [Vampire]. In earlier work we have used [Alloy] and [Prolog] for similar tasks. One of the reasons we chose Vampire this time, is that it operates on fairly standard FOL notation, whereas other tools use more idiosyncratic notations. Moreover, Vampire is less apt to crash on the relatively large amounts of data involved in detailed representations of documents.
Explanation by way of a toy example
An exemplar and a transcript
Let us consider the following artificial example of an
exemplar e:
Figure 1: Exemplar Figure 2: Transcript

In order to justify the plausibility of t, we may assume
that the practice of the relevant community (or of this
particular project) is to include insertions and silently omit
deletions in a different hand (in this case, in red), to
normalize spelling (in this case, of Essexe
to
Essex
), and to add disambiguating remarks between
square brackets (in this case, [the Earl
of]
).
XML representations
Our first step is to create an XML representation of e. We propose:
<doc> Elizabeth went <del>with</del> <add>to</add> Essexe </doc>
Our next step is to create an XML representation of t. We propose:
<doc> Elizabeth went to <supplied>the Earl of</supplied> <choice><orig>Essexe</orig><reg>Essex</reg></choice>. </doc>
We observe that t both adds to and leaves things out of
e, in accordance with the transcribal practice suggested.
, must be counted as special
transcript tokens. Furthermore, we understand
<supplied>the Earl
of</supplied>
to indicate that the tokens <choice><orig>Essexe</orig><reg>Essex</reg></choice>Essexe
and
Essex
are typesimilar.
When it comes to e,
must be
considered a special exemplar token in relation to t.
<del>with</del>
Translation to FOL, and proof of T-similarity
Together, the XML representations of e and t, as input to an appropriately devised XSLT stylesheet (see [Appendix E]), yield the following output:
fof(case_specific_facts, axiom,
typeof(e1, t_Elizabeth) &
typeof(e2, t_went) &
typeof(e3, t_with) &
typeof(e4, t_to) &
typeof(e5, t_Essexe) &
typeof(t1, t_Elizabeth) &
typeof(t2, t_went) &
typeof(t3, t_to) &
typeof(t4, t_the) &
typeof(t5, t_Earl) &
typeof(t6, t_of) &
typeof(t7, t_Essex) &
typesimilar(t_Essexe, t_Essex) &
transcript(e1,t1) &
transcript(e2,t2) &
transcript(e4,t3) &
transcript(e5,t7) &
! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z)
) &
$distinct(
e1,e2,e3,e4,e5,
t1,t2,t3,t4,t5,t6,t7,
t_Elizabeth,t_went,t_with,t_to,t_Essexe,t_the,t_Earl,t_of,t_Essex
) &
! [X] : ((e_token(X) <=> (X=e1 | X=e2 | X=e4 | X=e5)) &
(t_token(X) <=> (X=t1 | X=t2 | X=t3 | X=t7)))
).
The Case specific facts are formulated as an axiom in the form of one conjunction with many conjuncts, representing the facts of e and t and the relations between them.
The listing above is in the so-called FOF (First-Order Form) notation used by Vampire. It is perhaps similar enough to the standard FOL notation used elsewhere in this paper as it is. For convenience, however, we include a translation to the usual notation, and group the various conjuncts into numbered groups for ease of reference:
Case specific facts:
-
typeof(e1,t_Elizabeth) ∧ typeof(e2,t_went) ∧ typeof(e3,t_with) ∧typeof(e4,t_to) ∧ typeof(e5,t_Essexe) ∧ -
typeof(t1,t_Elizabeth) ∧ typeof(t2,t_went) ∧ typeof(t3,t_to) ∧typeof(t4,t_the) ∧ typeof(t5,t_Earl) ∧ typeof(t6,t_of) ∧ typeof(t7,t_Essex) ∧ -
typesimilar(t_Essexe,t_Essex) ∧ -
transcript(e1,t1) & transcript(e2,t2) & transcript(e4,t3) & transcript(e5,t7) & -
∀(X,Y,Z) ( ((transcript(X,Y) & transcript(X,Z)) → Y=Z) & ((exemplar(X,Y) & exemplar(X,Z)) → Y=Z) ) & -
$distinct(e1,e2,e3,e4,e5, t1,t2,t3,t4,t5,t6,t7,t_Elizabeth,t_went,t_with,t_to,t_Essexe,t_the,t_Earl,t_of,t_Essex ) ∧ -
∀(X)((e_token(X) ⇔ (X=e1 ∨ X=e2 ∨ X=e4 ∨ X=e5)) ∧(t_token(X) ⇔ (X=t1 ∨ X=t2 ∨ X=t3 ∨ X=t7)))
In group 1 the tokens in the exemplar (normal as well as
special), named e1..e5
, are associated with their
respective types. The names of the types are simply given as a
string consisting of the word string in question, prefixed with
t_
. In other words, the token e1
is of type t_Elisabeth
, and so on.
In group 2 the tokens in the transcript, named
t1..t7
, are associated with their respective
types, in the same manner.
The statement in Group 3 states that the type named
t_Essexe
is similar to the type named
t_Essex
.
The statements in Group 4 assigns each of the exemplar
tokens in e to their corresponding transcript tokens in t. It
may deserve special attention that the special exemplar token e3
(of type t_with) and the special transcript tokens
t4, t5, and t6 (of types t_the, t_Earl,
and t_of), though represented as tokens associated
with their respective types, are not mentioned in the Group 4
statements. They are part of the formal description of the two
documents, though they do not play any role in T-similarity
proofs.
The statements in Group 5, 6, and 7 are there to deal with
the open world assumption of Vampire, and may be primarily of
technical interest.[36] In group 5, we state that there are no other pairs of
individuals than those mentioned in group 4 which stand in a
transcript relation to each other. In Group 6, the
$distinct
predicate is a special Vampire
predicate that makes sure that each of the constants mentioned
refer uniquely, i.e., e1
names an individual not
named by any of the other constants mentioned, etc.[37] The statement in Group 7 makes sure that there are no
other e_tokens or t_tokens in the universe of discourse than
those listed in the right-hand clauses of the two biconditionals.
When we add the axioms 1-10 and the definitions a-d of transcriptional implicature formulated earlier to the axiom of Case specific facts above, the theorem prover confirms that e and t are T-similar.[38]
Proof of T-similarity
We have gone through the same steps with the exemplar and transcript [Appendix B] of Melville's notes as described in the previous section on the toy example.
We first made XML representations of the exemplar and the transcript, see [Appendix C]. In both, we used fairly straightforward TEI encoding.[39]
As a sanity check we made sure we could transform the XML files to HTML which came as close as we thought feasible to the originals in textual as well as visual aspects. We present these HTML files in [Appendix D]. For convenience, special exemplar and special transcript tokens are marked in red.
The stylesheet briefly mentioned earlier, genFof.xsl, is reproduced in [Appendix E]. This stylesheet takes the two XML files in [Appendix C] as input, compares them, and produces the output in Vampire FOF format to be found in [Appendix F]. (In its current form, special token GIs are hard-coded into the stylesheet itself. One might easily extend the stylesheet so as to give users control over these, and to distinguish between special exemplar and transcript GI's.)
Vampire confirms as a theorem that the XML representations of the exemplar and the transcript satisfies the T-similarity predicate, or, in other words, that e and t as represented are T-similar.
We believe this to show that our definition of T-similarity, taking exceptions from default transcriptional implicature into account, is at least feasible. One such test can of course not be conclusive. As mentioned earlier, our hypothesis that there is a common core of assumptions underlying transcription in general is one that can only be tested empirically. We do not have the resources necessary for extensive empirical work.
A normalized transcription
However, we have at least made one other, very different,
normalized
, transcript of the exemplar ([Appendix G]). Unlike the transcript
discussed above, this one does contain some type similarity
statements: &
is declared type similar to
and
, &c
to etc.
,
almanacks
to almanacs
,
nonsence
to nonsense
, and
nominee
to nomine
.[40]
We put the normalized transcript to the same test. Again, T-similarity as defined here, with exceptions from default assumptions made explicitly and formally, shows that also the normalized transcript is T-similar to the exemplar. See [Appendix G].
This illustrates an important and more general point: Very different transcripts can be T-similar to one and the same exemplar. (Since T-similarity is symmetric and transitive, this means that the diplomatic and normalized transcriptions should be T-similar, too. We do not prove this.)
Another important point illustrated by this exercise is that the representation of the exemplar, in the way things are set up here, is not necessarily (or usually cannot even be) independent of the way the transcript is represented. (For example, a deletion in the exemplar should be represented as a special transcript token if and only if it is omitted from the transcript.)
Discussion
One of the things we think that this project illustrates, is the complexity of document structures and of the relations between documents. Our formalization has been on a fairly shallow level. We have succeeded in bringing the axioms of transcriptional implicature down to only a dozen fairly simple FOL statements, but the formal representation of even the short, less than 300 word Melville document, requires a conjunct of thousands of statements. Processing of the representation of even this small document requires a fairly large amount of computing power and time — with a medium-range personal computer Vampire takes several minutes to process the Melville transcripts.
Therefore, we would also like to observe here that Vampire has impressed us. As mentioned, we have tried to do this work with the help of other tools in the past, but have had to give up because of insurmountable difficulties with formalization, user interface, or speed. This is the first time we have had the occasion to do real work with formalization and automated theorem proving, even though still on a fairly modest scale. We realize that some of our work in Vampire might have been done more efficiently with better knowledge of its possibilities. Other theorem provers exist, but we are not presently aware of any that are better suited to the task.
We regard the work presented here as merely a proof of the concept of formalizing the modeling of documents and document relations, and of the applicability of the notion of T-similarity to transcription. Any implementation for real work would have to use other, more user-friendly and efficient tools, tools which do not, for example, require thorough knowledge of standard first order logic.
In [Huitfeldt / Sperberg-McQueen 2008] we presented two models of text:
The Grapheme-sequence Model
and the Readings
model
. We found the Grapheme-sequence Model
unsatisfactory because of a number limitations:
-
Texts are not just sequences of graphemes, -- the are usually more complex graph-like structures with notes, variants, alternate readings etc.
-
The analysis of document types must reflect compositionality and levels of types: there are atomic as well as boolean and compound types, types at the levels of letters, words, sentences, paragraphs, etc.
-
there may be different readings of the same document
In [Huitfeldt / Marcoux / Sperberg-McQueen 2010] we tried to develop a formal account without these limitations. However, we never reached a full working formalization, partly because of the problems with finding suitable formalization tools mentioned above.
With what is presented here, we are in many ways back with the
Grapheme-sequence Model
: We represent documents as
sequences — not even of letters, but of word tokens. (We
could of course have chosen to model on character level in stead.
That would have increased the number of tokens in the model,
without making much interesting difference in principle.) The type
structure is also represented as entirely flat, without levels or
compositionality.
Still, with these limitations, we have a working model which lends itself to full formalization and formal proof. We can only hope that someone may some time extend it beyond its current limitations, e.g., along the lines of [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
Some readers may have been suspicious of our move, in the section on Melville, from discussing the original exemplar and typescript on paper to discussing XML-representations of them. It was this move, however, which made us aware of a (perhaps to other obvious) fact: Our model presupposes that the representation of the exemplar must reflect what is not recorded in the transcript.[41]
Moreover, we think that this move shows that it might be useful to think in terms of T-similarity or similar formalizations of similarity between digital documents in contexts well beyond transcription and scholarly editing.
Originally, this talk had the subtitle a contribution
to markup semantics
. We can only hope that it will
be.
Conclusions and Future Work
We have proposed the term transcriptional implicature to denote the rules of inference which should govern the interpretation of a transcript in the absence of explicit statements to the contrary. Informally, the rules of transcriptional implicature (are intended to) capture those facts about transcription which are too obvious to need saying. We expect that the rules of transcriptional implicature vary from community of practice to community of practice; their accurate identification will be a matter of the sociology of scholarship, philosophy of science, and close examination of transcription practice. We believe that the rules of what we call default transcriptional implicature provide a plausible example of what the transcriptional implicature for a given community of practice might look like, and we have shown how the practice of some concrete examples can be captured formally in ways which exhibit clearly both the general rules governing transcription and the specific practices of the individual project.
We hope that this work can provide a basis for a full model of the logical structure of transcription, of the formal relation between exemplar and transcript, and of the ways in which a transcript can be used to infer facts about its exemplar.
Appendix A. Penn: Facsimile of exemplar and transcript.
Facsimiles of the first and last page of a letter from Penn to Lady Conway from 1675, reproduced on the insides of the front and back cover of [The Papers of William Penn]:
Figure 3: Penn, letter to Lady Conway, first page

Figure 4: Penn, letter to Lady Conway, last page

Facsimiles of excerpt from the transcript of the corresponding pages in [The Papers of William Penn], pages 356 and 358:
Figure 5: Penn, letter to Lady Conway, first page

Figure 6: Penn, letter to Lady Conway, last page

Appendix B. Melville: Facsimile of exemplar and transcript
[Hayford et al. 1988] contain facsimiles of the two pages of Melville's exemplar on pages 968 and 970.
Facsimiles of the same two pages in better quality can be found here:
Facsimile of transcript, [Hayford et al. 1988] pages 969 and 970
Figure 7: Transcript, first page

Figure 8: Transcript, second page

Appendix C. Melville: XML representations
XML representation of the exemplar
<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="../../genFofXSLT/genFof.xsl"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->]>
<text>
<body>
<p>
<lb/>A seaman figures in The Canterbury Tales. <lb/>With
<del>a</del>many a tempest had his beard been <lb/>shook. –
<del>S</del>
<del>Deep g</del> Secret grief is a <lb/>cannibal of its own heart
– <emph>Bacon</emph>. <lb/>&vbar;<emph>Claudia</emph> of the
<emph>Appian</emph> family, <q>I wish <lb/>&vbar; some fight or
pestilence would thin out <lb/>&vbar; this crowd.</q>
<emph>Arrogance.</emph>
<lb/>
<emph>Roast beef in the pulpet.</emph>
</p>
<p>
<lb/>An animal of a man — <q>do eagles wear
<lb/>spectacles?</q>— Health. — <emph>Contrast</emph>:
an <lb/>over spiritual man. <lb/>—— </p>
<p>
<lb/>
<q>Yes, Madam, Cain was a godless froward boy, & <lb/>Reuben
(Gen:49) & Absalom</q> Many pious men <lb/>have impious
children — (Devil as a Quaker)
<lb/>————— </p>
<p>
<lb/>A formal compact – Imprimis – First –
Second. <lb/>The aforesaid soul. said soul &c –
Duplicates – <lb/>&triplebar;<q>How was it about the
temptation on the <lb/>hill?</q> &c – D begs the hero to
form <lb/>one of a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb/>&c –
Leaves a letter to the D – <q>My <lb/>
<emph>Dear D</emph>
</q>
<floatingText>
<body>
<ab> – Conversation upon Gabriel, Micheal &
<lb/> Raphel – gentlemanly &c</ab>
</body>
</floatingText>
</p>
<p>
<lb/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb/> making
<sic>almanacks</sic> – going to a ball takes a long
<lb/>time making toilette. – The Doctor's coach stops
<lb/>the way. – <q>Do you believe all that stuff? <lb/>
nonsence – the world was never made. – “But Is
not <lb/>this you mentioned <emph>here</emph> – in the
scriptures?</q>
<lb/>Receives visits from the principal d's –
<q>Gentlemen</q> &c. <lb/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb/>be below with kings than above with fools?</q>
</p>
<p>
<lb/>It is better to laugh & not sin than to <del>be</del> weep
& be <lb/>wicked. — Ten loads of coal to burn him.
— <lb/>Brought to the stake — warmed himself by the
fire. </p>
<p>
<lb/>Ego non baptizo te in nominee Patris et <lb/>Filii et Spiritus
Sancti – sed in nomine <lb/>Diaboli. — Madness is
undefinable — <lb/>It & right reasons extremes of one.
<lb/>–Not the <add place="sup">(black art)</add> Goetic but
Theurgic magic — <lb/>seeks converse with the Intelligence,
Power, the <lb/>Angel. </p>
</body>
</text>
XML representation of the transcript
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->]>
<text>
<front>
<docTitle>
<titlePart>NOTES IN A SHAKESPEARE VOLUME
<span>969</span></titlePart>
</docTitle>
</front>
<body>
<fw>[on verso of last leaf of Volume VII, page [524]]</fw>
<p>
<lb n="1"/>A seaman figures in The Canterbury Tales. <lb n="2"
/>With <del>a</del>many a tempest had his beard been <lb n="3"
/>shook. – <del>S</del>
<del>Deep g</del> Secret grief is a <lb n="4"/>cannibal of its own
heart – <emph>Bacon</emph>. <lb n="5"/>
<metamark>&vbar;</metamark>
<emph>Claudia</emph> of the <emph>Appian</emph> family, <q>I wish
<lb n="6"/>
<metamark>&vbar;</metamark> some fight or pestilence would thin
out <lb n="7"/>
<metamark>&vbar;</metamark> this crowd.</q>
<emph>Arrogance.</emph>
<lb n="8"/>
<emph>Roast beef in the pulpet.</emph>
</p>
<p>
<lb n="9"/>An animal of a man — <q>do eagles wear <lb n="10"
/>spectacles?</q>— Health. — <emph>Contrast</emph>: an
<lb n="11"/>over spiritual man. <lb rend="dunno"/>
<metamark>——</metamark>
</p>
<p>
<lb n="12"/>
<q>Yes, Madam, Cain was a godless froward boy, & <lb n="13"
/>Reuben (Gen:49) & Absalom</q> Many pious men <lb n="14"
/>have impious children — (Devil as a Quaker) <lb
rend="dunno"/>
<metamark>—————</metamark>
</p>
<p>
<lb n="15"/>A formal compact – Imprimis – First –
Second. <lb n="16"/>The aforesaid soul. said soul &c –
Duplicates – <lb n="17"/>&triplebar;<q>How was it about the
temptation on the <lb n="18"/>hill?</q> &c <supplied rend="[]"
>inserted later below in lines 21–21b after Dear D“
— and circled <lb rend="dunno"/> with guideline to caret
here</supplied>
<floatingText>
<body>
<ab>Conversation upon Gabriel, Micheal & / <lb rend="dunno"/>
Raphel – gentlemanly &c<supplied rend="]"/></ab>
</body>
</floatingText> – D begs the hero to form <lb n="19"/>one of
a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb n="20"/>&c
– Leaves a letter to the D – <q>My <lb n="21"/> Dear D
</q> – <supplied rend="[]">later insertion in lines
21–21b, reported in line 18</supplied>
</p>
<p>
<lb n="22"/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb n="23"/>
<unclear>making</unclear>
<unclear><supplied>alm<gap/></supplied></unclear> – going to
a ball takes a long <lb n="24"/>time making toilette. – The
Doctor's coach stops <lb n="25"/>the way. – <q>Do you
believe all that stuff? <lb n="26"/> nonsence – the world
was never made. – <supplied rend="[">add</supplied>
“But<supplied rend="]"/> Is not <lb n="27"/>this you
mentioned <emph>here</emph> – in the scriptures?</q>
<lb n="28"/>Receives visits from the principal d's –
<q>Gentlemen</q> &c. <lb n="29"/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb n="30"/>be below with kings than above with fools?</q>
</p>
<p>
<fw>[on recto of last blank leaf of Volume VII, page [523]]</fw>
<lb n="31"/>It is better to laugh & not sin than to
<del>be</del> weep & be <lb n="32"/>wicked. — Ten loads
of coal to burn him. — <lb n="33"/>Brought to the stake
— warmed himself by the fire. </p>
<p>
<lb n="34"/>Ego non baptizo te in nominee Patris et <lb n="35"
/>Filii et Spiritus Sancti – sed in nomine <lb n="36"
/>Diaboli. — Madness is undefinable — <lb n="37"/>It
& right reasons extremes of one. <lb n="38"/>–Not the
<supplied rend="[">inserted above line with caret below</supplied>
(black art)<supplied rend="]"/> Goetic but Theurgic magic —
<lb n="39"/>seeks converse with the Intelligence, Power, the <lb
n="40"/>Angel. </p>
</body>
</text>
Appendix D. Melville: HTML presentations
HTML presentation of the exemplar
Special exemplar tokens in red.
Figure 9: HTML presentation of the exemplar

HTML presentation of the transcript
Special transcript tokens in red.
Figure 10: HTML presentation of the transcript

Appendix E. genFoF stylesheet
<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
xmlns:ym="http://www.marcouxmedias.com"
xmlns:xs="http://www.w3.org/2001/XMLSchema"
xpath-default-namespace="" xml:lang="fr-CA" xml:space="ignore"
exclude-result-prefixes="#all">
<xsl:output method="text" indent="no" encoding="UTF-8" />
<xsl:function name="ym:lastIndexOf" as="xs:integer">
<xsl:param name="str" />
<xsl:param name="car" />
<xsl:sequence select="if (contains($str, $car)) then
max(for $i in (1 to string-length($str)) return
if (substring($str, $i, 1) = $car) then $i else 0)
else 0" />
</xsl:function>
<xsl:variable name="fpNoExt">
<xsl:variable name="temp" select="base-uri(/)" />
<xsl:value-of select="substring($temp, 1, ym:lastIndexOf($temp, '.'))" />
</xsl:variable>
<xsl:variable name="indent" select="' '"/>
<xsl:function name="ym:tokenizePlus" as="item()*">
<xsl:param name="text" as="node()" />
<xsl:choose>
<xsl:when test="not($text/ancestor::orig)">
<xsl:variable name="toks" select=
"tokenize($text,'[^a-zA-Z0-9]+')[.]" />
<xsl:choose>
<!-- The values of test specified below are project-specific -->
<xsl:when test="$text/ancestor::front
or $text/ancestor::fw
or $text/ancestor::supplied
or $text/ancestor::sic
or $text/ancestor::floatingText">
<xsl:sequence select="for $t in $toks return '*' || $t" />
</xsl:when>
<xsl:otherwise>
<xsl:sequence select="$toks" />
</xsl:otherwise>
</xsl:choose>
</xsl:when>
<xsl:otherwise />
</xsl:choose>
</xsl:function>
<xsl:function name="ym:addEffSeqNum" as="item()*">
<xsl:param name="toks" as="item()*" />
<xsl:sequence select="
for $n in 1 to count($toks) return
(if (not(starts-with($toks[$n],'*'))) then
count($toks[position() lt $n and not(starts-with(.,'*'))]) || '*' else '')
|| $toks[$n]
" />
</xsl:function>
<xsl:template match="/">
<xsl:variable name="doc1" select="/*"/>
<xsl:variable name="toks1" select=
"ym:addEffSeqNum($doc1//text()/ym:tokenizePlus(.))" />
<xsl:variable name="doc2" select="document($fpNoExt || 'transcript.xml')/*"/>
<xsl:variable name="toks2" select=
"ym:addEffSeqNum($doc2//text()/ym:tokenizePlus(.))" />
<xsl:variable name="nNormToks" select="count($toks1[not(starts-with(.,'*'))])" />
<xsl:if test="$nNormToks ne count($toks2[not(starts-with(.,'*'))])">
<xsl:message terminate="no"
select="string-join(($nNormToks, count($toks2[not(starts-with(.,'*'))]), ' '), ' ')"
>WARNING: Normal token counts differ between documents.</xsl:message>
</xsl:if>
<xsl:text>fof(case_specific_facts, axiom,
</xsl:text>
<xsl:call-template name="processDoc">
<xsl:with-param name="docPrefix" select="'e'" />
<xsl:with-param name="toks" select="$toks1" />
</xsl:call-template>
<xsl:call-template name="processDoc">
<xsl:with-param name="docPrefix" select="'t'" />
<xsl:with-param name="toks" select="$toks2" />
</xsl:call-template>
<xsl:for-each select="$doc2//choice">
<!-- The [.] in the following are to get rid of possible empty strings at the
beginning and end of the string: -->
<xsl:variable name="toksOrig" select="tokenize(orig,'[^a-zA-Z0-9]+')[.]"/>
<xsl:variable name="toksReg" select="tokenize(reg,'[^a-zA-Z0-9]+')[.]"/>
<xsl:if test="count($toksOrig) ne count($toksReg) or not(count($toksOrig))">
<xsl:message select="string-join((count($toksOrig), count($toksReg), ' '), ' ')"
>WARNING: Token count mismatch within a <choice> element.</xsl:message>
</xsl:if>
<xsl:for-each select="1 to count($toksOrig)">
<xsl:if test="$toksOrig[current()] ne $toksReg[current()]">
<xsl:value-of select="$indent" />
<xsl:text>typesimilar(t_</xsl:text>
<xsl:value-of select="$toksOrig[current()]" />
<xsl:text>, t_</xsl:text>
<xsl:value-of select="$toksReg[current()]" />
<xsl:text>) &
</xsl:text>
</xsl:if>
</xsl:for-each>
</xsl:for-each>
<xsl:for-each select="1 to count($toks1)">
<xsl:if test="not(starts-with($toks1[current()],'*'))">
<xsl:variable name="idNormTok" select="substring-before($toks1[current()],'*')" />
<xsl:value-of select="$indent" />
<xsl:text>transcript(e</xsl:text>
<xsl:value-of select="." />
<xsl:text>,t</xsl:text>
<xsl:value-of select="max(
for $i in 1 to count($toks2) return
if (starts-with($toks2[$i], $idNormTok || '*')) then $i else 0
)" />
<xsl:text>) &
</xsl:text>
</xsl:if>
</xsl:for-each>
<!--% No other pairs of corresponding tokens (gives us reciprocity for free):-->
<xsl:text>
! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z)
) &
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text>$distinct(
</xsl:text><xsl:value-of select="$indent" /><xsl:value-of select="$indent" />
<xsl:value-of select="
string-join(for $i in 1 to count($toks1) return ('e' || $i), ',')
" />
<xsl:text>,
</xsl:text><xsl:value-of select="$indent" /><xsl:value-of select="$indent" />
<xsl:value-of select="
string-join(for $i in 1 to count($toks2) return ('t' || $i), ',')
" />
<xsl:text>,
</xsl:text><xsl:value-of select="$indent" /><xsl:value-of select="$indent" />
<xsl:variable name="allTypes" select="
(for $i in 1 to count($toks1) return 't_' || substring-after($toks1[$i], '*')),
(for $i in 1 to count($toks2) return 't_' || substring-after($toks2[$i], '*'))
"/>
<xsl:value-of select="string-join(distinct-values($allTypes), ',')" />
<xsl:text>
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text>) &
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text>! [X] : ((e_token(X) <=> (X=e</xsl:text>
<xsl:value-of select="string-join(
(1 to count($toks1))[let $i := . return not(starts-with($toks1[$i],'*'))],
' | X=e')" />
<xsl:text>)) &
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text> (t_token(X) <=> (X=t</xsl:text>
<xsl:value-of select="string-join(
(1 to count($toks2))[let $i := . return not(starts-with($toks2[$i],'*'))],
' | X=t')" />
<xsl:text>)))
).</xsl:text>
</xsl:template>
<xsl:template name="processDoc">
<xsl:param name="toks" />
<xsl:param name="docPrefix" />
<!--<xsl:message select="$toks" />-->
<xsl:for-each select="$toks">
<!--
<xsl:value-of select="$indent" />
<xsl:value-of select="if (starts-with(.,'*')) then 's' else ''"/>
<xsl:value-of select="$docPrefix"/>
<xsl:text>_token(</xsl:text>
<xsl:value-of select="$docPrefix"/>
<xsl:value-of select="position()" />
<xsl:text>) &
</xsl:text>
^ 2026-03-31 Zoom meeting, decided we don’t need them. -->
<xsl:value-of select="$indent" />
<xsl:text>typeof(</xsl:text>
<xsl:value-of select="$docPrefix"/>
<xsl:value-of select="position()"/>
<xsl:text>, t_</xsl:text>
<xsl:value-of select="substring-after(.,'*')"/>
<xsl:text>) &
</xsl:text>
</xsl:for-each>
</xsl:template>
</xsl:stylesheet>
Appendix F. Melville: FOF representation of exemplar and transcript
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% General axioms
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(transcript_domain, axiom,
! [X,Y] : (transcript(X,Y) => (e_token(X) & t_token(Y)))).
fof(exemplar_domain, axiom,
! [X,Y] : (exemplar(X,Y) => (t_token(X) & e_token(Y)))).
%%%% Slightly modified from Default:
fof(typeof_domain, axiom,
! [X,Y] : (typeof(X,Y) =>
((t_token(X) | e_token(X) | st_token(X) | se_token(X))
& type(Y)))).
%%%% Slightly modified from Default:
fof(no_class_overlap, axiom,
! [X] : (
(e_token(X) <=> (~t_token(X) & ~type(X) & ~se_token(X) & ~st_token(X))) &
(t_token(X) <=> (~e_token(X) & ~type(X) & ~se_token(X) & ~st_token(X))) &
(type(X) <=> (~e_token(X) & ~t_token(X)) & ~se_token(X) & ~st_token(X)) &
(se_token(X) <=> (~t_token(X) & ~type(X) & ~e_token(X) & ~st_token(X))) &
(st_token(X) <=> (~t_token(X) & ~type(X) & ~se_token(X) & ~e_token(X)))
)).
fof(exemplar_and_transcript_inverse, axiom,
! [X,Y] : (exemplar(X,Y) <=> transcript(Y,X))).
fof(at_most_one_type_per_token, axiom,
! [X,Y,Z] : ((typeof(X,Y) & typeof(X,Z)) => Y=Z)).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Axioms not in Default
fof(typesimilar_domain, axiom,
! [X,Y] : (typesimilar(X,Y) => (type(X) & type(Y)))).
fof(typesimilar_reflexive, axiom,
! [X] : (type(X) => typesimilar(X,X))).
fof(typesimilar_symmetric, axiom,
! [X,Y] : (typesimilar(X,Y) => typesimilar(Y,X))).
fof(type_similar_transitive, axiom,
! [X,Y,Z] : ((typesimilar(X,Y) & typesimilar(Y,Z)) => typesimilar(X,Z))).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Definitions
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(reciprocity,definition,
reciprocity <=> ! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z))).
fof(completeness, definition,
completeness <=> ! [X] :
(e_token(X) => ? [Y] : transcript(X,Y))).
fof(purity, definition,
purity <=> ! [X] :
(t_token(X) => ? [Y] : exemplar(X,Y))).
%%%% Slightly modified from Default:
fof(type_similarity, definition,
type_similarity <=> ![X,Y] : (transcript(X,Y) =>
? [Z,V] : (typeof(X,Z) & typeof(Y,V) & typesimilar(Z,V)))).
%%%% Slightly modified from Default:
fof(t_similarity, definition,
t_similarity <=> (reciprocity & completeness & purity & type_similarity)).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Case-specific facts
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(case_specific_facts, axiom,
typeof(e1, t_A) &
typeof(e2, t_seaman) &
typeof(e3, t_figures) &
typeof(e4, t_in) &
typeof(e5, t_The) &
typeof(e6, t_Canterbury) &
typeof(e7, t_Tales) &
typeof(e8, t_With) &
typeof(e9, t_a) &
typeof(e10, t_many) &
typeof(e11, t_a) &
typeof(e12, t_tempest) &
typeof(e13, t_had) &
typeof(e14, t_his) &
typeof(e15, t_beard) &
typeof(e16, t_been) &
typeof(e17, t_shook) &
typeof(e18, t_S) &
typeof(e19, t_Deep) &
typeof(e20, t_g) &
%%% ...Approximately 250-350 lines omitted...
typeof(e270, t_right) &
typeof(e271, t_reasons) &
typeof(e272, t_extremes) &
typeof(e273, t_of) &
typeof(e274, t_one) &
typeof(e275, t_Not) &
typeof(e276, t_the) &
typeof(e277, t_black) &
typeof(e278, t_art) &
typeof(e279, t_Goetic) &
typeof(e280, t_but) &
typeof(e281, t_Theurgic) &
typeof(e282, t_magic) &
typeof(e283, t_seeks) &
typeof(e284, t_converse) &
typeof(e285, t_with) &
typeof(e286, t_the) &
typeof(e287, t_Intelligence) &
typeof(e288, t_Power) &
typeof(e289, t_the) &
%%% ...Approximately 250-350 lines omitted...
typeof(t332, t_inserted) &
typeof(t333, t_above) &
typeof(t334, t_line) &
typeof(t335, t_with) &
typeof(t336, t_caret) &
typeof(t337, t_below) &
typeof(t338, t_black) &
typeof(t339, t_art) &
typeof(t340, t_Goetic) &
typeof(t341, t_but) &
typeof(t342, t_Theurgic) &
typeof(t343, t_magic) &
typeof(t344, t_seeks) &
typeof(t345, t_converse) &
typeof(t346, t_with) &
typeof(t347, t_the) &
typeof(t348, t_Intelligence) &
typeof(t349, t_Power) &
typeof(t350, t_the) &
typeof(t351, t_Angel) &
%%% ...Approximately 250-350 lines omitted...
transcript(e271,t326) &
transcript(e272,t327) &
transcript(e273,t328) &
transcript(e274,t329) &
transcript(e275,t330) &
transcript(e276,t331) &
transcript(e277,t338) &
transcript(e278,t339) &
transcript(e279,t340) &
transcript(e280,t341) &
transcript(e281,t342) &
transcript(e282,t343) &
transcript(e283,t344) &
transcript(e284,t345) &
transcript(e285,t346) &
transcript(e286,t347) &
transcript(e287,t348) &
transcript(e288,t349) &
transcript(e289,t350) &
transcript(e290,t351) &
% No other pairs of corresponding tokens (gives us reciprocity for free):
! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z)
) &
$distinct(
e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13,e14,e15,e16,e17,e18,e19,e20,e21,e22,e23,e24,e25,e26,e27,e28,e29,e30,e31,e32,e33,e34,e35,e36,e37,e38,e39,e40,e41,e42,e43,e44,e45,e46,e47,e48,e49,e50,e51,e52,e53,e54,e55,e56,e57,e58,e59,e60,e61,e62,e63,e64,e65,e66,e67,e68,e69,e70,e71,e72,e73,e74,e75,e76,e77,e78,e79,e80,e81,e82,e83,e84,e85,e86,e87,e88,e89,e90,e91,e92,e93,e94,e95,e96,e97,e98,e99,e100,e101,e102,e103,e104,e105,e106,e107,e108,e109,e110,e111,e112,e113,e114,e115,e116,e117,e118,e119,e120,e121,e122,e123,e124,e125,e126,e127,e128,e129,e130,e131,e132,e133,e134,e135,e136,e137,e138,e139,e140,e141,e142,e143,e144,e145,e146,e147,e148,e149,e150,e151,e152,e153,e154,e155,e156,e157,e158,e159,e160,e161,e162,e163,e164,e165,e166,e167,e168,e169,e170,e171,e172,e173,e174,e175,e176,e177,e178,e179,e180,e181,e182,e183,e184,e185,e186,e187,e188,e189,e190,e191,e192,e193,e194,e195,e196,e197,e198,e199,e200,e201,e202,e203,e204,e205,e206,e207,e208,e209,e210,e211,e212,e213,e214,e215,e216,e217,e218,e219,e220,e221,e222,e223,e224,e225,e226,e227,e228,e229,e230,e231,e232,e233,e234,e235,e236,e237,e238,e239,e240,e241,e242,e243,e244,e245,e246,e247,e248,e249,e250,e251,e252,e253,e254,e255,e256,e257,e258,e259,e260,e261,e262,e263,e264,e265,e266,e267,e268,e269,e270,e271,e272,e273,e274,e275,e276,e277,e278,e279,e280,e281,e282,e283,e284,e285,e286,e287,e288,e289,e290,
t1,t2,t3,t4,t5,t6,t7,t8,t9,t10,t11,t12,t13,t14,t15,t16,t17,t18,t19,t20,t21,t22,t23,t24,t25,t26,t27,t28,t29,t30,t31,t32,t33,t34,t35,t36,t37,t38,t39,t40,t41,t42,t43,t44,t45,t46,t47,t48,t49,t50,t51,t52,t53,t54,t55,t56,t57,t58,t59,t60,t61,t62,t63,t64,t65,t66,t67,t68,t69,t70,t71,t72,t73,t74,t75,t76,t77,t78,t79,t80,t81,t82,t83,t84,t85,t86,t87,t88,t89,t90,t91,t92,t93,t94,t95,t96,t97,t98,t99,t100,t101,t102,t103,t104,t105,t106,t107,t108,t109,t110,t111,t112,t113,t114,t115,t116,t117,t118,t119,t120,t121,t122,t123,t124,t125,t126,t127,t128,t129,t130,t131,t132,t133,t134,t135,t136,t137,t138,t139,t140,t141,t142,t143,t144,t145,t146,t147,t148,t149,t150,t151,t152,t153,t154,t155,t156,t157,t158,t159,t160,t161,t162,t163,t164,t165,t166,t167,t168,t169,t170,t171,t172,t173,t174,t175,t176,t177,t178,t179,t180,t181,t182,t183,t184,t185,t186,t187,t188,t189,t190,t191,t192,t193,t194,t195,t196,t197,t198,t199,t200,t201,t202,t203,t204,t205,t206,t207,t208,t209,t210,t211,t212,t213,t214,t215,t216,t217,t218,t219,t220,t221,t222,t223,t224,t225,t226,t227,t228,t229,t230,t231,t232,t233,t234,t235,t236,t237,t238,t239,t240,t241,t242,t243,t244,t245,t246,t247,t248,t249,t250,t251,t252,t253,t254,t255,t256,t257,t258,t259,t260,t261,t262,t263,t264,t265,t266,t267,t268,t269,t270,t271,t272,t273,t274,t275,t276,t277,t278,t279,t280,t281,t282,t283,t284,t285,t286,t287,t288,t289,t290,t291,t292,t293,t294,t295,t296,t297,t298,t299,t300,t301,t302,t303,t304,t305,t306,t307,t308,t309,t310,t311,t312,t313,t314,t315,t316,t317,t318,t319,t320,t321,t322,t323,t324,t325,t326,t327,t328,t329,t330,t331,t332,t333,t334,t335,t336,t337,t338,t339,t340,t341,t342,t343,t344,t345,t346,t347,t348,t349,t350,t351,
t_A,t_seaman,t_figures,t_in,t_The,t_Canterbury,t_Tales,t_With,t_a,t_many,t_tempest,t_had,t_his,t_beard,t_been,t_shook,t_S,t_Deep,t_g,t_Secret,t_grief,t_is,t_cannibal,t_of,t_its,t_own,t_heart,t_Bacon,t_Claudia,t_the,t_Appian,t_family,t_I,t_wish,t_some,t_fight,t_or,t_pestilence,t_would,t_thin,t_out,t_this,t_crowd,t_Arrogance,t_Roast,t_beef,t_pulpet,t_An,t_animal,t_man,t_do,t_eagles,t_wear,t_spectacles,t_Health,t_Contrast,t_an,t_over,t_spiritual,t_Yes,t_Madam,t_Cain,t_was,t_godless,t_froward,t_boy,t_Reuben,t_Gen,t_49,t_Absalom,t_Many,t_pious,t_men,t_have,t_impious,t_children,t_Devil,t_as,t_Quaker,t_formal,t_compact,t_Imprimis,t_First,t_Second,t_aforesaid,t_soul,t_said,t_c,t_Duplicates,t_How,t_it,t_about,t_temptation,t_on,t_hill,t_D,t_begs,t_hero,t_to,t_form,t_one,t_Society,t_s,t_name,t_be,t_weighty,t_Leaves,t_letter,t_My,t_Dear,t_Conversation,t_upon,t_Gabriel,t_Micheal,t_Raphel,t_gentlemanly,t_Terra,t_Oblivionis,t_Hellites,t_At,t_Astor,t_find,t_him,t_making,t_almanacks,t_going,t_ball,t_takes,t_long,t_time,t_toilette,t_Doctor,t_coach,t_stops,t_way,t_Do,t_you,t_believe,t_all,t_that,t_stuff,t_nonsence,t_world,t_never,t_made,t_But,t_Is,t_not,t_mentioned,t_here,t_scriptures,t_Receives,t_visits,t_from,t_principal,t_d,t_Gentlemen,t_Arguments,t_persuade,t_Would,t_rather,t_below,t_with,t_kings,t_than,t_above,t_fools,t_It,t_better,t_laugh,t_sin,t_weep,t_wicked,t_Ten,t_loads,t_coal,t_burn,t_Brought,t_stake,t_warmed,t_himself,t_by,t_fire,t_Ego,t_non,t_baptizo,t_te,t_nominee,t_Patris,t_et,t_Filii,t_Spiritus,t_Sancti,t_sed,t_nomine,t_Diaboli,t_Madness,t_undefinable,t_right,t_reasons,t_extremes,t_Not,t_black,t_art,t_Goetic,t_but,t_Theurgic,t_magic,t_seeks,t_converse,t_Intelligence,t_Power,t_Angel,t_NOTES,t_IN,t_SHAKESPEARE,t_VOLUME,t_969,t_verso,t_last,t_leaf,t_Volume,t_VII,t_page,t_524,t_inserted,t_later,t_lines,t_21,t_21b,t_after,t_and,t_circled,t_guideline,t_caret,t_insertion,t_reported,t_line,t_18,t_alm,t_add,t_recto,t_blank,t_523
) &
! [X] : ((e_token(X) <=> (X=e1 | X=e2 | X=e3 | X=e4 | X=e5 | X=e6 | X=e7 | X=e8 | X=e9 | X=e10 | X=e11 | X=e12 | X=e13 | X=e14 | X=e15 | X=e16 | X=e17 | X=e18 | X=e19 | X=e20 | X=e21 | X=e22 | X=e23 | X=e24 | X=e25 | X=e26 | X=e27 | X=e28 | X=e29 | X=e30 | X=e31 | X=e32 | X=e33 | X=e34 | X=e35 | X=e36 | X=e37 | X=e38 | X=e39 | X=e40 | X=e41 | X=e42 | X=e43 | X=e44 | X=e45 | X=e46 | X=e47 | X=e48 | X=e49 | X=e50 | X=e51 | X=e52 | X=e53 | X=e54 | X=e55 | X=e56 | X=e57 | X=e58 | X=e59 | X=e60 | X=e61 | X=e62 | X=e63 | X=e64 | X=e65 | X=e66 | X=e67 | X=e68 | X=e69 | X=e70 | X=e71 | X=e72 | X=e73 | X=e74 | X=e75 | X=e76 | X=e77 | X=e78 | X=e79 | X=e80 | X=e81 | X=e82 | X=e83 | X=e84 | X=e85 | X=e86 | X=e87 | X=e88 | X=e89 | X=e90 | X=e91 | X=e92 | X=e93 | X=e94 | X=e95 | X=e96 | X=e97 | X=e98 | X=e99 | X=e100 | X=e101 | X=e102 | X=e103 | X=e104 | X=e105 | X=e106 | X=e107 | X=e108 | X=e109 | X=e110 | X=e111 | X=e112 | X=e113 | X=e114 | X=e115 | X=e116 | X=e117 | X=e118 | X=e119 | X=e120 | X=e121 | X=e122 | X=e123 | X=e124 | X=e125 | X=e126 | X=e127 | X=e128 | X=e129 | X=e130 | X=e131 | X=e132 | X=e133 | X=e134 | X=e135 | X=e136 | X=e137 | X=e138 | X=e139 | X=e140 | X=e148 | X=e149 | X=e150 | X=e151 | X=e152 | X=e153 | X=e154 | X=e155 | X=e156 | X=e158 | X=e159 | X=e160 | X=e161 | X=e162 | X=e163 | X=e164 | X=e165 | X=e166 | X=e167 | X=e168 | X=e169 | X=e170 | X=e171 | X=e172 | X=e173 | X=e174 | X=e175 | X=e176 | X=e177 | X=e178 | X=e179 | X=e180 | X=e181 | X=e182 | X=e183 | X=e184 | X=e185 | X=e186 | X=e187 | X=e188 | X=e189 | X=e190 | X=e191 | X=e192 | X=e193 | X=e194 | X=e195 | X=e196 | X=e197 | X=e198 | X=e199 | X=e200 | X=e201 | X=e202 | X=e203 | X=e204 | X=e205 | X=e206 | X=e207 | X=e208 | X=e209 | X=e210 | X=e211 | X=e212 | X=e213 | X=e214 | X=e215 | X=e216 | X=e217 | X=e218 | X=e219 | X=e220 | X=e221 | X=e222 | X=e223 | X=e224 | X=e225 | X=e226 | X=e227 | X=e228 | X=e229 | X=e230 | X=e231 | X=e232 | X=e233 | X=e234 | X=e235 | X=e236 | X=e237 | X=e238 | X=e239 | X=e240 | X=e241 | X=e242 | X=e243 | X=e244 | X=e245 | X=e246 | X=e247 | X=e248 | X=e249 | X=e250 | X=e251 | X=e252 | X=e253 | X=e254 | X=e255 | X=e256 | X=e257 | X=e258 | X=e259 | X=e260 | X=e261 | X=e262 | X=e263 | X=e264 | X=e265 | X=e266 | X=e267 | X=e268 | X=e269 | X=e270 | X=e271 | X=e272 | X=e273 | X=e274 | X=e275 | X=e276 | X=e277 | X=e278 | X=e279 | X=e280 | X=e281 | X=e282 | X=e283 | X=e284 | X=e285 | X=e286 | X=e287 | X=e288 | X=e289 | X=e290)) &
(t_token(X) <=> (X=t17 | X=t18 | X=t19 | X=t20 | X=t21 | X=t22 | X=t23 | X=t24 | X=t25 | X=t26 | X=t27 | X=t28 | X=t29 | X=t30 | X=t31 | X=t32 | X=t33 | X=t34 | X=t35 | X=t36 | X=t37 | X=t38 | X=t39 | X=t40 | X=t41 | X=t42 | X=t43 | X=t44 | X=t45 | X=t46 | X=t47 | X=t48 | X=t49 | X=t50 | X=t51 | X=t52 | X=t53 | X=t54 | X=t55 | X=t56 | X=t57 | X=t58 | X=t59 | X=t60 | X=t61 | X=t62 | X=t63 | X=t64 | X=t65 | X=t66 | X=t67 | X=t68 | X=t69 | X=t70 | X=t71 | X=t72 | X=t73 | X=t74 | X=t75 | X=t76 | X=t77 | X=t78 | X=t79 | X=t80 | X=t81 | X=t82 | X=t83 | X=t84 | X=t85 | X=t86 | X=t87 | X=t88 | X=t89 | X=t90 | X=t91 | X=t92 | X=t93 | X=t94 | X=t95 | X=t96 | X=t97 | X=t98 | X=t99 | X=t100 | X=t101 | X=t102 | X=t103 | X=t104 | X=t105 | X=t106 | X=t107 | X=t108 | X=t109 | X=t110 | X=t111 | X=t112 | X=t113 | X=t114 | X=t115 | X=t116 | X=t117 | X=t118 | X=t119 | X=t120 | X=t121 | X=t122 | X=t123 | X=t124 | X=t125 | X=t126 | X=t127 | X=t128 | X=t153 | X=t154 | X=t155 | X=t156 | X=t157 | X=t158 | X=t159 | X=t160 | X=t161 | X=t162 | X=t163 | X=t164 | X=t165 | X=t166 | X=t167 | X=t168 | X=t169 | X=t170 | X=t171 | X=t172 | X=t173 | X=t174 | X=t175 | X=t176 | X=t177 | X=t178 | X=t179 | X=t180 | X=t191 | X=t192 | X=t193 | X=t194 | X=t195 | X=t196 | X=t197 | X=t198 | X=t199 | X=t201 | X=t202 | X=t203 | X=t204 | X=t205 | X=t206 | X=t207 | X=t208 | X=t209 | X=t210 | X=t211 | X=t212 | X=t213 | X=t214 | X=t215 | X=t216 | X=t217 | X=t218 | X=t219 | X=t220 | X=t221 | X=t222 | X=t223 | X=t224 | X=t225 | X=t226 | X=t227 | X=t228 | X=t229 | X=t231 | X=t232 | X=t233 | X=t234 | X=t235 | X=t236 | X=t237 | X=t238 | X=t239 | X=t240 | X=t241 | X=t242 | X=t243 | X=t244 | X=t245 | X=t246 | X=t247 | X=t248 | X=t249 | X=t250 | X=t251 | X=t252 | X=t253 | X=t254 | X=t255 | X=t256 | X=t257 | X=t258 | X=t259 | X=t260 | X=t261 | X=t262 | X=t263 | X=t264 | X=t276 | X=t277 | X=t278 | X=t279 | X=t280 | X=t281 | X=t282 | X=t283 | X=t284 | X=t285 | X=t286 | X=t287 | X=t288 | X=t289 | X=t290 | X=t291 | X=t292 | X=t293 | X=t294 | X=t295 | X=t296 | X=t297 | X=t298 | X=t299 | X=t300 | X=t301 | X=t302 | X=t303 | X=t304 | X=t305 | X=t306 | X=t307 | X=t308 | X=t309 | X=t310 | X=t311 | X=t312 | X=t313 | X=t314 | X=t315 | X=t316 | X=t317 | X=t318 | X=t319 | X=t320 | X=t321 | X=t322 | X=t323 | X=t324 | X=t325 | X=t326 | X=t327 | X=t328 | X=t329 | X=t330 | X=t331 | X=t338 | X=t339 | X=t340 | X=t341 | X=t342 | X=t343 | X=t344 | X=t345 | X=t346 | X=t347 | X=t348 | X=t349 | X=t350 | X=t351)))
).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Conjecture
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(t_similar, conjecture,
completeness & purity & reciprocity & type_similarity
).
Appendix G. Melville: XML, HTML, and FOF of normalized transcript
XML representation of exemplar for normalized transcript
<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="../../genFofXSLT/genFof-Norm.xsl"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->
<!-- For HTML: -->
<!ENTITY and "&" >
<!ENTITY etc "&c" >
<!-- FOR Fof: -->
<!-- <!ENTITY and "amp" >
<!ENTITY etc "ampc" >-->
]>
<text>
<body>
<p>
<lb/>A seaman figures in The Canterbury Tales. <lb/>With
<del>a</del>many a tempest had his beard been <lb/>shook. –
<del>S</del>
<del>Deep g</del> Secret grief is a <lb/>cannibal of its own heart
– <emph>Bacon</emph>.
<lb/><metamark>&vbar;</metamark><emph>Claudia</emph> of the
<emph>Appian</emph> family, <q>I wish
<lb/><metamark>&vbar;</metamark> some fight or pestilence would
thin out <lb/><metamark>&vbar;</metamark> this crowd.</q>
<emph>Arrogance.</emph>
<lb/>
<emph>Roast beef in the pulpet.</emph>
</p>
<p>
<lb/>An animal of a man — <q>do eagles wear
<lb/>spectacles?</q>— Health. — <emph>Contrast</emph>:
an <lb/>over spiritual man. <lb/><metamark>——
</metamark></p>
<p>
<lb/>
<q>Yes, Madam, Cain was a godless froward boy, ∧ <lb/>Reuben
(Gen:49) ∧ Absalom</q> Many pious men <lb/>have impious
children — (Devil as a Quaker)
<lb/><metamark>—————</metamark>
</p>
<p>
<lb/>A formal compact – Imprimis – First –
Second. <lb/>The aforesaid soul<sic>. said soul</sic> &etc; –
Duplicates – <lb/><metamark>&triplebar;</metamark><q>How was
it about the temptation on the <lb/>hill?</q>&etc; – D begs
the hero to form <lb/>one of a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb/>&etc; – Leaves
a letter to the D – <q>My <lb/>
<emph>Dear D</emph>
</q>
<floatingText>
<body>
<ab> – Conversation upon Gabriel, Micheal ∧
<lb/> Raphel – gentlemanly &etc;</ab>
</body>
</floatingText>
</p>
<p>
<lb/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb/> making
almanacks – going to a ball takes a long <lb/>time making
toilette. – The Doctor's coach stops <lb/>the way.
– <q>Do you believe all that stuff? <lb/> nonsence –
the world was never made. – “But Is not <lb/>this you
mentioned <emph>here</emph> – in the scriptures?</q>
<lb/>Receives visits from the principal d's –
<q>Gentlemen</q> &etc;. <lb/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb/>be below with kings than above with fools?</q>
</p>
<p>
<lb/>It is better to laugh ∧ not sin than to <del>be</del> weep
∧ be <lb/>wicked. — Ten loads of coal to burn him.
— <lb/>Brought to the stake — warmed himself by the
fire. </p>
<p>
<lb/>Ego non baptizo te in nominee Patris et <lb/>Filii et Spiritus
Sancti – sed in nomine <lb/>Diaboli. — Madness is
undefinable — <lb/>It ∧ right reasons extremes of one.
<lb/>–Not the <add place="sup">(black art)</add> Goetic but
Theurgic magic — <lb/>seeks converse with the Intelligence,
Power, the <lb/>Angel. </p>
</body>
</text>
XML representation of normalized transcript
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->
<!-- For HTML: -->
<!ENTITY and "<reg>and</reg>" >
<!ENTITY etc "<reg>etc</reg>" >
<!-- FOR Fof: -->
<!-- <!ENTITY and "amp" >
<!ENTITY etc "ampc" >-->
]>
<text>
<body>
<p>
<lb/>A seaman figures in The Canterbury Tales. <lb/>With many a
tempest had his beard been <lb/>shook. – Secret grief is a
<lb/>cannibal of its own heart – <emph>Bacon</emph>. <lb/>
<emph>Claudia</emph> of the <emph>Appian</emph> family, <q>I wish
<lb/> some fight or pestilence would thin out <lb/> this
crowd.</q>
<emph>Arrogance.</emph>
<lb/>
<emph>Roast beef in the pulpet.</emph></p>
<p>
<lb/>An animal of a man — <q>do eagles wear
<lb/>spectacles?</q>— Health. — <emph>Contrast</emph>:
an <lb/>over spiritual man. <lb/>
</p>
<p>
<lb/>
<q>Yes, Madam, Cain was a godless froward boy, ∧ <lb/>Reuben
(Gen:49) ∧ Absalom</q> Many pious men <lb/>have impious
children — (Devil as a Quaker) <lb/>
</p>
<p>
<lb/>A formal compact – Imprimis – First –
Second. <lb/>The aforesaid soul &etc; – Duplicates –
<lb/><q>How was it about the temptation on the <lb/>hill?</q>
&etc; <supplied> Conversation upon Gabriel, Michael and <lb/>
Raphael – gentlemanly &etc;</supplied> – D begs the
hero to form <lb/>one of a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb/> &etc; –
Leaves a letter to the D – <q>My <lb/>
<emph>Dear D</emph>
</q> – </p>
<p>
<lb/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb/> making <choice>
<orig>almanacks</orig>
<reg>almanacs</reg>
</choice> – going to a ball takes a long <lb/>time making
toilette. – The Doctor's coach stops <lb/>the way.
– <q>Do you believe all that stuff? <lb/>
<choice>
<orig>nonsence</orig>
<reg>nonsense</reg>
</choice> – the world was never made. – “But <choice>
<orig>Is</orig>
<reg>is</reg>
</choice> not <lb/>this you mentioned <emph>here</emph> – in
the scriptures?</q>
<lb/>Receives visits from the principal d's –
<q>Gentlemen</q> &etc;. <lb/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb/>be below with kings than above with fools?</q>
</p>
<p>
<lb/>It is better to laugh ∧ not sin than to weep ∧ be
<lb/>wicked. — Ten loads of coal to burn him. —
<lb/>Brought to the stake — warmed himself by the fire. </p>
<p>
<lb/>Ego non baptizo te in <choice>
<orig>nominee</orig>
<reg>nomine</reg>
</choice> Patris et <lb/>Filii et Spiritus Sancti – sed in
nomine <lb/>Diaboli. — Madness is undefinable — <lb/>It
∧ right reasons extremes of one. <lb/>–Not the (black
art) Goetic but Theurgic magic — <lb/>seeks converse with the
Intelligence, Power, the <lb/>Angel. </p>
</body>
</text>
HTML presentation of exemplar for normalized transcript
Special exemplar tokens in red.
Figure 11: HTML presentation of exemplar for normalized transcript

HTML presentation of exemplar for normalized transcript
Special exemplar tokens in red, type similar tokens in blue.
Figure 12: HTML presentation of normalized transcript

FOF representation of normalized exemplar and transcript
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% General axioms
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(transcript_domain, axiom,
! [X,Y] : (transcript(X,Y) => (e_token(X) & t_token(Y)))).
fof(exemplar_domain, axiom,
! [X,Y] : (exemplar(X,Y) => (t_token(X) & e_token(Y)))).
%%%% Slightly modified from Default:
fof(typeof_domain, axiom,
! [X,Y] : (typeof(X,Y) =>
((t_token(X) | e_token(X) | st_token(X) | se_token(X))
& type(Y)))).
%%%% Slightly modified from Default:
fof(no_class_overlap, axiom,
! [X] : (
(e_token(X) <=> (~t_token(X) & ~type(X) & ~se_token(X) & ~st_token(X))) &
(t_token(X) <=> (~e_token(X) & ~type(X) & ~se_token(X) & ~st_token(X))) &
(type(X) <=> (~e_token(X) & ~t_token(X)) & ~se_token(X) & ~st_token(X)) &
(se_token(X) <=> (~t_token(X) & ~type(X) & ~e_token(X) & ~st_token(X))) &
(st_token(X) <=> (~t_token(X) & ~type(X) & ~se_token(X) & ~e_token(X)))
)).
fof(exemplar_and_transcript_inverse, axiom,
! [X,Y] : (exemplar(X,Y) <=> transcript(Y,X))).
fof(at_most_one_type_per_token, axiom,
! [X,Y,Z] : ((typeof(X,Y) & typeof(X,Z)) => Y=Z)).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Axioms not in Default
fof(typesimilar_domain, axiom,
! [X,Y] : (typesimilar(X,Y) => (type(X) & type(Y)))).
fof(typesimilar_reflexive, axiom,
! [X] : (type(X) => typesimilar(X,X))).
fof(typesimilar_symmetric, axiom,
! [X,Y] : (typesimilar(X,Y) => typesimilar(Y,X))).
fof(type_similar_transitive, axiom,
! [X,Y,Z] : ((typesimilar(X,Y) & typesimilar(Y,Z)) => typesimilar(X,Z))).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Definitions
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(reciprocity,definition,
reciprocity <=> ! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z))).
fof(completeness, definition,
completeness <=> ! [X] :
(e_token(X) => ? [Y] : transcript(X,Y))).
fof(purity, definition,
purity <=> ! [X] :
(t_token(X) => ? [Y] : exemplar(X,Y))).
%%%% Slightly modified from Default:
fof(type_similarity, definition,
type_similarity <=> ![X,Y] : (transcript(X,Y) =>
? [Z,V] : (typeof(X,Z) & typeof(Y,V) & typesimilar(Z,V)))).
fof(t_similarity, definition,
t_similarity <=> (reciprocity & completeness & purity & type_similarity)).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Case-specific facts
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(case_specific_facts, axiom,
typeof(e1, t_A) &
typeof(e2, t_seaman) &
typeof(e3, t_figures) &
typeof(e4, t_in) &
typeof(e5, t_The) &
typeof(e6, t_Canterbury) &
typeof(e7, t_Tales) &
typeof(e8, t_With) &
typeof(e9, t_a) &
typeof(e10, t_many) &
typeof(e11, t_a) &
typeof(e12, t_tempest) &
typeof(e13, t_had) &
typeof(e14, t_his) &
typeof(e15, t_beard) &
typeof(e16, t_been) &
typeof(e17, t_shook) &
typeof(e18, t_S) &
typeof(e19, t_Deep) &
%%% ...Approximately 250-350 lines omitted...
typeof(e278, t_extremes) &
typeof(e279, t_of) &
typeof(e280, t_one) &
typeof(e281, t_Not) &
typeof(e282, t_the) &
typeof(e283, t_black) &
typeof(e284, t_art) &
typeof(e285, t_Goetic) &
typeof(e286, t_but) &
typeof(e287, t_Theurgic) &
typeof(e288, t_magic) &
typeof(e289, t_seeks) &
typeof(e290, t_converse) &
typeof(e291, t_with) &
typeof(e292, t_the) &
typeof(e293, t_Intelligence) &
typeof(e294, t_Power) &
typeof(e295, t_the) &
typeof(e296, t_Angel) &
typeof(t1, t_A) &
typeof(t2, t_seaman) &
typeof(t3, t_figures) &
typeof(t4, t_in) &
typeof(t5, t_The) &
typeof(t6, t_Canterbury) &
typeof(t7, t_Tales) &
typeof(t8, t_With) &
typeof(t9, t_many) &
typeof(t10, t_a) &
typeof(t11, t_tempest) &
typeof(t12, t_had) &
typeof(t13, t_his) &
typeof(t14, t_beard) &
typeof(t15, t_been) &
typeof(t16, t_shook) &
typeof(t17, t_Secret) &
typeof(t18, t_grief) &
typeof(t19, t_is) &
typeof(t20, t_a) &
%%% ...Approximately 250-350 lines omitted...
typeof(t270, t_reasons) &
typeof(t271, t_extremes) &
typeof(t272, t_of) &
typeof(t273, t_one) &
typeof(t274, t_Not) &
typeof(t275, t_the) &
typeof(t276, t_black) &
typeof(t277, t_art) &
typeof(t278, t_Goetic) &
typeof(t279, t_but) &
typeof(t280, t_Theurgic) &
typeof(t281, t_magic) &
typeof(t282, t_seeks) &
typeof(t283, t_converse) &
typeof(t284, t_with) &
typeof(t285, t_the) &
typeof(t286, t_Intelligence) &
typeof(t287, t_Power) &
typeof(t288, t_the) &
typeof(t289, t_Angel) &
typesimilar(t_almanacks, t_almanacs) &
typesimilar(t_nonsence, t_nonsense) &
typesimilar(t_Is, t_is) &
typesimilar(t_nominee, t_nomine) &
transcript(e1,t1) &
transcript(e2,t2) &
transcript(e3,t3) &
transcript(e4,t4) &
transcript(e5,t5) &
transcript(e6,t6) &
transcript(e7,t7) &
transcript(e8,t8) &
transcript(e10,t9) &
transcript(e11,t10) &
transcript(e12,t11) &
transcript(e13,t12) &
transcript(e14,t13) &
transcript(e15,t14) &
transcript(e16,t15) &
transcript(e17,t16) &
transcript(e21,t17) &
transcript(e22,t18) &
transcript(e23,t19) &
transcript(e24,t20) &
%%% ...Approximately 250-350 lines omitted...
transcript(e276,t269) &
transcript(e277,t270) &
transcript(e278,t271) &
transcript(e279,t272) &
transcript(e280,t273) &
transcript(e281,t274) &
transcript(e282,t275) &
transcript(e283,t276) &
transcript(e284,t277) &
transcript(e285,t278) &
transcript(e286,t279) &
transcript(e287,t280) &
transcript(e288,t281) &
transcript(e289,t282) &
transcript(e290,t283) &
transcript(e291,t284) &
transcript(e292,t285) &
transcript(e293,t286) &
transcript(e294,t287) &
transcript(e295,t288) &
transcript(e296,t289) &
! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z)
) &
$distinct(
e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13,e14,e15,e16,e17,e18,e19,e20,e21,e22,e23,e24,e25,e26,e27,e28,e29,e30,e31,e32,e33,e34,e35,e36,e37,e38,e39,e40,e41,e42,e43,e44,e45,e46,e47,e48,e49,e50,e51,e52,e53,e54,e55,e56,e57,e58,e59,e60,e61,e62,e63,e64,e65,e66,e67,e68,e69,e70,e71,e72,e73,e74,e75,e76,e77,e78,e79,e80,e81,e82,e83,e84,e85,e86,e87,e88,e89,e90,e91,e92,e93,e94,e95,e96,e97,e98,e99,e100,e101,e102,e103,e104,e105,e106,e107,e108,e109,e110,e111,e112,e113,e114,e115,e116,e117,e118,e119,e120,e121,e122,e123,e124,e125,e126,e127,e128,e129,e130,e131,e132,e133,e134,e135,e136,e137,e138,e139,e140,e141,e142,e143,e144,e145,e146,e147,e148,e149,e150,e151,e152,e153,e154,e155,e156,e157,e158,e159,e160,e161,e162,e163,e164,e165,e166,e167,e168,e169,e170,e171,e172,e173,e174,e175,e176,e177,e178,e179,e180,e181,e182,e183,e184,e185,e186,e187,e188,e189,e190,e191,e192,e193,e194,e195,e196,e197,e198,e199,e200,e201,e202,e203,e204,e205,e206,e207,e208,e209,e210,e211,e212,e213,e214,e215,e216,e217,e218,e219,e220,e221,e222,e223,e224,e225,e226,e227,e228,e229,e230,e231,e232,e233,e234,e235,e236,e237,e238,e239,e240,e241,e242,e243,e244,e245,e246,e247,e248,e249,e250,e251,e252,e253,e254,e255,e256,e257,e258,e259,e260,e261,e262,e263,e264,e265,e266,e267,e268,e269,e270,e271,e272,e273,e274,e275,e276,e277,e278,e279,e280,e281,e282,e283,e284,e285,e286,e287,e288,e289,e290,e291,e292,e293,e294,e295,e296,
t1,t2,t3,t4,t5,t6,t7,t8,t9,t10,t11,t12,t13,t14,t15,t16,t17,t18,t19,t20,t21,t22,t23,t24,t25,t26,t27,t28,t29,t30,t31,t32,t33,t34,t35,t36,t37,t38,t39,t40,t41,t42,t43,t44,t45,t46,t47,t48,t49,t50,t51,t52,t53,t54,t55,t56,t57,t58,t59,t60,t61,t62,t63,t64,t65,t66,t67,t68,t69,t70,t71,t72,t73,t74,t75,t76,t77,t78,t79,t80,t81,t82,t83,t84,t85,t86,t87,t88,t89,t90,t91,t92,t93,t94,t95,t96,t97,t98,t99,t100,t101,t102,t103,t104,t105,t106,t107,t108,t109,t110,t111,t112,t113,t114,t115,t116,t117,t118,t119,t120,t121,t122,t123,t124,t125,t126,t127,t128,t129,t130,t131,t132,t133,t134,t135,t136,t137,t138,t139,t140,t141,t142,t143,t144,t145,t146,t147,t148,t149,t150,t151,t152,t153,t154,t155,t156,t157,t158,t159,t160,t161,t162,t163,t164,t165,t166,t167,t168,t169,t170,t171,t172,t173,t174,t175,t176,t177,t178,t179,t180,t181,t182,t183,t184,t185,t186,t187,t188,t189,t190,t191,t192,t193,t194,t195,t196,t197,t198,t199,t200,t201,t202,t203,t204,t205,t206,t207,t208,t209,t210,t211,t212,t213,t214,t215,t216,t217,t218,t219,t220,t221,t222,t223,t224,t225,t226,t227,t228,t229,t230,t231,t232,t233,t234,t235,t236,t237,t238,t239,t240,t241,t242,t243,t244,t245,t246,t247,t248,t249,t250,t251,t252,t253,t254,t255,t256,t257,t258,t259,t260,t261,t262,t263,t264,t265,t266,t267,t268,t269,t270,t271,t272,t273,t274,t275,t276,t277,t278,t279,t280,t281,t282,t283,t284,t285,t286,t287,t288,t289,
t_A,t_seaman,t_figures,t_in,t_The,t_Canterbury,t_Tales,t_With,t_a,t_many,t_tempest,t_had,t_his,t_beard,t_been,t_shook,t_S,t_Deep,t_g,t_Secret,t_grief,t_is,t_cannibal,t_of,t_its,t_own,t_heart,t_Bacon,t_Claudia,t_the,t_Appian,t_family,t_I,t_wish,t_some,t_fight,t_or,t_pestilence,t_would,t_thin,t_out,t_this,t_crowd,t_Arrogance,t_Roast,t_beef,t_pulpet,t_An,t_animal,t_man,t_do,t_eagles,t_wear,t_spectacles,t_Health,t_Contrast,t_an,t_over,t_spiritual,t_Yes,t_Madam,t_Cain,t_was,t_godless,t_froward,t_boy,t_amp,t_Reuben,t_Gen,t_49,t_Absalom,t_Many,t_pious,t_men,t_have,t_impious,t_children,t_Devil,t_as,t_Quaker,t_formal,t_compact,t_Imprimis,t_First,t_Second,t_aforesaid,t_soul,t_said,t_ampc,t_Duplicates,t_How,t_it,t_about,t_temptation,t_on,t_hill,t_D,t_begs,t_hero,t_to,t_form,t_one,t_Society,t_s,t_name,t_be,t_weighty,t_Leaves,t_letter,t_My,t_Dear,t_Conversation,t_upon,t_Gabriel,t_Micheal,t_Raphel,t_gentlemanly,t_Terra,t_Oblivionis,t_Hellites,t_At,t_Astor,t_find,t_him,t_making,t_almanacks,t_going,t_ball,t_takes,t_long,t_time,t_toilette,t_Doctor,t_coach,t_stops,t_way,t_Do,t_you,t_believe,t_all,t_that,t_stuff,t_nonsence,t_world,t_never,t_made,t_But,t_Is,t_not,t_mentioned,t_here,t_scriptures,t_Receives,t_visits,t_from,t_principal,t_d,t_Gentlemen,t_Arguments,t_persuade,t_Would,t_rather,t_below,t_with,t_kings,t_than,t_above,t_fools,t_It,t_better,t_laugh,t_sin,t_weep,t_wicked,t_Ten,t_loads,t_coal,t_burn,t_Brought,t_stake,t_warmed,t_himself,t_by,t_fire,t_Ego,t_non,t_baptizo,t_te,t_nominee,t_Patris,t_et,t_Filii,t_Spiritus,t_Sancti,t_sed,t_nomine,t_Diaboli,t_Madness,t_undefinable,t_right,t_reasons,t_extremes,t_Not,t_black,t_art,t_Goetic,t_but,t_Theurgic,t_magic,t_seeks,t_converse,t_Intelligence,t_Power,t_Angel,t_Michael,t_and,t_Raphael,t_almanacs,t_nonsense
) &
! [X] : ((e_token(X) <=> (X=e1 | X=e2 | X=e3 | X=e4 | X=e5 | X=e6 | X=e7 | X=e8 | X=e10 | X=e11 | X=e12 | X=e13 | X=e14 | X=e15 | X=e16 | X=e17 | X=e21 | X=e22 | X=e23 | X=e24 | X=e25 | X=e26 | X=e27 | X=e28 | X=e29 | X=e30 | X=e31 | X=e32 | X=e33 | X=e34 | X=e35 | X=e36 | X=e37 | X=e38 | X=e39 | X=e40 | X=e41 | X=e42 | X=e43 | X=e44 | X=e45 | X=e46 | X=e47 | X=e48 | X=e49 | X=e50 | X=e51 | X=e52 | X=e53 | X=e54 | X=e55 | X=e56 | X=e57 | X=e58 | X=e59 | X=e60 | X=e61 | X=e62 | X=e63 | X=e64 | X=e65 | X=e66 | X=e67 | X=e68 | X=e69 | X=e70 | X=e71 | X=e72 | X=e73 | X=e74 | X=e75 | X=e76 | X=e77 | X=e78 | X=e79 | X=e80 | X=e81 | X=e82 | X=e83 | X=e84 | X=e85 | X=e86 | X=e87 | X=e88 | X=e89 | X=e90 | X=e91 | X=e92 | X=e93 | X=e94 | X=e95 | X=e96 | X=e97 | X=e98 | X=e99 | X=e100 | X=e103 | X=e104 | X=e105 | X=e106 | X=e107 | X=e108 | X=e109 | X=e110 | X=e111 | X=e112 | X=e113 | X=e114 | X=e115 | X=e116 | X=e117 | X=e118 | X=e119 | X=e120 | X=e121 | X=e122 | X=e123 | X=e124 | X=e125 | X=e126 | X=e127 | X=e128 | X=e129 | X=e130 | X=e131 | X=e132 | X=e133 | X=e134 | X=e135 | X=e136 | X=e137 | X=e138 | X=e139 | X=e140 | X=e141 | X=e142 | X=e151 | X=e152 | X=e153 | X=e154 | X=e155 | X=e156 | X=e157 | X=e158 | X=e159 | X=e160 | X=e161 | X=e162 | X=e163 | X=e164 | X=e165 | X=e166 | X=e167 | X=e168 | X=e169 | X=e170 | X=e171 | X=e172 | X=e173 | X=e174 | X=e175 | X=e176 | X=e177 | X=e178 | X=e179 | X=e180 | X=e181 | X=e182 | X=e183 | X=e184 | X=e185 | X=e186 | X=e187 | X=e188 | X=e189 | X=e190 | X=e191 | X=e192 | X=e193 | X=e194 | X=e195 | X=e196 | X=e197 | X=e198 | X=e199 | X=e200 | X=e201 | X=e202 | X=e203 | X=e204 | X=e205 | X=e206 | X=e207 | X=e208 | X=e209 | X=e210 | X=e211 | X=e212 | X=e213 | X=e214 | X=e215 | X=e216 | X=e217 | X=e218 | X=e219 | X=e220 | X=e221 | X=e222 | X=e223 | X=e224 | X=e225 | X=e226 | X=e227 | X=e228 | X=e229 | X=e230 | X=e231 | X=e232 | X=e233 | X=e235 | X=e236 | X=e237 | X=e238 | X=e239 | X=e240 | X=e241 | X=e242 | X=e243 | X=e244 | X=e245 | X=e246 | X=e247 | X=e248 | X=e249 | X=e250 | X=e251 | X=e252 | X=e253 | X=e254 | X=e255 | X=e256 | X=e257 | X=e258 | X=e259 | X=e260 | X=e261 | X=e262 | X=e263 | X=e264 | X=e265 | X=e266 | X=e267 | X=e268 | X=e269 | X=e270 | X=e271 | X=e272 | X=e273 | X=e274 | X=e275 | X=e276 | X=e277 | X=e278 | X=e279 | X=e280 | X=e281 | X=e282 | X=e283 | X=e284 | X=e285 | X=e286 | X=e287 | X=e288 | X=e289 | X=e290 | X=e291 | X=e292 | X=e293 | X=e294 | X=e295 | X=e296)) &
(t_token(X) <=> (X=t1 | X=t2 | X=t3 | X=t4 | X=t5 | X=t6 | X=t7 | X=t8 | X=t9 | X=t10 | X=t11 | X=t12 | X=t13 | X=t14 | X=t15 | X=t16 | X=t17 | X=t18 | X=t19 | X=t20 | X=t21 | X=t22 | X=t23 | X=t24 | X=t25 | X=t26 | X=t27 | X=t28 | X=t29 | X=t30 | X=t31 | X=t32 | X=t33 | X=t34 | X=t35 | X=t36 | X=t37 | X=t38 | X=t39 | X=t40 | X=t41 | X=t42 | X=t43 | X=t44 | X=t45 | X=t46 | X=t47 | X=t48 | X=t49 | X=t50 | X=t51 | X=t52 | X=t53 | X=t54 | X=t55 | X=t56 | X=t57 | X=t58 | X=t59 | X=t60 | X=t61 | X=t62 | X=t63 | X=t64 | X=t65 | X=t66 | X=t67 | X=t68 | X=t69 | X=t70 | X=t71 | X=t72 | X=t73 | X=t74 | X=t75 | X=t76 | X=t77 | X=t78 | X=t79 | X=t80 | X=t81 | X=t82 | X=t83 | X=t84 | X=t85 | X=t86 | X=t87 | X=t88 | X=t89 | X=t90 | X=t91 | X=t92 | X=t93 | X=t94 | X=t95 | X=t96 | X=t97 | X=t98 | X=t99 | X=t100 | X=t101 | X=t102 | X=t103 | X=t104 | X=t105 | X=t106 | X=t107 | X=t108 | X=t117 | X=t118 | X=t119 | X=t120 | X=t121 | X=t122 | X=t123 | X=t124 | X=t125 | X=t126 | X=t127 | X=t128 | X=t129 | X=t130 | X=t131 | X=t132 | X=t133 | X=t134 | X=t135 | X=t136 | X=t137 | X=t138 | X=t139 | X=t140 | X=t141 | X=t142 | X=t143 | X=t144 | X=t145 | X=t146 | X=t147 | X=t148 | X=t149 | X=t150 | X=t151 | X=t152 | X=t153 | X=t154 | X=t155 | X=t156 | X=t157 | X=t158 | X=t159 | X=t160 | X=t161 | X=t162 | X=t163 | X=t164 | X=t165 | X=t166 | X=t167 | X=t168 | X=t169 | X=t170 | X=t171 | X=t172 | X=t173 | X=t174 | X=t175 | X=t176 | X=t177 | X=t178 | X=t179 | X=t180 | X=t181 | X=t182 | X=t183 | X=t184 | X=t185 | X=t186 | X=t187 | X=t188 | X=t189 | X=t190 | X=t191 | X=t192 | X=t193 | X=t194 | X=t195 | X=t196 | X=t197 | X=t198 | X=t199 | X=t200 | X=t201 | X=t202 | X=t203 | X=t204 | X=t205 | X=t206 | X=t207 | X=t208 | X=t209 | X=t210 | X=t211 | X=t212 | X=t213 | X=t214 | X=t215 | X=t216 | X=t217 | X=t218 | X=t219 | X=t220 | X=t221 | X=t222 | X=t223 | X=t224 | X=t225 | X=t226 | X=t227 | X=t228 | X=t229 | X=t230 | X=t231 | X=t232 | X=t233 | X=t234 | X=t235 | X=t236 | X=t237 | X=t238 | X=t239 | X=t240 | X=t241 | X=t242 | X=t243 | X=t244 | X=t245 | X=t246 | X=t247 | X=t248 | X=t249 | X=t250 | X=t251 | X=t252 | X=t253 | X=t254 | X=t255 | X=t256 | X=t257 | X=t258 | X=t259 | X=t260 | X=t261 | X=t262 | X=t263 | X=t264 | X=t265 | X=t266 | X=t267 | X=t268 | X=t269 | X=t270 | X=t271 | X=t272 | X=t273 | X=t274 | X=t275 | X=t276 | X=t277 | X=t278 | X=t279 | X=t280 | X=t281 | X=t282 | X=t283 | X=t284 | X=t285 | X=t286 | X=t287 | X=t288 | X=t289)))
).
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Conjecture
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
fof(t_similar, conjecture,
completeness & purity & reciprocity & type_similarity
).
References
This list of references has not been updated since 2018. References added later have been provided in footnotes and in web links.
[André 1972] [André, Jacques.] Règles et recommandation pour les éditions critiques (Série latine). Paris: Société d'édition "Les belles lettres," 1972. Collection des universités de France, publiée sous le patronage de l'Association Guillaume Budé. [vi +] 48 pp.
[Carter 1952] Carter, Clarence E. Historical editing. Bulletins of the national archives, Number 7 [Washington, DC]: National Archives and Records Service, August 1952. National Archives publication number 53-4.
[Caton 2013] Caton, Paul.
Pure transcriptional encoding.
Paper given at
Digital Humanities 2013, Lincoln, Nebraska.
[The Papers of William Penn] Dunn, Mary Maples et al. The Papers of William Penn, 5 vols. (Philadelphia: University of Pennsylvania Press, 1981-1986).
[Grice 1975] Grice, H.P.
Logic and Conversation.
In Syntax and
Semantics, edited by P. Cole and J. Morgan, vol.3,
Speech Acts. New York: Academic Press, 1975.
Reprinted as chapter 2 of his Studies in the Way of
Words. Cambridge, Mass.: Harvard University Press,
1989, pp. 22–40.
[Goodman] Goodman, Nelson. Languages of Art. Hackett Publishing, 1976.
[Hayford et al. 1988] Hayford,
Harrison, Hershel Parker, and G. Thomas Tanselle, ed. Moby
Dick, or, The Whale.
Vol. 7 of The Writings of
Herman Melville The Northwestern–Newberry
Edition. Evanston [Ill.]: Northwestern University Press; Chicago :
Newberry Library, 1988, rpt. 1994, 1997.
[Huitfeldt / Sperberg-McQueen 2008] Huitfeldt, Claus,
and C. M. Sperberg-McQueen. What is transcription?
Literary & Linguistic Computing 23.3
(2008): 295-310. [doi:https://doi.org/10.1093/llc/fqn013].
[Huitfeldt / Marcoux / Sperberg-McQueen 2010] Huitfeldt,
Claus, Yves Marcoux, and C. M. Sperberg-McQueen. Extension
of the type/token distinction to document structure.
Paper
presented at Balisage: The Markup Conference 2010, Montréal,
Canada, August 3 - 6, 2010. In Proceedings of Balisage:
The Markup Conference 2010. Balisage Series on Markup
Technologies, vol. 5 (2010). [doi:https://doi.org/10.4242/BalisageVol5.Huitfeldt01]. On the Web at
[http://www.balisage.net/Proceedings/vol5/html/Huitfeldt01/BalisageVol5-Huitfeldt01.html].
[Marcoux 2006] Marcoux,
Yves. A natural-language approach to modeling: Why is some
XML so difficult to write?
Paper given at Extreme Markup
Languages®, Montréal, 2006. Proceedings of Extreme
Markup Languages® 2006. On the Web at [http://conferences.idealliance.org/extreme/html/2006/Marcoux01/EML2006Marcoux01.html].
[Marcoux / Rizkallah 2007]
Marcoux, Yves, and Élias Rizkallah. Exploring
intertextual semantics: A reflection on attributes and
optionality.
Paper given at Extreme Markup Languages®,
Montréal, 2007. Proceedings of Extreme Markup
Languages® 2007. On the Web at [http://conferences.idealliance.org/extreme/html/2007/Marcoux01/EML2007Marcoux01.html].
[Peano 1889] Peano, Ioseph.
Arithmetics principia nova methodo exposita.
Romae,
Florentiae: Bocca, 1889. English translation: Michael Nahas, [https://raw.githubusercontent.com/mdnahas/Peano_Book/master/Peano.pdf].
[Robinson / Solopova 2006]
Robinson, Peter, and Elizabeth Solopova. Guidelines for
Transcription of the Manuscripts of The Wife of Bath!s
Prologue.
18 March 2006. On the Web at [http://www.canterburytalesproject.org/pubs/transguide-MI.pdf].
[Simons et al. 2005] Simons,
Gary, Scott O. Farrar, Brian Fitzsimons, William D. Lewis, D.
Terence Langendoen, and Hector Gonzalez. The semantics of
markup: Mapping legacy markup schemas to a common
semantics.
In Proceedings of the 4th workshop
on NLP and XML (NLPXML-2004): held in cooperation with ACL- 04
Barcelona, Spain. Pp. 25-32. On the Web at [http://www.aclweb.org/anthology/W/W04/W04-0604.pdf] and
other locations. [doi:https://doi.org/10.3115/1621066.1621070].
[Sperberg-McQueen 2005]
Sperberg-McQueen, C. M. The meaning of OAI 2.0 Markup: An
exercise in markup interpretation.
Unpublished fragment,
December 2005. On the Web at [http://www.w3.org/2004/04/em-msm/ioai.html].
[Sperberg-McQueen et al. 2002]
Sperberg-McQueen, C. M., David Dubin, Claus Huitfeldt, and Allen
Renear. Drawing inferences on the basis of markup.
Paper given at Extreme Markup Languages®, Montréal, 2002.
>Proceedings of Extreme Markup Languages®
2002. On the Web at [http://conferences.idealliance.org/extreme/html/2002/CMSMcQ01/EML2002CMSMcQ01.html].
[Sperberg-McQueen / Huitfeldt / Marcoux 2009]
Sperberg-McQueen, C. M.. Claus Huitfeldt, and Yves Marcoux.
What is transcription? Part 2.
Talk given at
Digital Humanities 2009, College Park, Maryland. Slides on the Web
at [http://blackmesatech.com/2009/06/dh2009/]. Summary at
[http://www.mith2.umd.edu/dh09/wp-content/uploads/dh09_conferencepreceedings_final.pdf (pp. 257-260)].
[Sperberg-McQueen / Huitfeldt / Renear 2001]
Sperberg-McQueen, C. M., Claus Huitfeldt, and Allen Renear.
Meaning and interpretation of markup.
Markup Languages: Theory & Practice 2.3
(2001): 215–234. On the Web at [http://cmsmcq.com/2000/mim.html]. [doi:https://doi.org/10.1162/109966200750363599].
[Stevens and Burg 1997] Stevens, Michael E. and Steven B. Burg. Editing Historical Documents: A Handbook of Practice. Walnut Creek, London, New Delhi: Alta Mira Press, 1997.
[Tanselle 1989] Tanselle, G. Thomas. A Rationale of Textual Criticism. Philadelphia: University of Pennsylvania Press, 1989. 104 pp.
[Vander Meulen / Tanselle 1999] Vander Meulen, David,
and G. Thomas Tanselle, A system of manuscript
transcription.
Studies in Bibliography 52 (1999): 201-212.
[Wickett / Renear 2009]
Wickett, Karen M., and Allen Renear. A first order theory of
bibliographic objects.
Proceedings of the American Society for Information
Science and Technology 46.1 (2009): 1–8. On the
Web at [http://onlinelibrary.wiley.com/doi/10.1002/meet.2009.1450460378/full]
(subscription required). [doi:https://doi.org/10.1002/meet.2009.1450460378].
[Boolos et al. 2007] George Boolos, John P. Burgess, Richard C. Jeffrey. Computability and logic. 5th ed. Cambridge, New York: Cambridge University Press 2007. ISBN 9780521877527.
[1] See [2018 version of this paper].
[2] Except for abstracts and slides from various conferences, cf. below.
[3] We use the English term exemplar to denote the document from or of which a transcript is made, in preference to the term original, which evokes confusing associations when the exemplar is itself a transcript. In our usage, exemplar thus corresponds to what is referred to in German as the Vorlage or in French as the antigraphe.
[4] See, for example, [Carter 1952], [Tanselle 1989], [Stevens and Burg 1997], [Vander Meulen / Tanselle 1999], [Robinson / Solopova 2006].
[5] Our notion of transcriptional implicature is indeed inspired by Grice's theory, but we make no claim as to the similarities between our notion and Grice's theory of conversational and conventional implicature.
[6] In other contexts, such as linguistics, music, or genetics,
transcription
refers to different, though
related phenomena — see [Huitfeldt / Sperberg-McQueen 2008], pp
295-96.
[7] See also [Huitfeldt / Sperberg-McQueen 2008], p. 296.
[8] Many documents do of course contain non-textual material. It may be argued that at least some kinds of non-textual material lend themselves to a type-token analysis, but we do not believe that always to be the case. Our account is limited to the aspects of documents which do lend themselves to such analysis, and we do not argue that all aspects of documents do.
[9] The presupposition that exemplar as well as transcript lend themselves to analysis in terms of the concepts of tokens and types presented above implies that they must be notations as Goodman defines that term.
According to Goodman, the requirements of notational schemes are:
-
...[A] character in a notation is an abstraction class of character-indifference among inscriptions. As a result, no mark may belong to more than one character.
[Goodman] pp. 132-3. In our terminology: No token is an instantiation of more than one type. -
...the characters be finitely differentiated, or articulate.
[Goodman] p. 135. That is, for any mark it must be at least theoretically possible to determine whether or not it belongs to a certain sign (character). In our terminology: For any token, it must be at least theoretically possible to determine whether it instantiates a certain type or not.
These are what Goodman calls syntactic requirements that notational schemes have to satisfy. In order for a notational scheme to qualify as a notational system, it must also satisfy three semantic requirements: It must be unambiguous, disjoint and finitely differentiated.
More precisely:
-
Every notational system must be unambiguous. In other words, the extension (compliance-class) of a sign (character or inscription) must be the same for all occurrences of that sign – it cannot vary from case to case. [Goodman] p. 148.
-
Every extension (compliance-class) must be distinct (disjoint) from every other extension. In other words, no object in a domain can belong to two extensions. [Goodman] p. 150.
-
Every extension (compliance-class) must be finitely differentiated. In other words, for any object in the domain it must be at least theoretically possible to determine whether or not it belongs to a certain extension. [Goodman] p. 152.
It would be interesting to investigate to what extent our account of transcription implies that transcription is notational not only in terms of the syntactic, but also in terms of the semantic requirements. We think this may be the case, but do not further investigate the issue as it is not of direct relevance to the tasks we have set ourselves here.
[10] The reader will note that the notion of "mark" is very vague; in view of the many different ways in which written messages can be constructed or conveyed, we believe this to be unavoidable. The notion is not defined formally in the model.
[11] When the distinction is not clear from context, we use single quotes for types and double quotes for tokens.
[12] The notions of disjunctive and conjunctive types were introduced in [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
[13] The reader may feel free to read the T
in
T-similarity
to stand for text
,
type
, or transcriptional
.
[14] The reader may wonder how rules can be represented by a
mathematical relation. The name rules
suggests
imperative statements such as when a sentence is not
closed with a period, insert an extra space
or
whenever appropriate, render "u" as "v," or "v" as
"u"
, more than it suggests a relationship. The answer
is that T-similarity represents
transcription rules in a descriptive manner. It specifies how exemplar and
transcript are to stand relative to each other once the
transcription is completed, the implicit rules being: do
whatever it takes for that relation to hold when you are
done
.
Interestingly, in the statement of
practice accompanying some transcripts, the rules of
transcription are also formulated in descriptive
rather than imperative style: e.g., spelling is retained
as written
or the long "s" is presented as a
short "s"
, rather than retain spelling as
written
or convert any long "s" to a short
"s"
.
[15] For a systematic way of defining such a correspondence, see [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
[16] For simplicity, we consider here only cases with a single type repertoire.
[17] One might regard this as a principle of giving the author the benefit of the doubt. Or perhaps one might just as well regard it as an act of charity to the reader: Unless there is clear evidence of error, there is no reason to bother the reader with the mere possibility of error.
[19] Here and in the rest of this section, indented quotations are quotations from pp. 15-18 in [The Papers of William Penn].
[20] The quotation continues: , with the following two
exceptions. Endorsements are treated as dockets, and entered
into the provenance note (see below). If a document is
undated, an initial date line is supplied [within square
brackets]. If a document is dated at the close but not at the
opening, an initial date line is supplied [within square
brackets], and the closing date line is
retained.
[21] Unfortunately we have only had access in facsimile to the first and the last page of one letter, which are reproduced on the insides of the front and back cover. These two pages, and the corresponding transcript, are reproduced in appendix A.
[22] See, for example, the poem on pp 32-3.
[23] Except for such line breaks that occur in connection with the phenomena mentioned in the previous paragraph.
[24] It is a common typographic practice to use "sic" to mark misspellings and other irregularities which might otherwise be blamed on the typesetter, and in some contexts punctuation and paragraphing are routinely normalized, or introduced, by editors.
[25] We note in passing that the statement is silent on what happens in cases (if any) where the exemplar already contains braces, or an extra space after a sentence not closed with a period. If there are such cases, then their printed representation is ambiguous: it could represent the discreet editorial intervention described in the statement of practice, or it could be a strictly literal transcript of the paragraph. It seems likely that the editors know that there are no such cases, and trust the reader to infer it. But it might also be the case that the editors simply don't regard such ambiguities as interesting or important enough to be worth avoiding.
[26] However, this may also be regarded as an application of the rule of type similarity, cf. our discussion of statement 5.l below.
[27] If, on the other hand, statement 5.b were taken to mean that the writers of the manuscripts regarded the use of capitalization as a matter for individual judgment rather than orthographic system, resulting in usage that the modern reader perceives as arbitrary and inconsistent, then the matter would be similar to the case with u and v, discussed below. (Readers today do find the capitalization of seventeenth-century manuscripts capricious, but we do not believe that is the point being made in 5.b.)
[28] We take this to mean that "£" is represented as "l" (not as "l.") — i.e., that the fact that the full stop is placed before the closing quotation mark is a somewhat misleading effect of punctuation rules.
[29] A glyph shaped like A
is either uppercase
Latin letter A or uppercase Greek letter Alpha, depending on
context. A dot on the baseline may be a full stop (sentence
punctuation), an abbreviation marker, or a decimal point
depending on both the intratextual and the extratextual context.
In most European languages, ij
,
ch
, and ll
are each a
two-character string, but the first is a single character (a
single token instantiating a single type) in Dutch, and each of
the others is taken as a single character in traditional Spanish
orthography, even though i
, j
,
c
, h
, and l
also
occur individually in those languages.
[30] Once again, an alternative analysis might be that the different type systems are employed for the transcript and the exemplar. (In this case, the subset relationship would be the opposite of the one for "u" and "v".) And once again, rules of default transcriptional implicature would be broken in neither case.
[31] Numbering in square brackets added by us.
[32] The transcript appears to be complete with respect to Melville's notes, but omits the printed text of Shakespeare which appears in the same volume.
[33] An alternative interpretation would infer a type system in which for any letter such as e, there are companion types for inserted e, canceled e, overwritten e, and underlined [or italic] e. Occam's Razor and convenience in the formalization both lead us to prefer an account of this example in which insertions, etc., are compound tokens of corresponding compound types.
[34] The occurrence of a single x is in
fact perhaps strictly speaking ambiguous: statement [5] suggests
the interpretation one undeciphered letter
, but
the only occurrence of a single x in the
transcript (in the word extremes
on line 37)
seems more plausibly read as transcribing an
x in the exemplar.
[35] Based on [Melville's Marginalia Online]. See in particular [their transcript of 16530].
[36] In other systems, like Prolog or Alloy, they would not have been required, as these systems are based on a closed world assumption.
[37] This is a convenience feature of Vampire. We could have
done without it, but would then have had to include a statement
to which Vampire expands the $distinct statement, such as:
t_of ≠ t_Essex ∧ t_Earl ≠ t_Essex ∧ t_Earl
≠ t_of ∧ t_the ≠ t_Essex ∧ t_the ≠ t_of ∧
t_the ≠ t_Earl ∧ t_Essexe ≠ t_Essex ∧ t_Essexe
≠ t_of ∧ t_Essexe ≠ t_Earl ∧ t_Essexe ≠ t_the
∧ t_to ≠ t_Essex ∧ t_to ≠ t_of ∧ t_to ≠
t_Earl ∧ t_to ≠ t_the ∧ t_to ≠ t_Essexe ∧
t_with ≠ t_Essex ∧ t_with ≠ t_of ∧ t_with ≠
t_Earl ∧ t_with ≠ t_the ∧ t_with ≠ t_Essexe
∧ t_with ≠ t_to ∧ t_went ≠ t_Essex ∧ t_went
≠ t_of ∧ t_went ≠ t_Earl ∧ t_went ≠ t_the
∧ t_went ≠ t_Essexe ∧ t_went ≠ t_to ∧ t_went
≠ t_with ∧ t_Elizabeth ≠ t_Essex ∧ t_Elizabeth
≠ t_of ∧ t_Elizabeth ≠ t_Earl ∧ t_Elizabeth ≠
t_the ∧ t_Elizabeth ≠ t_Essexe ∧ t_Elizabeth ≠
t_to ∧ t_Elizabeth ≠ t_with ∧ t_Elizabeth ≠
t_went ∧ t7 ≠ t_Essex ∧ t_of ≠ t7 ∧ t_Earl
≠ t7 ∧ t_the ≠ t7 ∧ t_Essexe ≠ t7 ∧ t_to
≠ t7 ∧ t_with ≠ t7 ∧ t_went ≠ t7 ∧
t_Elizabeth ≠ t7 ∧ t6 ≠ t_Essex ∧ t6 ≠ t_of
∧ t_Earl ≠ t6 ∧ t_the ≠ t6 ∧ t_Essexe ≠
t6 ∧ t_to ≠ t6 ∧ t_with ≠ t6 ∧ t_went ≠
t6 ∧ t_Elizabeth ≠ t6 ∧ t6 ≠ t7 ∧ t5 ≠
t_Essex ∧ t5 ≠ t_of ∧ t5 ≠ t_Earl ∧ t_the
≠ t5 ∧ t_Essexe ≠ t5 ∧ t_to ≠ t5 ∧ t_with
≠ t5 ∧ t_went ≠ t5 ∧ t_Elizabeth ≠ t5 ∧
t5 ≠ t7 ∧ t5 ≠ t6 ∧ t4 ≠ t_Essex ∧ t4
≠ t_of ∧ t4 ≠ t_Earl ∧ t4 ≠ t_the ∧
t_Essexe ≠ t4 ∧ t_to ≠ t4 ∧ t_with ≠ t4 ∧
t_went ≠ t4 ∧ t_Elizabeth ≠ t4 ∧ t4 ≠ t7
∧ t4 ≠ t6 ∧ t4 ≠ t5 ∧ t3 ≠ t_Essex ∧
t3 ≠ t_of ∧ t3 ≠ t_Earl ∧ t3 ≠ t_the ∧
t_Essexe ≠ t3 ∧ t_to ≠ t3 ∧ t_with ≠ t3 ∧
t_went ≠ t3 ∧ t_Elizabeth ≠ t3 ∧ t3 ≠ t7
∧ t3 ≠ t6 ∧ t3 ≠ t5 ∧ t3 ≠ t4 ∧ t2
≠ t_Essex ∧ t2 ≠ t_of ∧ t2 ≠ t_Earl ∧ t2
≠ t_the ∧ t_Essexe ≠ t2 ∧ t_to ≠ t2 ∧
t_with ≠ t2 ∧ t_went ≠ t2 ∧ t_Elizabeth ≠ t2
∧ t2 ≠ t7 ∧ t2 ≠ t6 ∧ t2 ≠ t5 ∧ t2
≠ t4 ∧ t2 ≠ t3 ∧ t1 ≠ t_Essex ∧ t1 ≠
t_of ∧ t1 ≠ t_Earl ∧ t1 ≠ t_the ∧ t_Essexe
≠ t1 ∧ t_to ≠ t1 ∧ t_with ≠ t1 ∧ t_went
≠ t1 ∧ t_Elizabeth ≠ t1 ∧ t1 ≠ t7 ∧ t1
≠ t6 ∧ t1 ≠ t5 ∧ t1 ≠ t4 ∧ t1 ≠ t3
∧ t1 ≠ t2 ∧ e5 ≠ t_Essex ∧ e5 ≠ t_of
∧ e5 ≠ t_Earl ∧ e5 ≠ t_the ∧ e5 ≠
t_Essexe ∧ t_to ≠ e5 ∧ t_with ≠ e5 ∧ t_went
≠ e5 ∧ t_Elizabeth ≠ e5 ∧ e5 ≠ t7 ∧ e5
≠ t6 ∧ e5 ≠ t5 ∧ e5 ≠ t4 ∧ e5 ≠ t3
∧ e5 ≠ t2 ∧ e5 ≠ t1 ∧ e4 ≠ t_Essex ∧
e4 ≠ t_of ∧ e4 ≠ t_Earl ∧ e4 ≠ t_the ∧ e4
≠ t_Essexe ∧ e4 ≠ t_to ∧ t_with ≠ e4 ∧
t_went ≠ e4 ∧ t_Elizabeth ≠ e4 ∧ e4 ≠ t7
∧ e4 ≠ t6 ∧ e4 ≠ t5 ∧ e4 ≠ t4 ∧ e4
≠ t3 ∧ e4 ≠ t2 ∧ e4 ≠ t1 ∧ e4 ≠ e5
∧ e3 ≠ t_Essex ∧ e3 ≠ t_of ∧ e3 ≠ t_Earl
∧ e3 ≠ t_the ∧ e3 ≠ t_Essexe ∧ e3 ≠ t_to
∧ e3 ≠ t_with ∧ t_went ≠ e3 ∧ t_Elizabeth
≠ e3 ∧ e3 ≠ t7 ∧ e3 ≠ t6 ∧ e3 ≠ t5
∧ e3 ≠ t4 ∧ e3 ≠ t3 ∧ e3 ≠ t2 ∧ e3
≠ t1 ∧ e3 ≠ e5 ∧ e3 ≠ e4 ∧ e2 ≠
t_Essex ∧ e2 ≠ t_of ∧ e2 ≠ t_Earl ∧ e2 ≠
t_the ∧ e2 ≠ t_Essexe ∧ e2 ≠ t_to ∧ e2 ≠
t_with ∧ e2 ≠ t_went ∧ t_Elizabeth ≠ e2 ∧ e2
≠ t7 ∧ e2 ≠ t6 ∧ e2 ≠ t5 ∧ e2 ≠ t4
∧ e2 ≠ t3 ∧ e2 ≠ t2 ∧ e2 ≠ t1 ∧ e2
≠ e5 ∧ e2 ≠ e4 ∧ e2 ≠ e3 ∧ e1 ≠
t_Essex ∧ e1 ≠ t_of ∧ e1 ≠ t_Earl ∧ e1 ≠
t_the ∧ e1 ≠ t_Essexe ∧ e1 ≠ t_to ∧ e1 ≠
t_with ∧ e1 ≠ t_went ∧ e1 ≠ t_Elizabeth ∧ e1
≠ t7 ∧ e1 ≠ t6 ∧ e1 ≠ t5 ∧ e1 ≠ t4
∧ e1 ≠ t3 ∧ e1 ≠ t2 ∧ e1 ≠ t1 ∧ e1
≠ e5 ∧ e1 ≠ e4 ∧ e1 ≠ e3 ∧ e1 ≠
e2
[38] Technically, this is obtained by submitting to Vampire a FOF-representation of the axioms 1-10 and the definitions a-d, with T-similarity as a conjecture. Vampire's response is that the conjecture is a theorem, i.e., that it is logically implied by the other statements.
[39] Experts on TEI and critical editing will probably find our encoding idiosyncratic and/or unsatisfactory. Moreover, any serious TEI-based project today would probably make one base transcription including all the information contained in both the transcription discussed here (in Appendix C and the normalized transcription in Appendix G. It should be mentioned, though, that Michael made a more thorough TEI representation of the transcript for a related project in 2017, using a special TEI-based tag set for manuscript transcription ([http://mlcd.blackmesatech.com/mlcd/2015/W/tip/Melville/Melville-notes.sourcedoc.xml]).
[40] Technically, Micheal
and
Raphel
are not declared type similar to
Michael
and Raphael
, but that is
because they occur only as parts of a supplied
element in the transcript.
[41] It may be plausible to omit unreadable writing from a transcript. However, in order to leave, for example, text in a different hand or a different language, an addition or deletion out from a transcript, the transcriber needs to recognize it as such, and may thus very well be able to read it.