Sperberg-McQueen, C. M., Claus Huitfeldt and Yves Marcoux. “Transcriptional implicature.” Presented at Balisage: The Markup Conference 2026, Washington, DC, August 3 - 7, 2026. In Proceedings of Balisage: The Markup Conference 2026. Balisage Series on Markup Technologies, vol. 31 (2026). https://doi.org/10.4242/BalisageVol31.Huitfeldt01.
Balisage: The Markup Conference 2026 August 3 - 7, 2026
Balisage Paper: Transcriptional implicature
C. M. Sperberg-McQueen
Founder and principal
Black Mesa Technologies LLC
C. M. Sperberg-McQueen† 2024 was
the founder and principal of Black Mesa Technologies, a
consultancy specializing in helping memory institutions improve
the long term preservation of and access to the information for
which they are responsible.
He served as editor in chief of the TEI Guidelines from 1988
to 2000, and also served as co-editor of the World Wide Web
Consortium's XML 1.0 and XML Schema 1.1 specifications.
Claus Huitfeldt works at the Department of Philosophy of the
University of Bergen, Norway. He was founding Director
(1990-2000) of the Wittgenstein Archives at the University of
Bergen, for which he developed the text encoding system MECS as
well as the editorial methods for the publication of
Wittgenstein's Nachlass - The Bergen Electronic Edition (Oxford
University Press, 2000).
Yves Marcoux
Honorary Professor (Professeur honoraire)
École de bibliothéconomie et des sciences de
l’information, Université de Montréal
Yves Marcoux has been a faculty member at EBSI, University
of Montréal, from 1991 to 2018, then adjunct professor and
lecturer until 2025. He was mainly involved in teaching,
research, standardization, and international cooperation
activities in the field of document informatics. Prior to his
appointment at EBSI, Dr. Marcoux worked for 10 years in systems
maintenance and development, in Canada, the U.S., and Europe. He
obtained his Ph.D. in theoretical computer science from
Université de Montréal in 1991. His main research interests are
intertextual semantics, the design of communication, markup
languages and digital humanities.
The concept of transcriptional implicature is proposed as a
way of accounting for common practices in transcription while also
making sense of variations in practice. A formal model implicit in
default transcriptional implicature can be elaborated in different
ways to reflect a number of commonalities as well as variations in
transcription practice which have hitherto eluded precise
description in formal terms. We regard the formalization of
transcription as an essential step towards formalizing the
semantics of colloquial XML vocabularies like TEI.
The publication of this paper is meant as
a tribute to the
late C. Michael Sperberg-McQueen by the coauthors. In their mind,
the paper represents the best possible "graceful" conclusion they
managed to produce of the work on transcription they carried out
with Michael over a span of more than two decades (a large portion
of which the third author was absent from the team).
While we by no means consider ourselves in a position to do
justice to the numerous enticing ideas and practical developments
contributed by Michael on the topic, we have felt it important to
at least try our best to make known some of the insights, hopes,
and the enthusiasm he had regarding the possibilities of
formalization for furthering the understanding of theoretical and
practical aspects of transcription.
Whatever good ideas may be found in this account stem in one
way or another from Michael. Only we, however, must bear the
responsibility for the many shortcomings that may remain. Their
number would have been much smaller had we been able, as we
repeatedly found ourselves wishing, to rely on the wit and
benevolence of our friend to keep us from going astray.
Our first formal publication on this topic came out in 2008
[Huitfeldt / Sperberg-McQueen 2008], followed by a Balisage paper in 2010
[Huitfeldt / Marcoux / Sperberg-McQueen 2010]. The notion of transcriptional
implicature is partly based on (and what we saw as a
low-hanging fruit from) ideas from these earlier publications. What
is presented here was more or less finished some time around 2015,
and last touched in 2018.[1] However, it was put on the back burner and has never
been formally published until now.[2]
At the time, we regarded transcriptional implicature as only
one of the parts of a more ambitious and comprehensive "formal
account", or a theory of "the logic", of transcription.
Regrettably, we did not manage to accomplish this quite demanding
task before Michael's untimely death.
Materials from that later period, going beyond what we present
here, will hopefully be made available as part of the [Sperberg-McQueen Nachlass project].
We would like to thank participants of a number of conferences
through the years, such as DH2009, Balisage2010, and DH2014, for
their interest, questions, and suggestions.
In particular, we want to thank Michael's wife, Marian
Sperberg-McQueen, for encouraging the posthumous publication of
this paper.
— Claus Huitfeldt & Yves Marcoux
Introduction
The notion of transcriptional implicature
There is hardly any universal transcription practice: for
every generalization we find exceptions. Is everything in the exemplar[3] transcribed? Not when deletions and irrelevant
material are excluded. Does everything in the transcript reproduce
some word or character in the exemplar? Not when line breaks are
marked explicitly with vertical bars, or notes are added.[4] Many scholarly editions account for variations like
these in an explicit statement of transcription practice. Such
statements typically describe deviations from
the usual practice, but rarely the ways in which it
exemplifies usual practice. Common practice
may sometimes be felt to be so obvious that it needs no mention or
explanation.
By transcriptional implicature for a
given community, we mean informally the things, suggested
or entailed by the rules of transcription, that members of that
community may find it unnecessary to mention explicitly.
Different communities of transcription practice have
different sets of tacit assumptions and thus different rules of
transcriptional implicature. Is there a common core of
transcriptional practice shared by all communities? There may be;
it is an empirical question. A definitive answer would require
more detailed studies of a wider variety of communities of
practice than we can currently manage. Our hypothesis, however, is
that there is such a common core, at least in the sense that the
transcriptional implicature of any community of practice can be
described with reference to some common set of rules.
If (as we conjecture) there is such a thing as the
transcriptional implicature of a given community, then the
transcriptional practice of any project of that community could be
described by listing the ways in which it
deviates from the transcriptional implicature.
Such deviations can be addition, modification, or
withdrawal of rules. Thus, the rules of the transcriptional
implicature are defeasible for particular projects. They apply,
except if explicitly excluded.
If (as we conjecture) the transcriptional implicature of a
community can in turn be described as a set of deviations from
some default transcriptional implicature, then it follows that any
project's transcription practice could be described with reference
to the default transcriptional implicature, by merging the list of
ways in which the community’s implicature deviates from the
default implicature with the list of ways in which the project’s
practice deviates from the default community practice.
The (conjectured) default transcriptional
implicature is a set of rules which apply by default, but which
may be overridden in particular cases, analogous to the rules of
conversational implicature proposed by H. P. Grice as a way of
explicating the logic of everyday conversation [Grice 1975].[5]
Earlier work
Since most digital transcriptions are represented today using
markup (e.g., TEI), providing a formal model
of transcription might prove useful for work on the semantics of
markup. Indeed, prevalent approaches to markup semantics are based
on formal languages like first-order predicate logic [Sperberg-McQueen / Huitfeldt / Renear 2001], which would a priori blend
well with a formal model of transcription, yielding a framework in
which meaning can be assigned to markup constructs whose
colloquial definitions involve transcription, for example,
p elements defined as
transcription of a text block in a
manuscript.
Some have proposed to explicate the
meaning of markup by specifying, for each
construct in a markup vocabulary, a sentence schema in a natural
language, with blanks to be filled in with data from the document
[Marcoux 2006]; others make a similar proposal but
allow sentence schemata in formal languages like first-order
predicate logic as well [Sperberg-McQueen / Huitfeldt / Renear 2001]. This appears
straightforward, although far from trivial, for metadata [Wickett / Renear 2009] and perhaps even for born-digital texts, but
how shall the meaning of a p element be formalized in
a markup language which defines it as containing a
transcription of a text block in a
manuscript? What does it mean for
a document to be a transcription of another document? Laying a
formal ground on which such questions can be tackled, may be seen
as contributions to the semantics of markup, as applied in
transcription.
Early work [Huitfeldt / Sperberg-McQueen 2008] has explored the nature
of the similarity between transcripts and their exemplars. If
documents are defined as sequences of characters, then perhaps the
similarity consists simply in the exemplar and the transcript
containing the same sequence of characters? This can be formalized
but proves disappointing, partly because the definition of
documents as sequences of characters omits text structures like
division into paragraphs and partly because the model offers no
way of describing disagreements among transcribers about how to
read the exemplar, or about which character distinctions (e.g.
i/j, u/v, s/ſ) to retain and which to level. It is also
wrong for the reason that few transcripts have
exactly the same character sequence as their
exemplar.
Later work has extended the model by introducing boolean
(disjunctive and conjunctive) types, which allowed explicit
representation of ambiguity, as well as
readings, which allowed modeling transcriber
agreement and disagreement explicitly [Sperberg-McQueen / Huitfeldt / Marcoux 2009],
and lifting the analyses from characters to higher-level textual
structures [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
Formalization seems interesting to us for two reasons: First,
and most importantly, formalization encourages a level of
explicitness and rigor which may help us better understand what
transcription is, and thereby also inform our discussions and
judgments in concrete cases. Second, now that transcription is
mostly done digitally, a formal model may help us build better
applications and better understand their possibilities and
limitations.
Our goal is simply to describe formally the relationship that
exists between two documents when the one is said to be a
transcription of the other. It is not our goal
to establish normative criteria which may serve to distinguish good
from bad or right from wrong in transcription, nor to construct a
theory of the psychological or physical processes involved in
transcription.
The term transcription may be used
variously to refer to the act or
process of transcribing a document, to the
physical product of that act (that is, another document), or to
the relation between the two documents.[6]
Our formal approach focuses on the
relation that must obtain between two
documents for the one to be a transcription of the other. Where
such a relation obtains, we call one of the documents a
transcript, and the other an
exemplar.
We take as basic facts about this relation that:
any exemplar antedates its
transcript,
the transcript has been made by someone for certain
reasons and with certain
intentions,
it has been made on the basis of
access to the exemplar (or facsimile of
the exemplar), and
the text of the transcript is the same as, or at least
similar in some way to, the text of the exemplar.[7]
Of all those aspects of the relation between exemplar and
transcript, our formalization addresses only the last one, that
is, the similarity relation that is assumed
to obtain between the texts of a transcript and of its
exemplar.
The kinds, levels or degrees of text similarity required for
a document to be counted as a transcription of another obviously
vary from context to context, and it is not hard to find examples
of disagreement about the matter. The formal model must be
flexible enough to allow for such variations.
The basic notions
The main idea in our approach to formalization is that
documents are textual objects[8] and, as such, can be analyzed in terms of a theory of
types and tokens.[9] In turn, the similarity
relationship that must obtain between two documents for one to be
a transcript of the other (we will call it
T-similarity) is defined in terms of
type-token analyses of the documents.
Here are the basic notions on which the current account is based.
document: A document (in
particular, an exemplar or a
transcript) is a physical phenomenon
containing or exhibiting marks. (For
example, the copy of Moby Dick on the third author's
nightstand, the Wittgenstein notebook cataloged as MS 108 in
the Wren library.)
mark: A mark is a
perceptually discernible arrangement of physical reality in a
document; some (and perhaps, but not necessarily, all) marks
are tokens (for example, words written in
pen or pencil in a manuscript, or the printed words, fly specks
and pencil marks of an old book).[10]
token: A token is a
mark determined to instantiate a
particular type while reading the
document. (For example, a letter, punctuation mark, a word or a
sentence on a piece of paper or parchment.) Tokens can be basic
(atomic) or compound. A compound token is one that is read as
the composition of smaller tokens, called its subtokens. A
whole document is viewed as a single "highest-level" compound
token, i.e., one that is not part of a larger compound
token.
type: A
type is a character, letter, word,
sentence, chapter, or any other object instantiated by one or
more tokens in one or more documents. (For example, the
character 'A' of the Latin alphabet, or the sentence 'I know
Verona'.) Like tokens, types can be basic or compound. A type
entering in the composition of a compound type is said to be a
subtype of the latter.
For example, the written word "cat" might be read as a
compound token instantiating the compound word type 'cat', and
composed of the three basic subtokens "c", "a", and "t",
instantiating respectively the basic letter types 'c', 'a', and
't', which are subtypes of the word type 'cat'.[11]
Documents (including exemplars and transcripts), which are
highest-level compound tokens, each instantiate a
(highest-level) compound type.
Types can also be conjunctive or
disjunctive, corresponding to ambiguous
tokens. For example, in a manuscript where it is not possible
to determine whether some token is of type 'i' or 'j', yet
clear that it is not of any other type, the token can be said
to be of the disjunctive type 'i/j'. In some documents, as for
example so-called ambigrams, certain
tokens are meant to be read in more than one way. Those tokens
can then be regarded as conjunctive types.[12]
type system: A type
system is a set of type repertoires, which in turn are sets of
types. For example, a type system may contain:
a repertoire of characters used in 17th-century English
manuscripts,
a repertoire of words used in the English of that
period,
a repertoire of sentences formed using those words, and
repertoires of paragraphs, sections, and other textual
structures.
All tokens instantiate at least one type in at
least one repertoire. The same token cannot instantiate more
than one type in any given repertoire. Strictly speaking, a
token may represent different types in different repertoires.
For example, the token "I" in English may instantiate both a
character, a word, and a sentence, but not more than one of
each. However, for the sake of convenience, we will usually
speak of "the type" (singular) of any given token.
The model allows for constraints to be defined on which types
can be instantiated by a compound token, in terms of the subtypes
instantiated by its subtokens. For example, it could be that any
compound token must be assigned a compound type such that the
subtokens of the compound tokens instantiate the subtypes of the
compound type (as was the case in the 'cat' example above).
Identification of tokens and types
The process of reading of a document, in particular
identifying the tokens in the document and the type that each of
them instantiates, is outside the scope of our model. Whenever we
refer to a document, we will allow ourselves to refer also to its
tokens and to the type of each of these tokens, without further
explanation.
Like
most formal models, our work assumes that names can be assigned to
individuals in the universe of discourse, in particular tokens and
types. The practical issues involved in reaching agreement on the
identity of individuals and the names to be used for them are
outside the scope of the formalization. When individuals are as
numerous as the tokens in a novel, of course, agreement on names
for them all is likely to present practical difficulties.
The absence from our account of any description of a method
for assigning names to the tokens of a document reflects not the
belief that such an operation is easy, only the fact that it has
no bearing on the logical structure of transcription, or of
transcriptional implicature. Like the process of reading a
document in the first place, the process of identifying the tokens
in a document sufficiently well to enable productive disagreement
about how to read them is of great practical importance but
outside the focus of our work.
T-similarity
The model defines a relation between documents called
T-similarity.[13] Two documents are T-similar if and only if, as
high-level tokens, they are either of the same (high-level) type,
or of similar types, where similarity among
types is defined in a ad hoc manner, in the
context of some given transcription project or set of
projects.
T-similarity is the relation that captures
transcription. Its formal definition reflects the
rules of transcription.[14]
In the simplest case, T-similarity is based on type equality;
in other words, exemplar and transcript must be of
identical types, not just similar types.
However, the definition can be relaxed, and this is how we can
hope to capture reasonably realistic aspects of the similarity
between exemplar and transcript in actual projects. Allowing for
such loosening of the similarity conditions is an essential aspect
of our model. We will later further relax the similarity
conditions with the introduction of special
tokens.
The intent behind T-similarity (with suitable versions of
similarity among types) is that it corresponds to the kind of
similarity expected to obtain in practice between an exemplar and
a corresponding transcript, as discussed above. We do not expect,
nor claim, that T-similarity can capture all possible views of
exemplar-transcript similarity, but we suggest it represents a
minimal necessary condition. Also, as already
mentioned, similarity of text is but one aspect of the
relationship between exemplar and transcript in real life:
normally a transcript must follow temporally after its exemplar,
be made in the presence of the exemplar, and be the result of an
intentional effort.
As a mathematical relation, T-similarity is reflexive,
transitive, and symmetric. The irreflexive, intransitive, and
asymmetric nature of the factual relation observed between
transcripts and their exemplars appears to us to be a consequence
of contingent facts and not a property of the similarity relation
holding between the documents.
We have informally defined transcriptional implicature for a
given community as the things, suggested or entailed by the
rules of transcription, that members of that community may find
it unnecessary to mention explicitly. From the point of
view of our formal model, which focuses on the relation between
exemplar and transcript, this amounts to the things
about the T-similarity relation that
go without saying.
Document pairs
We propose to characterize T-similarity as a set of formal
statements expressed in first order logic (FOL),
with variables ranging over the types and tokens of the related
documents. "Things" about the T-similarity relation (including
those that purportedly go without saying) can then
be formally expressed as such statements.
Such a representation of T-similarity can be put to work in
various ways. For instance, given a pair of documents, we can seek
to formally establish whether they stand in a T-similarity
relationship or not. This will be done for a simple pair of
documents in [appendices].
Or, given one known document (say, a transcript) and assuming
that it is T-similar to some other document (an exemplar), we can
ask ourselves what can be inferred about that other document on
the basis of the known document. Exploratory work in this
direction have been described in Claus Huitfeldt and C. M.
Sperberg-McQueen, transcriptional Implicature: Using a
Transcript to Reason about an Exemplarin [Digital Humanities 2017 Conference Abstracts], p. 266,
([http://blackmesatech.com/2017/08/MLCD/).
To do those things, we need to be able to
represent pairs of documents in FOL. The
goal of this section is to introduce FOL elements – predicates and
constants – that can be used to "say things" about pairs of
documents, one exemplar and one transcript, in terms of their
types and tokens.
In the first-order formulae below, we use the following
notation: The variables X, Y, Z, etc. range over
all individuals in the universe of discourse (i.e.,
the types and tokens of the related documents). Each individual
carries one (and only one) of the predicates:
type(X): X is a type
e_token(X): X is a token in the exemplar
t_token(X): X is a token in the transcript
We define the following binary relations (predicates):
typeof(X,Y) ⇔ the type of X is Y.
transcript(X,Y) ⇔ token Y in the transcript
corresponds to token X in the exemplar.
exemplar(X,Y) ⇔ token Y in the exemplar
corresponds to token X in the transcript.
As already mentioned, we are silent on how tokens are
identified and assigned types. We are also silent on how the
correspondences between exemplar and transcript tokens represented
by the transcript() and exemplar() predicates are established.[15]
Finally, we introduce the following two individual constants:
e: the exemplar (as a single, usually compound,
token)
t: the transcript (as a single, usually compound,
token)
Note that we have:
e_token(e), and
t_token(t).
It should be clear that facts or hypotheses (to be proved or
disproved) about pairs of documents can be formulated using the
above apparatus. In particular, an exemplar-transcript pair can be
entirely represented, in terms of its types and tokens, as a
series of statements using the above predicates together with
additional constants standing for the involved types and tokens.
An example is given in [Appendices].
As previously mentioned, we do not prescribe nor describe any
method for naming the constants corresponding to types and tokens,
we simply take for granted the use of some adequate naming
scheme.
Axiomatization
The above definitions can be expressed as the following set
of axioms that document pairs must satisfy.
Axioms
transcript domain: For every X and
Y, if X is the transcript of Y, then X is a t_token
and Y is an e_token.
Initial formulation of the default transcriptional
implicature
The initial formulation of the conjectured
default transcriptional implicature which we
present here could be described informally as a transcript
contains the same text as its exemplar. Formally, this
means that the exemplar and the transcript are tokens instantiating
the same type: the transcript re-instantiates the type of the
exemplar. Thus, only documents of equal types
are considered T-similar; this corresponds to the simplest case of
T-similarity presented earlier.
T-similarity could then be
formalized in a single rule: ∃(X)
(typeof(e,X) ∧ typeof(t,X)).
In interesting cases, however, the type X (of which e and
t are both tokens) will be a compound type consisting of some
structure of smaller compound types (which in turn consist of
smaller ones still), instantiated by a compound token which
similarly consists of smaller tokens.
It is thus more appropriate to define
T-similarity as the conjunction of four conditions, as
follows:
Default transcriptional implicature definition of
T-similarity
Reciprocity: The
transcript() and exemplar() predicates map tokens in a
one-to-one fashion.
This proposition defines formally what it means for
e and t to be T-similar.
With the axioms given earlier and a FOL representation of any
pair of documents, we can verify whether T-similarity,
i.e., the conjunction of the four rules of reciprocity,
completeness, purity, and type identity, comes out as a theorem or
not. This is described in more detail below.
Discussion
It may be worth stressing that the rules just given do not
constitute a claim that every transcript is reciprocal, pure,
complete, or thoroughly type similar to its exemplar. They amount
to a claim that these things are true unless otherwise
stated or unless the standards of a given
community of practice dictate otherwise. That is, they
express general assumptions about transcripts which are defeasible in particular cases.
Consider different transcripts of the same exemplar. They may
vary for several reasons. As a first application of our model, let
us use it to describe some of those reasons.
Transcripts may disagree about which of the marks in the
exemplar instantiate types and are thus tokens (one transcript may
read a mark as a decorative pen stroke, the other as a letter in a
word). They may agree on the set of tokens found in the exemplar
but disagree on which types they instantiate. Or they may use
different type systems. One, for example, may distinguish the
allographs i/j, u/v, and ſ/s, while the other treats the
allographs as instantiations of the same type (as they are
instances of the same grapheme).
The use of different type systems can lead to the same kinds
of difference between transcripts as different understandings of
the exemplar; failure to understand the nature of the disagreement
(different reading of the exemplar? or different choice of type
system?) can lead to confusion and acrimony. Analysis of
conflicting transcripts through the common lens of a formal model
might reduce such confusion and acrimony.
We observe that a number of common variations in
transcription practice can be classified according to which rule of
the initial formulation of the default transcriptional implicature
they override.
Some transcripts omit deleted material, extraneous material,
illegible material, or material in specific writing systems
(mathematics, Greek, …). In other words, transcripts are not always
entirely complete.
Some transcripts mark lines with bars, sometimes also adding
line numbers. In some cases omissions may be marked by symbols or
standard phrases ("[Illegible]", etc.) In other words, transcripts
are not always entirely pure.
Some transcripts preserve allographic variations, while others
level those distinctions. For example, some transcripts preserve
the distinctions between long and short s, or vocalic and
consonantal i and u, others do not. This might be seen as a
limitation of the rule of thorough type similarity. We believe it
is more natural, however, to assume in such cases that the
transcripts employ different type systems, one in which the
allographs are considered tokens of the same type, and another one
in which they are not. In other words, different transcripts do not
necessarily employ identical type systems.
Some transcripts silently expand abbreviations, normalize
spelling and correct slips of the pen. At character level, the
rules of completeness and purity seem to be broken. Even so, such
transcripts may observe thorough type similarity at higher level
tokens, such as words. In other words, transcripts are not always
entirely type similar through and through.
Expansion of abbreviations in brackets or italics can preserve
the rules of default transcriptional implicature on word and higher
levels, but introduces characters in transcripts which lack
corresponding characters in the exemplar. Again, transcripts are
not always entirely pure and do not always preserve entirely
thorough type similarity.
In cases of doubt, as in the case of parts of manuscripts
which are hard to read because of wear, damage, or difficulties in
handwriting, transcribers tend to interpret words as correctly
spelled and sentences as grammatically well-formed, at least unless
there is evidence to the contrary. This is often referred to as the
principle of charity.[17] This tendency may be accounted for in at least two
ways: Either we can regard the token in e as instantiating a
disjunctive type and the token in t as instantiating one of the
unobjectionable disjuncts, or else we can regard the transcriber as
choosing to read the token in e as an instantiation of an
unobjectionable type. In the former case,
type similarity must be
defined to allow pairs of the form "x/y" in e and either 'x' or
'y' in t (so the types are similar but not identical). In the
latter case, type identity is preserved but the reading of e
simply ignores the fact that the passage in e might plausibly be
interpreted differently.
Reformulation of the default
transcriptional implicature to permit
exceptions
As we have just seen, in almost all actual transcription
projects there are at least a few exceptions to each of the four
rules a-d: material in the exemplar not transcribed (perhaps
because irrelevant to the purpose of the transcript), material in
the transcript not present in the original (if only page numbers
and footnotes), and transcription conventions which transcribe
selected tokens with tokens of similar, not identical type.
The examples which follow show a variety of such exceptions;
the formalizations given show one way in which the defeasibility of
the simple rules just given can be represented in logical formulae.
As will become clear below, concrete transcription practices
can be summarized as a set of exceptions to the rules of purity,
completeness, and type similarity. Since the rules given in the
preceding section do not countenance exceptions, it will prove
necessary, for reasoning about any actual transcript and its
exemplar, to reformulate those rules to capture the qualification
that they each apply except in special
cases.
We therefore extend the document-pair model with the following
predicates and relations:
se_token(X): X is a special exemplar
token
st_token(X): X is a special transcript
token
typesimilar(X,Y): types X and Y are
considered similar
In order to account for the inclusion of special exemplar and
special transcript tokens we modify the typeof domain
and the no class overlap axioms as
follows:
Note that in the absence of special tokens and of unequal
types that are typesimilar, the reformulation changes nothing to
T-similarity. Thus, the reformulation does not per
se change the default transcriptional
implicature.
However, the reformulation allows a concise description of any
transcription practice in terms of its
deviations from the rules of the default
transcriptional implicature. For any given project's transcription
practice, we can define the extension of the predicates se_token,
st_token, and typesimilar, and combine them with the rules just
given to provide a basis for inference.
The skeptical reader may object to the explanatory value of
this approach: In general, anything can be described as a deviation
from any set of rules. More specifically, the skeptical reader will
have observed that taken strictly, this reformulation amounts to
saying that the properties of completeness, purity, and
type-identity will apply in all cases, except when they do not; the
formalization just given has no way to express the expectation that
completeness, purity, and type-identity are the normal, expected,
or usual case, and the cases covered by the predicates se_token,
st_token, and typesimilar (when it does not coincide with type
identity) are special, unusual, and less frequent cases.
A logic designed to formalize defeasible reasoning would
perhaps capture that distinction better; testing the utility of
such formalisms for reasoning about transcription remains a
desideratum for the future. In the meantime, however, we hope that
this formulation in terms of standard first-order logic will
suffice for the purposes of our argument.
Readers may have been struck by some similarity (pun intended)
between our initial formulation of the default transcriptional
implicature and certain methods or criteria for identifying or
measuring document similarity, such as, for example, the so-called
Levensthein edit distance between strings of characters.[18] The edit distance between two such strings is measured
in terms of the minimal number of deletions, insertions or
substitutions of individual characters in the one string that is
required in order to make it identical to the other.
Analogously, we might think of special character tokens as
deletions, special transcript tokens as insertions, and type
similar tokens as substitutions. The difference, however, is that
while Levensthein similarity is well suited for operations on
character strings, it is less well suited for work on documents
with a non-trivial structure. The additions and modifications we
made in the reformulation of the default transcriptional
implicature to permit exceptions may be seen as an attempt to take
care of that difference.
Can our formulation of the default
transcriptional implicature be put to empirical test? In principle,
yes. At least we can see whether it makes good sense when applied
to the actual transcriptional practices of a wide range of real
projects. In practice, however, we lack the resources required to
establish any really broad empirical basis. In this paper, we have
limited ourselves to looking at two examples of descriptions of
transcription practice. First we discuss an example drawn from a
U.S.-based historical documentary edition (the papers of William
Penn), then the transcription in a literary edition of some
manuscript notes by Hermann Melville.
Example 1: The papers of William Penn
Statement of practice
In the chapter Editorial Method in [The Papers of William Penn], the editors provide what we regard as a
statement of practice, stating that:[19]
...In this edition, we aim to print a completely faithful
transcript of each original text, including blemishes and errors.
... In general, our editorial interpolations [enclosed in square
brackets] within the text are minimal...
We observe that the first sentence seems at least implicitly
to confirm our rules of completeness and cleanliness: Every token
of e is transcribed by exactly one token of the same type in
t, and every token in t is exemplified by exactly one token of
the same type in e. The second sentence introduces an exception:
transcripts also contain additional material in the form of
editorial interpolations (that is, transcripts are not entirely
clean). However, such added material is explicitly marked by
square brackets. The fact that the editors find it worth pointing
this out may be taken to suggest that it does not go without
saying in the relevant community that blemishes and errors are
retained, or that additions are always marked explicitly. The
editors continue:
Our editorial rules may be summarized as follows:
1. Each document selected for publication in The
Papers of William Penn is printed in full. ...
At first sight, this looks like a straightforward
confirmation of the rule of completeness. So why does it not go
without saying? Perhaps because it is not altogether common
practice of this community to print every document in full. Or
perhaps because the edition is a selection (that is, not complete
in the sense of containing all the papers of William Penn) the
editors found it important to make clear that although their
transcript is not complete with respect to the entire body of
material, it is complete with respect to each individual
document.
2. Each document is numbered, for convenient
cross-reference, and is supplied with a short title.
From this we can infer that transcription breaks the implicit
rule of purity by adding document numbers
and titles, which do not occur in the exemplar. It seems unlikely
that the editors find it necessary to make this explicit statement
on the assumption that otherwise readers would believe that
William Penn himself numbered his letters and provided them with
titles. It is more likely that the community in question does
expect to find some indication of where one document ends and the
next begins as well as some means of referring to individual
documents, but has not agreed on a uniform way of doing so.
3. The format of each document (including the salutation and
complimentary closure in letters) is rendered as in the original
or copy...[20]
It is not entirely clear from this, or from the rest of the
declaration of practice, what the editors mean by the
format of a document. Comparison of a transcript
to a facsimile of its exemplar,[21] however, suggests that what they mean is page layout.
At least the transcripts seem to preserve such features as blocks
or lines of text flushed to the right or left or indented, blank
lines, and so on. (We also note that, elsewhere in the edition,
poems are printed as lines of verse.[22])
It is possible that the editors' practice is based on an
explicit or implicit distinction between phenomena such as
paragraphs, salutations, signatures, date lines, poems, verse
lines, etc. In either case, they seem to be referring to what we
would call compound types, and to try to preserve T-similarity
between transcript and exemplar also in this respect.
It is therefore perhaps striking to observe that the
transcripts contain no indication of line or page breaks[23] in the exemplar, and that the editors do not mention
this fact. Should we consider this omission a violation of the
rule of completeness? (That is, are line and page breaks not
considered tokens of some type?) If so, is the reason why the
editors do not mention this omission that it is part of the
transcriptional implicature of the relevant community of practice?
Or do they simply rely on the appearance of the transcript on the
printed page with regular, running lines etc. to make it too
obvious to deserve mention that line and page breaks of the
exemplar must be different?
The statement numbered 4 in [The Papers of William Penn] p. 17
deals with the rendering of datings according to Julian,
Gregorian, and Quaker calendars, a complicated issue the details
of which we do not go into here.
5. The text of each document is rendered as follows:
a. Spelling is retained as written. Misspelled words are not
marked with an editorial [sic]. If the sense of a word is
obscured through misspelling, its meaning is clarified in a
footnote.
b. Capitalization is retained as written. In
seventeenth-century manuscripts, the capitalization of such
letters as "c," "k," "p," "s," and "w" is often a matter of
judgment, and we cannot claim that our readings are definitive.
Whenever it is clear to us that the initial letter in a sentence
has not been capitalized, it is left lower case.
c. Punctuation and paragraphing are retained as written.
When a sentence is not closed with a period, we have inserted an
extra space.
d. Words or phrases inserted into the text are placed
{within braces}.
e. Words or phrases deleted from the text are crossed
through.
f. Slips of the pen are retained as written, and are not
marked by [sic].
g. Contractions, abbreviations, superscript letters, and
ampersands are retained as written. When a contraction is marked
by a tilde, it is expanded.
5.a suggests that this transcript, by retaining original
spelling, does not deviate from general transcriptional
implicature, but also that it is normal practice in the relevant
community to do so, i.e., by silently normalizing spelling or
marking misspelling with [sic].[24] We take 5.f to mean that the editors make a
distinction between slips of the pen (as in 5.a) from
misspellings, but that they treat them the same way.
The insertion of footnotes to clarify the meaning of
misspelled words mentioned in 5.a, however, does represent an
exception to the default rule of purity, as does the expansion of
contractions marked by tilde mentioned in 5.g. Other deviations
from the rule of purity are identified in 5.c for the insertion of
an extra space when a sentence is not closed with a period, and
5.d for the use of braces to surround editorial insertions into
the text.[25]
5.e simply seems to confirm the default rules of
transcriptional implicature in stating that deleted words or
phrases in the exemplar are crossed through in the transcript.
That the editors are explicit about this may suggest that normal
practice in the community is different. (Perhaps inserting special
markers for deleted text (breaking the rule of purity) or leaving
deleted text out (breaking the rule of completeness)).[26]
Statement 5.b seems to be a paradigmatic application of what
we referred to above as the principle of charity. We take it to be
saying that deciding whether a given character in the manuscript
is uppercase or lowercase requires judgment on the part of the
transcriber.[27]
Thus, to a large extent the statements above explicitly
confirm rules which form part of the default transcriptional
implicature; that they are stated explicitly may suggest that in
the relevant community of practice (or among the expected readers
of the edition), it might be common to deviate from the default
rules in these cases.
h. The thorn is rendered as "th," and superscript
contractions attached to the thorn are brought down to the line
and expanded: as "the," "them," or "that." Our justification for
this procedure is that we no longer have a thorn, and modern
readers mistake it for "y." Likewise, since modern readers do not
recognize that "u" and "v" were used interchangeably in the
seventeenth century, we have rendered "u" as "v," or "v" as "u,"
whenever appropriate.
i. The £ sign in superscript is rendered as "l."[28]
j. The tailed "p" is expanded into "per," "pro," or "pre,"
as indicated by the rest of the word.
k. The long "s" is presented as a short "s." The double "ff"
is presented as a capital "F."
The statement contained in the first clause of the first
sentence of 5.h, i.e., rendering the letter "Þ" as two
characters, "t" and "h", may seem to deviate from all the four
rules of our default transcriptional implicature: 1) it breaks the
one-to-one correspondence between the tokens of the exemplar and
the transcript (no reciprocity), 2) there are tokens (thorns) in
the exemplar which do not appear in the transcript
(incompleteness), 3) there are tokens ("th"-sequences) in the
transcript which do not occur in the exemplar (impurity), and 4)
none of the characters in question are of the same type (no type
similarity). We observe that the statement may suggest that the
normal practice of this community would be to represent "thorn"
with tokens of a distinct type, which would not deviate from
default transcriptional practice at all.
However, if we interpret the statement to the effect that the
project employs a type system in which "Þ" and "th" are
tokens of the same type, the practice is entirely in accordance
with the rules of our default transcriptional implicature. On the
one hand, therefore, such an interpretation seems attractive. On
the other hand, it may seem objectionable: In all other contexts
than where they occur in sequence and stand for a thorn, the
tokens "t" and "h" will still be taken as tokens of two different
types.
To this objection we may answer that it is not at all unusual
for the type of tokens, or even the decision about where to draw
boundaries between tokens, to be context dependent.[29] Also, we do have here a principled justification for
the context-dependent treatment of the thorn: A convention already
exists for representing Þ as
th in English spelling; "Þ" and "th" are
traditionally pronounced identically by modern readers; and, as
the statement itself argues, modern readers tend to mistake the
thorn for a "y". (The last argument seems to appeal to a principle
of charity with the reader.)
If we are not convinced by these arguments and thus reluctant
to accept that "Þ" and "th" may be regarded as tokens of the
same type, one alternative might be to interpret the first clause
of 5.h to the effect that the thorn is regarded as a contraction
which is expanded to "th". In that case, the practice constitutes
a deviation from default rules on a par with other contractions
— see the discussion of tilde in 5.g above. In any case, the
same goes for the statement made in the second clause of the first
sentence concerning expansion of superscript contractions attached
to the thorn.
Now to the second sentence of 5.h, which observes that "u"
and "v" were used interchangeably in the seventeenth century, and
therefore the one is rendered as the other (or vice versa)
"whenever appropriate" — presumably this means according to
the expectations of modern readers (and thus this practice may be
seen as another instance of "charity to the reader"). This may be
a case for assuming that the project employs a type system
according to which "u" and "v" are tokens of the same type, and in
that case no deviation from default transcriptional implicature is
implied.
What may seem awkward about this analysis is that even if "u"
and "v" are in free distribution in the seventeenth century texts
(alternate realizations of the same type, just as allographs are
alternate realizations of the same grapheme), they are certainly
distinct types (and indeed distinct graphemes) in modern English.
An alternative analysis, therefore, may be that the type system of
the exemplar subsumes that of the transcript: every type in the
transcript corresponds to exactly one type in the exemplar, but
more than one type in the transcript ( here, u and
v) may correspond to the same type in the exemplar
(here, the type instantiated by the tokens we read as
u and v in the exemplar). Either
way, no deviance from the default transcriptional implicature is
implied.
Statements 5.i and 5.k also describe the type system used:
superscript £ is treated as a token of type "l",[30] and long and short "s" are regarded as tokens of the
same type. This will produce less ambiguity in the transcript than
was the case for "Þ" and "th" or for "v" and "u", since long
"s" will not occur at all in the transcript. Regarding "ff" as a
token of capital "F", however, could in principle lead to
ambiguity.
Statement 5.j may seem to break the rule of purity for the
expansion of the abbreviations mentioned. However, if tailed "p"s
are regarded as tokens (depending on context) of the types "per,"
"pro," or "pre," the rule is preserved.
The statements 5.h-k may or may not be understood as
representing deviations from the rules of default transcriptional
implicature, deviations which are either not generally shared, or
dealt with in other ways, by the community of practice. In either
case, giving a formal account of all details involved here will
lead to a certain amount of complication. As a first simplifying
approximation, we may suggest that the statements imply that the
project at hand employs a type system in which the following
tokens are regarded as tokens of the same type:
"Þ" and "th"
"Þ" with superscripts and "the", "them", or "that"
(depending on context and/or the nature of the superscript,
— the statement is silent on this point)
"u" and "v"
"U" and "V"
Superscript "£" and "l"
tailed "p" and "per," "pro," or "pre," as indicated by the
rest of the word.
long and short s
"ff" and "F"
Under this assumption, the statements 5.h-5.k do not
represent any deviation from default transcriptional implicature:
reciprocity, purity, completeness and thorough type similarity are
all preserved.
l. Words underlined in manuscript are
italicized.
m. Blanks in the manuscript, missing words, and illegible
words are rendered as [blank] or [missing word] or [illegible
word] or [illegible deletion]. If a missing word can be supplied,
it is inserted [within square brackets]. If the supplied word is
conjectural, it is followed by a question mark.
Statement 5.l, that words underlined in the exemplar are
rendered in italics in the transcript, may be seen in different
ways within our framework of rules of transcriptional implicature.
(The following remarks are relevant also to statement 5.e above,
where deleted words or phrases in the exemplar are rendered as
crossed through in the transcript.) We might simply assume that
the editors regard underlined and italicized tokens as typesimilar
(like allographs of the same grapheme). If so, however, why do
they preserve the distinction in the transcript? After all, the
distinction between other allographs of the same grapheme, like
long and short s, are not preserved.
It may seem more natural to assume that the editors
understand underlining in the exemplar as signaling a higher-level
feature, let us call it emphasis, which
pertains not to the atomic letter tokens individually, but to the
entire underlined word, and that they have chosen to signal this
feature by other means, i.e., italics, in the transcript. On this
assumption, an underlined word in the transcript is treated as a
compound token of a type different from the same word without
underlining. Also on this account, however, we might regard this
as a way of preserving type similarity, and thus in accordance
with the rules of transcriptional implicature.
Statement 5.m, that information about blanks, missing words,
illegible words, and conjectural readings are represented by words
or phrases within square brackets, signals a break with the rules
of both completeness and purity. Again, we may interpret the
editors' statement as a confirmation that these rules are
otherwise followed. The reason why they see a need to make this
explicit may be either that these rules are normally followed in
such cases, or that the phenomena in question are normally
signaled in other ways.
Interpretation in terms of T-similarity
Given the reformulation of the default transcriptional
implicature offered above, we can interpret the statement of
practice in terms of the three predicates
se_token (special exemplar token),
se_token (special transcript token), and
typesimilar (similar types).
As far as we can tell, the only tokens in e that are not
present in t are the tildes marking contraction. However, these
contractions are expanded in t, so one might argue that this is
a case of type similarity rather than special exemplar
tokens.
The discussion above indicates that the following tokens in
t are special transcript tokens, i.e., that they correspond to
nothing in e:
footnotes (and their markers)
extra space at end of sentence lacking final
punctuation
braces marking inserted material
expansions of contractions (a) marked by tilde, (b) using
thorn + superscripts as the,
them, that, (c)
using tailed "p" as per,
pro, pre.
constituent "t" and "h" of "th" = thorn.
constituent "l" and "." of "l." = superscript £.
We have concluded that the type system of t may be assumed
to be the same as that of e, but also noted that the transcript
seems to assume several cases of type similarity. Types t1 and
t2 are similar iff:
t1 = t2
or t1 = thorn, t2 = 〈 "t", "h" 〉
or t1 = superscript £, t2 = "l"
or t1 = "ff", t2 = "F"
or (disjunctive(t1) ∧ t2 ∈ disjuncts(t1))
We conclude that together, the statement of practice and our
study of the relation between the exemplar and the transcript, may
confirm our hypothesis that the underlying idea of transcription
in this edition is that of T-similarity.
Example 2: Melville's notes in an edition of
Shakespeare
Statement of practice
On pages 955 to 970 of [Hayford et al. 1988], the editors
discuss notes made by the American author Herman Melville in an
edition of Shakespeare. On pages 967-970 they transcribe the notes
and provide facsimile images of the pages. (Facsimiles of
Melville's notes and the editor's transcript, taken from [Hayford et al. 1988], can be found in [Appendix B].) On page
967 they provide a guide to Symbols used, which we
quote in full:[31]
[1] [...] revision or
insertion enclosed in square brackets was made later than
initial inscription of leaf
[2] <...> letters or
words enclosed in diamond brackets were canceled by lining out
[3] <...>word
letter(s) or word(s) written over are enclosed in diamond
brackets closed up to the following word or letter that was
superimposed
[4] ?word prefixed by a
question mark indicates conjectural reading
[5] xxxx undeciphered
letters (number of x's approximates numbers of letters
involved)
[6] all words in roman are Melville's
[7] all words in italics outside
brackets are words Melville underlined
[8] all words in italics inside
brackets are editorial
There is no explicit statement, in this short text, that the
exemplar has been transcribed in full or that nothing appears in
the transcript that is not transcribing something in the exemplar.
On the other hand, statement [6], that all words in roman are
Melville's, and statements [4] and [8], which explain how
editorial conjectures and additions are marked, indicate that
great care has been taken to let the reader know which parts of
the document are added by the editors. One way of making sense of
this is to infer that for these transcribers it goes
without saying that unless otherwise indicated, the
transcript is complete and pure in our sense.[32]
Statements [1], [2], [3], and [7]
indicate that the transcribers re-instantiate all insertions,
cancellations, overwritings, and underlined text in the original,
using the notations indicated. We model this by taking insertion,
cancellation, overwriting, and underlining as compound types,
which can like other types be recognized in the exemplar and
re-instantiated in the transcript.[33] These statements have a further implication for the
type system to be used in reading the transcript: the charitable
reader will infer that the various forms of brackets used to
signal insertions, cancellations, and overwriting do not appear in
the exemplar, since if they did, any occurrence of them in the
transcript would become ambiguous.[34]
Statement [5] describes an exception to the usual rule of
type-identity between tokens in e and tokens in t, and also to
the default 1:1 mapping between tokens in the two documents.
Undeciphered tokens are transcribed by an
approximate and not an
exact number of x's, because tokens cannot
be counted reliably until they have been identified as tokens of
specific types. (To be a token is to be a token of a specific
type.) And yet, if the occurrences of x in
the transcript are tokens, then they must surely be tokens of some
type.
This would call for treating a sequence of n occurrences of
x as a single token of the type undeciphered sequence of letters about n characters
wide. In the exemplar, the undeciphered sequence would
be an atomic (or basic) token, not a compound one; in the
transcript, it would be a compound token composed of an
appropriate number of occurrences of x. The
undeciphered token in the exemplar would then illustrate the
principle that different readers may plausibly read an exemplar in
different ways: a reader who manages to decipher a word will
assign it and its characters to the appropriate types, while a
reader who finds the word illegible will assign it to an undeciphered letter-sequence type of
appropriate length.
(In the formalizations of the exemplar and the transcript
below, however, we rely on a source of information according to
which the sequence of letters in question, which are transcribed
as ?almxxxx has been deciphered as
almanacks. In this situation, the sequence in the
exemplar must be considered as a special exemplar token, while the
sequence in the transcript is a special transcript token.)
In contrast to the Penn Papers, page and line breaks are
preserved in this edition. We observe that on this point none of
the editions make any explicit statement on their choice. On our
account, this suggests that the transcriptional implicatures of
the two communities of practice are different: In this case it
apparently goes without saying that line breaks are preserved, in
the other it apparently goes without saying that they are not.
It is interesting to observe that the statement of practice
in the case of Melville's notes is formulated in terms of what the
reader will see in t and what it means, whereas the Penn Paper's
statement of practice starts from phenomena in e and describes
the rules the transcribers have followed in creating t. The
statement of practice for Melville's notes, that is, specifies
more or less directly what inferences the reader can make about
e, given what is in t, whereas the Penn Papers rules specify
more or less directly what the transcriber must do in t, given
what is in e. To use the Penn transcripts to make inferences
about e requires that the rules be applied backwards, so to
speak.
Interpretation in terms of
T-similarity
As before, we can interpret the statement of practice in
terms of the three predicates se_token
(special exemplar token), se_token (special
transcript token), and typesimilar (similar
types).
Among tokens present in e and not present in t are marks
(arrows, circles, and carets to mark the desired point of
insertion) which indicate where inserted material belongs. (Note
that these are not mentioned explicitly in the list of symbols,
perhaps because they do not appear as symbols in the transcript;
their existence can be inferred only from the editorial remarks in
the transcript, at line 18 of page 969.) None of these types are
in fact transcribed in t. So, they can all be considered special
exemplar tokens in e, having no corresponding tokens in
t.
The phrase Conversation upon Gabriel, Micheal &
Raphel – gentlemanly &c occurs in line 18 of t, but
only below in e. The occurrence in e may be considered a
special exemplar token.
Similarly, if one believes (as one probably should[35]) that the string represented as
?almxxxx in line 23 of t can be read as
almanacks, then almanacks (or the
substring anacks) of e should be considered a
special exemplar token.
The List of Symbols makes clear that italicized words within
angle brackets are editorial, not authorial, i.e., they are special
exemplar tokens.
Similarly, a question mark prefixed to a word whose reading
is uncertain is a special transcript token.
Square brackets, angle brackets, angle brackets with
immediately following text, question marks prefixed to words, and
sequences of the character x are all
identified by the List of Symbols as having special meaning to
which no token in the exemplar directly corresponds, i.e., they are
special transcript tokens.
We may take as a given that in normal cases it will be clear
to a competent reader on inspection whether a token in t is or
is not an instance of one of these types. However, one may also
easily imagine cases in which it is not clear; these are not
different in kind from legibility issues for the reader of t,
both consisting in uncertainty as to the type instantiated by a
token (or uncertainty as to whether a given mark is a token or
not). As already pointed out, the process of reading lies outside
the scope of our work.
There are no clear examples of type similarity between tokens in
e and t. However, see our remarks on a normalized transcript
further below.
Proof strategy
We have now discussed the Melville edition's statement of
practice in terms of the rules of transcriptional implicature. By
doing so we have formulated an account in more or less concise
prose of which parts of e and t should be assumed special
exemplar or transcript tokens, and which types should be assumed
type-similar in order for e and t to be considered T-similar.
How can we establish a formal proof to show that, given these
assumptions, t and e are T-similar?
After all, e and t are concrete, visual objects. They
cannot directly be subjected to the formal procedures required for
such proofs, and definitely not to the digital tools we will use
to test such proofs.
What we can do, however, is to let digital representations
of the documents stand in for e and t (and all the tokens
contained therein), and then perform the formal proof on these
representations.
There are many ways this can be done. The way we have
chosen, is to create TEI-XML documents representing t and e.
(As a sanity test, we also produce HTML visual replicas of e and
t from the XML.) From these XML documents we create one FOL
representation of both documents and of the relations between
them. By feeding this representation alongside our transcriptional
implicature axioms to a theorem prover we check whether
T-similarity between the two documents can be proven.
The theorem prover we have used is [Vampire]. In
earlier work we have used [Alloy] and [Prolog] for
similar tasks. One of the reasons we chose Vampire this time, is
that it operates on fairly standard FOL notation, whereas other
tools use more idiosyncratic notations. Moreover, Vampire is less
apt to crash on the relatively large amounts of data involved in
detailed representations of documents.
Explanation by way of a toy example
An exemplar and a transcript
Let us consider the following artificial example of an
exemplar e:
Figure 1: Exemplar
and the following, equally artificial, transcript of
e, t :
Figure 2: Transcript
In order to justify the plausibility of t, we may assume
that the practice of the relevant community (or of this
particular project) is to include insertions and silently omit
deletions in a different hand (in this case, in red), to
normalize spelling (in this case, of Essexe to
Essex), and to add disambiguating remarks between
square brackets (in this case, [the Earl
of]).
XML representations
Our first step is to create an XML representation of e. We
propose:
<doc>
Elizabeth went <del>with</del> <add>to</add> Essexe
</doc>
Our next step is to create an XML representation of t. We
propose:
<doc>
Elizabeth went to <supplied>the Earl of</supplied> <choice><orig>Essexe</orig><reg>Essex</reg></choice>.
</doc>
We observe that t both adds to and leaves things out of
e, in accordance with the transcribal practice suggested.
<supplied>the Earl
of</supplied>, must be counted as special
transcript tokens. Furthermore, we understand
<choice><orig>Essexe</orig><reg>Essex</reg></choice>
to indicate that the tokens Essexe and
Essex are typesimilar.
When it comes to e,
<del>with</del> must be
considered a special exemplar token in relation to t.
Translation to FOL, and proof of T-similarity
Together, the XML representations of e and t, as input
to an appropriately devised XSLT stylesheet (see [Appendix E]), yield the following
output:
The Case specific facts are formulated
as an axiom in the form of one conjunction with many conjuncts,
representing the facts of e and t and the relations between
them.
The listing above is in the so-called FOF (First-Order Form)
notation used by Vampire. It is perhaps similar enough to the
standard FOL notation used elsewhere in this paper as it is. For
convenience, however, we include a translation to the usual
notation, and group the various conjuncts into numbered groups
for ease of reference:
In group 1 the tokens in the exemplar (normal as well as
special), named e1..e5, are associated with their
respective types. The names of the types are simply given as a
string consisting of the word string in question, prefixed with
t_. In other words, the token e1
is of type t_Elisabeth, and so on.
In group 2 the tokens in the transcript, named
t1..t7, are associated with their respective
types, in the same manner.
The statement in Group 3 states that the type named
t_Essexe is similar to the type named
t_Essex.
The statements in Group 4 assigns each of the exemplar
tokens in e to their corresponding transcript tokens in t. It
may deserve special attention that the special exemplar token e3
(of type t_with) and the special transcript tokens
t4, t5, and t6 (of types t_the, t_Earl,
and t_of), though represented as tokens associated
with their respective types, are not mentioned in the Group 4
statements. They are part of the formal description of the two
documents, though they do not play any role in T-similarity
proofs.
The statements in Group 5, 6, and 7 are there to deal with
the open world assumption of Vampire, and may be primarily of
technical interest.[36] In group 5, we state that there are no other pairs of
individuals than those mentioned in group 4 which stand in a
transcript relation to each other. In Group 6, the
$distinct predicate is a special Vampire
predicate that makes sure that each of the constants mentioned
refer uniquely, i.e., e1 names an individual not
named by any of the other constants mentioned, etc.[37] The statement in Group 7 makes sure that there are no
other e_tokens or t_tokens in the universe of discourse than
those listed in the right-hand clauses of the two biconditionals.
When we add the axioms 1-10 and the definitions a-d of
transcriptional implicature formulated earlier to the axiom of
Case specific facts above, the theorem
prover confirms that e and t are T-similar.[38]
Proof of T-similarity
We have gone through the same steps with the exemplar and
transcript [Appendix B] of Melville's notes as described
in the previous section on the toy example.
We first made XML representations of the exemplar and the
transcript, see [Appendix C]. In both, we used fairly
straightforward TEI encoding.[39]
As a sanity check we made sure we could transform the XML
files to HTML which came as close as we thought feasible to the
originals in textual as well as visual aspects. We present these
HTML files in [Appendix D]. For convenience, special
exemplar and special transcript tokens are marked in red.
The stylesheet briefly mentioned earlier,
genFof.xsl, is reproduced in [Appendix E]. This stylesheet takes the two XML files in [Appendix C] as input, compares them, and produces the output
in Vampire FOF format to be found in [Appendix F]. (In
its current form, special token GIs are hard-coded into the
stylesheet itself. One might easily extend the stylesheet so as to
give users control over these, and to distinguish between special
exemplar and transcript GI's.)
Vampire confirms as a theorem that the XML representations of
the exemplar and the transcript satisfies the T-similarity
predicate, or, in other words, that e and t as represented are
T-similar.
We believe this to show that our definition of T-similarity,
taking exceptions from default transcriptional implicature into
account, is at least feasible. One such test can of course not be
conclusive. As mentioned earlier, our hypothesis that there is a
common core of assumptions underlying transcription in general is
one that can only be tested empirically. We do not have the
resources necessary for extensive empirical work.
A normalized transcription
However, we have at least made one other, very different,
normalized, transcript of the exemplar ([Appendix G]). Unlike the transcript
discussed above, this one does contain some type similarity
statements: & is declared type similar to
and, &c to etc.,
almanacks to almanacs,
nonsence to nonsense, and
nominee to nomine.[40]
We put the normalized transcript to the same test. Again,
T-similarity as defined here, with exceptions from default
assumptions made explicitly and formally, shows that also the
normalized transcript is T-similar to the exemplar. See [Appendix G].
This illustrates an important and more general point: Very
different transcripts can be T-similar to one and the same
exemplar. (Since T-similarity is symmetric and transitive, this
means that the diplomatic and normalized transcriptions should be
T-similar, too. We do not prove this.)
Another important point illustrated by this exercise is that
the representation of the exemplar, in the way things are set up
here, is not necessarily (or usually cannot even be) independent
of the way the transcript is represented. (For example, a deletion
in the exemplar should be represented as a special transcript
token if and only if it is omitted from the transcript.)
Discussion
One of the things we think that this project illustrates, is
the complexity of document structures and of the relations between
documents. Our formalization has been on a fairly shallow level. We
have succeeded in bringing the axioms of transcriptional
implicature down to only a dozen fairly simple FOL statements, but
the formal representation of even the short, less than 300 word
Melville document, requires a conjunct of thousands of statements.
Processing of the representation of even this small document
requires a fairly large amount of computing power and time —
with a medium-range personal computer Vampire takes several minutes
to process the Melville transcripts.
Therefore, we would also like to observe here that Vampire has
impressed us. As mentioned, we have tried to do this work with the
help of other tools in the past, but have had to give up because of
insurmountable difficulties with formalization, user interface, or
speed. This is the first time we have had the occasion to do real
work with formalization and automated theorem proving, even though
still on a fairly modest scale. We realize that some of our work in
Vampire might have been done more efficiently with better knowledge
of its possibilities. Other theorem provers exist, but we are not
presently aware of any that are better suited to the task.
We regard the work presented here as merely a proof of the
concept of formalizing the modeling of documents and document
relations, and of the applicability of the notion of T-similarity
to transcription. Any implementation for real work would
have to use other, more user-friendly and efficient tools, tools
which do not, for example, require thorough knowledge of standard
first order logic.
In [Huitfeldt / Sperberg-McQueen 2008] we presented two models of text:
The Grapheme-sequence Model and the Readings
model. We found the Grapheme-sequence Model
unsatisfactory because of a number limitations:
Texts are not just sequences of graphemes, -- the are
usually more complex graph-like structures with notes, variants,
alternate readings etc.
The analysis of document types must reflect
compositionality and levels of types: there are atomic as well
as boolean and compound types, types at the levels of letters,
words, sentences, paragraphs, etc.
there may be different readings of the same document
In [Huitfeldt / Marcoux / Sperberg-McQueen 2010] we tried to develop a formal
account without these limitations. However, we never reached a full
working formalization, partly because of the problems with finding
suitable formalization tools mentioned above.
With what is presented here, we are in many ways back with the
Grapheme-sequence Model: We represent documents as
sequences — not even of letters, but of word tokens. (We
could of course have chosen to model on character level in stead.
That would have increased the number of tokens in the model,
without making much interesting difference in principle.) The type
structure is also represented as entirely flat, without levels or
compositionality.
Still, with these limitations, we have a working model which
lends itself to full formalization and formal proof. We can only
hope that someone may some time extend it beyond its current
limitations, e.g., along the lines of [Huitfeldt / Marcoux / Sperberg-McQueen 2010].
Some readers may have been suspicious of our move, in the
section on Melville, from discussing the original exemplar and
typescript on paper to discussing XML-representations of them. It
was this move, however, which made us aware of a (perhaps to other
obvious) fact: Our model presupposes that the representation of the
exemplar must reflect what is not recorded in
the transcript.[41]
Moreover, we think that this move shows that it might be
useful to think in terms of T-similarity or similar formalizations
of similarity between digital documents in contexts well beyond
transcription and scholarly editing.
Originally, this talk had the subtitle a contribution
to markup semantics. We can only hope that it will
be.
Conclusions and Future Work
We have proposed the term transcriptional
implicature to denote the rules of inference which
should govern the interpretation of a transcript in the absence of
explicit statements to the contrary. Informally, the rules of
transcriptional implicature (are intended to) capture those facts
about transcription which are too obvious to need saying. We expect
that the rules of transcriptional implicature vary from community
of practice to community of practice; their accurate identification
will be a matter of the sociology of scholarship, philosophy of
science, and close examination of transcription practice. We
believe that the rules of what we call default
transcriptional implicature provide a plausible example
of what the transcriptional implicature for a given community of
practice might look like, and we have shown how the practice of
some concrete examples can be captured formally in ways which
exhibit clearly both the general rules governing transcription and
the specific practices of the individual project.
We hope that this work can provide a basis for a full model of
the logical structure of transcription, of the formal relation
between exemplar and transcript, and of the ways in which a
transcript can be used to infer facts about its exemplar.
Appendix A. Penn: Facsimile of exemplar and transcript.
Facsimiles of the first and last page of a letter from Penn to
Lady Conway from 1675, reproduced on the insides of the front and
back cover of [The Papers of William Penn]:
Figure 3: Penn, letter to Lady Conway, first page
Figure 4: Penn, letter to Lady Conway, last page
Facsimiles of excerpt from the transcript of the corresponding
pages in [The Papers of William Penn], pages 356 and 358:
Figure 5: Penn, letter to Lady Conway, first page
Figure 6: Penn, letter to Lady Conway, last page
Appendix B. Melville: Facsimile of exemplar and transcript
[Hayford et al. 1988] contain facsimiles of the two pages of
Melville's exemplar on pages 968 and 970.
Facsimiles of the same two pages in better quality can be
found here:
<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="../../genFofXSLT/genFof.xsl"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->]>
<text>
<body>
<p>
<lb/>A seaman figures in The Canterbury Tales. <lb/>With
<del>a</del>many a tempest had his beard been <lb/>shook. –
<del>S</del>
<del>Deep g</del> Secret grief is a <lb/>cannibal of its own heart
– <emph>Bacon</emph>. <lb/>&vbar;<emph>Claudia</emph> of the
<emph>Appian</emph> family, <q>I wish <lb/>&vbar; some fight or
pestilence would thin out <lb/>&vbar; this crowd.</q>
<emph>Arrogance.</emph>
<lb/>
<emph>Roast beef in the pulpet.</emph>
</p>
<p>
<lb/>An animal of a man — <q>do eagles wear
<lb/>spectacles?</q>— Health. — <emph>Contrast</emph>:
an <lb/>over spiritual man. <lb/>—— </p>
<p>
<lb/>
<q>Yes, Madam, Cain was a godless froward boy, & <lb/>Reuben
(Gen:49) & Absalom</q> Many pious men <lb/>have impious
children — (Devil as a Quaker)
<lb/>————— </p>
<p>
<lb/>A formal compact – Imprimis – First –
Second. <lb/>The aforesaid soul. said soul &c –
Duplicates – <lb/>&triplebar;<q>How was it about the
temptation on the <lb/>hill?</q> &c – D begs the hero to
form <lb/>one of a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb/>&c –
Leaves a letter to the D – <q>My <lb/>
<emph>Dear D</emph>
</q>
<floatingText>
<body>
<ab> – Conversation upon Gabriel, Micheal &
<lb/> Raphel – gentlemanly &c</ab>
</body>
</floatingText>
</p>
<p>
<lb/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb/> making
<sic>almanacks</sic> – going to a ball takes a long
<lb/>time making toilette. – The Doctor's coach stops
<lb/>the way. – <q>Do you believe all that stuff? <lb/>
nonsence – the world was never made. – “But Is
not <lb/>this you mentioned <emph>here</emph> – in the
scriptures?</q>
<lb/>Receives visits from the principal d's –
<q>Gentlemen</q> &c. <lb/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb/>be below with kings than above with fools?</q>
</p>
<p>
<lb/>It is better to laugh & not sin than to <del>be</del> weep
& be <lb/>wicked. — Ten loads of coal to burn him.
— <lb/>Brought to the stake — warmed himself by the
fire. </p>
<p>
<lb/>Ego non baptizo te in nominee Patris et <lb/>Filii et Spiritus
Sancti – sed in nomine <lb/>Diaboli. — Madness is
undefinable — <lb/>It & right reasons extremes of one.
<lb/>–Not the <add place="sup">(black art)</add> Goetic but
Theurgic magic — <lb/>seeks converse with the Intelligence,
Power, the <lb/>Angel. </p>
</body>
</text>
XML representation of the transcript
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->]>
<text>
<front>
<docTitle>
<titlePart>NOTES IN A SHAKESPEARE VOLUME
<span>969</span></titlePart>
</docTitle>
</front>
<body>
<fw>[on verso of last leaf of Volume VII, page [524]]</fw>
<p>
<lb n="1"/>A seaman figures in The Canterbury Tales. <lb n="2"
/>With <del>a</del>many a tempest had his beard been <lb n="3"
/>shook. – <del>S</del>
<del>Deep g</del> Secret grief is a <lb n="4"/>cannibal of its own
heart – <emph>Bacon</emph>. <lb n="5"/>
<metamark>&vbar;</metamark>
<emph>Claudia</emph> of the <emph>Appian</emph> family, <q>I wish
<lb n="6"/>
<metamark>&vbar;</metamark> some fight or pestilence would thin
out <lb n="7"/>
<metamark>&vbar;</metamark> this crowd.</q>
<emph>Arrogance.</emph>
<lb n="8"/>
<emph>Roast beef in the pulpet.</emph>
</p>
<p>
<lb n="9"/>An animal of a man — <q>do eagles wear <lb n="10"
/>spectacles?</q>— Health. — <emph>Contrast</emph>: an
<lb n="11"/>over spiritual man. <lb rend="dunno"/>
<metamark>——</metamark>
</p>
<p>
<lb n="12"/>
<q>Yes, Madam, Cain was a godless froward boy, & <lb n="13"
/>Reuben (Gen:49) & Absalom</q> Many pious men <lb n="14"
/>have impious children — (Devil as a Quaker) <lb
rend="dunno"/>
<metamark>—————</metamark>
</p>
<p>
<lb n="15"/>A formal compact – Imprimis – First –
Second. <lb n="16"/>The aforesaid soul. said soul &c –
Duplicates – <lb n="17"/>&triplebar;<q>How was it about the
temptation on the <lb n="18"/>hill?</q> &c <supplied rend="[]"
>inserted later below in lines 21–21b after Dear D“
— and circled <lb rend="dunno"/> with guideline to caret
here</supplied>
<floatingText>
<body>
<ab>Conversation upon Gabriel, Micheal & / <lb rend="dunno"/>
Raphel – gentlemanly &c<supplied rend="]"/></ab>
</body>
</floatingText> – D begs the hero to form <lb n="19"/>one of
a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb n="20"/>&c
– Leaves a letter to the D – <q>My <lb n="21"/> Dear D
</q> – <supplied rend="[]">later insertion in lines
21–21b, reported in line 18</supplied>
</p>
<p>
<lb n="22"/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb n="23"/>
<unclear>making</unclear>
<unclear><supplied>alm<gap/></supplied></unclear> – going to
a ball takes a long <lb n="24"/>time making toilette. – The
Doctor's coach stops <lb n="25"/>the way. – <q>Do you
believe all that stuff? <lb n="26"/> nonsence – the world
was never made. – <supplied rend="[">add</supplied>
“But<supplied rend="]"/> Is not <lb n="27"/>this you
mentioned <emph>here</emph> – in the scriptures?</q>
<lb n="28"/>Receives visits from the principal d's –
<q>Gentlemen</q> &c. <lb n="29"/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb n="30"/>be below with kings than above with fools?</q>
</p>
<p>
<fw>[on recto of last blank leaf of Volume VII, page [523]]</fw>
<lb n="31"/>It is better to laugh & not sin than to
<del>be</del> weep & be <lb n="32"/>wicked. — Ten loads
of coal to burn him. — <lb n="33"/>Brought to the stake
— warmed himself by the fire. </p>
<p>
<lb n="34"/>Ego non baptizo te in nominee Patris et <lb n="35"
/>Filii et Spiritus Sancti – sed in nomine <lb n="36"
/>Diaboli. — Madness is undefinable — <lb n="37"/>It
& right reasons extremes of one. <lb n="38"/>–Not the
<supplied rend="[">inserted above line with caret below</supplied>
(black art)<supplied rend="]"/> Goetic but Theurgic magic —
<lb n="39"/>seeks converse with the Intelligence, Power, the <lb
n="40"/>Angel. </p>
</body>
</text>
Appendix D. Melville: HTML presentations
HTML presentation of the exemplar
Special exemplar tokens in red.
Figure 9: HTML presentation of the exemplar
HTML presentation of the transcript
Special transcript tokens in red.
Figure 10: HTML presentation of the transcript
Appendix E. genFoF stylesheet
<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
xmlns:ym="http://www.marcouxmedias.com"
xmlns:xs="http://www.w3.org/2001/XMLSchema"
xpath-default-namespace="" xml:lang="fr-CA" xml:space="ignore"
exclude-result-prefixes="#all">
<xsl:output method="text" indent="no" encoding="UTF-8" />
<xsl:function name="ym:lastIndexOf" as="xs:integer">
<xsl:param name="str" />
<xsl:param name="car" />
<xsl:sequence select="if (contains($str, $car)) then
max(for $i in (1 to string-length($str)) return
if (substring($str, $i, 1) = $car) then $i else 0)
else 0" />
</xsl:function>
<xsl:variable name="fpNoExt">
<xsl:variable name="temp" select="base-uri(/)" />
<xsl:value-of select="substring($temp, 1, ym:lastIndexOf($temp, '.'))" />
</xsl:variable>
<xsl:variable name="indent" select="' '"/>
<xsl:function name="ym:tokenizePlus" as="item()*">
<xsl:param name="text" as="node()" />
<xsl:choose>
<xsl:when test="not($text/ancestor::orig)">
<xsl:variable name="toks" select=
"tokenize($text,'[^a-zA-Z0-9]+')[.]" />
<xsl:choose>
<!-- The values of test specified below are project-specific -->
<xsl:when test="$text/ancestor::front
or $text/ancestor::fw
or $text/ancestor::supplied
or $text/ancestor::sic
or $text/ancestor::floatingText">
<xsl:sequence select="for $t in $toks return '*' || $t" />
</xsl:when>
<xsl:otherwise>
<xsl:sequence select="$toks" />
</xsl:otherwise>
</xsl:choose>
</xsl:when>
<xsl:otherwise />
</xsl:choose>
</xsl:function>
<xsl:function name="ym:addEffSeqNum" as="item()*">
<xsl:param name="toks" as="item()*" />
<xsl:sequence select="
for $n in 1 to count($toks) return
(if (not(starts-with($toks[$n],'*'))) then
count($toks[position() lt $n and not(starts-with(.,'*'))]) || '*' else '')
|| $toks[$n]
" />
</xsl:function>
<xsl:template match="/">
<xsl:variable name="doc1" select="/*"/>
<xsl:variable name="toks1" select=
"ym:addEffSeqNum($doc1//text()/ym:tokenizePlus(.))" />
<xsl:variable name="doc2" select="document($fpNoExt || 'transcript.xml')/*"/>
<xsl:variable name="toks2" select=
"ym:addEffSeqNum($doc2//text()/ym:tokenizePlus(.))" />
<xsl:variable name="nNormToks" select="count($toks1[not(starts-with(.,'*'))])" />
<xsl:if test="$nNormToks ne count($toks2[not(starts-with(.,'*'))])">
<xsl:message terminate="no"
select="string-join(($nNormToks, count($toks2[not(starts-with(.,'*'))]), ' '), ' ')"
>WARNING: Normal token counts differ between documents.</xsl:message>
</xsl:if>
<xsl:text>fof(case_specific_facts, axiom,
</xsl:text>
<xsl:call-template name="processDoc">
<xsl:with-param name="docPrefix" select="'e'" />
<xsl:with-param name="toks" select="$toks1" />
</xsl:call-template>
<xsl:call-template name="processDoc">
<xsl:with-param name="docPrefix" select="'t'" />
<xsl:with-param name="toks" select="$toks2" />
</xsl:call-template>
<xsl:for-each select="$doc2//choice">
<!-- The [.] in the following are to get rid of possible empty strings at the
beginning and end of the string: -->
<xsl:variable name="toksOrig" select="tokenize(orig,'[^a-zA-Z0-9]+')[.]"/>
<xsl:variable name="toksReg" select="tokenize(reg,'[^a-zA-Z0-9]+')[.]"/>
<xsl:if test="count($toksOrig) ne count($toksReg) or not(count($toksOrig))">
<xsl:message select="string-join((count($toksOrig), count($toksReg), ' '), ' ')"
>WARNING: Token count mismatch within a <choice> element.</xsl:message>
</xsl:if>
<xsl:for-each select="1 to count($toksOrig)">
<xsl:if test="$toksOrig[current()] ne $toksReg[current()]">
<xsl:value-of select="$indent" />
<xsl:text>typesimilar(t_</xsl:text>
<xsl:value-of select="$toksOrig[current()]" />
<xsl:text>, t_</xsl:text>
<xsl:value-of select="$toksReg[current()]" />
<xsl:text>) &
</xsl:text>
</xsl:if>
</xsl:for-each>
</xsl:for-each>
<xsl:for-each select="1 to count($toks1)">
<xsl:if test="not(starts-with($toks1[current()],'*'))">
<xsl:variable name="idNormTok" select="substring-before($toks1[current()],'*')" />
<xsl:value-of select="$indent" />
<xsl:text>transcript(e</xsl:text>
<xsl:value-of select="." />
<xsl:text>,t</xsl:text>
<xsl:value-of select="max(
for $i in 1 to count($toks2) return
if (starts-with($toks2[$i], $idNormTok || '*')) then $i else 0
)" />
<xsl:text>) &
</xsl:text>
</xsl:if>
</xsl:for-each>
<!--% No other pairs of corresponding tokens (gives us reciprocity for free):-->
<xsl:text>
! [X,Y,Z] : (
((transcript(X,Y) & transcript(X,Z)) => Y=Z) &
((exemplar(X,Y) & exemplar(X,Z)) => Y=Z)
) &
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text>$distinct(
</xsl:text><xsl:value-of select="$indent" /><xsl:value-of select="$indent" />
<xsl:value-of select="
string-join(for $i in 1 to count($toks1) return ('e' || $i), ',')
" />
<xsl:text>,
</xsl:text><xsl:value-of select="$indent" /><xsl:value-of select="$indent" />
<xsl:value-of select="
string-join(for $i in 1 to count($toks2) return ('t' || $i), ',')
" />
<xsl:text>,
</xsl:text><xsl:value-of select="$indent" /><xsl:value-of select="$indent" />
<xsl:variable name="allTypes" select="
(for $i in 1 to count($toks1) return 't_' || substring-after($toks1[$i], '*')),
(for $i in 1 to count($toks2) return 't_' || substring-after($toks2[$i], '*'))
"/>
<xsl:value-of select="string-join(distinct-values($allTypes), ',')" />
<xsl:text>
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text>) &
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text>! [X] : ((e_token(X) <=> (X=e</xsl:text>
<xsl:value-of select="string-join(
(1 to count($toks1))[let $i := . return not(starts-with($toks1[$i],'*'))],
' | X=e')" />
<xsl:text>)) &
</xsl:text>
<xsl:value-of select="$indent" />
<xsl:text> (t_token(X) <=> (X=t</xsl:text>
<xsl:value-of select="string-join(
(1 to count($toks2))[let $i := . return not(starts-with($toks2[$i],'*'))],
' | X=t')" />
<xsl:text>)))
).</xsl:text>
</xsl:template>
<xsl:template name="processDoc">
<xsl:param name="toks" />
<xsl:param name="docPrefix" />
<!--<xsl:message select="$toks" />-->
<xsl:for-each select="$toks">
<!--
<xsl:value-of select="$indent" />
<xsl:value-of select="if (starts-with(.,'*')) then 's' else ''"/>
<xsl:value-of select="$docPrefix"/>
<xsl:text>_token(</xsl:text>
<xsl:value-of select="$docPrefix"/>
<xsl:value-of select="position()" />
<xsl:text>) &
</xsl:text>
^ 2026-03-31 Zoom meeting, decided we don’t need them. -->
<xsl:value-of select="$indent" />
<xsl:text>typeof(</xsl:text>
<xsl:value-of select="$docPrefix"/>
<xsl:value-of select="position()"/>
<xsl:text>, t_</xsl:text>
<xsl:value-of select="substring-after(.,'*')"/>
<xsl:text>) &
</xsl:text>
</xsl:for-each>
</xsl:template>
</xsl:stylesheet>
Appendix F. Melville: FOF representation of exemplar and transcript
Appendix G. Melville: XML, HTML, and FOF of normalized transcript
XML representation of exemplar for normalized
transcript
<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="../../genFofXSLT/genFof-Norm.xsl"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->
<!-- For HTML: -->
<!ENTITY and "&" >
<!ENTITY etc "&c" >
<!-- FOR Fof: -->
<!-- <!ENTITY and "amp" >
<!ENTITY etc "ampc" >-->
]>
<text>
<body>
<p>
<lb/>A seaman figures in The Canterbury Tales. <lb/>With
<del>a</del>many a tempest had his beard been <lb/>shook. –
<del>S</del>
<del>Deep g</del> Secret grief is a <lb/>cannibal of its own heart
– <emph>Bacon</emph>.
<lb/><metamark>&vbar;</metamark><emph>Claudia</emph> of the
<emph>Appian</emph> family, <q>I wish
<lb/><metamark>&vbar;</metamark> some fight or pestilence would
thin out <lb/><metamark>&vbar;</metamark> this crowd.</q>
<emph>Arrogance.</emph>
<lb/>
<emph>Roast beef in the pulpet.</emph>
</p>
<p>
<lb/>An animal of a man — <q>do eagles wear
<lb/>spectacles?</q>— Health. — <emph>Contrast</emph>:
an <lb/>over spiritual man. <lb/><metamark>——
</metamark></p>
<p>
<lb/>
<q>Yes, Madam, Cain was a godless froward boy, ∧ <lb/>Reuben
(Gen:49) ∧ Absalom</q> Many pious men <lb/>have impious
children — (Devil as a Quaker)
<lb/><metamark>—————</metamark>
</p>
<p>
<lb/>A formal compact – Imprimis – First –
Second. <lb/>The aforesaid soul<sic>. said soul</sic> &etc; –
Duplicates – <lb/><metamark>&triplebar;</metamark><q>How was
it about the temptation on the <lb/>hill?</q>&etc; – D begs
the hero to form <lb/>one of a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb/>&etc; – Leaves
a letter to the D – <q>My <lb/>
<emph>Dear D</emph>
</q>
<floatingText>
<body>
<ab> – Conversation upon Gabriel, Micheal ∧
<lb/> Raphel – gentlemanly &etc;</ab>
</body>
</floatingText>
</p>
<p>
<lb/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb/> making
almanacks – going to a ball takes a long <lb/>time making
toilette. – The Doctor's coach stops <lb/>the way.
– <q>Do you believe all that stuff? <lb/> nonsence –
the world was never made. – “But Is not <lb/>this you
mentioned <emph>here</emph> – in the scriptures?</q>
<lb/>Receives visits from the principal d's –
<q>Gentlemen</q> &etc;. <lb/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb/>be below with kings than above with fools?</q>
</p>
<p>
<lb/>It is better to laugh ∧ not sin than to <del>be</del> weep
∧ be <lb/>wicked. — Ten loads of coal to burn him.
— <lb/>Brought to the stake — warmed himself by the
fire. </p>
<p>
<lb/>Ego non baptizo te in nominee Patris et <lb/>Filii et Spiritus
Sancti – sed in nomine <lb/>Diaboli. — Madness is
undefinable — <lb/>It ∧ right reasons extremes of one.
<lb/>–Not the <add place="sup">(black art)</add> Goetic but
Theurgic magic — <lb/>seeks converse with the Intelligence,
Power, the <lb/>Angel. </p>
</body>
</text>
XML representation of normalized transcript
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
<!ENTITY mdash "—" ><!--=em dash-->
<!ENTITY ndash "–" ><!--=en dash-->
<!ENTITY vbar "|" ><!--=vertical bar-->
<!ENTITY apos "—" ><!--=apostrophe-->
<!ENTITY triplebar "|" ><!--/triple horizontal bars-->
<!ENTITY ldquo "“" ><!--=double quotation mark, left-->
<!ENTITY uarr "↑" ><!--/uparrow A: =upward arrow-->
<!-- For HTML: -->
<!ENTITY and "<reg>and</reg>" >
<!ENTITY etc "<reg>etc</reg>" >
<!-- FOR Fof: -->
<!-- <!ENTITY and "amp" >
<!ENTITY etc "ampc" >-->
]>
<text>
<body>
<p>
<lb/>A seaman figures in The Canterbury Tales. <lb/>With many a
tempest had his beard been <lb/>shook. – Secret grief is a
<lb/>cannibal of its own heart – <emph>Bacon</emph>. <lb/>
<emph>Claudia</emph> of the <emph>Appian</emph> family, <q>I wish
<lb/> some fight or pestilence would thin out <lb/> this
crowd.</q>
<emph>Arrogance.</emph>
<lb/>
<emph>Roast beef in the pulpet.</emph></p>
<p>
<lb/>An animal of a man — <q>do eagles wear
<lb/>spectacles?</q>— Health. — <emph>Contrast</emph>:
an <lb/>over spiritual man. <lb/>
</p>
<p>
<lb/>
<q>Yes, Madam, Cain was a godless froward boy, ∧ <lb/>Reuben
(Gen:49) ∧ Absalom</q> Many pious men <lb/>have impious
children — (Devil as a Quaker) <lb/>
</p>
<p>
<lb/>A formal compact – Imprimis – First –
Second. <lb/>The aforesaid soul &etc; – Duplicates –
<lb/><q>How was it about the temptation on the <lb/>hill?</q>
&etc; <supplied> Conversation upon Gabriel, Michael and <lb/>
Raphael – gentlemanly &etc;</supplied> – D begs the
hero to form <lb/>one of a <emph>
<q>Society of D's</q>
</emph> – his name would be weighty <lb/> &etc; –
Leaves a letter to the D – <q>My <lb/>
<emph>Dear D</emph>
</q> – </p>
<p>
<lb/>
<q>Terra Oblivionis</q>
<q>Hellites</q> – At the Astor find him <lb/> making <choice>
<orig>almanacks</orig>
<reg>almanacs</reg>
</choice> – going to a ball takes a long <lb/>time making
toilette. – The Doctor's coach stops <lb/>the way.
– <q>Do you believe all that stuff? <lb/>
<choice>
<orig>nonsence</orig>
<reg>nonsense</reg>
</choice> – the world was never made. – “But <choice>
<orig>Is</orig>
<reg>is</reg>
</choice> not <lb/>this you mentioned <emph>here</emph> – in
the scriptures?</q>
<lb/>Receives visits from the principal d's –
<q>Gentlemen</q> &etc;. <lb/>
<emph>Arguments</emph> to persuade – <q>Would you not rather
<lb/>be below with kings than above with fools?</q>
</p>
<p>
<lb/>It is better to laugh ∧ not sin than to weep ∧ be
<lb/>wicked. — Ten loads of coal to burn him. —
<lb/>Brought to the stake — warmed himself by the fire. </p>
<p>
<lb/>Ego non baptizo te in <choice>
<orig>nominee</orig>
<reg>nomine</reg>
</choice> Patris et <lb/>Filii et Spiritus Sancti – sed in
nomine <lb/>Diaboli. — Madness is undefinable — <lb/>It
∧ right reasons extremes of one. <lb/>–Not the (black
art) Goetic but Theurgic magic — <lb/>seeks converse with the
Intelligence, Power, the <lb/>Angel. </p>
</body>
</text>
HTML presentation of exemplar for normalized transcript
Special exemplar tokens in red.
Figure 11: HTML presentation of exemplar for normalized
transcript
HTML presentation of exemplar for normalized transcript
Special exemplar tokens in red, type similar tokens in
blue.
Figure 12: HTML presentation of normalized transcript
FOF representation of normalized exemplar and
transcript
This list of references has not been updated
since 2018. References added later have been provided in
footnotes and in web links.
[André 1972] [André,
Jacques.] Règles et recommandation pour les éditions
critiques (Série latine). Paris: Société d'édition "Les
belles lettres," 1972. Collection des universités de France,
publiée sous le patronage de l'Association Guillaume Budé. [vi +]
48 pp.
[Carter 1952] Carter,
Clarence E. Historical editing. Bulletins of
the national archives, Number 7 [Washington, DC]: National Archives
and Records Service, August 1952. National Archives publication
number 53-4.
[Caton 2013] Caton, Paul.
Pure transcriptional encoding. Paper given at
Digital Humanities 2013, Lincoln, Nebraska.
[The Papers of William Penn] Dunn,
Mary Maples et al. The Papers of William
Penn, 5 vols. (Philadelphia: University of Pennsylvania
Press, 1981-1986).
[Grice 1975] Grice, H.P.
Logic and Conversation. In Syntax and
Semantics, edited by P. Cole and J. Morgan, vol.3,
Speech Acts. New York: Academic Press, 1975.
Reprinted as chapter 2 of his Studies in the Way of
Words. Cambridge, Mass.: Harvard University Press,
1989, pp. 22–40.
[Goodman] Goodman, Nelson.
Languages of Art. Hackett Publishing,
1976.
[Hayford et al. 1988] Hayford,
Harrison, Hershel Parker, and G. Thomas Tanselle, ed. Moby
Dick, or, The Whale. Vol. 7 of The Writings of
Herman Melville The Northwestern–Newberry
Edition. Evanston [Ill.]: Northwestern University Press; Chicago :
Newberry Library, 1988, rpt. 1994, 1997.
[Marcoux / Rizkallah 2007]
Marcoux, Yves, and Élias Rizkallah. Exploring
intertextual semantics: A reflection on attributes and
optionality. Paper given at Extreme Markup Languages®,
Montréal, 2007. Proceedings of Extreme Markup
Languages® 2007. On the Web at [http://conferences.idealliance.org/extreme/html/2007/Marcoux01/EML2007Marcoux01.html].
[Simons et al. 2005] Simons,
Gary, Scott O. Farrar, Brian Fitzsimons, William D. Lewis, D.
Terence Langendoen, and Hector Gonzalez. The semantics of
markup: Mapping legacy markup schemas to a common
semantics. In Proceedings of the 4th workshop
on NLP and XML (NLPXML-2004): held in cooperation with ACL- 04
Barcelona, Spain. Pp. 25-32. On the Web at [http://www.aclweb.org/anthology/W/W04/W04-0604.pdf] and
other locations. [doi:https://doi.org/10.3115/1621066.1621070].
[Sperberg-McQueen 2005]
Sperberg-McQueen, C. M. The meaning of OAI 2.0 Markup: An
exercise in markup interpretation. Unpublished fragment,
December 2005. On the Web at [http://www.w3.org/2004/04/em-msm/ioai.html].
[Sperberg-McQueen et al. 2002]
Sperberg-McQueen, C. M., David Dubin, Claus Huitfeldt, and Allen
Renear. Drawing inferences on the basis of markup.
Paper given at Extreme Markup Languages®, Montréal, 2002.
>Proceedings of Extreme Markup Languages®
2002. On the Web at [http://conferences.idealliance.org/extreme/html/2002/CMSMcQ01/EML2002CMSMcQ01.html].
[Stevens and Burg 1997] Stevens, Michael E. and Steven B.
Burg. Editing Historical Documents: A Handbook of
Practice. Walnut Creek, London, New Delhi: Alta Mira
Press, 1997.
[Tanselle 1989]
Tanselle, G. Thomas. A Rationale of Textual
Criticism. Philadelphia: University of Pennsylvania
Press, 1989. 104 pp.
[Vander Meulen / Tanselle 1999] Vander Meulen, David,
and G. Thomas Tanselle, A system of manuscript
transcription.Studies in Bibliography 52 (1999): 201-212.
[Boolos et al. 2007] George
Boolos, John P. Burgess, Richard C. Jeffrey. Computability
and logic. 5th ed. Cambridge, New York: Cambridge
University Press 2007. ISBN 9780521877527.
[2] Except for abstracts and slides from various conferences,
cf. below.
[3] We use the English term exemplar to
denote the document from or of which a transcript is made, in
preference to the term original, which
evokes confusing associations when the exemplar is itself a
transcript. In our usage, exemplar thus corresponds to what is
referred to in German as the Vorlage or in
French as the antigraphe.
[5] Our notion of transcriptional implicature is indeed
inspired by Grice's theory, but we make no claim as to the
similarities between our notion and Grice's theory of
conversational and conventional implicature.
[6] In other contexts, such as linguistics, music, or genetics,
transcription refers to different, though
related phenomena — see [Huitfeldt / Sperberg-McQueen 2008], pp
295-96.
[8] Many documents do of course contain non-textual material.
It may be argued that at least some kinds of non-textual
material lend themselves to a type-token analysis, but we do not
believe that always to be the case. Our account is limited to
the aspects of documents which do lend themselves to such
analysis, and we do not argue that all aspects of documents
do.
[9] The presupposition that exemplar as well as transcript lend
themselves to analysis in terms of the concepts of tokens and
types presented above implies that they must be
notations as Goodman defines that
term.
According to Goodman, the requirements of notational
schemes are:
...[A] character in a notation is an abstraction
class of character-indifference among inscriptions. As a
result, no mark may belong to more than one
character. [Goodman] pp. 132-3. In
our terminology: No token is an instantiation of more than
one type.
...the characters be finitely differentiated, or
articulate. [Goodman] p. 135. That
is, for any mark it must be at least theoretically possible
to determine whether or not it belongs to a certain sign
(character). In our terminology: For any token, it must be at
least theoretically possible to determine whether it
instantiates a certain type or not.
These are what Goodman calls syntactic
requirements that notational schemes have
to satisfy. In order for a notational scheme to qualify as a
notational system, it must also satisfy
three semantic requirements: It must be
unambiguous,disjoint and finitely
differentiated.
More precisely:
Every notational system must be unambiguous. In other
words, the extension (compliance-class) of a sign (character
or inscription) must be the same for all occurrences of that
sign – it cannot vary from case to case. [Goodman] p. 148.
Every extension (compliance-class) must be distinct
(disjoint) from every other extension. In other words, no
object in a domain can belong to two extensions. [Goodman] p. 150.
Every extension (compliance-class) must be finitely
differentiated. In other words, for any object in the domain
it must be at least theoretically possible to determine
whether or not it belongs to a certain extension. [Goodman] p. 152.
It would be interesting to investigate to what extent our
account of transcription implies that transcription is
notational not only in terms of the syntactic, but also in terms
of the semantic requirements. We think this may be the case, but
do not further investigate the issue as it is not of direct
relevance to the tasks we have set ourselves here.
[10] The reader will note that the notion of "mark" is very
vague; in view of the many different ways in which written
messages can be constructed or conveyed, we believe this to
be unavoidable. The notion is not defined formally in the
model.
[11] When the distinction is not clear from context, we use
single quotes for types and double quotes for tokens.
[13] The reader may feel free to read the T in
T-similarity to stand for text,
type, or transcriptional.
[14] The reader may wonder how rules can be represented by a
mathematical relation. The name rules suggests
imperative statements such as when a sentence is not
closed with a period, insert an extra space or
whenever appropriate, render "u" as "v," or "v" as
"u", more than it suggests a relationship. The answer
is that T-similarity represents
transcription rules in a descriptive manner. It specifies how exemplar and
transcript are to stand relative to each other once the
transcription is completed, the implicit rules being: do
whatever it takes for that relation to hold when you are
done.
Interestingly, in the statement of
practice accompanying some transcripts, the rules of
transcription are also formulated in descriptive
rather than imperative style: e.g., spelling is retained
as written or the long "s" is presented as a
short "s", rather than retain spelling as
written or convert any long "s" to a short
"s".
[16] For simplicity, we consider here only cases with a single
type repertoire.
[17] One might regard this as a principle of giving the author
the benefit of the doubt. Or perhaps one might just as well
regard it as an act of charity to the reader: Unless there is
clear evidence of error, there is no reason to bother the reader
with the mere possibility of error.
[19] Here and in the rest of this section, indented quotations
are quotations from pp. 15-18 in [The Papers of William Penn].
[20] The quotation continues: , with the following two
exceptions. Endorsements are treated as dockets, and entered
into the provenance note (see below). If a document is
undated, an initial date line is supplied [within square
brackets]. If a document is dated at the close but not at the
opening, an initial date line is supplied [within square
brackets], and the closing date line is
retained.
[21] Unfortunately we have only had access in facsimile to the
first and the last page of one letter, which are reproduced on
the insides of the front and back cover. These two pages, and
the corresponding transcript, are reproduced in appendix A.
[23] Except for such line breaks that occur in connection with
the phenomena mentioned in the previous paragraph.
[24] It is a common typographic practice to use "sic" to mark
misspellings and other irregularities which might otherwise be
blamed on the typesetter, and in some contexts punctuation and
paragraphing are routinely normalized, or introduced, by
editors.
[25] We note in passing that the statement is silent on what
happens in cases (if any) where the exemplar already contains
braces, or an extra space after a sentence not closed with a
period. If there are such cases, then their printed
representation is ambiguous: it could represent the discreet
editorial intervention described in the statement of practice,
or it could be a strictly literal transcript of the paragraph.
It seems likely that the editors know that there are no such
cases, and trust the reader to infer it. But it might also be
the case that the editors simply don't regard such ambiguities
as interesting or important enough to be worth avoiding.
[26] However, this may also be regarded as an application of the
rule of type similarity, cf. our discussion of statement 5.l
below.
[27] If, on the other hand, statement 5.b were taken to mean
that the writers of the manuscripts regarded the use of
capitalization as a matter for individual judgment rather than
orthographic system, resulting in usage that the modern reader
perceives as arbitrary and inconsistent, then the matter would
be similar to the case with u and
v, discussed below. (Readers today do find
the capitalization of seventeenth-century manuscripts
capricious, but we do not believe that is the point being made
in 5.b.)
[28] We take this to mean that "£" is
represented as "l" (not as "l.") — i.e., that the fact
that the full stop is placed before the closing quotation mark
is a somewhat misleading effect of punctuation rules.
[29] A glyph shaped like A is either uppercase
Latin letter A or uppercase Greek letter Alpha, depending on
context. A dot on the baseline may be a full stop (sentence
punctuation), an abbreviation marker, or a decimal point
depending on both the intratextual and the extratextual context.
In most European languages, ij,
ch, and ll are each a
two-character string, but the first is a single character (a
single token instantiating a single type) in Dutch, and each of
the others is taken as a single character in traditional Spanish
orthography, even though i, j,
c, h, and l also
occur individually in those languages.
[30] Once again, an alternative analysis might be that the
different type systems are employed for the transcript and the
exemplar. (In this case, the subset relationship would be the
opposite of the one for "u" and "v".) And once again, rules of
default transcriptional implicature would be broken in neither
case.
[32] The transcript appears to be complete with respect to
Melville's notes, but omits the printed text of Shakespeare
which appears in the same volume.
[33] An alternative interpretation would infer a type system in
which for any letter such as e, there are
companion types for inserted e,
canceled e, overwritten
e, and underlined [or italic]
e. Occam's Razor and convenience in the formalization
both lead us to prefer an account of this example in which
insertions, etc., are compound tokens of corresponding compound
types.
[34] The occurrence of a single x is in
fact perhaps strictly speaking ambiguous: statement [5] suggests
the interpretation one undeciphered letter, but
the only occurrence of a single x in the
transcript (in the word extremes on line 37)
seems more plausibly read as transcribing an
x in the exemplar.
[38] Technically, this is obtained by submitting to Vampire a
FOF-representation of the axioms 1-10 and the definitions a-d,
with T-similarity as a conjecture. Vampire's response is that
the conjecture is a theorem, i.e., that it is logically implied
by the other statements.
[39] Experts on TEI and critical editing will probably find our
encoding idiosyncratic and/or unsatisfactory. Moreover, any
serious TEI-based project today would probably make one base
transcription including all the information contained in both
the transcription discussed here (in Appendix C and the
normalized transcription in Appendix G. It should be
mentioned, though, that Michael made a more thorough TEI
representation of the transcript for a related project in 2017,
using a special TEI-based tag set for manuscript transcription
([http://mlcd.blackmesatech.com/mlcd/2015/W/tip/Melville/Melville-notes.sourcedoc.xml]).
[40] Technically, Micheal and
Raphel are not declared type similar to
Michael and Raphael, but that is
because they occur only as parts of a supplied
element in the transcript.
[41] It may be plausible to omit unreadable writing from a
transcript. However, in order to leave, for example, text in a
different hand or a different language, an addition or deletion
out from a transcript, the transcriber needs to recognize it as
such, and may thus very well be able to read it.
[André,
Jacques.] Règles et recommandation pour les éditions
critiques (Série latine). Paris: Société d'édition "Les
belles lettres," 1972. Collection des universités de France,
publiée sous le patronage de l'Association Guillaume Budé. [vi +]
48 pp.
Carter,
Clarence E. Historical editing. Bulletins of
the national archives, Number 7 [Washington, DC]: National Archives
and Records Service, August 1952. National Archives publication
number 53-4.
Grice, H.P.
Logic and Conversation. In Syntax and
Semantics, edited by P. Cole and J. Morgan, vol.3,
Speech Acts. New York: Academic Press, 1975.
Reprinted as chapter 2 of his Studies in the Way of
Words. Cambridge, Mass.: Harvard University Press,
1989, pp. 22–40.
Hayford,
Harrison, Hershel Parker, and G. Thomas Tanselle, ed. Moby
Dick, or, The Whale. Vol. 7 of The Writings of
Herman Melville The Northwestern–Newberry
Edition. Evanston [Ill.]: Northwestern University Press; Chicago :
Newberry Library, 1988, rpt. 1994, 1997.
Huitfeldt, Claus,
and C. M. Sperberg-McQueen. What is transcription?Literary & Linguistic Computing 23.3
(2008): 295-310. [doi:https://doi.org/10.1093/llc/fqn013].
Simons,
Gary, Scott O. Farrar, Brian Fitzsimons, William D. Lewis, D.
Terence Langendoen, and Hector Gonzalez. The semantics of
markup: Mapping legacy markup schemas to a common
semantics. In Proceedings of the 4th workshop
on NLP and XML (NLPXML-2004): held in cooperation with ACL- 04
Barcelona, Spain. Pp. 25-32. On the Web at [http://www.aclweb.org/anthology/W/W04/W04-0604.pdf] and
other locations. [doi:https://doi.org/10.3115/1621066.1621070].
Sperberg-McQueen, C. M. The meaning of OAI 2.0 Markup: An
exercise in markup interpretation. Unpublished fragment,
December 2005. On the Web at [http://www.w3.org/2004/04/em-msm/ioai.html].
George
Boolos, John P. Burgess, Richard C. Jeffrey. Computability
and logic. 5th ed. Cambridge, New York: Cambridge
University Press 2007. ISBN 9780521877527.