<?xml version="1.0" encoding="utf-8"?><article xmlns="http://docbook.org/ns/docbook" xmlns:xlink="http://www.w3.org/1999/xlink" version="5.0-subset Balisage-1.5">
 <title>Transcriptional implicature</title>
 <info>
    <confgroup>
      <conftitle>Balisage: The Markup Conference 2026</conftitle>
      <confdates>August 3-7, 2026</confdates>
   </confgroup>
  <abstract>
   <para>The concept of transcriptional implicature is proposed as a
    way of accounting for common practices in transcription while also
    making sense of variations in practice. A formal model implicit in
    default transcriptional implicature can be elaborated in different
    ways to reflect a number of commonalities as well as variations in
    transcription practice which have hitherto eluded precise
    description in formal terms. We regard the formalization of
    transcription as an essential step towards formalizing the
    semantics of colloquial XML vocabularies like TEI.</para>
  </abstract>
  <author>
   <personname>
    <firstname>C. M.</firstname>
    <surname>Sperberg-McQueen</surname>
   </personname>
   <personblurb>
    <para>C. M. Sperberg-McQueen<superscript>† 2024</superscript> was
     the founder and principal of Black Mesa Technologies, a
     consultancy specializing in helping memory institutions improve
     the long term preservation of and access to the information for
     which they are responsible.</para>
    <para>He served as editor in chief of the TEI Guidelines from 1988
     to 2000, and also served as co-editor of the World Wide Web
     Consortium's XML 1.0 and XML Schema 1.1 specifications. </para>
   </personblurb>
   <affiliation>
    <jobtitle>Founder and principal</jobtitle>
    <orgname>Black Mesa Technologies LLC</orgname>
   </affiliation>
  </author>

  <author>
   <personname>
    <firstname>Claus</firstname>
    <surname>Huitfeldt</surname>
   </personname>
   <personblurb>
    <para>Claus Huitfeldt works at the Department of Philosophy of the
     University of Bergen, Norway. He was founding Director
     (1990-2000) of the Wittgenstein Archives at the University of
     Bergen, for which he developed the text encoding system MECS as
     well as the editorial methods for the publication of
     Wittgenstein's Nachlass - The Bergen Electronic Edition (Oxford
     University Press, 2000).</para>
   </personblurb>
   <affiliation>
    <jobtitle>Professor</jobtitle>
    <orgname>University of Bergen</orgname>
   </affiliation>
   <email>Claus.Huitfeldt@uib.no</email>
  </author>
  <author>
   <personname>
    <firstname>Yves</firstname>
    <surname>Marcoux</surname>
   </personname>
   <personblurb>
    <para>Yves Marcoux has been a faculty member at EBSI, University
     of Montréal, from 1991 to 2018, then adjunct professor and
     lecturer until 2025. He was mainly involved in teaching,
     research, standardization, and international cooperation
     activities in the field of document informatics. Prior to his
     appointment at EBSI, Dr. Marcoux worked for 10 years in systems
     maintenance and development, in Canada, the U.S., and Europe. He
     obtained his Ph.D. in theoretical computer science from
     Université de Montréal in 1991. His main research interests are
     intertextual semantics, the design of communication, markup
     languages and digital humanities.</para>
   </personblurb>
   <affiliation>
    <jobtitle>Honorary Professor (Professeur honoraire)</jobtitle>
    <orgname>École de bibliothéconomie et des sciences de
     l’information, Université de Montréal</orgname>
   </affiliation>
  </author>
<legalnotice><para>Copyright © 2026 by the authors. Used with permission.</para></legalnotice>
 </info>
 <blockquote>
  <title>Foreword</title>
  <para>The publication of this paper is meant as
   <!--must be understood first and foremost as--> a tribute to the
   late C. Michael Sperberg-McQueen by the coauthors. In their mind,
   the paper represents the best possible "graceful" conclusion they
   managed to produce of the work on transcription they carried out
   with Michael over a span of more than two decades (a large portion
   of which the third author was absent from the team).</para>
  <para>While we by no means consider ourselves in a position to do
   justice to the numerous enticing ideas and practical developments
   contributed by Michael on the topic, we have felt it important to
   at least try our best to make known some of the insights, hopes,
   and the enthusiasm he had regarding the possibilities of
   formalization for furthering the understanding of theoretical and
   practical aspects of transcription.</para>
  <para>Whatever good ideas may be found in this account stem in one
   way or another from Michael. Only we, however, must bear the
   responsibility for the many shortcomings that may remain. Their
   number would have been much smaller had we been able, as we
   repeatedly found ourselves wishing, to rely on the wit and
   benevolence of our friend to keep us from going astray.</para>

  <para> Our first formal publication on this topic came out in 2008
    [<xref linkend="wit1"/>], followed by a Balisage paper in 2010
    [<xref linkend="ettdds"/>]. The notion of <quote>transcriptional
    implicature</quote> is partly based on (and what we saw as a
   low-hanging fruit from) ideas from these earlier publications. What
   is presented here was more or less finished some time around 2015,
   and last touched in 2018.<footnote>
    <para>See [<link xlink:href="http://mlcd.blackmesatech.com/mlcd/2015/W/tip/transcriptional_Implicature_2018.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">2018 version of this paper</link>].</para>
   </footnote> However, it was put on the back burner and has never
   been formally published until now.<footnote>
    <para>Except for abstracts and slides from various conferences,
     cf. below.</para>
   </footnote>
  </para>
  <para>At the time, we regarded transcriptional implicature as only
   one of the parts of a more ambitious and comprehensive "formal
   account", or a theory of "the logic", of transcription.
   Regrettably, we did not manage to accomplish this quite demanding
   task before Michael's untimely death.
   <!--Without him, we realize we
   will not be able to finish.-->
  </para>

  <para>Materials from that later period, going beyond what we present
   here, will hopefully be made available as part of the [<link xlink:href=" https://dh.phil-fak.uni-koeln.de/en/research/cmsmcq" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Sperberg-McQueen Nachlass project</link>].</para>

  <para>We would like to thank participants of a number of conferences
   through the years, such as DH2009, Balisage2010, and DH2014, for
   their interest, questions, and suggestions.</para>

  <para>In particular, we want to thank Michael's wife, Marian
   Sperberg-McQueen, for encouraging the posthumous publication of
   this paper.</para>
<attribution>Claus Huitfeldt &amp; Yves Marcoux</attribution>
  
 </blockquote>
 <section>
  <title>Introduction</title>
  <section>
   <title>The notion of transcriptional implicature</title>
   <para>There is hardly any universal transcription practice: for
    every generalization we find exceptions. Is everything in the exemplar<footnote>
     <para>We use the English term <emphasis>exemplar</emphasis> to
      denote the document from or of which a transcript is made, in
      preference to the term <emphasis>original</emphasis>, which
      evokes confusing associations when the exemplar is itself a
      transcript. In our usage, exemplar thus corresponds to what is
      referred to in German as the <emphasis>Vorlage</emphasis> or in
      French as the <emphasis>antigraphe</emphasis>.</para>
    </footnote> transcribed? Not when deletions and irrelevant
    material are excluded. Does everything in the transcript reproduce
    some word or character in the exemplar? Not when line breaks are
    marked explicitly with vertical bars, or notes are added.<footnote>
     <para>See, for example, [<xref linkend="Carter"/>], [<xref linkend="tanselle1989"/>], [<xref linkend="StevensAndBurg"/>],
       [<xref linkend="vmt1999"/>], [<xref linkend="bath"/>].</para>
    </footnote> Many scholarly editions account for variations like
    these in an explicit statement of transcription practice. Such
    statements typically describe <emphasis>deviations</emphasis> from
    the usual practice, but rarely the ways in which it
     <emphasis>exemplifies</emphasis> usual practice. Common practice
    may sometimes be felt to be so obvious that it needs no mention or
    explanation.</para>
   <para>By <emphasis>transcriptional implicature</emphasis> for a
    given community, we mean informally <quote>the things, suggested
     or entailed by the rules of transcription, that members of that
     community may find it unnecessary to mention explicitly</quote>. </para>
   <para>Different communities of transcription practice have
    different sets of tacit assumptions and thus different rules of
    transcriptional implicature. Is there a common core of
    transcriptional practice shared by all communities? There may be;
    it is an empirical question. A definitive answer would require
    more detailed studies of a wider variety of communities of
    practice than we can currently manage. Our hypothesis, however, is
    that there is such a common core, at least in the sense that the
    transcriptional implicature of any community of practice can be
    described with reference to some common set of rules.</para>
   <para>If (as we conjecture) there is such a thing as the
    transcriptional implicature of a given community, then the
    transcriptional practice of any project of that community could be
    described <!--documented--> by listing the ways in which it
     <quote>deviates</quote> from the transcriptional implicature.
    Such <quote>deviations</quote> can be addition, modification, or
    withdrawal of rules. Thus, the rules of the transcriptional
    implicature are defeasible for particular projects. They apply,
    except if explicitly excluded.</para>
   <para>If (as we conjecture) the transcriptional implicature of a
    community can in turn be described as a set of deviations from
    some default transcriptional implicature, then it follows that any
    project's transcription practice could be described with reference
    to the default transcriptional implicature, by merging the list of
    ways in which the community’s implicature deviates from the
    default implicature with the list of ways in which the project’s
    practice deviates from the default community practice. </para>
   <para><!--YMA -->The (conjectured) default transcriptional
    implicature is a set of rules which apply by default, but which
    may be overridden in particular cases, analogous to the rules of
    conversational implicature proposed by H. P. Grice as a way of
    explicating the logic of everyday conversation [<xref linkend="grice"/>].<footnote>
     <para>Our notion of transcriptional implicature is indeed
      inspired by Grice's theory, but we make no claim as to the
      similarities between our notion and Grice's theory of
      conversational and conventional implicature.</para>
    </footnote></para>
  </section>
  <section>
   <title>Earlier work</title>
   <para>Since most digital transcriptions are represented today using
    markup (e.g., TEI), providing a <emphasis>formal</emphasis> model
    of transcription might prove useful for work on the semantics of
    markup. Indeed, prevalent approaches to markup semantics are based
    on formal languages like first-order predicate logic [<xref linkend="mim"/>], which would <emphasis>a priori</emphasis> blend
    well with a formal model of transcription, yielding a framework in
    which meaning can be assigned to markup constructs whose
    colloquial definitions involve transcription, for example, 
     <emphasis>p</emphasis> elements defined as
     <emphasis>transcription of a text block in a
     manuscript</emphasis>.</para>
   <para>Some have proposed to explicate the
     <emphasis>meaning</emphasis> of markup by specifying, for each
    construct in a markup vocabulary, a sentence schema in a natural
    language, with blanks to be filled in with data from the document
     [<xref linkend="is2006"/>]; others make a similar proposal but
    allow sentence schemata in formal languages like first-order
    predicate logic as well [<xref linkend="mim"/>]. This appears
    straightforward, although far from trivial, for metadata [<xref linkend="fotbo"/>] and perhaps even for born-digital texts, but
    how shall the meaning of a <code>p</code> element be formalized in
    a markup language which defines it as containing a
     <emphasis>transcription of a text block in a
     manuscript</emphasis>? What does it <emphasis>mean</emphasis> for
    a document to be a transcription of another document? Laying a
    formal ground on which such questions can be tackled, may be seen
    as contributions to the semantics of markup, as applied in
    transcription.</para>
   <para>Early work [<xref linkend="wit1"/>] has explored the nature
    of the similarity between transcripts and their exemplars. If
    documents are defined as sequences of characters, then perhaps the
    similarity consists simply in the exemplar and the transcript
    containing the same sequence of characters? This can be formalized
    but proves disappointing, partly because the definition of
    documents as sequences of characters omits text structures like
    division into paragraphs and partly because the model offers no
    way of describing disagreements among transcribers about how to
    read the exemplar, or about which character distinctions (e.g.
    i/j, u/v, s/ſ) to retain and which to level. It is also
    wrong for the reason that few transcripts have
     <emphasis>exactly</emphasis> the same character sequence as their
    exemplar. </para>
   <para>Later work has extended the model by introducing boolean
    (disjunctive and conjunctive) types, which allowed explicit
    representation of ambiguity, as well as
     <emphasis>readings</emphasis>, which allowed modeling transcriber
    agreement and disagreement explicitly [<xref linkend="wit2"/>],
    and lifting the analyses from characters to higher-level textual
    structures [<xref linkend="ettdds"/>].</para>
   <para>In this paper, we propose to assess the notion of
    transcriptional implicature using a formal approach, in the line
    of our earlier work on transcription [<xref linkend="wit1"/>] and
     [<xref linkend="ettdds"/>].</para>
  </section>
 </section>

 <section>
  <title>Outline of the formal model</title>
  <para>Formalization seems interesting to us for two reasons: First,
   and most importantly, formalization encourages a level of
   explicitness and rigor which may help us better understand what
   transcription is, and thereby also inform our discussions and
   judgments in concrete cases. Second, now that transcription is
   mostly done digitally, a formal model may help us build better
   applications and better understand their possibilities and
   limitations. </para>
  <para>Our goal is simply to describe formally the relationship that
   exists between two documents when the one is said to be a
   transcription of the other. It is <emphasis>not</emphasis> our goal
   to establish normative criteria which may serve to distinguish good
   from bad or right from wrong in transcription, nor to construct a
   theory of the psychological or physical processes involved in
   transcription. </para>


  <para>The formal model we use is that of [<xref linkend="ettdds"/>],
   but we present just enough technical details to be able to flesh
   out our discussions of transcriptional implicature. For further
   details, see [<xref linkend="wit1"/>] and [<xref linkend="ettdds"/>].</para>

  <section>
   <title>What is formalized</title>
   <para>The term <emphasis>transcription</emphasis> may be used
    variously to refer to the <emphasis>act</emphasis> or
     <emphasis>process</emphasis> of transcribing a document, to the
    physical product of that act (that is, another document), or to
    the <emphasis>relation</emphasis> between the two documents.<footnote>
     <para>In other contexts, such as linguistics, music, or genetics,
       <quote>transcription</quote> refers to different, though
      related phenomena — see [<xref linkend="wit1"/>], pp
      295-96.</para>
    </footnote>
   </para>
   <para>Our formal approach focuses on the
     <emphasis>relation</emphasis> that must obtain between two
    documents for the one to be a transcription of the other. Where
    such a relation obtains, we call one of the documents a
     <emphasis>transcript</emphasis>, and the other an
     <emphasis>exemplar</emphasis>. </para>
   <para>We take as basic facts about this relation that:<itemizedlist>
     <listitem>
      <para>any exemplar <emphasis>antedates</emphasis> its
       transcript,</para>
     </listitem>
     <listitem>
      <para>the transcript has been made by someone for certain
       reasons and with certain
       <emphasis>intentions</emphasis>,</para>
     </listitem>
     <listitem>
      <para>it has been made on the basis of
        <emphasis>access</emphasis> to the exemplar (or facsimile of
       the exemplar), and</para>
     </listitem>
     <listitem>
      <para>the text of the transcript is the same as, or at least
        <emphasis>similar</emphasis> in some way to, the text of the exemplar.<footnote>
        <para>See also [<xref linkend="wit1"/>], p. 296.</para>
       </footnote></para>
     </listitem>
    </itemizedlist></para>
   <para>Of all those aspects of the relation between exemplar and
    transcript, our formalization addresses only the last one, that
    is, the <emphasis>similarity relation</emphasis> that is assumed
    to obtain between the texts of a transcript and of its
    exemplar.</para>
   <para>The kinds, levels or degrees of text similarity required for
    a document to be counted as a transcription of another obviously
    vary from context to context, and it is not hard to find examples
    of disagreement about the matter. The formal model must be
    flexible enough to allow for such variations.</para>
  </section>
  <section>
   <title>The basic notions</title>
   <para>The main idea in our approach to formalization is that
    documents are textual objects<footnote>
     <para>Many documents do of course contain non-textual material.
      It may be argued that at least some kinds of non-textual
      material lend themselves to a type-token analysis, but we do not
      believe that always to be the case. Our account is limited to
      the aspects of documents which do lend themselves to such
      analysis, and we do not argue that all aspects of documents
      do.</para>
    </footnote> and, as such, can be analyzed in terms of a theory of
    types and tokens.<footnote>
     <para>The presupposition that exemplar as well as transcript lend
      themselves to analysis in terms of the concepts of tokens and
      types presented above implies that they must be
       <emphasis>notations</emphasis> as Goodman defines that
      term.</para>
     <para>According to Goodman, the requirements of notational
      schemes are: <itemizedlist>
       <listitem>
        <para><quote>...[A] character in a notation is an abstraction
          class of character-indifference among inscriptions. As a
          result, no mark may belong to more than one
          character.</quote> [<xref linkend="Goodman"/>] pp. 132-3. In
         our terminology: No token is an instantiation of more than
         one type.</para>
        <!--<note>
                  <para>A sign (character) consists of a maximal class of sign-equivalent
                    (character-indifferent) inscriptions. This means that no inscription
                      (<quote>mark</quote>) can belong to two signs (characters), or in
                    other words, that signs are distinct. [<xref linkend="Goodman"/>] p.
                    132.</para>
                </note>-->
       </listitem>
       <listitem>
        <para><quote>...the characters be finitely differentiated, or
          articulate.</quote> [<xref linkend="Goodman"/>] p. 135. That
         is, for any mark it must be at least theoretically possible
         to determine whether or not it belongs to a certain sign
         (character). In our terminology: For any token, it must be at
         least theoretically possible to determine whether it
         instantiates a certain type or not.</para>
        <!--<note>
                  <para>Every sign (character) must be finitely differentiated, or
                    articulate. For any mark it must be at least theoretically possible to
                    determine whether or not it belongs to a certain sign
                    (character).</para>
                </note>-->
       </listitem>
      </itemizedlist>
     </para>
     <para>These are what Goodman calls <emphasis>syntactic</emphasis>
      requirements that notational <emphasis>schemes</emphasis> have
      to satisfy. In order for a notational scheme to qualify as a
      notational <emphasis>system</emphasis>, it must also satisfy
      three <emphasis>semantic</emphasis> requirements: It must be
       <emphasis>unambiguous,</emphasis>
      <emphasis>disjoint</emphasis> and <emphasis>finitely
       differentiated</emphasis>.</para>
     <para>More precisely: <itemizedlist>
       <listitem>
        <para>Every notational system must be unambiguous. In other
         words, the extension (compliance-class) of a sign (character
         or inscription) must be the same for all occurrences of that
         sign – it cannot vary from case to case. [<xref linkend="Goodman"/>] p. 148.</para>
       </listitem>
       <listitem>
        <para>Every extension (compliance-class) must be distinct
         (disjoint) from every other extension. In other words, no
         object in a domain can belong to two extensions. [<xref linkend="Goodman"/>] p. 150.</para>
       </listitem>
       <listitem>
        <para>Every extension (compliance-class) must be finitely
         differentiated. In other words, for any object in the domain
         it must be at least theoretically possible to determine
         whether or not it belongs to a certain extension. [<xref linkend="Goodman"/>] p. 152.</para>
       </listitem>
      </itemizedlist>
     </para>
     <para>It would be interesting to investigate to what extent our
      account of transcription implies that transcription is
      notational not only in terms of the syntactic, but also in terms
      of the semantic requirements. We think this may be the case, but
      do not further investigate the issue as it is not of direct
      relevance to the tasks we have set ourselves here.</para>
    </footnote> In turn, the <emphasis>similarity</emphasis>
    relationship that must obtain between two documents for one to be
    a transcript of the other (we will call it
     <emphasis>T-similarity</emphasis>) is defined in terms of
    type-token analyses of the documents.</para>
   <!--<note place="block"><para>Back to the original text.</para></note>-->
   <para>Here are the basic notions on which the current account is based.<itemizedlist>
     <listitem>
      <para><emphasis role="bold">document</emphasis>: A document (in
       particular, an <emphasis>exemplar</emphasis> or a
        <emphasis>transcript</emphasis>) is a physical phenomenon
       containing or exhibiting <emphasis>marks</emphasis>. (For
       example, the copy of Moby Dick on the third author's
       nightstand, the Wittgenstein notebook cataloged as MS 108 in
       the Wren library.)</para>
     </listitem>
     <listitem>
      <para><emphasis role="bold">mark</emphasis>: A mark is a
       perceptually discernible arrangement of physical reality in a
       document; some (and perhaps, but not necessarily, all) marks
       are <emphasis>tokens</emphasis> (for example, words written in
       pen or pencil in a manuscript, or the printed words, fly specks
       and pencil marks of an old book).<footnote>
        <para>The reader will note that the notion of "mark" is very
         vague; in view of the many different ways in which written
         messages can be constructed or conveyed, we believe this to
         be unavoidable. The notion is not defined formally in the
         model.</para>
       </footnote></para>
     </listitem>
     <listitem>
      <para><emphasis role="bold">token</emphasis>: A token is a
        <emphasis>mark</emphasis> determined to instantiate a
       particular <emphasis>type</emphasis> while reading the
       document. (For example, a letter, punctuation mark, a word or a
       sentence on a piece of paper or parchment.) Tokens can be basic
       (atomic) or compound. A compound token is one that is read as
       the composition of smaller tokens, called its subtokens. A
       whole document is viewed as a single "highest-level" compound
       token, i.e., one that is not part of a larger compound
       token.</para>
     </listitem>
     <listitem>
      <para><emphasis role="bold">type</emphasis>: A
        <emphasis>type</emphasis> is a character, letter, word,
       sentence, chapter, or any other object instantiated by one or
       more tokens in one or more documents. (For example, the
       character 'A' of the Latin alphabet, or the sentence 'I know
       Verona'.) Like tokens, types can be basic or compound. A type
       entering in the composition of a compound type is said to be a
        <emphasis>subtype</emphasis> of the latter.</para>
      <para>For example, the written word "cat" might be read as a
       compound token instantiating the compound word type 'cat', and
       composed of the three basic subtokens "c", "a", and "t",
       instantiating respectively the basic letter types 'c', 'a', and
       't', which are subtypes of the word type 'cat'.<footnote>
        <para>When the distinction is not clear from context, we use
         single quotes for types and double quotes for tokens.
         <!--The small number of examples in this paper did not
         seem to us to warrant the burden of establishing a formal
         notational convention.--></para>
       </footnote></para>
      <para>Documents (including exemplars and transcripts), which are
       highest-level compound tokens, each instantiate a
       (highest-level) compound type.</para>
      <para>Types can also be <emphasis>conjunctive</emphasis> or
        <emphasis>disjunctive</emphasis>, corresponding to ambiguous
       tokens. For example, in a manuscript where it is not possible
       to determine whether some token is of type 'i' or 'j', yet
       clear that it is not of any other type, the token can be said
       to be of the disjunctive type 'i/j'. In some documents, as for
       example so-called <emphasis>ambigrams</emphasis>, certain
       tokens are meant to be read in more than one way. Those tokens
       can then be regarded as conjunctive types.<footnote>
        <para>The notions of disjunctive and conjunctive types were
         introduced in [<xref linkend="ettdds"/>].</para>
       </footnote>
      </para>
     </listitem>
     <listitem>
      <para><emphasis role="bold">type system</emphasis>: A type
       system is a set of type repertoires, which in turn are sets of
       types. For example, a type system may contain:<itemizedlist>
        <listitem>
         <para>a repertoire of characters used in 17th-century English
          manuscripts, </para>
        </listitem>
        <listitem>
         <para>a repertoire of words used in the English of that
          period, </para>
        </listitem>
        <listitem>
         <para>a repertoire of sentences formed using those words, and
         </para>
        </listitem>
        <listitem>
         <para>repertoires of paragraphs, sections, and other textual
          structures.
          <!--recognized and reproduced by a given project--></para>
        </listitem>
       </itemizedlist>All tokens instantiate at least one type in at
       least one repertoire. The same token cannot instantiate more
       than one type in any given repertoire. Strictly speaking, a
       token may represent different types in different repertoires.
       For example, the token "I" in English may instantiate both a
       character, a word, and a sentence, but not more than one of
       each. However, for the sake of convenience, we will usually
       speak of "the type" (singular) of any given token.</para>
     </listitem>
    </itemizedlist></para>
   <para>The model allows for constraints to be defined on which types
    can be instantiated by a compound token, in terms of the subtypes
    instantiated by its subtokens. For example, it could be that any
    compound token must be assigned a compound type such that the
    subtokens of the compound tokens instantiate the subtypes of the
    compound type (as was the case in the 'cat' example above).</para>
  </section>

  <section>
   <title>Identification of tokens and types</title>
   <!--I prefer exclusions, because it also talks of naming stuff.-->
   <para>The process of reading of a document, in particular
    identifying the tokens in the document and the type that each of
    them instantiates, is outside the scope of our model. Whenever we
    refer to a document, we will allow ourselves to refer also to its
    tokens and to the type of each of these tokens, without further
    explanation.</para>
   <para><!--YMA 2026-05-31 I like this paragraph, so I retrieved it from 2025/W/TransImp/transcriptional_Implicature_2025_CH.xml, uncommented it, and placed it the best place I found (here!).-->Like
    most formal models, our work assumes that names can be assigned to
    individuals in the universe of discourse, in particular tokens and
    types. The practical issues involved in reaching agreement on the
    identity of individuals and the names to be used for them are
    outside the scope of the formalization. When individuals are as
    numerous as the tokens in a novel, of course, agreement on names
    for them all is likely to present practical difficulties.</para>
   <para>The absence from our account of any description of a method
    for assigning names to the tokens of a document reflects not the
    belief that such an operation is easy, only the fact that it has
    no bearing on the logical structure of transcription, or of
    transcriptional implicature. Like the process of reading a
    document in the first place, the process of identifying the tokens
    in a document sufficiently well to enable productive disagreement
    about how to read them is of great practical importance but
    outside the focus of our work.</para>
   <!--YMA 2026-06-15 I am moving a very long comment that used to be here to the end of the document.-->
  </section>

  <section xml:id="T-similarity">
   <title>T-similarity</title>
   <para>The model defines a relation between documents called
     <emphasis>T-similarity</emphasis>.<footnote>
     <para>The reader may feel free to read the <quote>T</quote> in
       <quote>T-similarity</quote> to stand for <quote>text</quote>,
       <quote>type</quote>, or <quote>transcriptional</quote>.</para>
    </footnote> Two documents are T-similar if and only if, as
    high-level tokens, they are either of the same (high-level) type,
    or of <emphasis>similar</emphasis> types, where similarity among
    types is defined in a <emphasis>ad hoc</emphasis> manner, in the
    context of some given transcription project or set of
    projects.</para>
   <para>T-similarity is the relation that captures <!--the-->
    transcription<!-- practice-->. Its formal definition reflects the
    rules of transcription.<footnote>
     <para>The reader may wonder how rules can be represented by a
      mathematical relation. The name <quote>rules</quote> suggests
      imperative statements such as <quote>when a sentence is not
       closed with a period, insert an extra space</quote> or
       <quote>whenever appropriate, render "u" as "v," or "v" as
       "u"</quote>, more than it suggests a relationship. The answer
      is that <emphasis role="ital">T-similarity</emphasis> represents
      transcription rules in a <emphasis role="ital">descriptive</emphasis> manner. It specifies how exemplar and
      transcript are to stand relative to each other once the
      transcription is completed, the implicit rules being: <quote>do
       whatever it takes for that relation to hold when you are
       done</quote>.</para>
     <para>Interestingly, in the <emphasis role="ital">statement of
       practice</emphasis> accompanying some transcripts, the rules of
      transcription are also <!--usually--> formulated in descriptive
      rather than imperative style: e.g., <quote>spelling is retained
       as written</quote> or <quote>the long "s" is presented as a
       short "s"</quote>, rather than <quote>retain spelling as
       written</quote> or <quote>convert any long "s" to a short
       "s"</quote>.</para>
    </footnote>
   </para>
   <para>In the simplest case, T-similarity is based on type equality;
    in other words, exemplar and transcript must be of
     <emphasis>identical</emphasis> types, not just similar types.
    However, the definition can be relaxed, and this is how we can
    hope to capture reasonably realistic aspects of the similarity
    between exemplar and transcript in actual projects. Allowing for
    such loosening of the similarity conditions is an essential aspect
    of our model. We will later further relax the similarity
    conditions with the introduction of <emphasis>special</emphasis>
    tokens.</para>
   <para>The intent behind T-similarity (with suitable versions of
    similarity among types) is that it corresponds to the kind of
    similarity expected to obtain in practice between an exemplar and
    a corresponding transcript, as discussed above. We do not expect,
    nor claim, that T-similarity can capture all possible views of
    exemplar-transcript similarity, but we suggest it represents a
    minimal <emphasis>necessary condition</emphasis>. Also, as already
    mentioned, similarity of text is but one aspect of the
    relationship between exemplar and transcript in real life:
    normally a transcript must follow temporally after its exemplar,
    be made in the presence of the exemplar, and be the result of an
    intentional effort.</para>
   <para>As a mathematical relation, T-similarity is reflexive,
    transitive, and symmetric. The irreflexive, intransitive, and
    asymmetric nature of the factual relation observed between
    transcripts and their exemplars appears to us to be a consequence
    of contingent facts and not a property of the similarity relation
    holding between the documents.</para>
   <para>We have informally defined transcriptional implicature for a
    given community as <quote>the things, suggested or entailed by the
     rules of transcription, that members of that community may find
     it unnecessary to mention explicitly</quote>. From the point of
    view of our formal model, which focuses on the relation between
    exemplar and transcript, this amounts to the things
     <emphasis>about the T-similarity relation</emphasis> that
     <quote>go without saying</quote>.</para>
  </section>
  <section>
   <title>Document pairs</title>
   <para>We propose to characterize T-similarity as a set of formal
    statements expressed in first order <!--predicate--> logic (FOL),
    with variables ranging over the types and tokens of the related
    documents. "Things" about the T-similarity relation (including
    those that purportedly <quote>go without saying</quote>) can then
    be formally expressed as such statements.</para>
   <para>Such a representation of T-similarity can be put to work in
    various ways. For instance, given a pair of documents, we can seek
    to formally establish whether they stand in a T-similarity
    relationship or not. This will be done for a simple pair of
    documents in [<link xlink:href="#B" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">appendices</link>].</para>
   <para>Or, given one known document (say, a transcript) and assuming
    that it is T-similar to some other document (an exemplar), we can
    ask ourselves what can be inferred about that other document on
    the basis of the known document. Exploratory work in this
    direction have been described in Claus Huitfeldt and C. M.
    Sperberg-McQueen, <quote>transcriptional Implicature: Using a
     Transcript to Reason about an Exemplar</quote>
    <emphasis>in </emphasis> [<link xlink:href="https://dh2017.adho.org/abstracts/DH2017-abstracts.pdf" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Digital Humanities 2017 Conference Abstracts</link>], p. 266,
     ([<link xlink:href="http://blackmesatech.com/2017/08/MLCD/" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest"/>).</para>
   <para>To do those things, we need to be able to
     <emphasis>represent</emphasis> pairs of documents in FOL. The
    goal of this section is to introduce FOL elements – predicates and
    constants – that can be used to "say things" about pairs of
    documents, one exemplar and one transcript, in terms of their
    types and tokens.</para>
   <para>In the first-order formulae below, we use the following
    notation:<!--<lb/>--> The variables X, Y, Z, etc. range over
    all individuals in the<!--domain--> universe of discourse (i.e., 
    the types and tokens of the related documents). Each individual
    carries one (and only one) of the predicates:<itemizedlist>
     <listitem>
      <para>type(X): X is a type</para>
     </listitem>
     <listitem>
      <para>e_token(X): X is a token in the exemplar</para>
     </listitem>
     <listitem>
      <para>t_token(X): X is a token in the transcript</para>
     </listitem>
    </itemizedlist></para>
   <para>We define the following binary relations (predicates):</para>
   <itemizedlist>
    <listitem>
     <para>typeof(X,Y) ⇔ the type of X is Y.</para>
     <!--<para>Note that: (&forall;&X;, &Y;)[&typeof;(&X;,&Y;) &implies; ((&e_token;(&X;)
            &or; &t_token;(&X;)) &and; &type;(&Y;)].</para>-->
    </listitem>
    <listitem>
     <para>transcript(X,Y) ⇔ token Y in the transcript
      corresponds to token X in the exemplar.</para>
     <!--<para>Note that: (&forall;&X;, &Y;)[&transcript;(&X;,&Y;) &implies;
            &e_token;(&X;) &and; &t_token;(&Y;)]</para>-->
    </listitem>
    <listitem>
     <para>exemplar(X,Y) ⇔ token Y in the exemplar
      corresponds to token X in the transcript.</para>
     <!--<para>Note that: (&forall;&X;, &Y;)[&exemplar;(&X;,&Y;) &implies; &t_token;(&X;)
            &and; &e_token;(&Y;)]</para>-->
    </listitem>
   </itemizedlist>
   <para>As already mentioned, we are silent on how tokens are
    identified and assigned types. We are also silent on how the
    correspondences between exemplar and transcript tokens represented
    by the transcript() and exemplar() predicates are established.<footnote>
     <para>For a systematic way of defining such a correspondence, see
       [<xref linkend="ettdds"/>].</para>
    </footnote></para>
   <para>Finally, we introduce the following two individual constants: <itemizedlist>
     <listitem>
      <para><emphasis role="ital">e</emphasis>: the exemplar (as a single, usually compound,
       token)</para>
     </listitem>
     <listitem>
      <para><emphasis role="ital">t</emphasis>: the transcript (as a single, usually compound,
       token)</para>
     </listitem>
    </itemizedlist></para>
   <para>Note that we have:<itemizedlist>
     <listitem>
      <para>e_token(<emphasis role="ital">e</emphasis>), and</para>
     </listitem>
     <listitem>
      <para>t_token(<emphasis role="ital">t</emphasis>).</para>
     </listitem>
    </itemizedlist></para>
   <para>It should be clear that facts or hypotheses (to be proved or
    disproved) about pairs of documents can be formulated using the
    above apparatus. In particular, an exemplar-transcript pair can be
    entirely represented, in terms of its types and tokens, as a
    series of statements using the above predicates together with
    additional constants standing for the involved types and tokens.
    An example is given in [<link xlink:href="#F" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Appendices</link>].</para>
   <para>As previously mentioned, we do not prescribe nor describe any
    method for naming the constants corresponding to types and tokens,
    we simply take for granted the use of some adequate naming
    scheme.</para>
   <!--<para>Also, by hypothesis:<itemizedlist><listitem><para>&transcript;(&e;,&t;), and thus</para></listitem><listitem><para>&exemplar;(&t;,&e;).</para></listitem></itemizedlist></para>-->
  </section>
  <section>
   <title>Axiomatization</title>
   <para>The above definitions can be expressed as the following set
    of axioms that document pairs must satisfy.</para>


   <orderedlist>
    <title>Axioms</title>
    <listitem>
     <para><emphasis>transcript domain:</emphasis> For every X and
      Y, if X is the transcript of Y, then X is a t_token
      and Y is an e_token. </para>
     <para><code> ∀(X,Y)(transcript(X,Y) →
       (e_token(X) ∧ t_token(Y)))</code></para>
    </listitem>
    <listitem>
     <para><emphasis>exemplar domain:</emphasis> For every X and
      Y, if X is the exemplar of Y, then X is an e_token
      and Y is a t_token. </para>
     <para><code> ∀(X,Y)(exemplar(X,Y) →
       (t_token(X) ∧ e_token(Y)))</code></para>
    </listitem>
    <listitem>
     <para><emphasis>typeof domain:</emphasis> For every X and Y,
      if the type of X is Y then X is a t_token or an
      e_token, and Y is a type. </para>
     <para><code> ∀(X,Y)(typeof(X,Y) →
       ((t_token(X) ∨ e_token(X)) ∧
       type(Y)))</code></para>
    </listitem>
    <listitem>
     <para><emphasis>no class overlap:</emphasis> Every individual X
      is either an e_token or a t_token or a type. </para>
     <para><code> ∀(X)( (e_token(X) → (
       ¬t_token(X) ∧ ¬type(X))) ∧
      </code></para>
     <para><code> (t_token(X) → ( ¬e_token(X) ∧
       ¬type(X))) ∧ </code></para>
     <para><code> (type(X) → ( ¬e_token(X) ∧
       ¬t_token(X))))</code></para>
    </listitem>
    <listitem>
     <para><emphasis>Exemplar and transcript
      inverse:</emphasis><!--<lb/>--> For every X and Y, if X is
      the exemplar of Y, then Y is the transcript of
      X.</para>
     <para><code> ∀(X,Y)(exemplar(X,Y) ⇔
       transcript(Y,X))</code></para>
    </listitem>
    <listitem>
     <para><emphasis>At most one type per
      token:</emphasis><!--<lb/>--> For every X and Y, if the type
      of X is Y then no other individual is the type of X. </para>
     <para><code> ∀(X,Y,Z)((typeof(X,Y) ∧
       typeof(X,Z)) → Y=Z)</code><footnote>
       <para>For simplicity, we consider here only cases with a single
        type repertoire.</para>
      </footnote></para>
    </listitem>
   </orderedlist>


  </section>
 </section>
 <section xml:id="DefaultImplicature">
  <title>Initial formulation of the default transcriptional
   implicature</title>
  <!--<para>The <emphasis>default</emphasis> transcriptional implicature with which we start our formalization may seem so embarrassingly trivial in any context of transcription that, indeed, few would find it necessary to state it explicitly. If one wanted to, one might say something like <quote>a transcript contains the same text as its exemplar</quote>. </para>-->
  <para>The initial formulation of the conjectured
    <emphasis>default</emphasis> transcriptional implicature which we
   present here could be described informally as <quote>a transcript
    contains the same text as its exemplar</quote>. Formally, this
   means that the exemplar and the transcript are tokens instantiating
   the same type: the transcript re-instantiates the type of the
   exemplar. Thus, only documents of <emphasis>equal</emphasis> types
   are considered T-similar; this corresponds to the simplest case of
   T-similarity presented earlier.</para>
  <para><emphasis role="ital">T-similarity</emphasis> could then be
   formalized in a single rule: <!--<lb/>--> ∃(X)
   (typeof(<emphasis role="ital">e</emphasis>,X) ∧ typeof(<emphasis role="ital">t</emphasis>,X)).</para>
  <para>In interesting cases, however, the type X (of which <emphasis role="ital">e</emphasis> and
   <emphasis role="ital">t</emphasis> are both tokens) will be a compound type consisting of some
   structure of smaller compound types (which in turn consist of
   smaller ones still), instantiated by a compound token which
   similarly consists of smaller tokens.</para>
  <!--<para>The statement above will thus amount to the conjunction of the following claims:</para>-->
  <!--<para>There is a one-to-one correspondence between the tokens of a transcript and the tokens of its exemplar, such that every pair of corresponding tokens have the same type. Formally, this is a second-order statement, but we can approximate it using four first-order sentences: </para>-->
  <para>It is thus more appropriate <!--or natural?--> to define
   T-similarity as the conjunction of four conditions, as
   follows:</para>
  <orderedlist numeration="loweralpha">
   <title>Default transcriptional implicature definition of
    T-similarity</title>
   <listitem>
    <para><emphasis role="bold">Reciprocity:</emphasis> The
     transcript() and exemplar() predicates map tokens in a
     one-to-one fashion<!--, and are inverse of each other-->.</para>
    <!--In other words, ...-->
    <para><code> reciprocity ⇔ ∀(X,Y,Z)(
      (((transcript(X,Y) ∧ transcript(X,Z)) →
      Y=Z) ∧ </code></para>
    <para><code> ((exemplar(X,Y) ∧ exemplar(X,Z)) →
      Y=Z)))</code></para>
   </listitem>
   <listitem>
    <para><emphasis role="bold">Completeness:</emphasis> There is
     nothing in the exemplar that is left out from the transcript. </para>
    <para>In other words, for each and every token X in the exemplar
     there is <!--one and only--> at least one corresponding token Y
     in the transcript. </para>
    <para><code> completeness ⇔ ∀(X)(e_token(X)
      → ∃(Y)(exemplar(Y,X)))</code></para>
   </listitem>
   <listitem>
    <para><emphasis role="bold">Purity:</emphasis> There is nothing in
     the transcript except what comes from the exemplar.</para>
    <para>In other words, for each and every token X in the
     transcript there is <!--one and only--> at least one
     corresponding token Y in the exemplar.</para>
    <para><code> purity ⇔ ∀(X)((t_token(X) →
      ∃(Y)(transcript(Y,X))))</code></para>
   </listitem>
   <listitem>
    <para><emphasis role="bold">Type identity:</emphasis> The exemplar
     and the transcript are type-identical through and through (not
     just the tokens <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis>).</para>
    <para>In other words, for every pair of corresponding tokens X
     in the exemplar and Y in the transcript, X and Y are of the
     same type.</para>
    <para><code> type_identity ⇔
      ∀(X,Y)((transcript(X,Y) →
      ∃(Z)(typeof(X,Z) ∧
     typeof(Y,Z))))</code></para>
   </listitem>
  </orderedlist>
  <para>We can now define T-similarity as follows: <itemizedlist>
    <listitem>
     <para><emphasis role="bold">T-similarity:</emphasis> The relation
      between exemplar and transcript is one of reciprocity,
      completeness, purity, and type identity.</para>
     <para><code> t_similarity ⇔ (reciprocity ∧
       completeness ∧ purity ∧ type_identity)
      </code></para>
    </listitem>
   </itemizedlist> This proposition defines formally what it means for
   <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> to be T-similar.</para>

  <para>With the axioms given earlier and a FOL representation of any
   pair of documents, we can verify whether <code>T-similarity</code>,
   i.e., the conjunction of the four rules of reciprocity,
   completeness, purity, and type identity, comes out as a theorem or
   not. This is described in more detail below.</para>
 </section>

 <section>
  <title>Discussion</title>
  <!--2026-06-22 YMA Discovered Claus’s suggestion to move the following paragraph from the end of the previous section to here, and agreed!-->
  <para>It may be worth stressing that the rules just given do not
   constitute a claim that every transcript is reciprocal, pure,
   complete, or thoroughly type similar to its exemplar. They amount
   to a claim that these things are true <emphasis>unless otherwise
    stated</emphasis> or <emphasis>unless the standards of a given
    community of practice dictate otherwise</emphasis>. That is, they
   express general assumptions about transcripts which are <emphasis role="ital">defeasible</emphasis> in particular cases. </para>
  <para>Consider different transcripts of the same exemplar. They may
   vary for several reasons. As a first application of our model, let
   us use it to describe some of those reasons.</para>
  <para>Transcripts may disagree about which of the marks in the
   exemplar instantiate types and are thus tokens (one transcript may
   read a mark as a decorative pen stroke, the other as a letter in a
   word). They may agree on the set of tokens found in the exemplar
   but disagree on which types they instantiate. Or they may use
   different type systems. One, for example, may distinguish the
   allographs i/j, u/v, and ſ/s, while the other treats the
   allographs as instantiations of the same type (as they are
   instances of the same grapheme).</para>
  <para>The use of different type systems can lead to the same kinds
   of difference between transcripts as different understandings of
   the exemplar; failure to understand the nature of the disagreement
   (different reading of the exemplar? or different choice of type
   system?) can lead to confusion and acrimony. Analysis of
   conflicting transcripts through the common lens of a formal model
   might reduce such confusion and acrimony.</para>

  <!--  <note>
   <para>CH 2026-06-11: The text below is taken from what was
    "Application to simple problems".</para>
  </note>-->
  <para> We observe that a number of common variations in
   transcription practice can be classified according to which rule of
   the initial formulation of the default transcriptional implicature
   they override. </para>

  <para>Some transcripts omit deleted material, extraneous material,
   illegible material, or material in specific writing systems
   (mathematics, Greek, …). In other words, transcripts are not always
   entirely complete.</para>

  <para>Some transcripts mark lines with bars, sometimes also adding
   line numbers. In some cases omissions may be marked by symbols or
   standard phrases ("[Illegible]", etc.) In other words, transcripts
   are not always entirely pure.</para>

  <para>Some transcripts preserve allographic variations, while others
   level those distinctions. For example, some transcripts preserve
   the distinctions between long and short s, or vocalic and
   consonantal i and u, others do not. This might be seen as a
   limitation of the rule of thorough type similarity. We believe it
   is more natural, however, to assume in such cases that the
   transcripts employ different type systems, one in which the
   allographs are considered tokens of the same type, and another one
   in which they are not. In other words, different transcripts do not
   necessarily employ identical type systems. </para>

  <para>Some transcripts silently expand abbreviations, normalize
   spelling and correct slips of the pen. At character level, the
   rules of completeness and purity seem to be broken. Even so, such
   transcripts may observe thorough type similarity at higher level
   tokens, such as words. In other words, transcripts are not always
   entirely type similar through and through.</para>

  <para>Expansion of abbreviations in brackets or italics can preserve
   the rules of default transcriptional implicature on word and higher
   levels, but introduces characters in transcripts which lack
   corresponding characters in the exemplar. Again, transcripts are
   not always entirely pure and do not always preserve entirely
   thorough type similarity.</para>

  <para>In cases of doubt, as in the case of parts of manuscripts
   which are hard to read because of wear, damage, or difficulties in
   handwriting, transcribers tend to interpret words as correctly
   spelled and sentences as grammatically well-formed, at least unless
   there is evidence to the contrary. This is often referred to as the
   principle of charity.<footnote>
    <para>One might regard this as a principle of giving the author
     the benefit of the doubt. Or perhaps one might just as well
     regard it as an act of charity to the reader: Unless there is
     clear evidence of error, there is no reason to bother the reader
     with the mere possibility of error.</para>
   </footnote> This tendency may be accounted for in at least two
   ways: Either we can regard the token in <emphasis role="ital">e</emphasis> as instantiating a
   disjunctive type and the token in <emphasis role="ital">t</emphasis> as instantiating one of the
   unobjectionable disjuncts, or else we can regard the transcriber as
   choosing to read the token in <emphasis role="ital">e</emphasis> as an instantiation of an
   unobjectionable type. In the former case,
   <!--the rule of thorough--> type similarity must be <!--weakened-->
   defined to allow pairs of the form "x/y" in <emphasis role="ital">e</emphasis> and either 'x' or
   'y' in <emphasis role="ital">t</emphasis> (so the types are similar but not identical). In the
   latter case, type identity is preserved but the reading of <emphasis role="ital">e</emphasis>
   simply ignores the fact that the passage in <emphasis role="ital">e</emphasis> might plausibly be
   interpreted differently.
   <!--YMA Probably true <note><para>The last two sentences need revision.</para></note>--></para>

 </section>
 <section xml:id="TI-defeasible">
  <title>Reformulation of the default
   <!--general rules of-->transcriptional implicature to permit
   exceptions</title>

  <para>As we have just seen, in almost all actual transcription
   projects there are at least a few exceptions to each of the four
   rules a-d: material in the exemplar not transcribed (perhaps
   because irrelevant to the purpose of the transcript), material in
   the transcript not present in the original (if only page numbers
   and footnotes), and transcription conventions which transcribe
   selected tokens with tokens of similar, not identical type. </para>
  <para>The examples which follow show a variety of such exceptions;
   the formalizations given show one way in which the defeasibility of
   the simple rules just given can be represented in logical formulae. </para>

  <para>As will become clear below, concrete transcription practices
   can be summarized as a set of exceptions to the rules of purity,
   completeness, and type similarity. Since the rules given in the
   preceding section do not countenance exceptions, it will prove
   necessary, for reasoning about any actual transcript and its
   exemplar, to reformulate those rules to capture the qualification
   that they each apply <emphasis>except in special
   cases</emphasis>.</para>

  <para>We therefore extend the document-pair model with the following
   predicates and relations: <itemizedlist>
    <listitem>
     <para><code>se_token(X): X is a special exemplar
       token</code></para>
    </listitem>
    <listitem>
     <para><code>st_token(X): X is a special transcript
       token</code></para>
    </listitem>
    <listitem>
     <para><code>typesimilar(X,Y): types X and Y are
       considered similar</code></para>
    </listitem>
   </itemizedlist></para>
  <para>In order to account for the inclusion of special exemplar and
   special transcript tokens we modify the <code>typeof domain</code>
   and the <code>no class overlap</code> axioms as
    follows:<orderedlist startingnumber="3">
    <listitem>
     <para><emphasis>typeof domain:</emphasis></para>
     <para><code> ∀(X,Y)(typeof(X,Y) →
      </code></para>
     <para><code>((t_token(X) ∨ e_token(X) ∨
       st_token(X) ∨ se_token(X)) ∧
       type(Y)))</code></para>
    </listitem>
    <listitem>
     <para><emphasis>no class overlap:</emphasis></para>
     <para><code> ∀(X) (</code></para>
     <para><code> (e_token(X) ⇔ (¬t_token(X) ∧
       ¬type(X) ∧ ¬se_token(X) ∧
       ¬st_token(X))) ∧</code></para>
     <para><code> (t_token(X) ⇔ (¬e_token(X) ∧
       ¬type(X) ∧ ¬se_token(X) ∧
       ¬st_token(X))) ∧</code></para>
     <para><code> (type(X); ⇔ (¬e_token(X) ∧
       ~t_t_token(X)) ∧ ¬se_token(X) ∧
       ¬st_token(X)) ∧</code></para>
     <para><code> (se_token(X) ⇔ (¬t_token(X) ∧
       ¬type(X) ∧ ¬e_token(X) ∧
       ¬st_token(X))) ∧</code></para>
     <para><code> (st_token(X) ⇔ (¬t_token(X) ∧
       ¬type(X) ∧ ¬se_token(X) ∧
       ¬e_token(X)))</code></para>

    </listitem>
   </orderedlist></para>

  <para>We add the following axioms:<orderedlist startingnumber="7">
    <listitem>
     <para><emphasis>typesimilar domain:</emphasis>
     </para>
     <para><code> ∀(X,Y)(type similar(X,Y) →
       (type(X) ∧ type(Y)))</code></para>
    </listitem>

    <listitem>
     <para><emphasis>typesimilar reflexive:</emphasis>
     </para>
     <para><code> ∀(X)(type(X) →
       typesimilar(X,X))</code></para>
    </listitem>

    <listitem>
     <para><emphasis>typesimilar symmetric:</emphasis>
     </para>
     <para><code> ∀(X,Y)(typesimilar(X,Y) →
       typesimilar(Y,X))</code></para>
    </listitem>

    <listitem>
     <para><emphasis>typesimilar transitive:</emphasis>
     </para>
     <para><code> ∀(X,Y,Z)((typesimilar(X,Y) ∧
       typesimilar(Y,Z)) → typesimilar(X,Z))</code></para>
    </listitem>


   </orderedlist></para>

  <para>In the definition of T-similarity, we replace the
   type-identity rule with the following:</para>
  <orderedlist numeration="loweralpha" startingnumber="4">
   <listitem>
    <para><emphasis role="bold">Type similarity:</emphasis> The
     exemplar and the transcript are type-similar through and
     through.</para>
    <para>In other words, for every pair of corresponding tokens X
     in the exemplar and Y in the transcript, X and Y are of
     similar types.</para>
    <para><code> type_similarity ⇔
      ∀(X,Y)((transcript(X,Y) → </code></para>
    <para><code>( ∃(Z,V)(typeof(X,Z) ∧
      typeof(Y,V) ∧ typesimilar(Z,V)))))</code></para>
   </listitem>
  </orderedlist>
  <para>Then, naturally, we change the definition of
    <code>T-similarity</code> to require <code>type_similarity</code>
   instead of <code>type_identity</code>:</para>
  <itemizedlist>
   <listitem>
    <para><emphasis role="bold">T-similarity:</emphasis>
    </para>
    <para><code> t_similarity ⇔ (reciprocity ∧
      completeness ∧ purity ∧ type_similarity)
     </code></para>
   </listitem>
  </itemizedlist>

  <para>Note that in the absence of special tokens and of unequal
   types that are typesimilar, the reformulation changes nothing to
   T-similarity. Thus, the reformulation does not <emphasis>per
    se</emphasis> change the default transcriptional
   implicature.</para>
  <para>However, the reformulation allows a concise description of any
   transcription practice in terms of its
    <emphasis>deviations</emphasis> from the rules of the default
   transcriptional implicature. For any given project's transcription
   practice, we can define the extension of the predicates se_token,
   st_token, and typesimilar, and combine them with the rules just
   given to provide a basis for inference.</para>

  <para>The skeptical reader may object to the explanatory value of
   this approach: In general, anything can be described as a deviation
   from any set of rules. More specifically, the skeptical reader will
   have observed that taken strictly, this reformulation amounts to
   saying that the properties of completeness, purity, and
   type-identity will apply in all cases, except when they do not; the
   formalization just given has no way to express the expectation that
   completeness, purity, and type-identity are the normal, expected,
   or usual case, and the cases covered by the predicates se_token,
   st_token, and typesimilar (when it does not coincide with type
   identity) are special, unusual, and less frequent cases.</para>

  <para>A logic designed to formalize defeasible reasoning would
   perhaps capture that distinction better; testing the utility of
   such formalisms for reasoning about transcription remains a
   desideratum for the future. In the meantime, however, we hope that
   this formulation in terms of standard first-order logic will
   suffice for the purposes of our argument. </para>
  <para>Readers may have been struck by some similarity (pun intended)
   between our initial formulation of the default transcriptional
   implicature and certain methods or criteria for identifying or
   measuring document similarity, such as, for example, the so-called
   Levensthein edit distance between strings of characters.<footnote>
    <para>See [<link xlink:href="https://www.balisage.net/Proceedings/vol25/print/Huitfeldt01/BalisageVol25-Huitfeldt01.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Balisage paper on Document similarity</link>].</para>
   </footnote> The edit distance between two such strings is measured
   in terms of the minimal number of deletions, insertions or
   substitutions of individual characters in the one string that is
   required in order to make it identical to the other.</para>
  <para>Analogously, we might think of special character tokens as
   deletions, special transcript tokens as insertions, and type
   similar tokens as substitutions. The difference, however, is that
   while Levensthein similarity is well suited for operations on
   character strings, it is less well suited for work on documents
   with a non-trivial structure. The additions and modifications we
   made in the reformulation of the default transcriptional
   implicature to permit exceptions may be seen as an attempt to take
   care of that difference. </para>
  <para>Can our <!--hypothesis about--> formulation of the default
   transcriptional implicature be put to empirical test? In principle,
   yes. At least we can see whether it makes good sense when applied
   to the actual transcriptional practices of a wide range of real
   projects. In practice, however, we lack the resources required to
   establish any really broad empirical basis. In this paper, we have
   limited ourselves to looking at two examples of descriptions of
   transcription practice. First we discuss an example drawn from a
   U.S.-based historical documentary edition (the papers of William
   Penn), then the transcription in a literary edition of some
   manuscript notes by Hermann Melville. </para>
  <!--<note><para>CH: Somewhere, a note about simplicity: We want a small, economical (lean) set of axioms, and we want one that captures common intuitions.<!-\-<lb/>-\-></para></note>-->
  <!--<note><para>CH: Revision of &date.last.revised; ends here.<!-\-<lb/>-\-></para></note>-->
 </section>

 <section xml:id="Penn">
  <title>Example 1: The papers of William Penn</title>

  <section xml:id="PennStatement">
   <title>Statement of practice</title>
   <para>In the chapter <quote>Editorial Method</quote> in [<xref linkend="Dunn"/>], the editors provide what we regard as a
     <quote>statement of practice</quote>, stating that:<footnote>
     <para>Here and in the rest of this section, indented quotations
      are quotations from pp. 15-18 in [<xref linkend="Dunn"/>].</para>
    </footnote></para>
   <blockquote>
    <para>...In this edition, we aim to print a completely faithful
     transcript of each original text, including blemishes and errors.
     ... In general, our editorial interpolations [enclosed in square
     brackets] within the text are minimal...</para>
   </blockquote>
   <para>We observe that the first sentence seems at least implicitly
    to confirm our rules of completeness and cleanliness: Every token
    of <emphasis role="ital">e</emphasis> is transcribed by exactly one token of the same type in
    <emphasis role="ital">t</emphasis>, and every token in <emphasis role="ital">t</emphasis> is exemplified by exactly one token of
    the same type in <emphasis role="ital">e</emphasis>. The second sentence introduces an exception:
    transcripts also contain additional material in the form of
    editorial interpolations (that is, transcripts are not entirely
    clean). However, such added material is explicitly marked by
    square brackets. The fact that the editors find it worth pointing
    this out may be taken to suggest that it does not go without
    saying in the relevant community that blemishes and errors are
    retained, or that additions are always marked explicitly. The
    editors continue: </para>
   <blockquote>
    <para>Our editorial rules may be summarized as follows:</para>
    <para>1. Each document selected for publication in <emphasis>The
      Papers of William Penn</emphasis> is printed in full. ...</para>
   </blockquote>
   <para>At first sight, this looks like a straightforward
    confirmation of the rule of completeness. So why does it not go
    without saying? Perhaps because it is not altogether common
    practice of this community to print every document in full. Or
    perhaps because the edition is a selection (that is, not complete
    in the sense of containing all the papers of William Penn) the
    editors found it important to make clear that although their
    transcript is not complete with respect to the entire body of
    material, it is complete with respect to each individual
    document.</para>
   <blockquote>
    <para>2. Each document is numbered, for convenient
     cross-reference, and is supplied with a short title.</para>
   </blockquote>
   <para>From this we can infer that transcription breaks the implicit
    rule of <!--YMA cleanliness--> purity by adding document numbers
    and titles, which do not occur in the exemplar. It seems unlikely
    that the editors find it necessary to make this explicit statement
    on the assumption that otherwise readers would believe that
    William Penn himself numbered his letters and provided them with
    titles. It is more likely that the community in question does
    expect to find some indication of where one document ends and the
    next begins as well as some means of referring to individual
    documents, but has not agreed on a uniform way of doing so. </para>
   <blockquote>
    <para>3. The format of each document (including the salutation and
     complimentary closure in letters) is rendered as in the original
     or copy...<footnote>
      <para>The quotation continues: <quote>, with the following two
        exceptions. Endorsements are treated as dockets, and entered
        into the provenance note (see below). If a document is
        undated, an initial date line is supplied [within square
        brackets]. If a document is dated at the close but not at the
        opening, an initial date line is supplied [within square
        brackets], and the closing date line is
       retained.</quote></para>
     </footnote></para>
   </blockquote>
   <para>It is not entirely clear from this, or from the rest of the
    declaration of practice, what the editors mean by the
     <quote>format</quote> of a document. Comparison of a transcript
    to a facsimile of its exemplar,<footnote>
     <para>Unfortunately we have only had access in facsimile to the
      first and the last page of one letter, which are reproduced on
      the insides of the front and back cover. These two pages, and
      the corresponding transcript, are reproduced in <link xlink:href="#A" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">appendix A</link>.</para>
    </footnote> however, suggests that what they mean is page layout.
    At least the transcripts seem to preserve such features as blocks
    or lines of text flushed to the right or left or indented, blank
    lines, and so on. (We also note that, elsewhere in the edition,
    poems are printed as lines of verse.<footnote>
     <para>See, for example, the poem on pp 32-3.</para>
    </footnote>)</para>
   <para>It is possible that the editors' practice is based on an
    explicit or implicit distinction between phenomena such as
    paragraphs, salutations, signatures, date lines, poems, verse
    lines, etc. In either case, they seem to be referring to what we
    would call compound types, and to try to preserve T-similarity
    between transcript and exemplar also in this respect.</para>
   <para>It is therefore perhaps striking to observe that the
    transcripts contain no indication of line or page breaks<footnote>
     <para>Except for such line breaks that occur in connection with
      the phenomena mentioned in the previous paragraph.</para>
    </footnote> in the exemplar, and that the editors do not mention
    this fact. Should we consider this omission a violation of the
    rule of completeness? (That is, are line and page breaks not
    considered tokens of some type?) If so, is the reason why the
    editors do not mention this omission that it is part of the
    transcriptional implicature of the relevant community of practice?
    Or do they simply rely on the appearance of the transcript on the
    printed page with regular, running lines etc. to make it too
    obvious to deserve mention that line and page breaks of the
    exemplar must be different?</para>
   <para>The statement numbered 4 in [<xref linkend="Dunn"/>] p. 17
    deals with the rendering of datings according to Julian,
    Gregorian, and Quaker calendars, a complicated issue the details
    of which we do not go into here.</para>
   <blockquote>
    <para>5. The text of each document is rendered as follows: </para>
    <para>a. Spelling is retained as written. Misspelled words are not
     marked with an editorial [sic]. If the sense of a word is
     obscured through misspelling, its meaning is clarified in a
     footnote. </para>
    <para>b. Capitalization is retained as written. In
     seventeenth-century manuscripts, the capitalization of such
     letters as "c," "k," "p," "s," and "w" is often a matter of
     judgment, and we cannot claim that our readings are definitive.
     Whenever it is clear to us that the initial letter in a sentence
     has not been capitalized, it is left lower case.</para>
    <para>c. Punctuation and paragraphing are retained as written.
     When a sentence is not closed with a period, we have inserted an
     extra space. </para>
    <para>d. Words or phrases inserted into the text are placed
     {within braces}. </para>
    <para>e. Words or phrases deleted from the text are crossed
     through.<!--YMA Don’t change: this is a quotation! <footnote><para>Concerning this and the previous point: inserted or deleted <emphasis>by Penn</emphasis>.</para></footnote>--></para>
    <para>f. Slips of the pen are retained as written, and are not
     marked by [sic].</para>
    <para>g. Contractions, abbreviations, superscript letters, and
     ampersands are retained as written. When a contraction is marked
     by a tilde, it is expanded.</para>
   </blockquote>
   <para>5.a suggests that this transcript, by retaining original
    spelling, does not deviate from general transcriptional
    implicature, but also that it is normal practice in the relevant
    community to do so, i.e., by silently normalizing spelling or
    marking misspelling with [sic].<footnote>
     <para>It is a common typographic practice to use "sic" to mark
      misspellings and other irregularities which might otherwise be
      blamed on the typesetter, and in some contexts punctuation and
      paragraphing are routinely normalized, or introduced, by
      editors.</para>
    </footnote> We take 5.f to mean that the editors make a
    distinction between slips of the pen (as in 5.a) from
    misspellings, but that they treat them the same way. </para>
   <para> The insertion of footnotes to clarify the meaning of
    misspelled words mentioned in 5.a, however, does represent an
    exception to the default rule of purity, as does the expansion of
    contractions marked by tilde mentioned in 5.g. Other deviations
    from the rule of purity are identified in 5.c for the insertion of
    an extra space when a sentence is not closed with a period, and
    5.d for the use of braces to surround editorial insertions into
    the text.<footnote>
     <para>We note in passing that the statement is silent on what
      happens in cases (if any) where the exemplar already contains
      braces, or an extra space after a sentence not closed with a
      period. If there are such cases, then their printed
      representation is ambiguous: it could represent the discreet
      editorial intervention described in the statement of practice,
      or it could be a strictly literal transcript of the paragraph.
      It seems likely that the editors know that there are no such
      cases, and trust the reader to infer it. But it might also be
      the case that the editors simply don't regard such ambiguities
      as interesting or important enough to be worth avoiding.</para>
    </footnote></para>
   <para>5.e simply seems to confirm the default rules of
    transcriptional implicature in stating that deleted words or
    phrases in the exemplar are crossed through in the transcript.
    That the editors are explicit about this may suggest that normal
    practice in the community is different. (Perhaps inserting special
    markers for deleted text (breaking the rule of purity) or leaving
    deleted text out (breaking the rule of completeness)).<footnote>
     <para>However, this may also be regarded as an application of the
      rule of type similarity, cf. our discussion of statement 5.l
      below.</para>
    </footnote></para>
   <para>Statement 5.b seems to be a paradigmatic application of what
    we referred to above as the principle of charity. We take it to be
    saying that deciding whether a given character in the manuscript
    is uppercase or lowercase requires judgment on the part of the
    transcriber.<footnote>
     <para>If, on the other hand, statement 5.b were taken to mean
      that the writers of the manuscripts regarded the use of
      capitalization as a matter for individual judgment rather than
      orthographic system, resulting in usage that the modern reader
      perceives as arbitrary and inconsistent, then the matter would
      be similar to the case with <emphasis>u</emphasis> and
       <emphasis>v</emphasis>, discussed below. (Readers today do find
      the capitalization of seventeenth-century manuscripts
      capricious, but we do not believe that is the point being made
      in 5.b.)</para>
    </footnote></para>
   <para>Thus, to a large extent the statements above explicitly
    confirm rules which form part of the default transcriptional
    implicature; that they are stated explicitly may suggest that in
    the relevant community of practice (or among the expected readers
    of the edition), it might be common to deviate from the default
    rules in these cases.</para>
   <blockquote>
    <para>h. The thorn is rendered as "th," and superscript
     contractions attached to the thorn are brought down to the line
     and expanded: as "the," "them," or "that." Our justification for
     this procedure is that we no longer have a thorn, and modern
     readers mistake it for "y." Likewise, since modern readers do not
     recognize that "u" and "v" were used interchangeably in the
     seventeenth century, we have rendered "u" as "v," or "v" as "u,"
     whenever appropriate.</para>
    <para>i. The £ sign in superscript is rendered as "l."<footnote>
      <para>We take this to mean that "£" is
       represented as "l" (not as "l.") — i.e., that the fact
       that the full stop is placed before the closing quotation mark
       is a somewhat misleading effect of punctuation rules.</para>
     </footnote></para>
    <para>j. The tailed "p" is expanded into "per," "pro," or "pre,"
     as indicated by the rest of the word.</para>
    <para>k. The long "s" is presented as a short "s." The double "ff"
     is presented as a capital "F."</para>
   </blockquote>
   <para>The statement contained in the first clause of the first
    sentence of 5.h, i.e., rendering the letter "Þ" as two
    characters, "t" and "h", may seem to deviate from all the four
    rules of our default transcriptional implicature: 1) it breaks the
    one-to-one correspondence between the tokens of the exemplar and
    the transcript (no reciprocity), 2) there are tokens (thorns) in
    the exemplar which do not appear in the transcript
    (incompleteness), 3) there are tokens ("th"-sequences) in the
    transcript which do not occur in the exemplar (impurity), and 4)
    none of the characters in question are of the same type (no type
    similarity). We observe that the statement may suggest that the
    normal practice of this community would be to represent "thorn"
    with tokens of a distinct type, which would not deviate from
    default transcriptional practice at all.</para>
   <para>However, if we interpret the statement to the effect that the
    project employs a type system in which "Þ" and "th" are
    tokens of the same type, the practice is entirely in accordance
    with the rules of our default transcriptional implicature. On the
    one hand, therefore, such an interpretation seems attractive. On
    the other hand, it may seem objectionable: In all other contexts
    than where they occur in sequence and stand for a thorn, the
    tokens "t" and "h" will still be taken as tokens of two different
    types.
    <!--<note>
     <para>If we allow this kind of move, even a document in which
       <emphasis>every</emphasis> token of the exemplar is represented by a
      string of characters in another alphabet (like Pinyin transliterations of
      Chinese in Latin script) would count as a transcription. Unless we would
      like to allow for transliterations to count as transcriptions, we would
      seem to need some kind of principled justification for this way of dealing
      with the thorn.</para>
    </note> We might therefore want some kind of more principled justification for this way of dealing with the thorn. --></para>
   <para>To this objection we may answer that it is not at all unusual
    for the type of tokens, or even the decision about where to draw
    boundaries between tokens, to be context dependent.<footnote>
     <para>A glyph shaped like <quote>A</quote> is either uppercase
      Latin letter A or uppercase Greek letter Alpha, depending on
      context. A dot on the baseline may be a full stop (sentence
      punctuation), an abbreviation marker, or a decimal point
      depending on both the intratextual and the extratextual context.
      In most European languages, <quote>ij</quote>,
      <quote>ch</quote>, and <quote>ll</quote> are each a
      two-character string, but the first is a single character (a
      single token instantiating a single type) in Dutch, and each of
      the others is taken as a single character in traditional Spanish
      orthography, even though <quote>i</quote>, <quote>j</quote>,
       <quote>c</quote>, <quote>h</quote>, and <quote>l</quote> also
      occur individually in those languages.</para>
    </footnote> Also, we do have here a principled justification for
    the context-dependent treatment of the thorn: A convention already
    exists for representing <quote>Þ</quote> as
     <quote>th</quote> in English spelling; "Þ" and "th" are
    traditionally pronounced identically by modern readers; and, as
    the statement itself argues, modern readers tend to mistake the
    thorn for a "y". (The last argument seems to appeal to a principle
    of charity with the <emphasis>reader</emphasis>.)</para>
   <para>If we are not convinced by these arguments and thus reluctant
    to accept that "Þ" and "th" may be regarded as tokens of the
    same type, one alternative might be to interpret the first clause
    of 5.h to the effect that the thorn is regarded as a contraction
    which is expanded to "th". In that case, the practice constitutes
    a deviation from default rules on a par with other contractions
    — see the discussion of tilde in 5.g above. In any case, the
    same goes for the statement made in the second clause of the first
    sentence concerning expansion of superscript contractions attached
    to the thorn.</para>
   <para>Now to the second sentence of 5.h, which observes that "u"
    and "v" were used interchangeably in the seventeenth century, and
    therefore the one is rendered as the other (or vice versa)
    "whenever appropriate" — presumably this means according to
    the expectations of modern readers (and thus this practice may be
    seen as another instance of "charity to the reader"). This may be
    a case for assuming that the project employs a type system
    according to which "u" and "v" are tokens of the same type, and in
    that case no deviation from default transcriptional implicature is
    implied. </para>
   <para>What may seem awkward about this analysis is that even if "u"
    and "v" are in free distribution in the seventeenth century texts
    (alternate realizations of the same type, just as allographs are
    alternate realizations of the same grapheme), they are certainly
    distinct types (and indeed distinct graphemes) in modern English.
    An alternative analysis, therefore, may be that the type system of
    the exemplar subsumes that of the transcript: every type in the
    transcript corresponds to exactly one type in the exemplar, but
    more than one type in the transcript ( here, <quote>u</quote> and
     <quote>v</quote>) may correspond to the same type in the exemplar
    (here, the type instantiated by the tokens we read as
     <quote>u</quote> and <quote>v</quote> in the exemplar). Either
    way, no deviance from the default transcriptional implicature is
    implied.</para>
   <!--<note><para>MSM 2016-01-04</para><para>If we wish to make the last sentence true, we will need to find a new formulation of the default transcriptional implicature, since our current formulation assumes the same type system in the readings of E and T.</para></note>-->
   <para>Statements 5.i and 5.k also describe the type system used:
    superscript £ is treated as a token of type "l",<footnote>
     <para>Once again, an alternative analysis might be that the
      different type systems are employed for the transcript and the
      exemplar. (In this case, the subset relationship would be the
      opposite of the one for "u" and "v".) And once again, rules of
      default transcriptional implicature would be broken in neither
      case.</para>
    </footnote> and long and short "s" are regarded as tokens of the
    same type. This will produce less ambiguity in the transcript than
    was the case for "Þ" and "th" or for "v" and "u", since long
    "s" will not occur at all in the transcript. Regarding "ff" as a
    token of capital "F", however, could in principle lead to
    ambiguity.</para>
   <para>Statement 5.j may seem to break the rule of purity for the
    expansion of the abbreviations mentioned. However, if tailed "p"s
    are regarded as tokens (depending on context) of the types "per,"
    "pro," or "pre," the rule is preserved.</para>
   <para>The statements 5.h-k may or may not be understood as
    representing deviations from the rules of default transcriptional
    implicature, deviations which are either not generally shared, or
    dealt with in other ways, by the community of practice. In either
    case, giving a formal account of all details involved here will
    lead to a certain amount of complication. As a first simplifying
    approximation, we may suggest that the statements imply that the
    project at hand employs a type system in which the following
    tokens are regarded as tokens of the same type: <orderedlist>
     <listitem>
      <para>"Þ" and "th"</para>
     </listitem>
     <listitem>
      <para>"Þ" with superscripts and "the", "them", or "that"
       (depending on context and/or the nature of the superscript,
       — the statement is silent on this point) </para>
     </listitem>
     <listitem>
      <para>"u" and "v"</para>
     </listitem>
     <listitem>
      <para>"U" and "V"</para>
     </listitem>
     <listitem>
      <para>Superscript "£" and "l"</para>
     </listitem>
     <listitem>
      <para>tailed "p" and "per," "pro," or "pre," as indicated by the
       rest of the word. </para>
     </listitem>
     <listitem>
      <para>long and short s</para>
     </listitem>
     <listitem>
      <para>"ff" and "F"</para>
     </listitem>
    </orderedlist>Under this assumption, the statements 5.h-5.k do not
    represent any deviation from default transcriptional implicature:
    reciprocity, purity, completeness and thorough type similarity are
    all preserved. </para>
   <blockquote>
    <para>l. Words underlined in manuscript are
      <emphasis>italicized</emphasis>.</para>
    <para>m. Blanks in the manuscript, missing words, and illegible
     words are rendered as [blank] or [missing word] or [illegible
     word] or [illegible deletion]. If a missing word can be supplied,
     it is inserted [within square brackets]. If the supplied word is
     conjectural, it is followed by a question mark.</para>
   </blockquote>
   <!--<note><para>Next two paras added by CH 07.01.25:</para></note>-->
   <para>Statement 5.l, that words underlined in the exemplar are
    rendered in italics in the transcript, may be seen in different
    ways within our framework of rules of transcriptional implicature.
    (The following remarks are relevant also to statement 5.e above,
    where deleted words or phrases in the exemplar are rendered as
    crossed through in the transcript.) We might simply assume that
    the editors regard underlined and italicized tokens as typesimilar
    (like allographs of the same grapheme). If so, however, why do
    they preserve the distinction in the transcript? After all, the
    distinction between other allographs of the same grapheme, like
    long and short s, are not preserved.</para>
   <para> It may seem more natural to assume that the editors
    understand underlining in the exemplar as signaling a higher-level
    feature, let us call it <emphasis>emphasis</emphasis>, which
    pertains not to the atomic letter tokens individually, but to the
    entire underlined word, and that they have chosen to signal this
    feature by other means, i.e., italics, in the transcript. On this
    assumption, an underlined word in the transcript is treated as a
    compound token of a type different from the same word without
    underlining. Also on this account, however, we might regard this
    as a way of preserving type similarity, and thus in accordance
    with the rules of transcriptional implicature.</para>
   <para>Statement 5.m, that information about blanks, missing words,
    illegible words, and conjectural readings are represented by words
    or phrases within square brackets, signals a break with the rules
    of both completeness and purity. Again, we may interpret the
    editors' statement as a confirmation that these rules are
    otherwise followed. The reason why they see a need to make this
    explicit may be either that these rules are normally followed in
    such cases, or that the phenomena in question are normally
    signaled in other ways.
    <!--<note><para>Note about rules 6 (concerning provenance) and 7 (special rules for two documents, and why we are not dealing with them here. </para></note>-->
   </para>
  </section>


  <section xml:id="PennReformulation">
   <title>Interpretation in terms of T-similarity</title>

   <para>Given the reformulation of the default transcriptional
    implicature offered above, we can interpret the statement of
    practice in terms of the three predicates
     <emphasis>se_token</emphasis> (special exemplar token),
     <emphasis>se_token</emphasis> (special transcript token), and
     <emphasis>typesimilar</emphasis> (similar types).</para>

   <!--   <section xml:id="Penn-formal-set">
    <title>Special exemplar tokens in &e;</title>-->

   <para>As far as we can tell, the only tokens in <emphasis role="ital">e</emphasis> that are not
    present in <emphasis role="ital">t</emphasis> are the tildes marking contraction. However, these
    contractions are expanded in <emphasis role="ital">t</emphasis>, so one might argue that this is
    a case of type similarity rather than special exemplar
    tokens.</para>
   <!--</section>-->


   <!--   <section>
    <title>Special transcript tokens in &t;</title>-->

   <para>The discussion above indicates that the following tokens in
    <emphasis role="ital">t</emphasis> are special transcript tokens, i.e., that they correspond to
    nothing in <emphasis role="ital">e</emphasis>: <itemizedlist>
     <listitem>
      <para>footnotes (and their markers)</para>
     </listitem>
     <listitem>
      <para>extra space at end of sentence lacking final
       punctuation</para>
     </listitem>
     <listitem>
      <para>braces marking inserted material</para>
     </listitem>
     <listitem>
      <para>expansions of contractions (a) marked by tilde, (b) using
       thorn + superscripts as <emphasis>the</emphasis>,
        <emphasis>them</emphasis>, <emphasis>that</emphasis>, (c)
       using tailed "p" as <emphasis>per</emphasis>,
        <emphasis>pro</emphasis>, <emphasis>pre</emphasis>.</para>
     </listitem>
     <listitem>
      <para>constituent "t" and "h" of "th" = thorn. </para>
     </listitem>
     <listitem>
      <para>constituent "l" and "." of "l." = superscript £.</para>
     </listitem>
    </itemizedlist></para>
   <!--</section>-->
   <!--   <section>
    <title>Type similarity</title>-->
   <para>We have concluded that the type system of <emphasis role="ital">t</emphasis> may be assumed
    to be the same as that of <emphasis role="ital">e</emphasis>, but also noted that the transcript
    seems to assume several cases of type similarity. Types <emphasis role="ital">t1</emphasis> and
    <emphasis role="ital">t2</emphasis> are similar iff: <itemizedlist>
     <listitem>
      <para><emphasis role="ital">t1</emphasis> = <emphasis role="ital">t2</emphasis></para>
     </listitem>
     <listitem>
      <para>or <emphasis role="ital">t1</emphasis> = thorn, <emphasis role="ital">t2</emphasis> = 〈 "t", "h" 〉</para>
     </listitem>
     <listitem>
      <para>or <emphasis role="ital">t1</emphasis> = superscript £, <emphasis role="ital">t2</emphasis> = "l"</para>
     </listitem>
     <listitem>
      <para>or <emphasis role="ital">t1</emphasis> = "ff", <emphasis role="ital">t2</emphasis> = "F"</para>
     </listitem>
     <listitem>
      <para>or (disjunctive(<emphasis role="ital">t1</emphasis>) ∧ <emphasis role="ital">t2</emphasis> ∈ disjuncts(<emphasis role="ital">t1</emphasis>))</para>
     </listitem>
    </itemizedlist></para>
   <!--</section>-->
   <!--  <section>
   <title>Conclusion</title>-->
   <para>We conclude that together, the statement of practice and our
    study of the relation between the exemplar and the transcript, may
    confirm our hypothesis that the underlying idea of transcription
    in this edition is that of T-similarity. </para>
   <!--</section>-->
  </section>
 </section>
 <!--   <para><emphasis role="bold">Purity</emphasis>: For every <emphasis
     role="bital">normal token</emphasis> in the transcript there is
    exactly one corresponding token in the exemplar. (Nothing in
     <emphasis>&t;</emphasis> except what comes from
     <emphasis>E</emphasis>, and some <emphasis role="bital">special
     cases</emphasis>.) Special cases include: 

   <para><emphasis role="bold">Completeness</emphasis>: For every
     <emphasis role="bital">normal token</emphasis> in the exemplar
    there is exactly one corresponding token in the transcript.
    (Everything in <emphasis>E</emphasis> is transcribed, except for
    some <emphasis role="bital">special cases</emphasis>.)</para>
   

   <para><emphasis role="bold">Reciprocity</emphasis>: The relations
    identified in the rules of purity and completeness are inverses. </para>
   <para>N.B. No special cases needed.</para>

   <para><emphasis role="bold">Type-similarity through and
     through</emphasis>: <emphasis>&t;</emphasis> and
     <emphasis>E</emphasis> are type-similar <emphasis role="bital"
     >through and through</emphasis>.</para>
   <para>I.e.: In every pair of corresponding tokens, the two tokens
    are tokens of similar types.</para>

 <section xml:id="PennFormalization">
   <title>Formalization of the rules for the Penn Papers</title>
   <note>
    <para>YMA 2026-05-30: This section seems to still use typed logic.
     Are we going to change that to Vampiresque notation?</para>
   </note>

   <para><emphasis role="bold">Purity</emphasis>: (&forall; t :
    tokens(T)) <!-\-<lb/>-\->(normal-transcript-token(t)
    &iff;<!-\-<lb/>-\->&exist-1; e : tokens(E)) (e = exemplar(t))</para>
   <para> (&forall; t : tokens(T))
    <!-\-<lb/>-\->(normal-transcript-token(t) &or;
    special-transcript-token(t))</para>

   <para><emphasis role="bold">Special tokens in T</emphasis>:
    (&forall; t) (special-transcript-token(t) &iff;<!-\-<lb/>-\->
    in-footnote(t) &or; footnote-marker(t) &or;
    sentence-extra-whitespace(t) &or; brace(t) &or;
    inserted-material(t) &or; expansion(t) )</para>
   <para><emphasis role="bold">Definition of in-footnote</emphasis>:
    (&forall; t) (in-footnote(t) &iff;<!-\-<lb/>-\->footnote-token(t)
    &or; (&exist; t2) (footnote-token(t2) &and;
    ancestor-descendant(t2, t)))</para>
   <para><emphasis role="bold">Mapping rule for footnotes</emphasis>:
    (&forall; t : tokens(T)) (in-footnote(t) &rArr;&not; (&exist; e :
    tokens(E)) (e = exemplar(t))</para>
   <para>Etc.</para>
   <note>
    <para>Predicates like inserted-material and footnote-token need to
     be defined (at least in prose).</para>
   </note>

   <para><emphasis role="bold">Completeness</emphasis>: (&forall; e :
    tokens(E)) (normal-exemplar-token(e) &rArr; (&exist-1; t :
    tokens(T)) (t = transcript(e)))</para>
   <para> (&forall; t : tokens(E)) (normal-exemplar-token(e) &or;
    special-exemplar-token(e))</para>

   <para><emphasis role="bold">Special tokens in E</emphasis>:
    (&forall; t) (special-exemplar-token(t) &iff; contraction-tilde(t)
    )</para>

   <para><emphasis role="bold">Thorough type similarity</emphasis>:
    (&forall; e : tokens(E)) <!-\-<lb/>-\->(typesimilar(type(e),
    type(transcript(e))))</para>
  </section>-->

 <section xml:id="Melville">

  <title>Example 2: Melville's notes in an edition of
   Shakespeare</title>

  <section xml:id="HM-prose">
   <title>Statement of practice</title>
   <para>On pages 955 to 970 of [<xref linkend="md"/>], the editors
    discuss notes made by the American author Herman Melville in an
    edition of Shakespeare. On pages 967-970 they transcribe the notes
    and provide facsimile images of the pages. (Facsimiles of
    Melville's notes and the editor's transcript, taken from [<xref linkend="md"/>], can be found in [<xref linkend="B"/>].) On page
    967 they provide a guide to <quote>Symbols used</quote>, which we
    quote in full:<footnote>
     <para>Numbering in square brackets added by us.</para>
    </footnote>
    <blockquote>
     <itemizedlist>
      <listitem>
       <!--* 1 *-->
       <para>[1] <emphasis role="bold">[...]</emphasis> revision or
        insertion enclosed in square brackets was made later than
        initial inscription of leaf </para>
      </listitem>
      <listitem>
       <!--* 2 *-->
       <para>[2] <emphasis role="bold">&lt;...&gt;</emphasis> letters or
        words enclosed in diamond brackets were canceled by lining out
       </para>
      </listitem>
      <listitem>
       <!--* 3 *-->
       <para>[3] <emphasis role="bold">&lt;...&gt;word</emphasis>
        letter(s) or word(s) written over are enclosed in diamond
        brackets closed up to the following word or letter that was
        superimposed </para>
      </listitem>
      <listitem>
       <!--* 4 *-->
       <para>[4] <emphasis role="bold">?word</emphasis> prefixed by a
        question mark indicates conjectural reading </para>
      </listitem>
      <listitem>
       <!--* 5 *-->
       <para>[5] <emphasis role="bold">xxxx</emphasis> undeciphered
        letters (number of x's approximates numbers of letters
        involved) </para>
      </listitem>
      <listitem>
       <!--* 6 *-->
       <para>[6] all words in roman are Melville's </para>
      </listitem>
      <listitem>
       <!--* 7 *-->
       <para>[7] all words in italics <emphasis role="ital">outside
         brackets</emphasis> are words Melville underlined </para>
      </listitem>
      <listitem>
       <!--* 8 *-->
       <para>[8] all words in italics <emphasis role="ital">inside
         brackets</emphasis> are editorial </para>
      </listitem>
     </itemizedlist>
    </blockquote></para>
   <para>There is no explicit statement, in this short text, that the
    exemplar has been transcribed in full or that nothing appears in
    the transcript that is not transcribing something in the exemplar.
    On the other hand, statement [6], that all words in roman are
    Melville's, and statements [4] and [8], which explain how
    editorial conjectures and additions are marked, indicate that
    great care has been taken to let the reader know which parts of
    the document are added by the editors. One way of making sense of
    this is to infer that for these transcribers <emphasis>it goes
     without saying</emphasis> that unless otherwise indicated, the
    transcript is complete and pure in our sense.<footnote>
     <para>The transcript appears to be complete with respect to
      Melville's notes, but omits the printed text of Shakespeare
      which appears in the same volume.</para>
    </footnote>
   </para>
   <para>Statements <!--* 1, 2, 3, and 7 *--> [1], [2], [3], and [7]
    indicate that the transcribers re-instantiate all insertions,
    cancellations, overwritings, and underlined text in the original,
    using the notations indicated. We model this by taking insertion,
    cancellation, overwriting, and underlining as compound types,
    which can like other types be recognized in the exemplar and
    re-instantiated in the transcript.<footnote>
     <para>An alternative interpretation would infer a type system in
      which for any letter such as <emphasis>e</emphasis>, there are
      companion types for <emphasis>inserted e</emphasis>,
       <emphasis>canceled e</emphasis>, <emphasis>overwritten
       e</emphasis>, and <emphasis>underlined [or italic]
      e</emphasis>. Occam's Razor and convenience in the formalization
      both lead us to prefer an account of this example in which
      insertions, etc., are compound tokens of corresponding compound
      types.</para>
    </footnote> These statements have a further implication for the
    type system to be used in reading the transcript: the charitable
    reader will infer that the various forms of brackets used to
    signal insertions, cancellations, and overwriting do not appear in
    the exemplar, since if they did, any occurrence of them in the
    transcript would become ambiguous.<footnote>
     <para>The occurrence of a single <emphasis>x</emphasis> is in
      fact perhaps strictly speaking ambiguous: statement [5] suggests
      the interpretation <quote>one undeciphered letter</quote>, but
      the only occurrence of a single <emphasis>x</emphasis> in the
      transcript (in the word <quote>extremes</quote> on line 37)
      seems more plausibly read as transcribing an
       <emphasis>x</emphasis> in the exemplar.</para>
    </footnote>
   </para>

   <para>Statement [5] describes an exception to the usual rule of
    type-identity between tokens in <emphasis role="ital">e</emphasis> and tokens in <emphasis role="ital">t</emphasis>, and also to
    the default 1:1 mapping between tokens in the two documents.
    Undeciphered tokens are transcribed by an
     <emphasis>approximate</emphasis> and not an
     <emphasis>exact</emphasis> number of x's, because tokens cannot
    be counted reliably until they have been identified as tokens of
    specific types. (To be a token is to be a token of a specific
    type.) And yet, if the occurrences of <emphasis>x</emphasis> in
    the transcript are tokens, then they must surely be tokens of some
    type. </para>
   <para>This would call for treating a sequence of n occurrences of
     <emphasis>x</emphasis> as a single token of the type <emphasis role="ital">undeciphered sequence of letters about n characters
     wide</emphasis>. In the exemplar, the undeciphered sequence would
    be an atomic (or basic) token, not a compound one; in the
    transcript, it would be a compound token composed of an
    appropriate number of occurrences of <emphasis>x</emphasis>. The
    undeciphered token in the exemplar would then illustrate the
    principle that different readers may plausibly read an exemplar in
    different ways: a reader who manages to decipher a word will
    assign it and its characters to the appropriate types, while a
    reader who finds the word illegible will assign it to an <emphasis role="ital">undeciphered letter-sequence</emphasis> type of
    appropriate length.</para>
   <para>(In the formalizations of the exemplar and the transcript
    below, however, we rely on a source of information according to
    which the sequence of letters in question, which are transcribed
    as <quote>?almxxxx</quote> has been deciphered as
     <quote>almanacks</quote>. In this situation, the sequence in the
    exemplar must be considered as a special exemplar token, while the
    sequence in the transcript is a special transcript token.)</para>

   <para>In contrast to the Penn Papers, page and line breaks are
    preserved in this edition. We observe that on this point none of
    the editions make any explicit statement on their choice. On our
    account, this suggests that the transcriptional implicatures of
    the two communities of practice are different: In this case it
    apparently goes without saying that line breaks are preserved, in
    the other it apparently goes without saying that they are not. </para>
   <para> It is interesting to observe that the statement of practice
    in the case of Melville's notes is formulated in terms of what the
    reader will see in <emphasis role="ital">t</emphasis> and what it means, whereas the Penn Paper's
    statement of practice starts from phenomena in <emphasis role="ital">e</emphasis> and describes
    the rules the transcribers have followed in creating <emphasis role="ital">t</emphasis>. The
    statement of practice for Melville's notes, that is, specifies
    more or less directly what inferences the reader can make about
    <emphasis role="ital">e</emphasis>, given what is in <emphasis role="ital">t</emphasis>, whereas the Penn Papers rules specify
    more or less directly what the transcriber must do in <emphasis role="ital">t</emphasis>, given
    what is in <emphasis role="ital">e</emphasis>. To use the Penn transcripts to make inferences
    about <emphasis role="ital">e</emphasis> requires that the rules be applied backwards, so to
    speak.</para>
  </section>

  <section xml:id="HM-formal">
   <title>Interpretation in terms of
    <!--transcriptional
    implicature/-->T-similarity</title>

   <para>As before, we can interpret the statement of practice in
    terms of the three predicates <emphasis>se_token</emphasis>
    (special exemplar token), <emphasis>se_token</emphasis> (special
    transcript token), and <emphasis>typesimilar</emphasis> (similar
    types).
    <!--and a description of the type system which the
    transcribers intend to be used.--></para>

   <!--   <section xml:id="HM-formal-set">
    <title>Special exemplar tokens in &e;</title>-->
   <para>Among tokens present in <emphasis role="ital">e</emphasis> and not present in <emphasis role="ital">t</emphasis> are marks
    (arrows, circles, and carets to mark the desired point of
    insertion) which indicate where inserted material belongs. (Note
    that these are not mentioned explicitly in the list of symbols,
    perhaps because they do not appear as symbols in the transcript;
    their existence can be inferred only from the editorial remarks in
    the transcript, at line 18 of page 969.) None of these types are
    in fact transcribed in <emphasis role="ital">t</emphasis>. So, they can all be considered special
    exemplar tokens in <emphasis role="ital">e</emphasis>, having no corresponding tokens in
    <emphasis role="ital">t</emphasis>.</para>
   <para>The phrase <quote>Conversation upon Gabriel, Micheal &amp;
     Raphel – gentlemanly &amp;c</quote> occurs in line 18 of <emphasis role="ital">t</emphasis>, but
    only below in <emphasis role="ital">e</emphasis>. The occurrence in <emphasis role="ital">e</emphasis> may be considered a
    special exemplar token.
    <!--<note>
      <para>Remark here that real projects and TEI would be more
       sophisticated then that.</para>
     </note>-->
   </para>
   <para>Similarly, if one believes (as one probably should<footnote>
     <para>Based on [<link xlink:href="https://melvillesmarginalia.org" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Melville's
       Marginalia Online</link>]. See in particular [<link xlink:href="https://melvillesmarginalia.org/Viewer.aspx?pid=16530" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">their transcript of 16530</link>].</para>
    </footnote>) that the string represented as
     <quote>?almxxxx</quote> in line 23 of <emphasis role="ital">t</emphasis> can be read as
     <quote>almanacks</quote>, then <quote>almanacks</quote> (or the
    substring <quote>anacks</quote>) of <emphasis role="ital">e</emphasis> should be considered a
    special exemplar token. </para>
   <!--<note>
     <para>CH 2026-06-24: If time allows: A remark on the overwriting
      of "ge" by "Gen" in line 13, cf the online Melville marginalia.
     </para>
    </note>-->
   <!--</section>-->
   <!--   <section xml:id="HM-formal-stt">
    <title>Special transcript tokens in &t;</title>-->
   <para>The List of Symbols makes clear that italicized words within
    angle brackets are editorial, not authorial, i.e., they are special
    exemplar tokens.</para>
   <para>Similarly, a question mark prefixed to a word whose reading
    is uncertain is a special transcript token.</para>
   <para>Square brackets, angle brackets, angle brackets with
    immediately following text, question marks prefixed to words, and
    sequences of the character <emphasis>x</emphasis> are all
    identified by the List of Symbols as having special meaning to
    which no token in the exemplar directly corresponds, i.e., they are
    special transcript tokens. </para>
   <para>We may take as a given that in normal cases it will be clear
    to a competent reader on inspection whether a token in <emphasis role="ital">t</emphasis> is or
    is not an instance of one of these types. However, one may also
    easily imagine cases in which it is not clear; these are not
    different in kind from legibility issues for the reader of <emphasis role="ital">t</emphasis>,
    both consisting in uncertainty as to the type instantiated by a
    token (or uncertainty as to whether a given mark is a token or
    not). As already pointed out, the process of reading lies outside
    the scope of our work. </para>
   <para>
    <!-- <note>
  <para>Section on type systems commented out, - come back to it
   under discussion later of simplifications. Perhaps simply say
   that we assume that the type systems of &e; and &t; are the
   same. </para></note>-->
    There are no clear examples of type similarity between tokens in
    <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis>. However, see our remarks on a normalized transcript
    further below.</para>




   <!--</section>-->

   <!--<section xml:id="HM-formal-ts">
    <title>Type similarity</title>

    <para>The transcription practice described above in section [<xref
      linkend="HM-prose"/>] relaxes in one case the normal expectation
     of type identity between tokens in &e; and corresponding tokens
     in &t;: <itemizedlist>
      <listitem>
       <para>A word token prefixed in &t; with <emphasis>?</emphasis>
        is a conjectural reading; for purposes of inference we take
        this to mean that the transcribers think that the word token
        so marked in &t; and its exemplar in &e;
         <emphasis>probably</emphasis> instantiate the same type, but
        that they are not certain.</para>
       <note>
        <para>This para may need to go, or else place it
         elsewhere:</para>
        <para>Such statements about likelihoods are not simple to
         represent accurately in conventional logic; we content
         ourselves here with introducing a predicate
         &probably-identical;, which holds in such cases, without
         attempting to specify exactly what inferences may be drawn
         from instances of that predicate. The predicate
         &probably-identical; will be asserted only in cases of
         conjectural readings, not in cases where the reading is
         certain. </para>
       </note>
      </listitem>
     </itemizedlist>
    </para>
    <para>Two other cases could be treated as relaxing the
     type-identity rule, though in fact we model those cases
     differently: <itemizedlist>
      <listitem>
       <para>A sequence of <emphasis>x</emphasis>s in &t; transcribes
        a word in &t; which the transcribers were unable to read. If
        one holds that the word in &t; instantiates a type even though
        we are unable to say which, then the two tokens will not
        instantiate the same type.</para>
       <para>In the formalization offered below, however, we take a
        different approach, assuming a set of types for
         <quote>undeciphered word about n characters long</quote>, for
        all required values of n (for this transcription, we need only
        n = 4). </para>
      </listitem>
      <listitem>
       <para>If one prefers a type system in which underlined words
        and italic words are necessarily of different types, then the
        correspondence between underlining in &e; and italics in &t;
        will involve a form of type similarity instead of type
        identity.</para>
       <para>In the formalization offered below, we postulate that
        underlining and italics are two graphetically distinct
        realizations of the same types, so that the tokens in &t; and
        &e; do instantiate the same types, though in different ways.
       </para>
      </listitem>
     </itemizedlist>
    </para>
    <note>
     <para>I suggest we omit this, we have carefully avoided
      thematizing differences in type systems everywhere else:</para>
     <!-\-<section xml:id="HM-formal-tsystem">-\->
     <para>The type system in the Melville transcription</para>
     <para>... must contain letters, words, punctuation marks
      ...</para>
     <para>... includes topographic lines as types ...</para>
     <para>... includes insertion, cancellation, overwriting, and
      underlining as (compound) types ...</para>
     <para> ... Brackets appear in T but not in E ... Types in T, but
      not reinstantiated ... </para>
     <!-\-</section>-\->
    </note>
   </section>-->
  </section>

  <section>
   <title>Proof strategy</title>
   <para>We have now discussed the Melville edition's statement of
    practice in terms of the rules of transcriptional implicature. By
    doing so we have formulated an account in more or less concise
    prose of which parts of <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> should be assumed special
    exemplar or transcript tokens, and which types should be assumed
    type-similar in order for <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> to be considered T-similar.
    How can we establish a formal proof to show that, given these
    assumptions, <emphasis role="ital">t</emphasis> and <emphasis role="ital">e</emphasis> are T-similar? </para>
   <para>After all, <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> are concrete, visual objects. They
    cannot directly be subjected to the formal procedures required for
    such proofs, and definitely not to the digital tools we will use
    to test such proofs. </para>
   <para> What we can do, however, is to let digital representations
    of the documents stand in for <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> (and all the tokens
    contained therein), and then perform the formal proof on these
    representations. </para>
   <para> There are many ways this can be done. The way we have
    chosen, is to create TEI-XML documents representing <emphasis role="ital">t</emphasis> and <emphasis role="ital">e</emphasis>.
    (As a sanity test, we also produce HTML visual replicas of <emphasis role="ital">e</emphasis> and
    <emphasis role="ital">t</emphasis> from the XML.) From these XML documents we create one FOL
    representation of both documents and of the relations between
    them. By feeding this representation alongside our transcriptional
    implicature axioms to a theorem prover we check whether
    T-similarity between the two documents can be proven. </para>
   <para>The theorem prover we have used is [<link xlink:href="https://vprover.github.io/" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Vampire</link>]. In
    earlier work we have used [<link xlink:href="https://alloytools.org/" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Alloy</link>] and [<link xlink:href="https://www.swi-prolog.org/" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Prolog</link>] for
    similar tasks. One of the reasons we chose Vampire this time, is
    that it operates on fairly standard FOL notation, whereas other
    tools use more idiosyncratic notations. Moreover, Vampire is less
    apt to crash on the relatively large amounts of data involved in
    detailed representations of documents. </para>

  </section>
  <section>
   <title>Explanation by way of a toy example</title>

   <section>
    <title>An exemplar and a transcript</title>


    <para> Let us consider the following artificial example of an
     exemplar <emphasis role="ital">e</emphasis>: <figure xml:id="fig-e">
      <title>Exemplar</title>
      <mediaobject>
       <imageobject>
        <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-001.jpg" format="jpg" width="50%"/>
       </imageobject>
      </mediaobject>
     </figure> and the following, equally artificial, transcript of
     <emphasis role="ital">e</emphasis>, <emphasis role="ital">t</emphasis> : <figure xml:id="fig-t1">
      <title>Transcript</title>
      <mediaobject>
       <imageobject>
        <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-002.jpg" format="jpg" width="50%"/>
       </imageobject>
      </mediaobject>
     </figure>
    </para>

    <para>In order to justify the plausibility of <emphasis role="ital">t</emphasis>, we may assume
     that the practice of the relevant community (or of this
     particular project) is to include insertions and silently omit
     deletions in a different hand (in this case, in red), to
     normalize spelling (in this case, of <quote>Essexe</quote> to
      <quote>Essex</quote>), and to add disambiguating remarks between
     square brackets (in this case, <quote>[the Earl
     of]</quote>).</para>
   </section>
   <section>
    <title>XML representations</title>

    <para>Our first step is to create an XML representation of <emphasis role="ital">e</emphasis>. We
     propose: 
<programlisting xml:space="preserve">
&lt;doc&gt;
 Elizabeth went &lt;del&gt;with&lt;/del&gt; &lt;add&gt;to&lt;/add&gt; Essexe
&lt;/doc&gt;
</programlisting>
    </para>

    <para>Our next step is to create an XML representation of <emphasis role="ital">t</emphasis>. We
     propose: 
<programlisting xml:space="preserve">
&lt;doc&gt;
 Elizabeth went to &lt;supplied&gt;the Earl of&lt;/supplied&gt; &lt;choice&gt;&lt;orig&gt;Essexe&lt;/orig&gt;&lt;reg&gt;Essex&lt;/reg&gt;&lt;/choice&gt;.
&lt;/doc&gt;
</programlisting>
    </para>

    <para>We observe that <emphasis role="ital">t</emphasis> both adds to and leaves things out of
     <emphasis role="ital">e</emphasis>, in accordance with the transcribal practice suggested.
       <quote><code>&lt;supplied&gt;the Earl
      of&lt;/supplied&gt;</code></quote>, must be counted as special
     transcript tokens. Furthermore, we understand
       <quote><code>&lt;choice&gt;&lt;orig&gt;Essexe&lt;/orig&gt;&lt;reg&gt;Essex&lt;/reg&gt;&lt;/choice&gt;</code></quote>
     to indicate that the tokens <quote>Essexe</quote> and
      <quote>Essex</quote> are typesimilar.</para>

    <para> When it comes to <emphasis role="ital">e</emphasis>,
      <quote><code>&lt;del&gt;with&lt;</code>/del&gt;</quote> must be
     considered a special exemplar token in relation to <emphasis role="ital">t</emphasis>.</para>
   </section>

   <section>
    <title>Translation to FOL, and proof of T-similarity</title>

    <para>Together, the XML representations of <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis>, as input
     to an appropriately devised XSLT stylesheet (see [<link xlink:href="#E" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Appendix E</link>]), yield the following
     output:</para>

    <para>
<programlisting xml:space="preserve">
fof(case_specific_facts, axiom,
  typeof(e1, t_Elizabeth) &amp;
  typeof(e2, t_went) &amp;
  typeof(e3, t_with) &amp;
  typeof(e4, t_to) &amp;
  typeof(e5, t_Essexe) &amp;
  typeof(t1, t_Elizabeth) &amp;
  typeof(t2, t_went) &amp;
  typeof(t3, t_to) &amp;
  typeof(t4, t_the) &amp;
  typeof(t5, t_Earl) &amp;
  typeof(t6, t_of) &amp;
  typeof(t7, t_Essex) &amp;
  typesimilar(t_Essexe, t_Essex) &amp;
  transcript(e1,t1) &amp;
  transcript(e2,t2) &amp;
  transcript(e4,t3) &amp;
  transcript(e5,t7) &amp;
  ! [X,Y,Z] : (
              ((transcript(X,Y) &amp; transcript(X,Z)) =&gt; Y=Z) &amp;
              ((exemplar(X,Y) &amp; exemplar(X,Z)) =&gt; Y=Z)
              ) &amp;
  $distinct(
    e1,e2,e3,e4,e5,
    t1,t2,t3,t4,t5,t6,t7,
    t_Elizabeth,t_went,t_with,t_to,t_Essexe,t_the,t_Earl,t_of,t_Essex
  ) &amp;
  ! [X] : ((e_token(X) &lt;=&gt; (X=e1 | X=e2 | X=e4 | X=e5)) &amp;
           (t_token(X) &lt;=&gt; (X=t1 | X=t2 | X=t3 | X=t7)))
).
</programlisting>
    </para>

    <para>The <emphasis>Case specific facts</emphasis> are formulated
     as an axiom in the form of one conjunction with many conjuncts,
     representing the facts of <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> and the relations between
     them. </para>
    <para>The listing above is in the so-called FOF (First-Order Form)
     notation used by Vampire. It is perhaps similar enough to the
     standard FOL notation used elsewhere in this paper as it is. For
     convenience, however, we include a translation to the usual
     notation, and group the various conjuncts into numbered groups
     for ease of reference: </para>

    <para>
     <emphasis>Case specific facts:</emphasis>
     <orderedlist>
      <listitem>
       <para>
        <code> typeof(e1,t_Elizabeth) ∧ typeof(e2,t_went)
         ∧ typeof(e3,t_with) ∧ </code>
       </para>
       <para>
        <code> typeof(e4,t_to) ∧ typeof(e5,t_Essexe) ∧
        </code>
       </para>
      </listitem>
      <listitem>
       <para>
        <code> typeof(t1,t_Elizabeth) ∧ typeof(t2,t_went)
         ∧ typeof(t3,t_to) ∧ </code>
       </para>
       <para>
        <code> typeof(t4,t_the) ∧ typeof(t5,t_Earl) ∧
         typeof(t6,t_of) ∧ typeof(t7,t_Essex) ∧ </code>
       </para>
      </listitem>
      <listitem>
       <para>
        <code> typesimilar(t_Essexe,t_Essex) ∧ </code>
       </para>
      </listitem>
      <listitem>
       <para>
        <code> transcript(e1,t1) &amp; transcript(e2,t2) &amp;
         transcript(e4,t3) &amp; transcript(e5,t7) &amp; </code>
       </para>
      </listitem>
      <listitem>
       <para>
        <code>∀(X,Y,Z) ( ((transcript(X,Y) &amp;
         transcript(X,Z)) → Y=Z) &amp; ((exemplar(X,Y) &amp;
         exemplar(X,Z)) → Y=Z) ) &amp; </code>
       </para>
      </listitem>

      <listitem>
       <para>
        <code> $distinct(</code>
       </para>
       <para>
        <code> e1,e2,e3,e4,e5, t1,t2,t3,t4,t5,t6,t7,</code>
       </para>
       <para>
        <code>
         t_Elizabeth,t_went,t_with,t_to,t_Essexe,t_the,t_Earl,t_of,t_Essex
         ) ∧ </code>
       </para>
      </listitem>
      <listitem>
       <para>
        <code> ∀(X)((e_token(X) ⇔ (X=e1 ∨ X=e2
         ∨ X=e4 ∨ X=e5)) ∧ </code>
       </para>
       <para>
        <code> (t_token(X) ⇔ (X=t1 ∨ X=t2 ∨ X=t3
         ∨ X=t7))) </code>
       </para>
      </listitem>
     </orderedlist>
    </para>


    <para>In group 1 the tokens in the exemplar (normal as well as
     special), named <quote>e1..e5</quote>, are associated with their
     respective types. The names of the types are simply given as a
     string consisting of the word string in question, prefixed with
      <quote>t_</quote>. In other words, the token <quote>e1</quote>
     is of type <quote>t_Elisabeth</quote>, and so on. </para>

    <para>In group 2 the tokens in the transcript, named
      <quote>t1..t7</quote>, are associated with their respective
     types, in the same manner.</para>

    <para>The statement in Group 3 states that the type named
      <quote>t_Essexe</quote> is similar to the type named
      <quote>t_Essex</quote>.</para>
    <para>The statements in Group 4 assigns each of the exemplar
     tokens in <emphasis role="ital">e</emphasis> to their corresponding transcript tokens in <emphasis role="ital">t</emphasis>. It
     may deserve special attention that the special exemplar token e3
     (of type <code>t_with</code>) and the special transcript tokens
     t4, t5, and t6 (of types <code>t_the</code>, <code>t_Earl</code>,
     and <code>t_of</code>), though represented as tokens associated
     with their respective types, are not mentioned in the Group 4
     statements. They are part of the formal description of the two
     documents, though they do not play any role in T-similarity
     proofs.</para>


    <para>The statements in Group 5, 6, and 7 are there to deal with
     the open world assumption of Vampire, and may be primarily of
     technical interest.<footnote>
      <para>In other systems, like Prolog or Alloy, they would not
       have been required, as these systems are based on a closed
       world assumption.</para>
     </footnote> In group 5, we state that there are no other pairs of
     individuals than those mentioned in group 4 which stand in a
     transcript relation to each other. In Group 6, the
      <quote>$distinct</quote> predicate is a special Vampire
     predicate that makes sure that each of the constants mentioned
     refer uniquely, i.e., <quote>e1</quote> names an individual not
     named by any of the other constants mentioned, etc.<footnote>
      <para>This is a convenience feature of Vampire. We could have
       done without it, but would then have had to include a statement
       to which Vampire expands the $distinct statement, such as:
        <code>t_of ≠ t_Essex ∧ t_Earl ≠ t_Essex ∧ t_Earl
        ≠ t_of ∧ t_the ≠ t_Essex ∧ t_the ≠ t_of ∧
        t_the ≠ t_Earl ∧ t_Essexe ≠ t_Essex ∧ t_Essexe
        ≠ t_of ∧ t_Essexe ≠ t_Earl ∧ t_Essexe ≠ t_the
        ∧ t_to ≠ t_Essex ∧ t_to ≠ t_of ∧ t_to ≠
        t_Earl ∧ t_to ≠ t_the ∧ t_to ≠ t_Essexe ∧
        t_with ≠ t_Essex ∧ t_with ≠ t_of ∧ t_with ≠
        t_Earl ∧ t_with ≠ t_the ∧ t_with ≠ t_Essexe
        ∧ t_with ≠ t_to ∧ t_went ≠ t_Essex ∧ t_went
        ≠ t_of ∧ t_went ≠ t_Earl ∧ t_went ≠ t_the
        ∧ t_went ≠ t_Essexe ∧ t_went ≠ t_to ∧ t_went
        ≠ t_with ∧ t_Elizabeth ≠ t_Essex ∧ t_Elizabeth
        ≠ t_of ∧ t_Elizabeth ≠ t_Earl ∧ t_Elizabeth ≠
        t_the ∧ t_Elizabeth ≠ t_Essexe ∧ t_Elizabeth ≠
        t_to ∧ t_Elizabeth ≠ t_with ∧ t_Elizabeth ≠
        t_went ∧ t7 ≠ t_Essex ∧ t_of ≠ t7 ∧ t_Earl
        ≠ t7 ∧ t_the ≠ t7 ∧ t_Essexe ≠ t7 ∧ t_to
        ≠ t7 ∧ t_with ≠ t7 ∧ t_went ≠ t7 ∧
        t_Elizabeth ≠ t7 ∧ t6 ≠ t_Essex ∧ t6 ≠ t_of
        ∧ t_Earl ≠ t6 ∧ t_the ≠ t6 ∧ t_Essexe ≠
        t6 ∧ t_to ≠ t6 ∧ t_with ≠ t6 ∧ t_went ≠
        t6 ∧ t_Elizabeth ≠ t6 ∧ t6 ≠ t7 ∧ t5 ≠
        t_Essex ∧ t5 ≠ t_of ∧ t5 ≠ t_Earl ∧ t_the
        ≠ t5 ∧ t_Essexe ≠ t5 ∧ t_to ≠ t5 ∧ t_with
        ≠ t5 ∧ t_went ≠ t5 ∧ t_Elizabeth ≠ t5 ∧
        t5 ≠ t7 ∧ t5 ≠ t6 ∧ t4 ≠ t_Essex ∧ t4
        ≠ t_of ∧ t4 ≠ t_Earl ∧ t4 ≠ t_the ∧
        t_Essexe ≠ t4 ∧ t_to ≠ t4 ∧ t_with ≠ t4 ∧
        t_went ≠ t4 ∧ t_Elizabeth ≠ t4 ∧ t4 ≠ t7
        ∧ t4 ≠ t6 ∧ t4 ≠ t5 ∧ t3 ≠ t_Essex ∧
        t3 ≠ t_of ∧ t3 ≠ t_Earl ∧ t3 ≠ t_the ∧
        t_Essexe ≠ t3 ∧ t_to ≠ t3 ∧ t_with ≠ t3 ∧
        t_went ≠ t3 ∧ t_Elizabeth ≠ t3 ∧ t3 ≠ t7
        ∧ t3 ≠ t6 ∧ t3 ≠ t5 ∧ t3 ≠ t4 ∧ t2
        ≠ t_Essex ∧ t2 ≠ t_of ∧ t2 ≠ t_Earl ∧ t2
        ≠ t_the ∧ t_Essexe ≠ t2 ∧ t_to ≠ t2 ∧
        t_with ≠ t2 ∧ t_went ≠ t2 ∧ t_Elizabeth ≠ t2
        ∧ t2 ≠ t7 ∧ t2 ≠ t6 ∧ t2 ≠ t5 ∧ t2
        ≠ t4 ∧ t2 ≠ t3 ∧ t1 ≠ t_Essex ∧ t1 ≠
        t_of ∧ t1 ≠ t_Earl ∧ t1 ≠ t_the ∧ t_Essexe
        ≠ t1 ∧ t_to ≠ t1 ∧ t_with ≠ t1 ∧ t_went
        ≠ t1 ∧ t_Elizabeth ≠ t1 ∧ t1 ≠ t7 ∧ t1
        ≠ t6 ∧ t1 ≠ t5 ∧ t1 ≠ t4 ∧ t1 ≠ t3
        ∧ t1 ≠ t2 ∧ e5 ≠ t_Essex ∧ e5 ≠ t_of
        ∧ e5 ≠ t_Earl ∧ e5 ≠ t_the ∧ e5 ≠
        t_Essexe ∧ t_to ≠ e5 ∧ t_with ≠ e5 ∧ t_went
        ≠ e5 ∧ t_Elizabeth ≠ e5 ∧ e5 ≠ t7 ∧ e5
        ≠ t6 ∧ e5 ≠ t5 ∧ e5 ≠ t4 ∧ e5 ≠ t3
        ∧ e5 ≠ t2 ∧ e5 ≠ t1 ∧ e4 ≠ t_Essex ∧
        e4 ≠ t_of ∧ e4 ≠ t_Earl ∧ e4 ≠ t_the ∧ e4
        ≠ t_Essexe ∧ e4 ≠ t_to ∧ t_with ≠ e4 ∧
        t_went ≠ e4 ∧ t_Elizabeth ≠ e4 ∧ e4 ≠ t7
        ∧ e4 ≠ t6 ∧ e4 ≠ t5 ∧ e4 ≠ t4 ∧ e4
        ≠ t3 ∧ e4 ≠ t2 ∧ e4 ≠ t1 ∧ e4 ≠ e5
        ∧ e3 ≠ t_Essex ∧ e3 ≠ t_of ∧ e3 ≠ t_Earl
        ∧ e3 ≠ t_the ∧ e3 ≠ t_Essexe ∧ e3 ≠ t_to
        ∧ e3 ≠ t_with ∧ t_went ≠ e3 ∧ t_Elizabeth
        ≠ e3 ∧ e3 ≠ t7 ∧ e3 ≠ t6 ∧ e3 ≠ t5
        ∧ e3 ≠ t4 ∧ e3 ≠ t3 ∧ e3 ≠ t2 ∧ e3
        ≠ t1 ∧ e3 ≠ e5 ∧ e3 ≠ e4 ∧ e2 ≠
        t_Essex ∧ e2 ≠ t_of ∧ e2 ≠ t_Earl ∧ e2 ≠
        t_the ∧ e2 ≠ t_Essexe ∧ e2 ≠ t_to ∧ e2 ≠
        t_with ∧ e2 ≠ t_went ∧ t_Elizabeth ≠ e2 ∧ e2
        ≠ t7 ∧ e2 ≠ t6 ∧ e2 ≠ t5 ∧ e2 ≠ t4
        ∧ e2 ≠ t3 ∧ e2 ≠ t2 ∧ e2 ≠ t1 ∧ e2
        ≠ e5 ∧ e2 ≠ e4 ∧ e2 ≠ e3 ∧ e1 ≠
        t_Essex ∧ e1 ≠ t_of ∧ e1 ≠ t_Earl ∧ e1 ≠
        t_the ∧ e1 ≠ t_Essexe ∧ e1 ≠ t_to ∧ e1 ≠
        t_with ∧ e1 ≠ t_went ∧ e1 ≠ t_Elizabeth ∧ e1
        ≠ t7 ∧ e1 ≠ t6 ∧ e1 ≠ t5 ∧ e1 ≠ t4
        ∧ e1 ≠ t3 ∧ e1 ≠ t2 ∧ e1 ≠ t1 ∧ e1
        ≠ e5 ∧ e1 ≠ e4 ∧ e1 ≠ e3 ∧ e1 ≠
        e2</code>
      </para>
     </footnote> The statement in Group 7 makes sure that there are no
     other e_tokens or t_tokens in the universe of discourse than
     those listed in the right-hand clauses of the two biconditionals. </para>

    <para>When we add the axioms 1-10 and the definitions a-d of
     transcriptional implicature formulated earlier to the axiom of
      <emphasis>Case specific facts</emphasis> above, the theorem
     prover confirms that <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> are T-similar.<footnote>
      <para>Technically, this is obtained by submitting to Vampire a
       FOF-representation of the axioms 1-10 and the definitions a-d,
       with T-similarity as a conjecture. Vampire's response is that
       the conjecture is a theorem, i.e., that it is logically implied
       by the other statements.</para>
     </footnote></para>
   </section>

  </section>
  <section>
   <title>Proof of T-similarity</title>
   <!--   <note xml:id="omit-special-tokens"
    xreflabel="Note on special tokens">
    <para>2026-06-21 YMA: Here (or somewhere near) might be the right
     place to say that we are omitting some (or all?) special tokens
     for our formal descriptions of documents, because they play no
     role in T-similarity proofs.</para>
    <para>2026-06-25 CH: Well, I just mentioned that, two paras up.
     OK?</para>
   </note>-->

   <para>We have gone through the same steps with the exemplar and
    transcript [<xref linkend="B"/>] of Melville's notes as described
    in the previous section on the toy example. </para>
   <para>We first made XML representations of the exemplar and the
    transcript, see [<xref linkend="C"/>]. In both, we used fairly
    straightforward TEI encoding.<footnote>
     <para>Experts on TEI and critical editing will probably find our
      encoding idiosyncratic and/or unsatisfactory. Moreover, any
      serious TEI-based project today would probably make one base
      transcription including all the information contained in both
      the transcription discussed here (in <xref linkend="C"/> and the
      normalized transcription in <xref linkend="G"/>. It should be
      mentioned, though, that Michael made a more thorough TEI
      representation of the transcript for a related project in 2017,
      using a special TEI-based tag set for manuscript transcription
       ([<link xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://mlcd.blackmesatech.com/mlcd/2015/W/tip/Melville/Melville-notes.sourcedoc.xml</link>]).</para>
    </footnote></para>
   <para>As a sanity check we made sure we could transform the XML
    files to HTML which came as close as we thought feasible to the
    originals in textual as well as visual aspects. We present these
    HTML files in [<xref linkend="D"/>]. For convenience, special
    exemplar and special transcript tokens are marked in red.</para>
   <para>The stylesheet briefly mentioned earlier,
     <emphasis>genFof.xsl</emphasis>, is reproduced in [<xref linkend="E"/>]. This stylesheet takes the two XML files in [<xref linkend="C"/>] as input, compares them, and produces the output
    in Vampire FOF format to be found in [<xref linkend="F"/>]. (In
    its current form, special token GIs are hard-coded into the
    stylesheet itself. One might easily extend the stylesheet so as to
    give users control over these, and to distinguish between special
    exemplar and transcript GI's.)</para>
   <para>Vampire confirms as a theorem that the XML representations of
    the exemplar and the transcript satisfies the T-similarity
    predicate, or, in other words, that <emphasis role="ital">e</emphasis> and <emphasis role="ital">t</emphasis> as represented are
    T-similar.</para>
   <para>We believe this to show that our definition of T-similarity,
    taking exceptions from default transcriptional implicature into
    account, is at least feasible. One such test can of course not be
    conclusive. As mentioned earlier, our hypothesis that there is a
    common core of assumptions underlying transcription in general is
    one that can only be tested empirically. We do not have the
    resources necessary for extensive empirical work.</para>
  </section>
  <section>
   <title>A normalized transcription</title>
   <para>However, we have at least made one other, very different,
     <quote>normalized</quote>, transcript of the exemplar ([<link xlink:href="#G" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Appendix G</link>]). Unlike the transcript
    discussed above, this one does contain some type similarity
    statements: <quote>&amp;</quote> is declared type similar to
     <quote>and</quote>, <quote>&amp;c</quote> to <quote>etc.</quote>,
     <quote>almanacks</quote> to <quote>almanacs</quote>,
     <quote>nonsence</quote> to <quote>nonsense</quote>, and
     <quote>nominee</quote> to <quote>nomine</quote>.<footnote>
     <para>Technically, <quote>Micheal</quote> and
       <quote>Raphel</quote> are not declared type similar to
       <quote>Michael</quote> and <quote>Raphael</quote>, but that is
      because they occur only as parts of a <quote>supplied</quote>
      element in the transcript.</para>
    </footnote>
   </para>




   <para>We put the normalized transcript to the same test. Again,
    T-similarity as defined here, with exceptions from default
    assumptions made explicitly and formally, shows that also the
    normalized transcript is T-similar to the exemplar. See [<xref linkend="G"/>]. </para>

   <para>This illustrates an important and more general point: Very
    different transcripts can be T-similar to one and the same
    exemplar. (Since T-similarity is symmetric and transitive, this
    means that the diplomatic and normalized transcriptions should be
    T-similar, too. We do not prove this.)</para>
   <para>Another important point illustrated by this exercise is that
    the representation of the exemplar, in the way things are set up
    here, is not necessarily (or usually cannot even be) independent
    of the way the transcript is represented. (For example, a deletion
    in the exemplar should be represented as a special transcript
    token if and only if it is omitted from the transcript.)</para>

  </section>


 </section>


 <section>
  <title>Discussion</title>
  <para>One of the things we think that this project illustrates, is
   the complexity of document structures and of the relations between
   documents. Our formalization has been on a fairly shallow level. We
   have succeeded in bringing the axioms of transcriptional
   implicature down to only a dozen fairly simple FOL statements, but
   the formal representation of even the short, less than 300 word
   Melville document, requires a conjunct of thousands of statements.
   Processing of the representation of even this small document
   requires a fairly large amount of computing power and time —
   with a medium-range personal computer Vampire takes several minutes
   to process the Melville transcripts. </para>
  <para>Therefore, we would also like to observe here that Vampire has
   impressed us. As mentioned, we have tried to do this work with the
   help of other tools in the past, but have had to give up because of
   insurmountable difficulties with formalization, user interface, or
   speed. This is the first time we have had the occasion to do real
   work with formalization and automated theorem proving, even though
   still on a fairly modest scale. We realize that some of our work in
   Vampire might have been done more efficiently with better knowledge
   of its possibilities. Other theorem provers exist, but we are not
   presently aware of any that are better suited to the task. </para>
  <para>We regard the work presented here as merely a proof of the
   concept of formalizing the modeling of documents and document
   relations, and of the applicability of the notion of T-similarity
   to transcription. Any implementation for real work <!--will-->would
   have to use other, more user-friendly and efficient tools, tools
   which do not, for example, require thorough knowledge of standard
   first order logic. </para>
  <para>In [<xref linkend="wit1"/>] we presented two models of text:
   The <quote>Grapheme-sequence Model</quote> and the <quote>Readings
    model</quote>. We found the <quote>Grapheme-sequence Model</quote>
   unsatisfactory because of a number limitations: <itemizedlist>
    <listitem>
     <para>Texts are not just sequences of graphemes, -- the are
      usually more complex graph-like structures with notes, variants,
      alternate readings etc.</para>
    </listitem>
    <listitem>
     <para>The analysis of document types must reflect
      compositionality and levels of types: there are atomic as well
      as boolean and compound types, types at the levels of letters,
      words, sentences, paragraphs, etc.</para>
    </listitem>
    <listitem>
     <para>there may be different readings of the same document</para>
    </listitem>
   </itemizedlist>
  </para>
  <para> In [<xref linkend="ettdds"/>] we tried to develop a formal
   account without these limitations. However, we never reached a full
   working formalization, partly because of the problems with finding
   suitable formalization tools mentioned above.</para>
  <para>With what is presented here, we are in many ways back with the
    <quote>Grapheme-sequence Model</quote>: We represent documents as
   sequences — not even of letters, but of word tokens. (We
   could of course have chosen to model on character level in stead.
   That would have increased the number of tokens in the model,
   without making much interesting difference in principle.) The type
   structure is also represented as entirely flat, without levels or
   compositionality. </para>

  <para>Still, with these limitations, we have a working model which
   lends itself to full formalization and formal proof. We can only
   hope that someone may some time extend it beyond its current
   limitations, e.g., along the lines of [<xref linkend="ettdds"/>].</para>

  <para>Some readers may have been suspicious of our move, in the
   section on Melville, from discussing the original exemplar and
   typescript on paper to discussing XML-representations of them. It
   was this move, however, which made us aware of a (perhaps to other
   obvious) fact: Our model presupposes that the representation of the
   exemplar must reflect what is <emphasis>not</emphasis> recorded in
   the transcript.<footnote>
    <para>It may be plausible to omit unreadable writing from a
     transcript. However, in order to leave, for example, text in a
     different hand or a different language, an addition or deletion
     out from a transcript, the transcriber needs to recognize it as
     such, and may thus very well be able to read it.</para>
   </footnote></para>

  <para>Moreover, we think that this move shows that it might be
   useful to think in terms of T-similarity or similar formalizations
   of similarity between digital documents in contexts well beyond
   transcription and scholarly editing.</para>

  <para>Originally, this talk had the subtitle <quote>a contribution
    to markup semantics</quote>. We can only hope that it will
   be.</para>
 </section>

 <section xml:id="Future">
  <title>Conclusions and Future Work</title>
  <para>We have proposed the term <emphasis>transcriptional
    implicature</emphasis> to denote the rules of inference which
   should govern the interpretation of a transcript in the absence of
   explicit statements to the contrary. Informally, the rules of
   transcriptional implicature (are intended to) capture those facts
   about transcription which are too obvious to need saying. We expect
   that the rules of transcriptional implicature vary from community
   of practice to community of practice; their accurate identification
   will be a matter of the sociology of scholarship, philosophy of
   science, and close examination of transcription practice. We
   believe that the rules of what we call <emphasis>default
    transcriptional implicature</emphasis> provide a plausible example
   of what the transcriptional implicature for a given community of
   practice might look like, and we have shown how the practice of
   some concrete examples can be captured formally in ways which
   exhibit clearly both the general rules governing transcription and
   the specific practices of the individual project.</para>
  <para>We hope that this work can provide a basis for a full model of
   the logical structure of transcription, of the formal relation
   between exemplar and transcript, and of the ways in which a
   transcript can be used to infer facts about its exemplar.</para>
 </section>

 <appendix xml:id="A">
  <title>Penn: Facsimile of exemplar and transcript.</title>
  <para>Facsimiles of the first and last page of a letter from Penn to
   Lady Conway from 1675, reproduced on the insides of the front and
   back cover of [<xref linkend="Dunn"/>]: </para>
  <figure xml:id="Penn-03">
   <title>Penn, letter to Lady Conway, first page</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-003.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>
  <figure xml:id="Penn-04">
   <title>Penn, letter to Lady Conway, last page</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-004.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>


  <para>Facsimiles of excerpt from the transcript of the corresponding
   pages in [<xref linkend="Dunn"/>], pages 356 and 358: </para>

  <figure xml:id="Penn-B356">
   <title>Penn, letter to Lady Conway, first page</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-005.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>

  <figure xml:id="Penn-B358">
   <title>Penn, letter to Lady Conway, last page</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-006.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>

 </appendix>

 <appendix xml:id="B">
  <title>Melville: Facsimile of exemplar and transcript</title>
  <para>[<xref linkend="md"/>] contain facsimiles of the two pages of
   Melville's exemplar on pages 968 and 970. </para>

  <para>Facsimiles of the same two pages in better quality can be
   found here: <itemizedlist>
    <listitem>
     <para>[<link xlink:href="https://melvillesmarginalia.org/Viewer.aspx?pid=16530" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Melville's Marginalia Online, 16530</link>]</para>
    </listitem>
    <listitem>
     <para>[<link xlink:href="https://melvillesmarginalia.org/Viewer.aspx?did=31&amp;pid=16529" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">Melville's Marginalia Online, 16529</link>]</para>
    </listitem>
   </itemizedlist>
  </para>


  <para>Facsimile of transcript, [<xref linkend="md"/>] pages 969 and
   970</para>
  <figure xml:id="Melville-transcript1">
   <title>Transcript, first page</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-007.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>
  <figure xml:id="Melville-transcript2">
   <title>Transcript, second page</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-008.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>
 </appendix>

 <appendix xml:id="C">
  <title>Melville: XML representations</title>
  <para>XML representation of the exemplar</para>
<programlisting xml:space="preserve">
&lt;?xml version="1.0" encoding="UTF-8"?&gt;
&lt;?xml-stylesheet type="text/xsl" href="../../genFofXSLT/genFof.xsl"?&gt;
&lt;!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
 &lt;!ENTITY mdash  "—" &gt;&lt;!--=em dash--&gt;
 &lt;!ENTITY ndash  "–" &gt;&lt;!--=en dash--&gt;
 &lt;!ENTITY vbar  "|" &gt;&lt;!--=vertical bar--&gt;
 &lt;!ENTITY apos  "—" &gt;&lt;!--=apostrophe--&gt;
 &lt;!ENTITY triplebar "|" &gt;&lt;!--/triple horizontal bars--&gt;
 &lt;!ENTITY ldquo  "“" &gt;&lt;!--=double quotation mark, left--&gt;
 &lt;!ENTITY uarr   "↑" &gt;&lt;!--/uparrow A: =upward arrow--&gt;]&gt;
&lt;text&gt;
 &lt;body&gt;
  &lt;p&gt;
   &lt;lb/&gt;A seaman figures in The Canterbury Tales. &lt;lb/&gt;With
    &lt;del&gt;a&lt;/del&gt;many a tempest had his beard been &lt;lb/&gt;shook. &amp;ndash;
    &lt;del&gt;S&lt;/del&gt;
   &lt;del&gt;Deep g&lt;/del&gt; Secret grief is a &lt;lb/&gt;cannibal of its own heart
   &amp;ndash; &lt;emph&gt;Bacon&lt;/emph&gt;. &lt;lb/&gt;&amp;vbar;&lt;emph&gt;Claudia&lt;/emph&gt; of the
    &lt;emph&gt;Appian&lt;/emph&gt; family, &lt;q&gt;I wish &lt;lb/&gt;&amp;vbar; some fight or
    pestilence would thin out &lt;lb/&gt;&amp;vbar; this crowd.&lt;/q&gt;
   &lt;emph&gt;Arrogance.&lt;/emph&gt;
   &lt;lb/&gt;
   &lt;emph&gt;Roast beef in the pulpet.&lt;/emph&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;An animal of a man &amp;mdash; &lt;q&gt;do eagles wear
    &lt;lb/&gt;spectacles?&lt;/q&gt;&amp;mdash; Health. &amp;mdash; &lt;emph&gt;Contrast&lt;/emph&gt;:
   an &lt;lb/&gt;over spiritual man. &lt;lb/&gt;&amp;mdash;&amp;mdash; &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;
   &lt;q&gt;Yes, Madam, Cain was a godless froward boy, &amp;amp; &lt;lb/&gt;Reuben
    (Gen:49) &amp;amp; Absalom&lt;/q&gt; Many pious men &lt;lb/&gt;have impious
   children &amp;mdash; (Devil as a Quaker)
   &lt;lb/&gt;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash; &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;A formal compact &amp;ndash; Imprimis &amp;ndash; First &amp;ndash;
   Second. &lt;lb/&gt;The aforesaid soul. said soul &amp;amp;c &amp;ndash;
   Duplicates &amp;ndash; &lt;lb/&gt;&amp;triplebar;&lt;q&gt;How was it about the
    temptation on the &lt;lb/&gt;hill?&lt;/q&gt; &amp;amp;c &amp;ndash; D begs the hero to
   form &lt;lb/&gt;one of a &lt;emph&gt;
    &lt;q&gt;Society of D's&lt;/q&gt;
   &lt;/emph&gt; &amp;ndash; his name would be weighty &lt;lb/&gt;&amp;amp;c &amp;ndash;
   Leaves a letter to the D &amp;ndash; &lt;q&gt;My &lt;lb/&gt;
    &lt;emph&gt;Dear D&lt;/emph&gt;
   &lt;/q&gt;
   &lt;floatingText&gt;
    &lt;body&gt;
     &lt;ab&gt; &amp;ndash;  Conversation upon Gabriel, Micheal &amp;amp;
      &lt;lb/&gt; Raphel &amp;ndash; gentlemanly &amp;amp;c&lt;/ab&gt;
    &lt;/body&gt;
   &lt;/floatingText&gt;
  &lt;/p&gt;

  &lt;p&gt;
   &lt;lb/&gt;
   &lt;q&gt;Terra Oblivionis&lt;/q&gt;
   &lt;q&gt;Hellites&lt;/q&gt; &amp;ndash; At the Astor find him &lt;lb/&gt; making
    &lt;sic&gt;almanacks&lt;/sic&gt; &amp;ndash; going to a ball takes a long
   &lt;lb/&gt;time making toilette. &amp;ndash; The Doctor's coach stops
   &lt;lb/&gt;the way. &amp;ndash; &lt;q&gt;Do you believe all that stuff? &lt;lb/&gt;
    nonsence &amp;ndash; the world was never made. &amp;ndash; &amp;ldquo;But Is
    not &lt;lb/&gt;this you mentioned &lt;emph&gt;here&lt;/emph&gt; &amp;ndash; in the
    scriptures?&lt;/q&gt;
   &lt;lb/&gt;Receives visits from the principal d's &amp;ndash;
    &lt;q&gt;Gentlemen&lt;/q&gt; &amp;amp;c. &lt;lb/&gt;
   &lt;emph&gt;Arguments&lt;/emph&gt; to persuade &amp;ndash; &lt;q&gt;Would you not rather
    &lt;lb/&gt;be below with kings than above with fools?&lt;/q&gt;
  &lt;/p&gt;

  &lt;p&gt;
   &lt;lb/&gt;It is better to laugh &amp;amp; not sin than to &lt;del&gt;be&lt;/del&gt; weep
   &amp;amp; be &lt;lb/&gt;wicked. &amp;mdash; Ten loads of coal to burn him.
   &amp;mdash; &lt;lb/&gt;Brought to the stake &amp;mdash; warmed himself by the
   fire. &lt;/p&gt;

  &lt;p&gt;
   &lt;lb/&gt;Ego non baptizo te in nominee Patris et &lt;lb/&gt;Filii et Spiritus
   Sancti &amp;ndash; sed in nomine &lt;lb/&gt;Diaboli. &amp;mdash; Madness is
   undefinable &amp;mdash; &lt;lb/&gt;It &amp;amp; right reasons extremes of one.
   &lt;lb/&gt;–Not the &lt;add place="sup"&gt;(black art)&lt;/add&gt; Goetic but
   Theurgic magic &amp;mdash; &lt;lb/&gt;seeks converse with the Intelligence,
   Power, the &lt;lb/&gt;Angel. &lt;/p&gt;
 &lt;/body&gt;
&lt;/text&gt;
</programlisting>
  <para>XML representation of the transcript</para>
<programlisting xml:space="preserve">
&lt;?xml version="1.0" encoding="UTF-8"?&gt;
&lt;!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
 &lt;!ENTITY mdash  "—" &gt;&lt;!--=em dash--&gt;
 &lt;!ENTITY ndash  "–" &gt;&lt;!--=en dash--&gt;
 &lt;!ENTITY vbar  "|" &gt;&lt;!--=vertical bar--&gt;
 &lt;!ENTITY apos  "—" &gt;&lt;!--=apostrophe--&gt;
 &lt;!ENTITY triplebar "|" &gt;&lt;!--/triple horizontal bars--&gt;
 &lt;!ENTITY ldquo  "“" &gt;&lt;!--=double quotation mark, left--&gt;
 &lt;!ENTITY uarr   "↑" &gt;&lt;!--/uparrow A: =upward arrow--&gt;]&gt;
&lt;text&gt;
 &lt;front&gt;
  &lt;docTitle&gt;
   &lt;titlePart&gt;NOTES IN A SHAKESPEARE VOLUME
    &lt;span&gt;969&lt;/span&gt;&lt;/titlePart&gt;
  &lt;/docTitle&gt;
 &lt;/front&gt;
 &lt;body&gt;
  &lt;fw&gt;[on verso of last leaf of Volume VII, page [524]]&lt;/fw&gt;
  &lt;p&gt;
   &lt;lb n="1"/&gt;A seaman figures in The Canterbury Tales. &lt;lb n="2"
   /&gt;With &lt;del&gt;a&lt;/del&gt;many a tempest had his beard been &lt;lb n="3"
   /&gt;shook. &amp;ndash; &lt;del&gt;S&lt;/del&gt;
   &lt;del&gt;Deep g&lt;/del&gt; Secret grief is a &lt;lb n="4"/&gt;cannibal of its own
   heart &amp;ndash; &lt;emph&gt;Bacon&lt;/emph&gt;. &lt;lb n="5"/&gt;
   &lt;metamark&gt;&amp;vbar;&lt;/metamark&gt;
   &lt;emph&gt;Claudia&lt;/emph&gt; of the &lt;emph&gt;Appian&lt;/emph&gt; family, &lt;q&gt;I wish
     &lt;lb n="6"/&gt;
    &lt;metamark&gt;&amp;vbar;&lt;/metamark&gt; some fight or pestilence would thin
    out &lt;lb n="7"/&gt;
    &lt;metamark&gt;&amp;vbar;&lt;/metamark&gt; this crowd.&lt;/q&gt;
   &lt;emph&gt;Arrogance.&lt;/emph&gt;
   &lt;lb n="8"/&gt;
   &lt;emph&gt;Roast beef in the pulpet.&lt;/emph&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb n="9"/&gt;An animal of a man — &lt;q&gt;do eagles wear &lt;lb n="10"
    /&gt;spectacles?&lt;/q&gt;— Health. — &lt;emph&gt;Contrast&lt;/emph&gt;: an
    &lt;lb n="11"/&gt;over spiritual man. &lt;lb rend="dunno"/&gt;
   &lt;metamark&gt;——&lt;/metamark&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb n="12"/&gt;
   &lt;q&gt;Yes, Madam, Cain was a godless froward boy, &amp;amp; &lt;lb n="13"
    /&gt;Reuben (Gen:49) &amp;amp; Absalom&lt;/q&gt; Many pious men &lt;lb n="14"
   /&gt;have impious children — (Devil as a Quaker) &lt;lb
    rend="dunno"/&gt;
   &lt;metamark&gt;—————&lt;/metamark&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb n="15"/&gt;A formal compact &amp;ndash; Imprimis &amp;ndash; First &amp;ndash;
   Second. &lt;lb n="16"/&gt;The aforesaid soul. said soul &amp;amp;c &amp;ndash;
   Duplicates &amp;ndash; &lt;lb n="17"/&gt;&amp;triplebar;&lt;q&gt;How was it about the
    temptation on the &lt;lb n="18"/&gt;hill?&lt;/q&gt; &amp;amp;c &lt;supplied rend="[]"
    &gt;inserted later below in lines 21–21b after Dear D&amp;ldquo;
    — and circled &lt;lb rend="dunno"/&gt; with guideline to caret
    here&lt;/supplied&gt;
   &lt;floatingText&gt;
    &lt;body&gt;
     &lt;ab&gt;Conversation upon Gabriel, Micheal &amp;amp; / &lt;lb rend="dunno"/&gt;
      Raphel &amp;ndash; gentlemanly &amp;amp;c&lt;supplied rend="]"/&gt;&lt;/ab&gt;
    &lt;/body&gt;
   &lt;/floatingText&gt; &amp;ndash; D begs the hero to form &lt;lb n="19"/&gt;one of
   a &lt;emph&gt;
    &lt;q&gt;Society of D&amp;apos;s&lt;/q&gt;
   &lt;/emph&gt; &amp;ndash; his name would be weighty &lt;lb n="20"/&gt;&amp;amp;c
   &amp;ndash; Leaves a letter to the D &amp;ndash; &lt;q&gt;My &lt;lb n="21"/&gt; Dear D
   &lt;/q&gt; &amp;ndash; &lt;supplied rend="[]"&gt;later insertion in lines
    21–21b, reported in line 18&lt;/supplied&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb n="22"/&gt;
   &lt;q&gt;Terra Oblivionis&lt;/q&gt;
   &lt;q&gt;Hellites&lt;/q&gt; &amp;ndash; At the Astor find him &lt;lb n="23"/&gt;
   &lt;unclear&gt;making&lt;/unclear&gt;
   &lt;unclear&gt;&lt;supplied&gt;alm&lt;gap/&gt;&lt;/supplied&gt;&lt;/unclear&gt; &amp;ndash; going to
   a ball takes a long &lt;lb n="24"/&gt;time making toilette. &amp;ndash; The
   Doctor&amp;apos;s coach stops &lt;lb n="25"/&gt;the way. &amp;ndash; &lt;q&gt;Do you
    believe all that stuff? &lt;lb n="26"/&gt; nonsence &amp;ndash; the world
    was never made. &amp;ndash; &lt;supplied rend="["&gt;add&lt;/supplied&gt;
     &amp;ldquo;But&lt;supplied rend="]"/&gt; Is not &lt;lb n="27"/&gt;this you
    mentioned &lt;emph&gt;here&lt;/emph&gt; &amp;ndash; in the scriptures?&lt;/q&gt;
   &lt;lb n="28"/&gt;Receives visits from the principal d&amp;apos;s &amp;ndash;
    &lt;q&gt;Gentlemen&lt;/q&gt; &amp;amp;c. &lt;lb n="29"/&gt;
   &lt;emph&gt;Arguments&lt;/emph&gt; to persuade &amp;ndash; &lt;q&gt;Would you not rather
     &lt;lb n="30"/&gt;be below with kings than above with fools?&lt;/q&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;fw&gt;[on recto of last blank leaf of Volume VII, page [523]]&lt;/fw&gt;
   &lt;lb n="31"/&gt;It is better to laugh &amp;amp; not sin than to
    &lt;del&gt;be&lt;/del&gt; weep &amp;amp; be &lt;lb n="32"/&gt;wicked. — Ten loads
   of coal to burn him. — &lt;lb n="33"/&gt;Brought to the stake
   — warmed himself by the fire. &lt;/p&gt;
  &lt;p&gt;
   &lt;lb n="34"/&gt;Ego non baptizo te in nominee Patris et &lt;lb n="35"
   /&gt;Filii et Spiritus Sancti &amp;ndash; sed in nomine &lt;lb n="36"
   /&gt;Diaboli. — Madness is undefinable — &lt;lb n="37"/&gt;It
   &amp;amp; right reasons extremes of one. &lt;lb n="38"/&gt;–Not the
    &lt;supplied rend="["&gt;inserted above line with caret below&lt;/supplied&gt;
   (black art)&lt;supplied rend="]"/&gt; Goetic but Theurgic magic —
    &lt;lb n="39"/&gt;seeks converse with the Intelligence, Power, the &lt;lb
    n="40"/&gt;Angel. &lt;/p&gt;
 &lt;/body&gt;
&lt;/text&gt;
</programlisting>
 </appendix>

 <appendix xml:id="D">
  <title>Melville: HTML presentations</title>
  <para>HTML presentation of the exemplar</para>
  <para>Special exemplar tokens in red.</para>

  <figure xml:id="HTML-NNE-exemplar">
   <title>HTML presentation of the exemplar</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-009.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>

  <para>HTML presentation of the transcript</para>
  <para>Special transcript tokens in red.</para>

  <figure xml:id="HTML-NNE-transcript">
   <title>HTML presentation of the transcript</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-010.jpg" format="jpg" width="70%"/>
    </imageobject>
   </mediaobject>
  </figure>
 </appendix>

 <appendix xml:id="E">
  <title>genFoF stylesheet</title>
  <para> </para>
<programlisting xml:space="preserve">
&amp;lt;?xml version="1.0" encoding="UTF-8"?&gt;
&amp;lt;xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
  xmlns:ym="http://www.marcouxmedias.com"
  xmlns:xs="http://www.w3.org/2001/XMLSchema"
  xpath-default-namespace="" xml:lang="fr-CA" xml:space="ignore"
  exclude-result-prefixes="#all"&gt;

  &amp;lt;xsl:output method="text" indent="no" encoding="UTF-8" /&gt;

  &amp;lt;xsl:function name="ym:lastIndexOf" as="xs:integer"&gt;
    &amp;lt;xsl:param name="str" /&gt;
    &amp;lt;xsl:param name="car" /&gt;
    &amp;lt;xsl:sequence select="if (contains($str, $car)) then
      max(for $i in (1 to string-length($str)) return
      if (substring($str, $i, 1) = $car) then $i else 0)
      else 0" /&gt;
  &amp;lt;/xsl:function&gt;

  &amp;lt;xsl:variable name="fpNoExt"&gt;
    &amp;lt;xsl:variable name="temp" select="base-uri(/)" /&gt;
    &amp;lt;xsl:value-of select="substring($temp, 1, ym:lastIndexOf($temp, '.'))" /&gt;
  &amp;lt;/xsl:variable&gt;

  &amp;lt;xsl:variable name="indent" select="'  '"/&gt;

  &amp;lt;xsl:function name="ym:tokenizePlus" as="item()*"&gt;
    &amp;lt;xsl:param name="text" as="node()" /&gt;
    &amp;lt;xsl:choose&gt;
      &amp;lt;xsl:when test="not($text/ancestor::orig)"&gt;
        &amp;lt;xsl:variable name="toks" select=
          "tokenize($text,'[^a-zA-Z0-9]+')[.]" /&gt;
        &amp;lt;xsl:choose&gt;
&amp;lt;!-- The values of test specified below are project-specific --&gt;
          &amp;lt;xsl:when test="$text/ancestor::front
           or $text/ancestor::fw
           or $text/ancestor::supplied
           or $text/ancestor::sic
           or $text/ancestor::floatingText"&gt;
            &amp;lt;xsl:sequence select="for $t in $toks return '*' || $t" /&gt;
          &amp;lt;/xsl:when&gt;
          &amp;lt;xsl:otherwise&gt;
            &amp;lt;xsl:sequence select="$toks" /&gt;
          &amp;lt;/xsl:otherwise&gt;
        &amp;lt;/xsl:choose&gt;
      &amp;lt;/xsl:when&gt;
      &amp;lt;xsl:otherwise /&gt;
    &amp;lt;/xsl:choose&gt;
  &amp;lt;/xsl:function&gt;

  &amp;lt;xsl:function name="ym:addEffSeqNum" as="item()*"&gt;
    &amp;lt;xsl:param name="toks" as="item()*" /&gt;
    &amp;lt;xsl:sequence select="
      for $n in 1 to count($toks) return
        (if (not(starts-with($toks[$n],'*'))) then
          count($toks[position() lt $n and not(starts-with(.,'*'))]) || '*' else '')
          || $toks[$n]
      " /&gt;
  &amp;lt;/xsl:function&gt;

  &amp;lt;xsl:template match="/"&gt;
    &amp;lt;xsl:variable name="doc1" select="/*"/&gt;
    &amp;lt;xsl:variable name="toks1" select=
      "ym:addEffSeqNum($doc1//text()/ym:tokenizePlus(.))" /&gt;
    &amp;lt;xsl:variable name="doc2" select="document($fpNoExt || 'transcript.xml')/*"/&gt;
    &amp;lt;xsl:variable name="toks2" select=
      "ym:addEffSeqNum($doc2//text()/ym:tokenizePlus(.))" /&gt;
    &amp;lt;xsl:variable name="nNormToks" select="count($toks1[not(starts-with(.,'*'))])" /&gt;
    &amp;lt;xsl:if test="$nNormToks ne count($toks2[not(starts-with(.,'*'))])"&gt;
      &amp;lt;xsl:message terminate="no"
    select="string-join(($nNormToks, count($toks2[not(starts-with(.,'*'))]), ' '), ' ')"
    &gt;WARNING: Normal token counts differ between documents.&amp;lt;/xsl:message&gt;
    &amp;lt;/xsl:if&gt;

    &amp;lt;xsl:text&gt;fof(case_specific_facts, axiom,
&amp;lt;/xsl:text&gt;

    &amp;lt;xsl:call-template name="processDoc"&gt;
      &amp;lt;xsl:with-param name="docPrefix" select="'e'" /&gt;
      &amp;lt;xsl:with-param name="toks" select="$toks1" /&gt;
    &amp;lt;/xsl:call-template&gt;

    &amp;lt;xsl:call-template name="processDoc"&gt;
      &amp;lt;xsl:with-param name="docPrefix" select="'t'" /&gt;
      &amp;lt;xsl:with-param name="toks" select="$toks2" /&gt;
    &amp;lt;/xsl:call-template&gt;

    &amp;lt;xsl:for-each select="$doc2//choice"&gt;
&amp;lt;!-- The [.] in the following are to get rid of possible empty strings at the
beginning and end of the string: --&gt;
      &amp;lt;xsl:variable name="toksOrig" select="tokenize(orig,'[^a-zA-Z0-9]+')[.]"/&gt;
      &amp;lt;xsl:variable name="toksReg" select="tokenize(reg,'[^a-zA-Z0-9]+')[.]"/&gt;
      &amp;lt;xsl:if test="count($toksOrig) ne count($toksReg) or not(count($toksOrig))"&gt;
        &amp;lt;xsl:message select="string-join((count($toksOrig), count($toksReg), ' '), ' ')"
          &gt;WARNING: Token count mismatch within a &amp;lt;choice&gt; element.&amp;lt;/xsl:message&gt;
      &amp;lt;/xsl:if&gt;
      &amp;lt;xsl:for-each select="1 to count($toksOrig)"&gt;
        &amp;lt;xsl:if test="$toksOrig[current()] ne $toksReg[current()]"&gt;
          &amp;lt;xsl:value-of select="$indent" /&gt;
          &amp;lt;xsl:text&gt;typesimilar(t_&amp;lt;/xsl:text&gt;
          &amp;lt;xsl:value-of select="$toksOrig[current()]" /&gt;
          &amp;lt;xsl:text&gt;, t_&amp;lt;/xsl:text&gt;
          &amp;lt;xsl:value-of select="$toksReg[current()]" /&gt;
          &amp;lt;xsl:text&gt;) &amp;amp;
&amp;lt;/xsl:text&gt;
        &amp;lt;/xsl:if&gt;
      &amp;lt;/xsl:for-each&gt;
    &amp;lt;/xsl:for-each&gt;

    &amp;lt;xsl:for-each select="1 to count($toks1)"&gt;
      &amp;lt;xsl:if test="not(starts-with($toks1[current()],'*'))"&gt;
        &amp;lt;xsl:variable name="idNormTok" select="substring-before($toks1[current()],'*')" /&gt;
        &amp;lt;xsl:value-of select="$indent" /&gt;
        &amp;lt;xsl:text&gt;transcript(e&amp;lt;/xsl:text&gt;
        &amp;lt;xsl:value-of select="." /&gt;
        &amp;lt;xsl:text&gt;,t&amp;lt;/xsl:text&gt;
        &amp;lt;xsl:value-of select="max(
          for $i in 1 to count($toks2) return
            if (starts-with($toks2[$i], $idNormTok || '*')) then $i else 0
          )" /&gt;
        &amp;lt;xsl:text&gt;) &amp;amp;
&amp;lt;/xsl:text&gt;
      &amp;lt;/xsl:if&gt;
    &amp;lt;/xsl:for-each&gt;

&amp;lt;!--% No other pairs of corresponding tokens (gives us reciprocity for free):--&gt;
   &amp;lt;xsl:text&gt;
  ! [X,Y,Z] : (
              ((transcript(X,Y) &amp;amp; transcript(X,Z)) =&gt; Y=Z) &amp;amp;
              ((exemplar(X,Y) &amp;amp; exemplar(X,Z)) =&gt; Y=Z)
              ) &amp;amp;

&amp;lt;/xsl:text&gt;

    &amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:text&gt;$distinct(
&amp;lt;/xsl:text&gt;&amp;lt;xsl:value-of select="$indent" /&gt;&amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:value-of select="
      string-join(for $i in 1 to count($toks1) return ('e' || $i), ',')
      " /&gt;
    &amp;lt;xsl:text&gt;,
&amp;lt;/xsl:text&gt;&amp;lt;xsl:value-of select="$indent" /&gt;&amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:value-of select="
      string-join(for $i in 1 to count($toks2) return ('t' || $i), ',')
      " /&gt;
    &amp;lt;xsl:text&gt;,
&amp;lt;/xsl:text&gt;&amp;lt;xsl:value-of select="$indent" /&gt;&amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:variable name="allTypes" select="
      (for $i in 1 to count($toks1) return 't_' || substring-after($toks1[$i], '*')),
      (for $i in 1 to count($toks2) return 't_' || substring-after($toks2[$i], '*'))
      "/&gt;
    &amp;lt;xsl:value-of select="string-join(distinct-values($allTypes), ',')" /&gt;
    &amp;lt;xsl:text&gt;
&amp;lt;/xsl:text&gt;
    &amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:text&gt;) &amp;amp;
&amp;lt;/xsl:text&gt;

    &amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:text&gt;! [X] : ((e_token(X) &amp;lt;=&gt; (X=e&amp;lt;/xsl:text&gt;
    &amp;lt;xsl:value-of select="string-join(
      (1 to count($toks1))[let $i := . return not(starts-with($toks1[$i],'*'))],
      ' | X=e')" /&gt;
    &amp;lt;xsl:text&gt;)) &amp;amp;
&amp;lt;/xsl:text&gt;
    &amp;lt;xsl:value-of select="$indent" /&gt;
    &amp;lt;xsl:text&gt;         (t_token(X) &amp;lt;=&gt; (X=t&amp;lt;/xsl:text&gt;
    &amp;lt;xsl:value-of select="string-join(
      (1 to count($toks2))[let $i := . return not(starts-with($toks2[$i],'*'))],
      ' | X=t')" /&gt;
    &amp;lt;xsl:text&gt;)))
).&amp;lt;/xsl:text&gt;

  &amp;lt;/xsl:template&gt;

  &amp;lt;xsl:template name="processDoc"&gt;
    &amp;lt;xsl:param name="toks" /&gt;
    &amp;lt;xsl:param name="docPrefix" /&gt;
&amp;lt;!--&amp;lt;xsl:message select="$toks" /&gt;--&gt;
    &amp;lt;xsl:for-each select="$toks"&gt;
&amp;lt;!--
    &amp;lt;xsl:value-of select="$indent" /&gt;
      &amp;lt;xsl:value-of select="if (starts-with(.,'*')) then 's' else ''"/&gt;
      &amp;lt;xsl:value-of select="$docPrefix"/&gt;
      &amp;lt;xsl:text&gt;_token(&amp;lt;/xsl:text&gt;
      &amp;lt;xsl:value-of select="$docPrefix"/&gt;
      &amp;lt;xsl:value-of select="position()" /&gt;
      &amp;lt;xsl:text&gt;) &amp;amp;
    &amp;lt;/xsl:text&gt;
^ 2026-03-31 Zoom meeting, decided we don’t need them. --&gt;
      &amp;lt;xsl:value-of select="$indent" /&gt;
      &amp;lt;xsl:text&gt;typeof(&amp;lt;/xsl:text&gt;
      &amp;lt;xsl:value-of select="$docPrefix"/&gt;
      &amp;lt;xsl:value-of select="position()"/&gt;
      &amp;lt;xsl:text&gt;, t_&amp;lt;/xsl:text&gt;
      &amp;lt;xsl:value-of select="substring-after(.,'*')"/&gt;
      &amp;lt;xsl:text&gt;) &amp;amp;
&amp;lt;/xsl:text&gt;
    &amp;lt;/xsl:for-each&gt;
  &amp;lt;/xsl:template&gt;

&amp;lt;/xsl:stylesheet&gt;
</programlisting>
 </appendix>

 <appendix xml:id="F">
  <title>Melville: FOF representation of exemplar and transcript</title>
  <para> </para>
<programlisting xml:space="preserve">
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% General axioms
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(transcript_domain, axiom,
  ! [X,Y] : (transcript(X,Y) =&gt; (e_token(X) &amp; t_token(Y)))).

fof(exemplar_domain, axiom,
  ! [X,Y] : (exemplar(X,Y) =&gt; (t_token(X) &amp; e_token(Y)))).

%%%% Slightly modified from Default:
fof(typeof_domain, axiom,
  ! [X,Y] : (typeof(X,Y) =&gt; 
            ((t_token(X) | e_token(X) | st_token(X) | se_token(X)) 
             &amp; type(Y)))).

%%%% Slightly modified from Default:
fof(no_class_overlap, axiom,
  ! [X] : (
  (e_token(X) &lt;=&gt; (~t_token(X) &amp; ~type(X) &amp; ~se_token(X) &amp; ~st_token(X))) &amp;
  (t_token(X) &lt;=&gt; (~e_token(X) &amp; ~type(X) &amp; ~se_token(X) &amp; ~st_token(X))) &amp;
  (type(X) &lt;=&gt; (~e_token(X) &amp; ~t_token(X)) &amp; ~se_token(X) &amp; ~st_token(X)) &amp;
  (se_token(X) &lt;=&gt; (~t_token(X) &amp; ~type(X) &amp; ~e_token(X) &amp; ~st_token(X))) &amp;
  (st_token(X) &lt;=&gt; (~t_token(X) &amp; ~type(X) &amp; ~se_token(X) &amp; ~e_token(X)))
  )).

fof(exemplar_and_transcript_inverse, axiom,
  ! [X,Y] : (exemplar(X,Y) &lt;=&gt; transcript(Y,X))).

fof(at_most_one_type_per_token, axiom,
  ! [X,Y,Z] : ((typeof(X,Y) &amp; typeof(X,Z)) =&gt; Y=Z)).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Axioms not in Default

fof(typesimilar_domain, axiom,
  ! [X,Y] : (typesimilar(X,Y) =&gt; (type(X) &amp; type(Y)))).

fof(typesimilar_reflexive, axiom,
  ! [X] : (type(X) =&gt; typesimilar(X,X))).

fof(typesimilar_symmetric, axiom,
  ! [X,Y] : (typesimilar(X,Y) =&gt; typesimilar(Y,X))).

fof(type_similar_transitive, axiom,
  ! [X,Y,Z] : ((typesimilar(X,Y) &amp; typesimilar(Y,Z)) =&gt; typesimilar(X,Z))).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Definitions
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(reciprocity,definition,
  reciprocity &lt;=&gt; ! [X,Y,Z] : (
              ((transcript(X,Y) &amp; transcript(X,Z)) =&gt; Y=Z) &amp;
              ((exemplar(X,Y) &amp; exemplar(X,Z)) =&gt; Y=Z))).

fof(completeness, definition,
  completeness &lt;=&gt; ! [X] : 
          (e_token(X) =&gt; ? [Y] : transcript(X,Y))).

fof(purity, definition,
  purity &lt;=&gt; ! [X] :
          (t_token(X) =&gt; ? [Y] : exemplar(X,Y))).

%%%% Slightly modified from Default:
fof(type_similarity, definition,
  type_similarity &lt;=&gt; ![X,Y] : (transcript(X,Y) =&gt;
    ? [Z,V] : (typeof(X,Z) &amp; typeof(Y,V) &amp; typesimilar(Z,V)))).

%%%% Slightly modified from Default:
fof(t_similarity, definition,
  t_similarity &lt;=&gt; (reciprocity &amp; completeness &amp; purity &amp; type_similarity)).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Case-specific facts
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(case_specific_facts, axiom,
  typeof(e1, t_A) &amp;
  typeof(e2, t_seaman) &amp;
  typeof(e3, t_figures) &amp;
  typeof(e4, t_in) &amp;
  typeof(e5, t_The) &amp;
  typeof(e6, t_Canterbury) &amp;
  typeof(e7, t_Tales) &amp;
  typeof(e8, t_With) &amp;
  typeof(e9, t_a) &amp;
  typeof(e10, t_many) &amp;
  typeof(e11, t_a) &amp;
  typeof(e12, t_tempest) &amp;
  typeof(e13, t_had) &amp;
  typeof(e14, t_his) &amp;
  typeof(e15, t_beard) &amp;
  typeof(e16, t_been) &amp;
  typeof(e17, t_shook) &amp;
  typeof(e18, t_S) &amp;
  typeof(e19, t_Deep) &amp;
  typeof(e20, t_g) &amp;
%%% ...Approximately 250-350 lines omitted...
  typeof(e270, t_right) &amp;
  typeof(e271, t_reasons) &amp;
  typeof(e272, t_extremes) &amp;
  typeof(e273, t_of) &amp;
  typeof(e274, t_one) &amp;
  typeof(e275, t_Not) &amp;
  typeof(e276, t_the) &amp;
  typeof(e277, t_black) &amp;
  typeof(e278, t_art) &amp;
  typeof(e279, t_Goetic) &amp;
  typeof(e280, t_but) &amp;
  typeof(e281, t_Theurgic) &amp;
  typeof(e282, t_magic) &amp;
  typeof(e283, t_seeks) &amp;
  typeof(e284, t_converse) &amp;
  typeof(e285, t_with) &amp;
  typeof(e286, t_the) &amp;
  typeof(e287, t_Intelligence) &amp;
  typeof(e288, t_Power) &amp;
  typeof(e289, t_the) &amp;
%%% ...Approximately 250-350 lines omitted...
  typeof(t332, t_inserted) &amp;
  typeof(t333, t_above) &amp;
  typeof(t334, t_line) &amp;
  typeof(t335, t_with) &amp;
  typeof(t336, t_caret) &amp;
  typeof(t337, t_below) &amp;
  typeof(t338, t_black) &amp;
  typeof(t339, t_art) &amp;
  typeof(t340, t_Goetic) &amp;
  typeof(t341, t_but) &amp;
  typeof(t342, t_Theurgic) &amp;
  typeof(t343, t_magic) &amp;
  typeof(t344, t_seeks) &amp;
  typeof(t345, t_converse) &amp;
  typeof(t346, t_with) &amp;
  typeof(t347, t_the) &amp;
  typeof(t348, t_Intelligence) &amp;
  typeof(t349, t_Power) &amp;
  typeof(t350, t_the) &amp;
  typeof(t351, t_Angel) &amp;
%%% ...Approximately 250-350 lines omitted...
  transcript(e271,t326) &amp;
  transcript(e272,t327) &amp;
  transcript(e273,t328) &amp;
  transcript(e274,t329) &amp;
  transcript(e275,t330) &amp;
  transcript(e276,t331) &amp;
  transcript(e277,t338) &amp;
  transcript(e278,t339) &amp;
  transcript(e279,t340) &amp;
  transcript(e280,t341) &amp;
  transcript(e281,t342) &amp;
  transcript(e282,t343) &amp;
  transcript(e283,t344) &amp;
  transcript(e284,t345) &amp;
  transcript(e285,t346) &amp;
  transcript(e286,t347) &amp;
  transcript(e287,t348) &amp;
  transcript(e288,t349) &amp;
  transcript(e289,t350) &amp;
  transcript(e290,t351) &amp;

% No other pairs of corresponding tokens (gives us reciprocity for free):
  ! [X,Y,Z] : (
              ((transcript(X,Y) &amp; transcript(X,Z)) =&gt; Y=Z) &amp;
              ((exemplar(X,Y) &amp; exemplar(X,Z)) =&gt; Y=Z)
              ) &amp;

  $distinct(
    e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13,e14,e15,e16,e17,e18,e19,e20,e21,e22,e23,e24,e25,e26,e27,e28,e29,e30,e31,e32,e33,e34,e35,e36,e37,e38,e39,e40,e41,e42,e43,e44,e45,e46,e47,e48,e49,e50,e51,e52,e53,e54,e55,e56,e57,e58,e59,e60,e61,e62,e63,e64,e65,e66,e67,e68,e69,e70,e71,e72,e73,e74,e75,e76,e77,e78,e79,e80,e81,e82,e83,e84,e85,e86,e87,e88,e89,e90,e91,e92,e93,e94,e95,e96,e97,e98,e99,e100,e101,e102,e103,e104,e105,e106,e107,e108,e109,e110,e111,e112,e113,e114,e115,e116,e117,e118,e119,e120,e121,e122,e123,e124,e125,e126,e127,e128,e129,e130,e131,e132,e133,e134,e135,e136,e137,e138,e139,e140,e141,e142,e143,e144,e145,e146,e147,e148,e149,e150,e151,e152,e153,e154,e155,e156,e157,e158,e159,e160,e161,e162,e163,e164,e165,e166,e167,e168,e169,e170,e171,e172,e173,e174,e175,e176,e177,e178,e179,e180,e181,e182,e183,e184,e185,e186,e187,e188,e189,e190,e191,e192,e193,e194,e195,e196,e197,e198,e199,e200,e201,e202,e203,e204,e205,e206,e207,e208,e209,e210,e211,e212,e213,e214,e215,e216,e217,e218,e219,e220,e221,e222,e223,e224,e225,e226,e227,e228,e229,e230,e231,e232,e233,e234,e235,e236,e237,e238,e239,e240,e241,e242,e243,e244,e245,e246,e247,e248,e249,e250,e251,e252,e253,e254,e255,e256,e257,e258,e259,e260,e261,e262,e263,e264,e265,e266,e267,e268,e269,e270,e271,e272,e273,e274,e275,e276,e277,e278,e279,e280,e281,e282,e283,e284,e285,e286,e287,e288,e289,e290,
    t1,t2,t3,t4,t5,t6,t7,t8,t9,t10,t11,t12,t13,t14,t15,t16,t17,t18,t19,t20,t21,t22,t23,t24,t25,t26,t27,t28,t29,t30,t31,t32,t33,t34,t35,t36,t37,t38,t39,t40,t41,t42,t43,t44,t45,t46,t47,t48,t49,t50,t51,t52,t53,t54,t55,t56,t57,t58,t59,t60,t61,t62,t63,t64,t65,t66,t67,t68,t69,t70,t71,t72,t73,t74,t75,t76,t77,t78,t79,t80,t81,t82,t83,t84,t85,t86,t87,t88,t89,t90,t91,t92,t93,t94,t95,t96,t97,t98,t99,t100,t101,t102,t103,t104,t105,t106,t107,t108,t109,t110,t111,t112,t113,t114,t115,t116,t117,t118,t119,t120,t121,t122,t123,t124,t125,t126,t127,t128,t129,t130,t131,t132,t133,t134,t135,t136,t137,t138,t139,t140,t141,t142,t143,t144,t145,t146,t147,t148,t149,t150,t151,t152,t153,t154,t155,t156,t157,t158,t159,t160,t161,t162,t163,t164,t165,t166,t167,t168,t169,t170,t171,t172,t173,t174,t175,t176,t177,t178,t179,t180,t181,t182,t183,t184,t185,t186,t187,t188,t189,t190,t191,t192,t193,t194,t195,t196,t197,t198,t199,t200,t201,t202,t203,t204,t205,t206,t207,t208,t209,t210,t211,t212,t213,t214,t215,t216,t217,t218,t219,t220,t221,t222,t223,t224,t225,t226,t227,t228,t229,t230,t231,t232,t233,t234,t235,t236,t237,t238,t239,t240,t241,t242,t243,t244,t245,t246,t247,t248,t249,t250,t251,t252,t253,t254,t255,t256,t257,t258,t259,t260,t261,t262,t263,t264,t265,t266,t267,t268,t269,t270,t271,t272,t273,t274,t275,t276,t277,t278,t279,t280,t281,t282,t283,t284,t285,t286,t287,t288,t289,t290,t291,t292,t293,t294,t295,t296,t297,t298,t299,t300,t301,t302,t303,t304,t305,t306,t307,t308,t309,t310,t311,t312,t313,t314,t315,t316,t317,t318,t319,t320,t321,t322,t323,t324,t325,t326,t327,t328,t329,t330,t331,t332,t333,t334,t335,t336,t337,t338,t339,t340,t341,t342,t343,t344,t345,t346,t347,t348,t349,t350,t351,
    t_A,t_seaman,t_figures,t_in,t_The,t_Canterbury,t_Tales,t_With,t_a,t_many,t_tempest,t_had,t_his,t_beard,t_been,t_shook,t_S,t_Deep,t_g,t_Secret,t_grief,t_is,t_cannibal,t_of,t_its,t_own,t_heart,t_Bacon,t_Claudia,t_the,t_Appian,t_family,t_I,t_wish,t_some,t_fight,t_or,t_pestilence,t_would,t_thin,t_out,t_this,t_crowd,t_Arrogance,t_Roast,t_beef,t_pulpet,t_An,t_animal,t_man,t_do,t_eagles,t_wear,t_spectacles,t_Health,t_Contrast,t_an,t_over,t_spiritual,t_Yes,t_Madam,t_Cain,t_was,t_godless,t_froward,t_boy,t_Reuben,t_Gen,t_49,t_Absalom,t_Many,t_pious,t_men,t_have,t_impious,t_children,t_Devil,t_as,t_Quaker,t_formal,t_compact,t_Imprimis,t_First,t_Second,t_aforesaid,t_soul,t_said,t_c,t_Duplicates,t_How,t_it,t_about,t_temptation,t_on,t_hill,t_D,t_begs,t_hero,t_to,t_form,t_one,t_Society,t_s,t_name,t_be,t_weighty,t_Leaves,t_letter,t_My,t_Dear,t_Conversation,t_upon,t_Gabriel,t_Micheal,t_Raphel,t_gentlemanly,t_Terra,t_Oblivionis,t_Hellites,t_At,t_Astor,t_find,t_him,t_making,t_almanacks,t_going,t_ball,t_takes,t_long,t_time,t_toilette,t_Doctor,t_coach,t_stops,t_way,t_Do,t_you,t_believe,t_all,t_that,t_stuff,t_nonsence,t_world,t_never,t_made,t_But,t_Is,t_not,t_mentioned,t_here,t_scriptures,t_Receives,t_visits,t_from,t_principal,t_d,t_Gentlemen,t_Arguments,t_persuade,t_Would,t_rather,t_below,t_with,t_kings,t_than,t_above,t_fools,t_It,t_better,t_laugh,t_sin,t_weep,t_wicked,t_Ten,t_loads,t_coal,t_burn,t_Brought,t_stake,t_warmed,t_himself,t_by,t_fire,t_Ego,t_non,t_baptizo,t_te,t_nominee,t_Patris,t_et,t_Filii,t_Spiritus,t_Sancti,t_sed,t_nomine,t_Diaboli,t_Madness,t_undefinable,t_right,t_reasons,t_extremes,t_Not,t_black,t_art,t_Goetic,t_but,t_Theurgic,t_magic,t_seeks,t_converse,t_Intelligence,t_Power,t_Angel,t_NOTES,t_IN,t_SHAKESPEARE,t_VOLUME,t_969,t_verso,t_last,t_leaf,t_Volume,t_VII,t_page,t_524,t_inserted,t_later,t_lines,t_21,t_21b,t_after,t_and,t_circled,t_guideline,t_caret,t_insertion,t_reported,t_line,t_18,t_alm,t_add,t_recto,t_blank,t_523
  ) &amp;
  ! [X] : ((e_token(X) &lt;=&gt; (X=e1 | X=e2 | X=e3 | X=e4 | X=e5 | X=e6 | X=e7 | X=e8 | X=e9 | X=e10 | X=e11 | X=e12 | X=e13 | X=e14 | X=e15 | X=e16 | X=e17 | X=e18 | X=e19 | X=e20 | X=e21 | X=e22 | X=e23 | X=e24 | X=e25 | X=e26 | X=e27 | X=e28 | X=e29 | X=e30 | X=e31 | X=e32 | X=e33 | X=e34 | X=e35 | X=e36 | X=e37 | X=e38 | X=e39 | X=e40 | X=e41 | X=e42 | X=e43 | X=e44 | X=e45 | X=e46 | X=e47 | X=e48 | X=e49 | X=e50 | X=e51 | X=e52 | X=e53 | X=e54 | X=e55 | X=e56 | X=e57 | X=e58 | X=e59 | X=e60 | X=e61 | X=e62 | X=e63 | X=e64 | X=e65 | X=e66 | X=e67 | X=e68 | X=e69 | X=e70 | X=e71 | X=e72 | X=e73 | X=e74 | X=e75 | X=e76 | X=e77 | X=e78 | X=e79 | X=e80 | X=e81 | X=e82 | X=e83 | X=e84 | X=e85 | X=e86 | X=e87 | X=e88 | X=e89 | X=e90 | X=e91 | X=e92 | X=e93 | X=e94 | X=e95 | X=e96 | X=e97 | X=e98 | X=e99 | X=e100 | X=e101 | X=e102 | X=e103 | X=e104 | X=e105 | X=e106 | X=e107 | X=e108 | X=e109 | X=e110 | X=e111 | X=e112 | X=e113 | X=e114 | X=e115 | X=e116 | X=e117 | X=e118 | X=e119 | X=e120 | X=e121 | X=e122 | X=e123 | X=e124 | X=e125 | X=e126 | X=e127 | X=e128 | X=e129 | X=e130 | X=e131 | X=e132 | X=e133 | X=e134 | X=e135 | X=e136 | X=e137 | X=e138 | X=e139 | X=e140 | X=e148 | X=e149 | X=e150 | X=e151 | X=e152 | X=e153 | X=e154 | X=e155 | X=e156 | X=e158 | X=e159 | X=e160 | X=e161 | X=e162 | X=e163 | X=e164 | X=e165 | X=e166 | X=e167 | X=e168 | X=e169 | X=e170 | X=e171 | X=e172 | X=e173 | X=e174 | X=e175 | X=e176 | X=e177 | X=e178 | X=e179 | X=e180 | X=e181 | X=e182 | X=e183 | X=e184 | X=e185 | X=e186 | X=e187 | X=e188 | X=e189 | X=e190 | X=e191 | X=e192 | X=e193 | X=e194 | X=e195 | X=e196 | X=e197 | X=e198 | X=e199 | X=e200 | X=e201 | X=e202 | X=e203 | X=e204 | X=e205 | X=e206 | X=e207 | X=e208 | X=e209 | X=e210 | X=e211 | X=e212 | X=e213 | X=e214 | X=e215 | X=e216 | X=e217 | X=e218 | X=e219 | X=e220 | X=e221 | X=e222 | X=e223 | X=e224 | X=e225 | X=e226 | X=e227 | X=e228 | X=e229 | X=e230 | X=e231 | X=e232 | X=e233 | X=e234 | X=e235 | X=e236 | X=e237 | X=e238 | X=e239 | X=e240 | X=e241 | X=e242 | X=e243 | X=e244 | X=e245 | X=e246 | X=e247 | X=e248 | X=e249 | X=e250 | X=e251 | X=e252 | X=e253 | X=e254 | X=e255 | X=e256 | X=e257 | X=e258 | X=e259 | X=e260 | X=e261 | X=e262 | X=e263 | X=e264 | X=e265 | X=e266 | X=e267 | X=e268 | X=e269 | X=e270 | X=e271 | X=e272 | X=e273 | X=e274 | X=e275 | X=e276 | X=e277 | X=e278 | X=e279 | X=e280 | X=e281 | X=e282 | X=e283 | X=e284 | X=e285 | X=e286 | X=e287 | X=e288 | X=e289 | X=e290)) &amp;
           (t_token(X) &lt;=&gt; (X=t17 | X=t18 | X=t19 | X=t20 | X=t21 | X=t22 | X=t23 | X=t24 | X=t25 | X=t26 | X=t27 | X=t28 | X=t29 | X=t30 | X=t31 | X=t32 | X=t33 | X=t34 | X=t35 | X=t36 | X=t37 | X=t38 | X=t39 | X=t40 | X=t41 | X=t42 | X=t43 | X=t44 | X=t45 | X=t46 | X=t47 | X=t48 | X=t49 | X=t50 | X=t51 | X=t52 | X=t53 | X=t54 | X=t55 | X=t56 | X=t57 | X=t58 | X=t59 | X=t60 | X=t61 | X=t62 | X=t63 | X=t64 | X=t65 | X=t66 | X=t67 | X=t68 | X=t69 | X=t70 | X=t71 | X=t72 | X=t73 | X=t74 | X=t75 | X=t76 | X=t77 | X=t78 | X=t79 | X=t80 | X=t81 | X=t82 | X=t83 | X=t84 | X=t85 | X=t86 | X=t87 | X=t88 | X=t89 | X=t90 | X=t91 | X=t92 | X=t93 | X=t94 | X=t95 | X=t96 | X=t97 | X=t98 | X=t99 | X=t100 | X=t101 | X=t102 | X=t103 | X=t104 | X=t105 | X=t106 | X=t107 | X=t108 | X=t109 | X=t110 | X=t111 | X=t112 | X=t113 | X=t114 | X=t115 | X=t116 | X=t117 | X=t118 | X=t119 | X=t120 | X=t121 | X=t122 | X=t123 | X=t124 | X=t125 | X=t126 | X=t127 | X=t128 | X=t153 | X=t154 | X=t155 | X=t156 | X=t157 | X=t158 | X=t159 | X=t160 | X=t161 | X=t162 | X=t163 | X=t164 | X=t165 | X=t166 | X=t167 | X=t168 | X=t169 | X=t170 | X=t171 | X=t172 | X=t173 | X=t174 | X=t175 | X=t176 | X=t177 | X=t178 | X=t179 | X=t180 | X=t191 | X=t192 | X=t193 | X=t194 | X=t195 | X=t196 | X=t197 | X=t198 | X=t199 | X=t201 | X=t202 | X=t203 | X=t204 | X=t205 | X=t206 | X=t207 | X=t208 | X=t209 | X=t210 | X=t211 | X=t212 | X=t213 | X=t214 | X=t215 | X=t216 | X=t217 | X=t218 | X=t219 | X=t220 | X=t221 | X=t222 | X=t223 | X=t224 | X=t225 | X=t226 | X=t227 | X=t228 | X=t229 | X=t231 | X=t232 | X=t233 | X=t234 | X=t235 | X=t236 | X=t237 | X=t238 | X=t239 | X=t240 | X=t241 | X=t242 | X=t243 | X=t244 | X=t245 | X=t246 | X=t247 | X=t248 | X=t249 | X=t250 | X=t251 | X=t252 | X=t253 | X=t254 | X=t255 | X=t256 | X=t257 | X=t258 | X=t259 | X=t260 | X=t261 | X=t262 | X=t263 | X=t264 | X=t276 | X=t277 | X=t278 | X=t279 | X=t280 | X=t281 | X=t282 | X=t283 | X=t284 | X=t285 | X=t286 | X=t287 | X=t288 | X=t289 | X=t290 | X=t291 | X=t292 | X=t293 | X=t294 | X=t295 | X=t296 | X=t297 | X=t298 | X=t299 | X=t300 | X=t301 | X=t302 | X=t303 | X=t304 | X=t305 | X=t306 | X=t307 | X=t308 | X=t309 | X=t310 | X=t311 | X=t312 | X=t313 | X=t314 | X=t315 | X=t316 | X=t317 | X=t318 | X=t319 | X=t320 | X=t321 | X=t322 | X=t323 | X=t324 | X=t325 | X=t326 | X=t327 | X=t328 | X=t329 | X=t330 | X=t331 | X=t338 | X=t339 | X=t340 | X=t341 | X=t342 | X=t343 | X=t344 | X=t345 | X=t346 | X=t347 | X=t348 | X=t349 | X=t350 | X=t351)))
).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Conjecture
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(t_similar, conjecture,
 completeness &amp; purity  &amp;  reciprocity  &amp; type_similarity 
).

</programlisting>
 </appendix>

 <appendix xml:id="G">
  <title>Melville: XML, HTML, and FOF of normalized transcript</title>
  <para>XML representation of exemplar for normalized
   transcript</para>
<programlisting xml:space="preserve">
&lt;?xml version="1.0" encoding="UTF-8"?&gt;
&lt;?xml-stylesheet type="text/xsl" href="../../genFofXSLT/genFof-Norm.xsl"?&gt;
&lt;!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
 &lt;!ENTITY mdash  "—" &gt;&lt;!--=em dash--&gt;
 &lt;!ENTITY ndash  "–" &gt;&lt;!--=en dash--&gt;
 &lt;!ENTITY vbar  "|" &gt;&lt;!--=vertical bar--&gt;
 &lt;!ENTITY apos  "—" &gt;&lt;!--=apostrophe--&gt;
 &lt;!ENTITY triplebar "|" &gt;&lt;!--/triple horizontal bars--&gt;
 &lt;!ENTITY ldquo  "“" &gt;&lt;!--=double quotation mark, left--&gt;
 &lt;!ENTITY uarr   "↑" &gt;&lt;!--/uparrow A: =upward arrow--&gt;
&lt;!-- For HTML: --&gt;
 &lt;!ENTITY and "&amp;amp;" &gt;
 &lt;!ENTITY etc "&amp;amp;c" &gt;
&lt;!-- FOR Fof: --&gt;
 &lt;!-- &lt;!ENTITY and "amp" &gt;
 &lt;!ENTITY etc "ampc" &gt;--&gt;
]&gt;
&lt;text&gt;
 &lt;body&gt;
  &lt;p&gt;
   &lt;lb/&gt;A seaman figures in The Canterbury Tales. &lt;lb/&gt;With
    &lt;del&gt;a&lt;/del&gt;many a tempest had his beard been &lt;lb/&gt;shook. &amp;ndash;
    &lt;del&gt;S&lt;/del&gt;
   &lt;del&gt;Deep g&lt;/del&gt; Secret grief is a &lt;lb/&gt;cannibal of its own heart
   &amp;ndash; &lt;emph&gt;Bacon&lt;/emph&gt;.
    &lt;lb/&gt;&lt;metamark&gt;&amp;vbar;&lt;/metamark&gt;&lt;emph&gt;Claudia&lt;/emph&gt; of the
    &lt;emph&gt;Appian&lt;/emph&gt; family, &lt;q&gt;I wish
     &lt;lb/&gt;&lt;metamark&gt;&amp;vbar;&lt;/metamark&gt; some fight or pestilence would
    thin out &lt;lb/&gt;&lt;metamark&gt;&amp;vbar;&lt;/metamark&gt; this crowd.&lt;/q&gt;
   &lt;emph&gt;Arrogance.&lt;/emph&gt;
   &lt;lb/&gt;
   &lt;emph&gt;Roast beef in the pulpet.&lt;/emph&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;An animal of a man — &lt;q&gt;do eagles wear
    &lt;lb/&gt;spectacles?&lt;/q&gt;— Health. — &lt;emph&gt;Contrast&lt;/emph&gt;:
   an &lt;lb/&gt;over spiritual man. &lt;lb/&gt;&lt;metamark&gt;——
   &lt;/metamark&gt;&lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;
   &lt;q&gt;Yes, Madam, Cain was a godless froward boy, ∧ &lt;lb/&gt;Reuben
    (Gen:49) ∧ Absalom&lt;/q&gt; Many pious men &lt;lb/&gt;have impious
   children — (Devil as a Quaker)
    &lt;lb/&gt;&lt;metamark&gt;—————&lt;/metamark&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;A formal compact &amp;ndash; Imprimis &amp;ndash; First &amp;ndash;
   Second. &lt;lb/&gt;The aforesaid soul&lt;sic&gt;. said soul&lt;/sic&gt; &amp;etc; &amp;ndash;
   Duplicates &amp;ndash; &lt;lb/&gt;&lt;metamark&gt;&amp;triplebar;&lt;/metamark&gt;&lt;q&gt;How was
    it about the temptation on the &lt;lb/&gt;hill?&lt;/q&gt;&amp;etc; &amp;ndash; D begs
   the hero to form &lt;lb/&gt;one of a &lt;emph&gt;
    &lt;q&gt;Society of D&amp;apos;s&lt;/q&gt;
   &lt;/emph&gt; &amp;ndash; his name would be weighty &lt;lb/&gt;&amp;etc; &amp;ndash; Leaves
   a letter to the D &amp;ndash; &lt;q&gt;My &lt;lb/&gt;
    &lt;emph&gt;Dear D&lt;/emph&gt;
   &lt;/q&gt;
   &lt;floatingText&gt;
    &lt;body&gt;
     &lt;ab&gt; &amp;ndash;  Conversation upon Gabriel, Micheal ∧
      &lt;lb/&gt; Raphel &amp;ndash; gentlemanly &amp;etc;&lt;/ab&gt;
    &lt;/body&gt;
   &lt;/floatingText&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;
   &lt;q&gt;Terra Oblivionis&lt;/q&gt;
   &lt;q&gt;Hellites&lt;/q&gt; &amp;ndash; At the Astor find him &lt;lb/&gt; making
   almanacks &amp;ndash; going to a ball takes a long &lt;lb/&gt;time making
   toilette. &amp;ndash; The Doctor&amp;apos;s coach stops &lt;lb/&gt;the way.
   &amp;ndash; &lt;q&gt;Do you believe all that stuff? &lt;lb/&gt; nonsence &amp;ndash;
    the world was never made. &amp;ndash; &amp;ldquo;But Is not &lt;lb/&gt;this you
    mentioned &lt;emph&gt;here&lt;/emph&gt; &amp;ndash; in the scriptures?&lt;/q&gt;
   &lt;lb/&gt;Receives visits from the principal d&amp;apos;s &amp;ndash;
    &lt;q&gt;Gentlemen&lt;/q&gt; &amp;etc;. &lt;lb/&gt;
   &lt;emph&gt;Arguments&lt;/emph&gt; to persuade &amp;ndash; &lt;q&gt;Would you not rather
    &lt;lb/&gt;be below with kings than above with fools?&lt;/q&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;It is better to laugh ∧ not sin than to &lt;del&gt;be&lt;/del&gt; weep
   ∧ be &lt;lb/&gt;wicked. — Ten loads of coal to burn him.
   — &lt;lb/&gt;Brought to the stake — warmed himself by the
   fire. &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;Ego non baptizo te in nominee Patris et &lt;lb/&gt;Filii et Spiritus
   Sancti &amp;ndash; sed in nomine &lt;lb/&gt;Diaboli. — Madness is
   undefinable — &lt;lb/&gt;It ∧ right reasons extremes of one.
   &lt;lb/&gt;–Not the &lt;add place="sup"&gt;(black art)&lt;/add&gt; Goetic but
   Theurgic magic — &lt;lb/&gt;seeks converse with the Intelligence,
   Power, the &lt;lb/&gt;Angel. &lt;/p&gt;
 &lt;/body&gt;
&lt;/text&gt;
</programlisting>

  <para>XML representation of normalized transcript</para>
<programlisting xml:space="preserve">
&lt;?xml version="1.0" encoding="UTF-8"?&gt;
&lt;!DOCTYPE text SYSTEM "../../lib/Tei.dtd"[
 &lt;!ENTITY mdash  "—" &gt;&lt;!--=em dash--&gt;
 &lt;!ENTITY ndash  "–" &gt;&lt;!--=en dash--&gt;
 &lt;!ENTITY vbar  "|" &gt;&lt;!--=vertical bar--&gt;
 &lt;!ENTITY apos  "—" &gt;&lt;!--=apostrophe--&gt;
 &lt;!ENTITY triplebar "|" &gt;&lt;!--/triple horizontal bars--&gt;
 &lt;!ENTITY ldquo  "“" &gt;&lt;!--=double quotation mark, left--&gt;
 &lt;!ENTITY uarr   "↑" &gt;&lt;!--/uparrow A: =upward arrow--&gt;
&lt;!-- For HTML: --&gt;
 &lt;!ENTITY and "&lt;reg&gt;and&lt;/reg&gt;" &gt;
 &lt;!ENTITY etc "&lt;reg&gt;etc&lt;/reg&gt;" &gt;
&lt;!-- FOR Fof: --&gt;
 &lt;!-- &lt;!ENTITY and "amp" &gt;
 &lt;!ENTITY etc "ampc" &gt;--&gt;
]&gt;
&lt;text&gt;
 &lt;body&gt;
  &lt;p&gt;
   &lt;lb/&gt;A seaman figures in The Canterbury Tales. &lt;lb/&gt;With many a
   tempest had his beard been &lt;lb/&gt;shook. &amp;ndash; Secret grief is a
   &lt;lb/&gt;cannibal of its own heart &amp;ndash; &lt;emph&gt;Bacon&lt;/emph&gt;. &lt;lb/&gt;
   &lt;emph&gt;Claudia&lt;/emph&gt; of the &lt;emph&gt;Appian&lt;/emph&gt; family, &lt;q&gt;I wish
    &lt;lb/&gt; some fight or pestilence would thin out &lt;lb/&gt; this
    crowd.&lt;/q&gt;
   &lt;emph&gt;Arrogance.&lt;/emph&gt;
   &lt;lb/&gt;
   &lt;emph&gt;Roast beef in the pulpet.&lt;/emph&gt;&lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;An animal of a man — &lt;q&gt;do eagles wear
    &lt;lb/&gt;spectacles?&lt;/q&gt;— Health. — &lt;emph&gt;Contrast&lt;/emph&gt;:
   an &lt;lb/&gt;over spiritual man. &lt;lb/&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;
   &lt;q&gt;Yes, Madam, Cain was a godless froward boy, ∧ &lt;lb/&gt;Reuben
    (Gen:49) ∧ Absalom&lt;/q&gt; Many pious men &lt;lb/&gt;have impious
   children — (Devil as a Quaker) &lt;lb/&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;A formal compact &amp;ndash; Imprimis &amp;ndash; First &amp;ndash;
   Second. &lt;lb/&gt;The aforesaid soul &amp;etc; &amp;ndash; Duplicates &amp;ndash;
    &lt;lb/&gt;&lt;q&gt;How was it about the temptation on the &lt;lb/&gt;hill?&lt;/q&gt;
   &amp;etc; &lt;supplied&gt; Conversation upon Gabriel, Michael and &lt;lb/&gt;
    Raphael &amp;ndash; gentlemanly &amp;etc;&lt;/supplied&gt; &amp;ndash; D begs the
   hero to form &lt;lb/&gt;one of a &lt;emph&gt;
    &lt;q&gt;Society of D&amp;apos;s&lt;/q&gt;
   &lt;/emph&gt; &amp;ndash; his name would be weighty &lt;lb/&gt; &amp;etc; &amp;ndash;
   Leaves a letter to the D &amp;ndash; &lt;q&gt;My &lt;lb/&gt;
    &lt;emph&gt;Dear D&lt;/emph&gt;
   &lt;/q&gt; &amp;ndash; &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;
   &lt;q&gt;Terra Oblivionis&lt;/q&gt;
   &lt;q&gt;Hellites&lt;/q&gt; &amp;ndash; At the Astor find him &lt;lb/&gt; making &lt;choice&gt;
    &lt;orig&gt;almanacks&lt;/orig&gt;
    &lt;reg&gt;almanacs&lt;/reg&gt;
   &lt;/choice&gt; &amp;ndash; going to a ball takes a long &lt;lb/&gt;time making
   toilette. &amp;ndash; The Doctor&amp;apos;s coach stops &lt;lb/&gt;the way.
   &amp;ndash; &lt;q&gt;Do you believe all that stuff? &lt;lb/&gt;
    &lt;choice&gt;
     &lt;orig&gt;nonsence&lt;/orig&gt;
     &lt;reg&gt;nonsense&lt;/reg&gt;
    &lt;/choice&gt; &amp;ndash; the world was never made. &amp;ndash; &amp;ldquo;But &lt;choice&gt;
     &lt;orig&gt;Is&lt;/orig&gt;
     &lt;reg&gt;is&lt;/reg&gt;
    &lt;/choice&gt; not &lt;lb/&gt;this you mentioned &lt;emph&gt;here&lt;/emph&gt; &amp;ndash; in
    the scriptures?&lt;/q&gt;
   &lt;lb/&gt;Receives visits from the principal d&amp;apos;s &amp;ndash;
    &lt;q&gt;Gentlemen&lt;/q&gt; &amp;etc;. &lt;lb/&gt;
   &lt;emph&gt;Arguments&lt;/emph&gt; to persuade &amp;ndash; &lt;q&gt;Would you not rather
    &lt;lb/&gt;be below with kings than above with fools?&lt;/q&gt;
  &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;It is better to laugh ∧ not sin than to weep ∧ be
   &lt;lb/&gt;wicked. — Ten loads of coal to burn him. —
   &lt;lb/&gt;Brought to the stake — warmed himself by the fire. &lt;/p&gt;
  &lt;p&gt;
   &lt;lb/&gt;Ego non baptizo te in &lt;choice&gt;
    &lt;orig&gt;nominee&lt;/orig&gt;
    &lt;reg&gt;nomine&lt;/reg&gt;
   &lt;/choice&gt; Patris et &lt;lb/&gt;Filii et Spiritus Sancti &amp;ndash; sed in
   nomine &lt;lb/&gt;Diaboli. — Madness is undefinable — &lt;lb/&gt;It
   ∧ right reasons extremes of one. &lt;lb/&gt;–Not the (black
   art) Goetic but Theurgic magic — &lt;lb/&gt;seeks converse with the
   Intelligence, Power, the &lt;lb/&gt;Angel. &lt;/p&gt;
 &lt;/body&gt;
&lt;/text&gt;
</programlisting>



  <para>HTML presentation of exemplar for normalized transcript</para>
  <para>Special exemplar tokens in red.</para>

  <figure xml:id="HTML-Norm-exemplar">
   <title>HTML presentation of exemplar for normalized
    transcript</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-011.jpg" format="jpg" width="70%"/>
    </imageobject>
   </mediaobject>
  </figure>

  <para>HTML presentation of exemplar for normalized transcript</para>
  <para>Special exemplar tokens in red, type similar tokens in
   blue.</para>

  <figure xml:id="HTML-Norm-transcript">
   <title>HTML presentation of normalized transcript</title>
   <mediaobject>
    <imageobject>
     <imagedata fileref="../../../vol31/graphics/Huitfeldt01/Huitfeldt01-012.jpg" format="jpg" width="50%"/>
    </imageobject>
   </mediaobject>
  </figure>

  <para>FOF representation of normalized exemplar and
   transcript</para>
<programlisting xml:space="preserve">
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% General axioms
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(transcript_domain, axiom,
  ! [X,Y] : (transcript(X,Y) =&gt; (e_token(X) &amp; t_token(Y)))).

fof(exemplar_domain, axiom,
  ! [X,Y] : (exemplar(X,Y) =&gt; (t_token(X) &amp; e_token(Y)))).

%%%% Slightly modified from Default:
fof(typeof_domain, axiom,
  ! [X,Y] : (typeof(X,Y) =&gt; 
            ((t_token(X) | e_token(X) | st_token(X) | se_token(X)) 
             &amp; type(Y)))).

%%%% Slightly modified from Default:
fof(no_class_overlap, axiom,
  ! [X] : (
  (e_token(X) &lt;=&gt; (~t_token(X) &amp; ~type(X) &amp; ~se_token(X) &amp; ~st_token(X))) &amp;
  (t_token(X) &lt;=&gt; (~e_token(X) &amp; ~type(X) &amp; ~se_token(X) &amp; ~st_token(X))) &amp;
  (type(X) &lt;=&gt; (~e_token(X) &amp; ~t_token(X)) &amp; ~se_token(X) &amp; ~st_token(X)) &amp;
  (se_token(X) &lt;=&gt; (~t_token(X) &amp; ~type(X) &amp; ~e_token(X) &amp; ~st_token(X))) &amp;
  (st_token(X) &lt;=&gt; (~t_token(X) &amp; ~type(X) &amp; ~se_token(X) &amp; ~e_token(X)))
  )).

fof(exemplar_and_transcript_inverse, axiom,
  ! [X,Y] : (exemplar(X,Y) &lt;=&gt; transcript(Y,X))).

fof(at_most_one_type_per_token, axiom,
  ! [X,Y,Z] : ((typeof(X,Y) &amp; typeof(X,Z)) =&gt; Y=Z)).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Axioms not in Default

fof(typesimilar_domain, axiom,
  ! [X,Y] : (typesimilar(X,Y) =&gt; (type(X) &amp; type(Y)))).

fof(typesimilar_reflexive, axiom,
  ! [X] : (type(X) =&gt; typesimilar(X,X))).

fof(typesimilar_symmetric, axiom,
  ! [X,Y] : (typesimilar(X,Y) =&gt; typesimilar(Y,X))).

fof(type_similar_transitive, axiom,
  ! [X,Y,Z] : ((typesimilar(X,Y) &amp; typesimilar(Y,Z)) =&gt; typesimilar(X,Z))).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Definitions
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(reciprocity,definition,
  reciprocity &lt;=&gt; ! [X,Y,Z] : (
              ((transcript(X,Y) &amp; transcript(X,Z)) =&gt; Y=Z) &amp;
              ((exemplar(X,Y) &amp; exemplar(X,Z)) =&gt; Y=Z))).

fof(completeness, definition,
  completeness &lt;=&gt; ! [X] : 
          (e_token(X) =&gt; ? [Y] : transcript(X,Y))).

fof(purity, definition,
  purity &lt;=&gt; ! [X] :
          (t_token(X) =&gt; ? [Y] : exemplar(X,Y))).

%%%% Slightly modified from Default:
fof(type_similarity, definition,
  type_similarity &lt;=&gt; ![X,Y] : (transcript(X,Y) =&gt;
    ? [Z,V] : (typeof(X,Z) &amp; typeof(Y,V) &amp; typesimilar(Z,V)))).

fof(t_similarity, definition,
  t_similarity &lt;=&gt; (reciprocity &amp; completeness &amp; purity &amp; type_similarity)).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Case-specific facts
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(case_specific_facts, axiom,
  typeof(e1, t_A) &amp;
  typeof(e2, t_seaman) &amp;
  typeof(e3, t_figures) &amp;
  typeof(e4, t_in) &amp;
  typeof(e5, t_The) &amp;
  typeof(e6, t_Canterbury) &amp;
  typeof(e7, t_Tales) &amp;
  typeof(e8, t_With) &amp;
  typeof(e9, t_a) &amp;
  typeof(e10, t_many) &amp;
  typeof(e11, t_a) &amp;
  typeof(e12, t_tempest) &amp;
  typeof(e13, t_had) &amp;
  typeof(e14, t_his) &amp;
  typeof(e15, t_beard) &amp;
  typeof(e16, t_been) &amp;
  typeof(e17, t_shook) &amp;
  typeof(e18, t_S) &amp;
  typeof(e19, t_Deep) &amp;
%%% ...Approximately 250-350 lines omitted...
  typeof(e278, t_extremes) &amp;
  typeof(e279, t_of) &amp;
  typeof(e280, t_one) &amp;
  typeof(e281, t_Not) &amp;
  typeof(e282, t_the) &amp;
  typeof(e283, t_black) &amp;
  typeof(e284, t_art) &amp;
  typeof(e285, t_Goetic) &amp;
  typeof(e286, t_but) &amp;
  typeof(e287, t_Theurgic) &amp;
  typeof(e288, t_magic) &amp;
  typeof(e289, t_seeks) &amp;
  typeof(e290, t_converse) &amp;
  typeof(e291, t_with) &amp;
  typeof(e292, t_the) &amp;
  typeof(e293, t_Intelligence) &amp;
  typeof(e294, t_Power) &amp;
  typeof(e295, t_the) &amp;
  typeof(e296, t_Angel) &amp;
  typeof(t1, t_A) &amp;
  typeof(t2, t_seaman) &amp;
  typeof(t3, t_figures) &amp;
  typeof(t4, t_in) &amp;
  typeof(t5, t_The) &amp;
  typeof(t6, t_Canterbury) &amp;
  typeof(t7, t_Tales) &amp;
  typeof(t8, t_With) &amp;
  typeof(t9, t_many) &amp;
  typeof(t10, t_a) &amp;
  typeof(t11, t_tempest) &amp;
  typeof(t12, t_had) &amp;
  typeof(t13, t_his) &amp;
  typeof(t14, t_beard) &amp;
  typeof(t15, t_been) &amp;
  typeof(t16, t_shook) &amp;
  typeof(t17, t_Secret) &amp;
  typeof(t18, t_grief) &amp;
  typeof(t19, t_is) &amp;
  typeof(t20, t_a) &amp;
%%% ...Approximately 250-350 lines omitted...
  typeof(t270, t_reasons) &amp;
  typeof(t271, t_extremes) &amp;
  typeof(t272, t_of) &amp;
  typeof(t273, t_one) &amp;
  typeof(t274, t_Not) &amp;
  typeof(t275, t_the) &amp;
  typeof(t276, t_black) &amp;
  typeof(t277, t_art) &amp;
  typeof(t278, t_Goetic) &amp;
  typeof(t279, t_but) &amp;
  typeof(t280, t_Theurgic) &amp;
  typeof(t281, t_magic) &amp;
  typeof(t282, t_seeks) &amp;
  typeof(t283, t_converse) &amp;
  typeof(t284, t_with) &amp;
  typeof(t285, t_the) &amp;
  typeof(t286, t_Intelligence) &amp;
  typeof(t287, t_Power) &amp;
  typeof(t288, t_the) &amp;
  typeof(t289, t_Angel) &amp;

  typesimilar(t_almanacks, t_almanacs) &amp;
  typesimilar(t_nonsence, t_nonsense) &amp;
  typesimilar(t_Is, t_is) &amp;
  typesimilar(t_nominee, t_nomine) &amp;

  transcript(e1,t1) &amp;
  transcript(e2,t2) &amp;
  transcript(e3,t3) &amp;
  transcript(e4,t4) &amp;
  transcript(e5,t5) &amp;
  transcript(e6,t6) &amp;
  transcript(e7,t7) &amp;
  transcript(e8,t8) &amp;
  transcript(e10,t9) &amp;
  transcript(e11,t10) &amp;
  transcript(e12,t11) &amp;
  transcript(e13,t12) &amp;
  transcript(e14,t13) &amp;
  transcript(e15,t14) &amp;
  transcript(e16,t15) &amp;
  transcript(e17,t16) &amp;
  transcript(e21,t17) &amp;
  transcript(e22,t18) &amp;
  transcript(e23,t19) &amp;
  transcript(e24,t20) &amp;
%%% ...Approximately 250-350 lines omitted...
  transcript(e276,t269) &amp;
  transcript(e277,t270) &amp;
  transcript(e278,t271) &amp;
  transcript(e279,t272) &amp;
  transcript(e280,t273) &amp;
  transcript(e281,t274) &amp;
  transcript(e282,t275) &amp;
  transcript(e283,t276) &amp;
  transcript(e284,t277) &amp;
  transcript(e285,t278) &amp;
  transcript(e286,t279) &amp;
  transcript(e287,t280) &amp;
  transcript(e288,t281) &amp;
  transcript(e289,t282) &amp;
  transcript(e290,t283) &amp;
  transcript(e291,t284) &amp;
  transcript(e292,t285) &amp;
  transcript(e293,t286) &amp;
  transcript(e294,t287) &amp;
  transcript(e295,t288) &amp;
  transcript(e296,t289) &amp;

  ! [X,Y,Z] : (
              ((transcript(X,Y) &amp; transcript(X,Z)) =&gt; Y=Z) &amp;
              ((exemplar(X,Y) &amp; exemplar(X,Z)) =&gt; Y=Z)
              ) &amp;

  $distinct(
    e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13,e14,e15,e16,e17,e18,e19,e20,e21,e22,e23,e24,e25,e26,e27,e28,e29,e30,e31,e32,e33,e34,e35,e36,e37,e38,e39,e40,e41,e42,e43,e44,e45,e46,e47,e48,e49,e50,e51,e52,e53,e54,e55,e56,e57,e58,e59,e60,e61,e62,e63,e64,e65,e66,e67,e68,e69,e70,e71,e72,e73,e74,e75,e76,e77,e78,e79,e80,e81,e82,e83,e84,e85,e86,e87,e88,e89,e90,e91,e92,e93,e94,e95,e96,e97,e98,e99,e100,e101,e102,e103,e104,e105,e106,e107,e108,e109,e110,e111,e112,e113,e114,e115,e116,e117,e118,e119,e120,e121,e122,e123,e124,e125,e126,e127,e128,e129,e130,e131,e132,e133,e134,e135,e136,e137,e138,e139,e140,e141,e142,e143,e144,e145,e146,e147,e148,e149,e150,e151,e152,e153,e154,e155,e156,e157,e158,e159,e160,e161,e162,e163,e164,e165,e166,e167,e168,e169,e170,e171,e172,e173,e174,e175,e176,e177,e178,e179,e180,e181,e182,e183,e184,e185,e186,e187,e188,e189,e190,e191,e192,e193,e194,e195,e196,e197,e198,e199,e200,e201,e202,e203,e204,e205,e206,e207,e208,e209,e210,e211,e212,e213,e214,e215,e216,e217,e218,e219,e220,e221,e222,e223,e224,e225,e226,e227,e228,e229,e230,e231,e232,e233,e234,e235,e236,e237,e238,e239,e240,e241,e242,e243,e244,e245,e246,e247,e248,e249,e250,e251,e252,e253,e254,e255,e256,e257,e258,e259,e260,e261,e262,e263,e264,e265,e266,e267,e268,e269,e270,e271,e272,e273,e274,e275,e276,e277,e278,e279,e280,e281,e282,e283,e284,e285,e286,e287,e288,e289,e290,e291,e292,e293,e294,e295,e296,
    t1,t2,t3,t4,t5,t6,t7,t8,t9,t10,t11,t12,t13,t14,t15,t16,t17,t18,t19,t20,t21,t22,t23,t24,t25,t26,t27,t28,t29,t30,t31,t32,t33,t34,t35,t36,t37,t38,t39,t40,t41,t42,t43,t44,t45,t46,t47,t48,t49,t50,t51,t52,t53,t54,t55,t56,t57,t58,t59,t60,t61,t62,t63,t64,t65,t66,t67,t68,t69,t70,t71,t72,t73,t74,t75,t76,t77,t78,t79,t80,t81,t82,t83,t84,t85,t86,t87,t88,t89,t90,t91,t92,t93,t94,t95,t96,t97,t98,t99,t100,t101,t102,t103,t104,t105,t106,t107,t108,t109,t110,t111,t112,t113,t114,t115,t116,t117,t118,t119,t120,t121,t122,t123,t124,t125,t126,t127,t128,t129,t130,t131,t132,t133,t134,t135,t136,t137,t138,t139,t140,t141,t142,t143,t144,t145,t146,t147,t148,t149,t150,t151,t152,t153,t154,t155,t156,t157,t158,t159,t160,t161,t162,t163,t164,t165,t166,t167,t168,t169,t170,t171,t172,t173,t174,t175,t176,t177,t178,t179,t180,t181,t182,t183,t184,t185,t186,t187,t188,t189,t190,t191,t192,t193,t194,t195,t196,t197,t198,t199,t200,t201,t202,t203,t204,t205,t206,t207,t208,t209,t210,t211,t212,t213,t214,t215,t216,t217,t218,t219,t220,t221,t222,t223,t224,t225,t226,t227,t228,t229,t230,t231,t232,t233,t234,t235,t236,t237,t238,t239,t240,t241,t242,t243,t244,t245,t246,t247,t248,t249,t250,t251,t252,t253,t254,t255,t256,t257,t258,t259,t260,t261,t262,t263,t264,t265,t266,t267,t268,t269,t270,t271,t272,t273,t274,t275,t276,t277,t278,t279,t280,t281,t282,t283,t284,t285,t286,t287,t288,t289,
    t_A,t_seaman,t_figures,t_in,t_The,t_Canterbury,t_Tales,t_With,t_a,t_many,t_tempest,t_had,t_his,t_beard,t_been,t_shook,t_S,t_Deep,t_g,t_Secret,t_grief,t_is,t_cannibal,t_of,t_its,t_own,t_heart,t_Bacon,t_Claudia,t_the,t_Appian,t_family,t_I,t_wish,t_some,t_fight,t_or,t_pestilence,t_would,t_thin,t_out,t_this,t_crowd,t_Arrogance,t_Roast,t_beef,t_pulpet,t_An,t_animal,t_man,t_do,t_eagles,t_wear,t_spectacles,t_Health,t_Contrast,t_an,t_over,t_spiritual,t_Yes,t_Madam,t_Cain,t_was,t_godless,t_froward,t_boy,t_amp,t_Reuben,t_Gen,t_49,t_Absalom,t_Many,t_pious,t_men,t_have,t_impious,t_children,t_Devil,t_as,t_Quaker,t_formal,t_compact,t_Imprimis,t_First,t_Second,t_aforesaid,t_soul,t_said,t_ampc,t_Duplicates,t_How,t_it,t_about,t_temptation,t_on,t_hill,t_D,t_begs,t_hero,t_to,t_form,t_one,t_Society,t_s,t_name,t_be,t_weighty,t_Leaves,t_letter,t_My,t_Dear,t_Conversation,t_upon,t_Gabriel,t_Micheal,t_Raphel,t_gentlemanly,t_Terra,t_Oblivionis,t_Hellites,t_At,t_Astor,t_find,t_him,t_making,t_almanacks,t_going,t_ball,t_takes,t_long,t_time,t_toilette,t_Doctor,t_coach,t_stops,t_way,t_Do,t_you,t_believe,t_all,t_that,t_stuff,t_nonsence,t_world,t_never,t_made,t_But,t_Is,t_not,t_mentioned,t_here,t_scriptures,t_Receives,t_visits,t_from,t_principal,t_d,t_Gentlemen,t_Arguments,t_persuade,t_Would,t_rather,t_below,t_with,t_kings,t_than,t_above,t_fools,t_It,t_better,t_laugh,t_sin,t_weep,t_wicked,t_Ten,t_loads,t_coal,t_burn,t_Brought,t_stake,t_warmed,t_himself,t_by,t_fire,t_Ego,t_non,t_baptizo,t_te,t_nominee,t_Patris,t_et,t_Filii,t_Spiritus,t_Sancti,t_sed,t_nomine,t_Diaboli,t_Madness,t_undefinable,t_right,t_reasons,t_extremes,t_Not,t_black,t_art,t_Goetic,t_but,t_Theurgic,t_magic,t_seeks,t_converse,t_Intelligence,t_Power,t_Angel,t_Michael,t_and,t_Raphael,t_almanacs,t_nonsense
  ) &amp;
  ! [X] : ((e_token(X) &lt;=&gt; (X=e1 | X=e2 | X=e3 | X=e4 | X=e5 | X=e6 | X=e7 | X=e8 | X=e10 | X=e11 | X=e12 | X=e13 | X=e14 | X=e15 | X=e16 | X=e17 | X=e21 | X=e22 | X=e23 | X=e24 | X=e25 | X=e26 | X=e27 | X=e28 | X=e29 | X=e30 | X=e31 | X=e32 | X=e33 | X=e34 | X=e35 | X=e36 | X=e37 | X=e38 | X=e39 | X=e40 | X=e41 | X=e42 | X=e43 | X=e44 | X=e45 | X=e46 | X=e47 | X=e48 | X=e49 | X=e50 | X=e51 | X=e52 | X=e53 | X=e54 | X=e55 | X=e56 | X=e57 | X=e58 | X=e59 | X=e60 | X=e61 | X=e62 | X=e63 | X=e64 | X=e65 | X=e66 | X=e67 | X=e68 | X=e69 | X=e70 | X=e71 | X=e72 | X=e73 | X=e74 | X=e75 | X=e76 | X=e77 | X=e78 | X=e79 | X=e80 | X=e81 | X=e82 | X=e83 | X=e84 | X=e85 | X=e86 | X=e87 | X=e88 | X=e89 | X=e90 | X=e91 | X=e92 | X=e93 | X=e94 | X=e95 | X=e96 | X=e97 | X=e98 | X=e99 | X=e100 | X=e103 | X=e104 | X=e105 | X=e106 | X=e107 | X=e108 | X=e109 | X=e110 | X=e111 | X=e112 | X=e113 | X=e114 | X=e115 | X=e116 | X=e117 | X=e118 | X=e119 | X=e120 | X=e121 | X=e122 | X=e123 | X=e124 | X=e125 | X=e126 | X=e127 | X=e128 | X=e129 | X=e130 | X=e131 | X=e132 | X=e133 | X=e134 | X=e135 | X=e136 | X=e137 | X=e138 | X=e139 | X=e140 | X=e141 | X=e142 | X=e151 | X=e152 | X=e153 | X=e154 | X=e155 | X=e156 | X=e157 | X=e158 | X=e159 | X=e160 | X=e161 | X=e162 | X=e163 | X=e164 | X=e165 | X=e166 | X=e167 | X=e168 | X=e169 | X=e170 | X=e171 | X=e172 | X=e173 | X=e174 | X=e175 | X=e176 | X=e177 | X=e178 | X=e179 | X=e180 | X=e181 | X=e182 | X=e183 | X=e184 | X=e185 | X=e186 | X=e187 | X=e188 | X=e189 | X=e190 | X=e191 | X=e192 | X=e193 | X=e194 | X=e195 | X=e196 | X=e197 | X=e198 | X=e199 | X=e200 | X=e201 | X=e202 | X=e203 | X=e204 | X=e205 | X=e206 | X=e207 | X=e208 | X=e209 | X=e210 | X=e211 | X=e212 | X=e213 | X=e214 | X=e215 | X=e216 | X=e217 | X=e218 | X=e219 | X=e220 | X=e221 | X=e222 | X=e223 | X=e224 | X=e225 | X=e226 | X=e227 | X=e228 | X=e229 | X=e230 | X=e231 | X=e232 | X=e233 | X=e235 | X=e236 | X=e237 | X=e238 | X=e239 | X=e240 | X=e241 | X=e242 | X=e243 | X=e244 | X=e245 | X=e246 | X=e247 | X=e248 | X=e249 | X=e250 | X=e251 | X=e252 | X=e253 | X=e254 | X=e255 | X=e256 | X=e257 | X=e258 | X=e259 | X=e260 | X=e261 | X=e262 | X=e263 | X=e264 | X=e265 | X=e266 | X=e267 | X=e268 | X=e269 | X=e270 | X=e271 | X=e272 | X=e273 | X=e274 | X=e275 | X=e276 | X=e277 | X=e278 | X=e279 | X=e280 | X=e281 | X=e282 | X=e283 | X=e284 | X=e285 | X=e286 | X=e287 | X=e288 | X=e289 | X=e290 | X=e291 | X=e292 | X=e293 | X=e294 | X=e295 | X=e296)) &amp;
           (t_token(X) &lt;=&gt; (X=t1 | X=t2 | X=t3 | X=t4 | X=t5 | X=t6 | X=t7 | X=t8 | X=t9 | X=t10 | X=t11 | X=t12 | X=t13 | X=t14 | X=t15 | X=t16 | X=t17 | X=t18 | X=t19 | X=t20 | X=t21 | X=t22 | X=t23 | X=t24 | X=t25 | X=t26 | X=t27 | X=t28 | X=t29 | X=t30 | X=t31 | X=t32 | X=t33 | X=t34 | X=t35 | X=t36 | X=t37 | X=t38 | X=t39 | X=t40 | X=t41 | X=t42 | X=t43 | X=t44 | X=t45 | X=t46 | X=t47 | X=t48 | X=t49 | X=t50 | X=t51 | X=t52 | X=t53 | X=t54 | X=t55 | X=t56 | X=t57 | X=t58 | X=t59 | X=t60 | X=t61 | X=t62 | X=t63 | X=t64 | X=t65 | X=t66 | X=t67 | X=t68 | X=t69 | X=t70 | X=t71 | X=t72 | X=t73 | X=t74 | X=t75 | X=t76 | X=t77 | X=t78 | X=t79 | X=t80 | X=t81 | X=t82 | X=t83 | X=t84 | X=t85 | X=t86 | X=t87 | X=t88 | X=t89 | X=t90 | X=t91 | X=t92 | X=t93 | X=t94 | X=t95 | X=t96 | X=t97 | X=t98 | X=t99 | X=t100 | X=t101 | X=t102 | X=t103 | X=t104 | X=t105 | X=t106 | X=t107 | X=t108 | X=t117 | X=t118 | X=t119 | X=t120 | X=t121 | X=t122 | X=t123 | X=t124 | X=t125 | X=t126 | X=t127 | X=t128 | X=t129 | X=t130 | X=t131 | X=t132 | X=t133 | X=t134 | X=t135 | X=t136 | X=t137 | X=t138 | X=t139 | X=t140 | X=t141 | X=t142 | X=t143 | X=t144 | X=t145 | X=t146 | X=t147 | X=t148 | X=t149 | X=t150 | X=t151 | X=t152 | X=t153 | X=t154 | X=t155 | X=t156 | X=t157 | X=t158 | X=t159 | X=t160 | X=t161 | X=t162 | X=t163 | X=t164 | X=t165 | X=t166 | X=t167 | X=t168 | X=t169 | X=t170 | X=t171 | X=t172 | X=t173 | X=t174 | X=t175 | X=t176 | X=t177 | X=t178 | X=t179 | X=t180 | X=t181 | X=t182 | X=t183 | X=t184 | X=t185 | X=t186 | X=t187 | X=t188 | X=t189 | X=t190 | X=t191 | X=t192 | X=t193 | X=t194 | X=t195 | X=t196 | X=t197 | X=t198 | X=t199 | X=t200 | X=t201 | X=t202 | X=t203 | X=t204 | X=t205 | X=t206 | X=t207 | X=t208 | X=t209 | X=t210 | X=t211 | X=t212 | X=t213 | X=t214 | X=t215 | X=t216 | X=t217 | X=t218 | X=t219 | X=t220 | X=t221 | X=t222 | X=t223 | X=t224 | X=t225 | X=t226 | X=t227 | X=t228 | X=t229 | X=t230 | X=t231 | X=t232 | X=t233 | X=t234 | X=t235 | X=t236 | X=t237 | X=t238 | X=t239 | X=t240 | X=t241 | X=t242 | X=t243 | X=t244 | X=t245 | X=t246 | X=t247 | X=t248 | X=t249 | X=t250 | X=t251 | X=t252 | X=t253 | X=t254 | X=t255 | X=t256 | X=t257 | X=t258 | X=t259 | X=t260 | X=t261 | X=t262 | X=t263 | X=t264 | X=t265 | X=t266 | X=t267 | X=t268 | X=t269 | X=t270 | X=t271 | X=t272 | X=t273 | X=t274 | X=t275 | X=t276 | X=t277 | X=t278 | X=t279 | X=t280 | X=t281 | X=t282 | X=t283 | X=t284 | X=t285 | X=t286 | X=t287 | X=t288 | X=t289)))
).

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Conjecture
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

fof(t_similar, conjecture,
 completeness &amp; purity  &amp;  reciprocity  &amp; type_similarity 
).
</programlisting>

 </appendix>

 <bibliography>
  <title>References</title>
  <bibliomixed>
   <phrase>
    <emphasis role="bold">This list of references has not been updated
     since 2018. References added later have been provided in
     footnotes and in web links. </emphasis></phrase>
  </bibliomixed>

  <bibliomixed xml:id="Andre" xreflabel="André 1972">[André,
   Jacques.] <emphasis>Règles et recommandation pour les éditions
    critiques (Série latine)</emphasis>. Paris: Société d'édition "Les
   belles lettres," 1972. Collection des universités de France,
   publiée sous le patronage de l'Association Guillaume Budé. [vi +]
   48 pp.</bibliomixed>


  <bibliomixed xml:id="Carter" xreflabel="Carter 1952">Carter,
   Clarence E. <emphasis>Historical editing</emphasis>. Bulletins of
   the national archives, Number 7 [Washington, DC]: National Archives
   and Records Service, August 1952. National Archives publication
   number 53-4.</bibliomixed>

  <bibliomixed xml:id="Caton" xreflabel="Caton 2013">Caton, Paul.
    <quote>Pure transcriptional encoding.</quote> Paper given at
   Digital Humanities 2013, Lincoln, Nebraska. </bibliomixed>

  <bibliomixed xml:id="Dunn" xreflabel="The Papers of William Penn">Dunn, 
Mary Maples et al. <emphasis>The Papers of William
    Penn</emphasis>, 5 vols. (Philadelphia: University of Pennsylvania
   Press, 1981-1986).</bibliomixed>

  <bibliomixed xml:id="grice" xreflabel="Grice 1975">Grice, H.P.
    <quote>Logic and Conversation.</quote> In <emphasis>Syntax and
    Semantics</emphasis>, edited by P. Cole and J. Morgan, vol.3,
    <emphasis>Speech Acts</emphasis>. New York: Academic Press, 1975.
   Reprinted as chapter 2 of his <emphasis>Studies in the Way of
    Words</emphasis>. Cambridge, Mass.: Harvard University Press,
   1989, pp. 22–40.</bibliomixed>

  <bibliomixed xml:id="Goodman" xreflabel="Goodman">Goodman, Nelson.
    <emphasis>Languages of Art</emphasis>. Hackett Publishing,
   1976.</bibliomixed>

  <bibliomixed xml:id="md" xreflabel="Hayford et al. 1988">Hayford,
   Harrison, Hershel Parker, and G. Thomas Tanselle, ed. <quote>Moby
    Dick, or, The Whale.</quote> Vol. 7 of <emphasis>The Writings of
    Herman Melville</emphasis> The Northwestern–Newberry
   Edition. Evanston [Ill.]: Northwestern University Press; Chicago :
   Newberry Library, 1988, rpt. 1994, 1997. <!--* ISBN 0-8101-0268-4 (cloth), 0-8101-0269-2 (paper) *-->
   <!--* https://books.google.de/books?id=jnNBh61lpjUC *-->
  </bibliomixed>
  <bibliomixed xml:id="wit1" xreflabel="Huitfeldt / Sperberg-McQueen 2008">Huitfeldt, Claus,
   and C. M. Sperberg-McQueen. <quote>What is transcription?</quote>
   <emphasis>Literary &amp; Linguistic Computing</emphasis> 23.3
   (2008): 295-310. [doi:<biblioid class="doi">10.1093/llc/fqn013</biblioid>].</bibliomixed>
  <bibliomixed xml:id="ettdds" xreflabel="Huitfeldt / Marcoux / Sperberg-McQueen 2010">Huitfeldt,
   Claus, Yves Marcoux, and C. M. Sperberg-McQueen. <quote>Extension
    of the type/token distinction to document structure.</quote> Paper
   presented at Balisage: The Markup Conference 2010, Montréal,
   Canada, August 3 - 6, 2010. In <emphasis>Proceedings of Balisage:
    The Markup Conference 2010</emphasis>. Balisage Series on Markup
   Technologies, vol. 5 (2010). [doi:<biblioid class="doi">10.4242/BalisageVol5.Huitfeldt01</biblioid>]. On the Web at
    [<link xlink:href="http://www.balisage.net/Proceedings/vol5/html/Huitfeldt01/BalisageVol5-Huitfeldt01.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://www.balisage.net/Proceedings/vol5/html/Huitfeldt01/BalisageVol5-Huitfeldt01.html</link>]. </bibliomixed>
  <bibliomixed xml:id="is2006" xreflabel="Marcoux 2006">Marcoux,
   Yves. <quote>A natural-language approach to modeling: Why is some
    XML so difficult to write?</quote> Paper given at Extreme Markup
   Languages®, Montréal, 2006. <emphasis>Proceedings of Extreme
    Markup Languages® 2006</emphasis>. On the Web at [<link xlink:href="http://conferences.idealliance.org/extreme/html/2006/Marcoux01/EML2006Marcoux01.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://conferences.idealliance.org/extreme/html/2006/Marcoux01/EML2006Marcoux01.html</link>]. </bibliomixed>
  <bibliomixed xml:id="is2007" xreflabel="Marcoux / Rizkallah 2007">
   Marcoux, Yves, and Élias Rizkallah. <quote>Exploring
    intertextual semantics: A reflection on attributes and
    optionality.</quote> Paper given at Extreme Markup Languages®,
   Montréal, 2007. <emphasis>Proceedings of Extreme Markup
    Languages® 2007</emphasis>. On the Web at [<link xlink:href="http://conferences.idealliance.org/extreme/html/2007/Marcoux01/EML2007Marcoux01.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://conferences.idealliance.org/extreme/html/2007/Marcoux01/EML2007Marcoux01.html</link>].</bibliomixed>
  <bibliomixed xml:id="peano" xreflabel="Peano 1889">Peano, Ioseph.
    <quote>Arithmetics principia nova methodo exposita.</quote> Romae,
   Florentiae: Bocca, 1889. English translation: Michael Nahas, [<link xlink:href="https://raw.githubusercontent.com/mdnahas/Peano_Book/master/Peano.pdf" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">https://raw.githubusercontent.com/mdnahas/Peano_Book/master/Peano.pdf</link>].</bibliomixed>
  <bibliomixed xml:id="bath" xreflabel="Robinson / Solopova 2006">
   Robinson, Peter, and Elizabeth Solopova. <quote>Guidelines for
    Transcription of the Manuscripts of The Wife of Bath!s
    Prologue.</quote> 18 March 2006. On the Web at [<link xlink:href="http://www.canterburytalesproject.org/pubs/transguide-MI.pdf" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://www.canterburytalesproject.org/pubs/transguide-MI.pdf</link>]. </bibliomixed>
  <bibliomixed xml:id="simons" xreflabel="Simons et al. 2005">Simons,
   Gary, Scott O. Farrar, Brian Fitzsimons, William D. Lewis, D.
   Terence Langendoen, and Hector Gonzalez. <quote>The semantics of
    markup: Mapping legacy markup schemas to a common
    semantics.</quote> In <emphasis>Proceedings of the 4th workshop
    on NLP and XML (NLPXML-2004): held in cooperation with ACL- 04
    Barcelona, Spain</emphasis>. Pp. 25-32. On the Web at [<link xlink:href="http://www.aclweb.org/anthology/W/W04/W04-0604.pdf" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://www.aclweb.org/anthology/W/W04/W04-0604.pdf</link>] and
   other locations. [doi:<biblioid class="doi">10.3115/1621066.1621070</biblioid>].</bibliomixed>
  <bibliomixed xml:id="ioai" xreflabel="Sperberg-McQueen 2005">
   Sperberg-McQueen, C. M. <quote>The meaning of OAI 2.0 Markup: An
    exercise in markup interpretation.</quote> Unpublished fragment,
   December 2005. On the Web at [<link xlink:href="http://www.w3.org/2004/04/em-msm/ioai.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://www.w3.org/2004/04/em-msm/ioai.html</link>]. </bibliomixed>
  <bibliomixed xml:id="dibm" xreflabel="Sperberg-McQueen et al. 2002">
   Sperberg-McQueen, C. M., David Dubin, Claus Huitfeldt, and Allen
   Renear. <quote>Drawing inferences on the basis of markup.</quote>
   Paper given at Extreme Markup Languages®, Montréal, 2002.
    <emphasis> &gt;Proceedings of Extreme Markup Languages® 
    2002</emphasis>. On the Web at [<link xlink:href="http://conferences.idealliance.org/extreme/html/2002/CMSMcQ01/EML2002CMSMcQ01.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://conferences.idealliance.org/extreme/html/2002/CMSMcQ01/EML2002CMSMcQ01.html</link>]. </bibliomixed>
  <bibliomixed xml:id="wit2" xreflabel="Sperberg-McQueen / Huitfeldt / Marcoux 2009">
   Sperberg-McQueen, C. M.. Claus Huitfeldt, and Yves Marcoux.
    <quote>What is transcription? Part 2.</quote> Talk given at
   Digital Humanities 2009, College Park, Maryland. Slides on the Web
   at [<link xlink:href="http://blackmesatech.com/2009/06/dh2009/" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://blackmesatech.com/2009/06/dh2009/</link>]. Summary at
    [<link xlink:href="http://www.mith2.umd.edu/dh09/wp-content/uploads/dh09_conferencepreceedings_final.pdf" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest"/> (pp. 257-260)]. </bibliomixed>
  <bibliomixed xml:id="mim" xreflabel="Sperberg-McQueen / Huitfeldt / Renear 2001">
   Sperberg-McQueen, C. M., Claus Huitfeldt, and Allen Renear.
    <quote>Meaning and interpretation of markup.</quote>
   <emphasis>Markup Languages: Theory &amp; Practice</emphasis> 2.3
   (2001): 215–234. On the Web at [<link xlink:href="http://cmsmcq.com/2000/mim.html" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://cmsmcq.com/2000/mim.html</link>]. [doi:<biblioid class="doi">10.1162/109966200750363599</biblioid>].</bibliomixed>
  <bibliomixed xml:id="StevensAndBurg" xreflabel="Stevens and Burg 1997">Stevens, Michael E. and Steven B.
   Burg. <emphasis>Editing Historical Documents: A Handbook of
    Practice</emphasis>. Walnut Creek, London, New Delhi: Alta Mira
   Press, 1997.</bibliomixed>
  <bibliomixed xml:id="tanselle1989" xreflabel="Tanselle 1989">
   Tanselle, G. Thomas. <emphasis>A Rationale of Textual
    Criticism</emphasis>. Philadelphia: University of Pennsylvania
   Press, 1989. 104 pp.</bibliomixed>
  <bibliomixed xml:id="vmt1999" xreflabel="Vander Meulen / Tanselle 1999">Vander Meulen, David,
   and G. Thomas Tanselle, <quote>A system of manuscript
    transcription.</quote> 
   <emphasis>Studies in Bibliography</emphasis> 52 (1999): 201-212. </bibliomixed>
  <bibliomixed xml:id="fotbo" xreflabel="Wickett / Renear 2009">
   Wickett, Karen M., and Allen Renear. <quote>A first order theory of
    bibliographic objects.</quote>
   <emphasis>Proceedings of the American Society for Information
    Science and Technology</emphasis> 46.1 (2009): 1–8. On the
   Web at [<link xlink:href="http://onlinelibrary.wiley.com/doi/10.1002/meet.2009.1450460378/full" xlink:type="simple" xlink:show="new" xlink:actuate="onRequest">http://onlinelibrary.wiley.com/doi/10.1002/meet.2009.1450460378/full</link>]
   (subscription required). [doi:<biblioid class="doi">10.1002/meet.2009.1450460378</biblioid>]. </bibliomixed>
  <bibliomixed xml:id="Boolos" xreflabel="Boolos et al. 2007">George
   Boolos, John P. Burgess, Richard C. Jeffrey. <emphasis>Computability
    and logic</emphasis>. 5th ed. Cambridge, New York: Cambridge
   University Press 2007. ISBN 9780521877527.
   <!--<idno type="callNumber">QA9.59 .B66 2007</idno>--></bibliomixed>
 </bibliography>

 <!--PLEASE KEEP THIS COMMENT
<para><emphasis>YMA 2026-05-31 Various scraps follow:</emphasis></para><para>Before we introduce the basic notions of the model, we think it will be useful to present an informal blueprint scenario of transcription in which the various concepts of the model can be understood intuitively.</para><para>Here is the scenario: A transcriber (typically, a scholar) examines an existing document (the exemplar, a physical entity), with the purpose of creating another document (the transcript, also a physical entity) that has a textual contents similar to that of the exemplar, and is, for some reason, more accessible – or easier to read – than the original for the members of some target community of readers, who are, by hypothesis, interested in the textual contents of the exemplar.</para><para>As mentioned above, the examination performed by the transcriber is a type-toke analysis, in which the various marks in the exemplar are identified, and some of them (the <emphasis>tokens</emphasis>) are assigned a type from some predetermined set of possible types. Groups of tokens (<emphasis>compound tokens</emphasis>) are also identified and assigned a <emphasis>compound type</emphasis> from the same set of possible types. The grouping must be such that the exemplar as a whole corresponds to a single compound token, assigned to a single compound type, which in essence captures the complete textual contents of the exemplar.</para><para>For convenience, we call the breakdown of the exemplar into tokens with assigned types a <emphasis>reading</emphasis> of the exemplar (not defined formally in the present paper, see [<xref linkend="ettdds"/>] for a formal definition).</para><para>Establishing a reading of the exemplar is by no means clerical work. But the idea is to stay as close as possible to the physicality of the exemplar. For instance, if a token could plausibly be interpreted as an i or a j, it should be assigned the <emphasis>disjunctive</emphasis> type "i/j". The ambiguity is not lifted at this stage.</para><para>The true work of the scholar begins after the reading is established. Based on the rules of transcription in force, and perhaps using their experience, knowledge of the target community of readers, and other factors, they might decide to produce a transcript that has a <emphasis>different</emphasis> textual contents from that of the reading. For example, if a token has been assigned the disjunctive type "i/j", then, instead of writing a truly ambiguous token that can be read as i or j, they may choose to write either an i or a j, or add marks, or a note, indicating that an ambiguity is present in the exemplar.</para><para>Although this part of the transcriber’s work involves the physical writing of a document (the transcript), we find it more natural – and conceptually simpler – to think of it as involving only types. In this view, the transcribers "looks" at the textual contents established by the reading (the compound type assigned to the exemplar as a whole), then decides if and how that textual contents must be modified (if at all) to satisfy the transcription rules in force. Thus, they <emphasis>first</emphasis> decide on the textual contents of the transcript, before actually writing it.</para><para>In a way, we assume that, once the desired textual contents of the transcript is decided, it will be possible to write it out physically in such a way that any reading of it (by some member of the target community of readers) will indeed result in assigning the desired (compound) type to the transcript. This does not seem to us to be an unreasonable assumption, and allows us to envision and discuss the work of the transcriber as creating a textual contents (a compound type) rather than the physical transcript.</para>
-->

 <!--<para>We have said that the aspect of transcription and of transcription practices that we want to model is the <emphasis>relationship</emphasis> that must obtain between the text of a document and the text of some other document to warrant asserting that the latter is a transcript of the former according to the transcription practice in force. But we have not defined yet how we model the text of a document. We do this by way of the notion of <emphasis>reading</emphasis>.<footnote><para>The term reading here is related to, but not to be confused with, the term reading as typically used in textual scholarship, where readings refer to variants in or among textual witnesses.</para></footnote></para><para>A <emphasis>reading of a document</emphasis> identifies the tokens in the document and the type that each of them instantiates. In other words, it is the result of a type-token analysis of the document. As such, a reading consists in the (finite) set of all tokens identified in the document, together with a mapping that associates each token to some type in the underlying type system. Thus, a reading relates material phenomena (tokens) to abstract objects (types).</para><para>Our appeal to the notion of readings will most often be implicit. Whenever we refer to a document, we will allow ourselves to also refer to its tokens and to the type of each of these tokens, without further notice. In particular, <emphasis>T-similarity</emphasis>, the similarity relation meant to hold precisely between exemplars and transcripts, will be defined simply in terms of the tokens of each document and of the type of each token.</para><para>Is is important to realize that identifying the tokens of a document and associating a type to each of them is not a trivial task, and its outcome is likely to vary with the <emphasis>agent</emphasis> (human or not) performing it. Although strictly speaking there could be as many readings of a document as there are "readers" (reading agents), the following argument might help to convince oneself that, in practice, the variation is limited; at least limited enough to make one feel comfortable with a formal model based on type-token analyses. <!-\-All these awkward formulations to avoid speaking of "the reader" of the paper.-\-></para><para>We will be dealing with two sorts of documents: exemplars and transcripts. In any actual transcription project, exemplars, on the one hand, are analyzed by transcribers, so we can simply decide that "the" reading of an exemplar is that of some chosen transcriber. Transcripts, on the other hand, mostly tend to be univocal, being typically more recent, and written with a preoccupation for clarity and lack of ambiguity. Thus, assuming a unique reading of each transcript is not too daunting.</para><para>It is also possible not to assume uniqueness of readings, and simply accept that any formal statement involving readings is "reader-dependent". Even from that stance, we believe that usual transcription contexts are such that, in practice, our formal model can yield valuable insights in a wide range of situations.</para>-->
 <!--<note>
        <para>YMA 2026-06-08: I am putting the following stuff in a note… I am not at ease
          with it, I get a feeling that it is trying to bundle the main argument of the
          whole paper as a sort of aside. See if maybe you are more comfortable with my
          reworking of the presentation of readings above. If so, we might perhaps just
          remove the rest of this note.</para>
        <para>this implicit reference to readings hides a conceptual difficulty:
          <!-\-CH deleted because of the unintended(?) pun: ", which the careful reader may 
          have spotted." // YM: This is very funny; I hadn’t thought of the possible pun, 
          so it was totally unintended. So, I get from this that you are an extra-careful 
          reader!-\->
          there can be more than one reading of one and the same document.</para>
        <para>This is often particularly clear with the kind of documents which are
          usually subject to transcription: Exemplars are often old, worn, handwritten
          documents, they may be difficult to read, unclear or ambiguous. It may be hard
          to tell, for example, which letter or word (if any) some handwritten mark on the
          surface of a worn piece of paper is. </para>
        <para>It is mostly less clear that documents which serve as the
            <emphasis>transcripts</emphasis>, normally in the form of digital or printed
          texts, lend themselves to more than one reading. While exemplars often tend
          towards the polyvocal, their transcripts mostly tend towards the univocal. This
          is natural, as the aim of transcription is typically either to remove, clarify,
          or explicitly expose difficulties in the exemplar. Even so, it is easy to see
          that transcripts as well may have more than one reading.<footnote>
            <para>At least because the conventions (the transcriptional implicatures at
              play in) transcripts may be unclear or misunderstood.</para>
          </footnote>
        </para>
        <para>The choice (or identification) of a reading for any particular document in
          any particular case is of course decided by the <emphasis>agent</emphasis>
          (human or not), the practice of the relevant community, psychological facts,
          historical or cultural conventions etc. It lies outside our scope to account for
          these factors, but we will suggest how the multiplicity of readings may come
          into play in practice in the context of transcription projects: </para>
        <para>We may think of a transcriber as reading a document, and
            <quote>selecting</quote> one among several possible readings of the exemplar,
          i.e., to identify the tokens in the document and the type that each of them
          instantiates. His task is to reinstantiate this reading in another document, the
          transcript. One complication is that the transcriber will not always
          reinstantiate his reading in its entirety &ndash; he may be guided by
          conventions to make certain omissions, changes or additions in the
          reinstantiation. Next, the reader of the transcription will establish his own
          reading of <emphasis>that</emphasis> document. A complication here is that the
          transcript will usually contain material (notes, critical marks etc.) that the
          reader does (or should) not take to reinstantiate any part of the exemplar. At
          best, the reader's reading of the transcript may be partially identical to the
          transcriber's entire reading of the document.<footnote>
            <para>Michael Sperberg-McQueen worked hard to establish a formal account of
              readings and their interrelations. (See
              e.g.http://mlcd.blackmesatech.com/mlcd/2018/Talks/London-201802/slides.html
              [Can't get xref to work here]), characterizing the task of the reader of a
              transcript as that of <quote>reconstructing</quote> or
                <quote>infering</quote> (parts of) the transcriber's reading.</para>
          </footnote>
        </para>
        <para>We may seem to be in muddy waters. A formalization of the notion of readings
          is beyond the scope of this paper. </para>
      </note>-->
 <!--  <para> Indeed, establishing a reading of a document is by no means automatic
    or exclusively clerical work. Even if the idea is to stay as close as
    possible to the physicality of the document, there is considerable leeway in
    how the identification of tokens and types can be accomplished. In
    particular, it is likely to vary with the <emphasis>agent</emphasis> (human
    or not) performing the task.</para>
   <para>For the exemplar, the choice of an appropriate reading is simple: we
    take the transcriber’s reading, including their choice of underlying type
    system. However, for the transcript, meant to be read by numerous – and
    possibly vastly different – readers, there is no natural "one" reading.
    Thus, it must born in mind that, strictly speaking, any formal statement
    involving readings is fundamentally reader-dependent.</para>-->




</article>