Prologue
I have presented many of the themes below in papers or conversations at Balisage over the course of many years. And the current elaborations and entailments drawn will be recognized as characteristic (no eye-rolling please!). This is fitting. For good or for ill almost nothing that I have written, said, or thought, would have happened without Balisage, this truly extraordinary community which has shaped (and tolerated, or more importantly, not tolerated) my odd meditations over the last 35 years. Balisage made everything possible.
Introduction
Abstraction is the unifying ideology for the triumphalist story of the information sciences generally, and of computer science in particular. It plays a leading role in how we tell the story of programming languages, software engineering, data management systems, and document text encoding. It also exemplifies the search for generality, efficiency, and elegance that characterizes these fields.
That the abstraction agenda may be problematic has long been noticed by others, both
generally (Our meddling intellect / Mis-shapes the beauteous forms of things:— / We murder to
dissect.
) and specifically with respect to formal systems for information management. With
regard to the latter, a special place must be given to William Kent’s renowned Data and Reality (Kent 1978). In addition, the applied ontology community, cited frequently below, has focused
on how formal ontologies often mishandle notions such as identity, unity, time, and
space (Guarino & Welty 2002, Borgo et al. 2022). And some writers, like the artificial intelligence theorist Brian Cantwell Smith,
have argued that no formalization can ever carve the world at its joints (Smith 1996, Smith 2019).
We continue that conversation here, noting that these difficulties are not merely local defects that just require better abstractions than ones currently presenting as problematic, but recurrences of paradoxes that have accompanied abstraction since antiquity, and that the resolutions adopted in working systems are not really resolutions, but in some sense, and not a good sense, workarounds. They replace serviceable common-sense concepts with awkward surrogates that, when embedded in the relevant social contexts where systems are being used, could create real problems. Most importantly, perhaps, we note that the stakes have changed: our flawed models provide the foundation for the lights-out automatic processing that runs the world we live in.
Abstraction I: The Logical Level
Consider the familiar periodized history of data management systems. Practices before
the emergence of the relational model in the 1970s are typically characterized as
flawed because of their direct interaction with physical
storage structures, an intrinsic connection that limited the functionality of these
systems and generated various liabilities. The most obvious of these were failures
of physical data independence
. Codd discusses these limitations in his classic paper introducing the relational
model (Codd 1970), and they are institutionalized in the 1975 Interim Report of the ANSI/X3/SPARC
Study Group on Data Base Management Systems (ANSI/X3/SPARC 1975).
The relational model addresses this problem by presenting a simple abstraction that
represents data independently of how that data is stored. This representation is a
table – or, mathematically, a relation
, a set of tuples. Relations are then mapped to the underlying storage structures
and software applications and end users then interact with stored data via the familiar
language of columns, rows, and values.
This provides physical and logical data independence: changes in storage methods may be made without any effect on subroutines or preexisting queries, and data attributes may be added or removed without affecting them. In addition, it is no longer necessary for end users and application programmers to understand how data is stored; inconsistencies can be more easily avoided; inferencing and validation are easier; and system design, documentation, reformatting, and interoperability are simplified and disciplined.
The descriptive markup story runs along similar lines, although here the focus is primarily on abstracting away from intended processing rather than from storage methods (Goldfarb 1981, ISO 8879:1986). A logical-level model is defined – in this case an ordered rooted tree with an accompanying formal grammar as a schema, and the editorial (or, more generally, communicative) components of the document represented as nodes. Processing instructions are mapped to node types. As with the relational model the advantages, grounded in abstraction and indirection, are substantial and varied (Coombs et al. 1987, DeRose et al. 1990).
In the beginning, abstract logical models such as the relational model and markup
grammars seemed to be not only better than any other approach to data management,
but the best imaginable approach
(Coombs et al. 1987). In fact, these models seemed to function at the level of the information itself,
rather than at the level of storing and processing data representations. Alternatively,
one might say that they functioned at the level of human understanding of the problem
domain, establishing a sort of conceptual impedance match between the formal model
and human understanding.
Abstraction II: The Conceptual Level
However, to some in the database community it appeared that relational databases still recorded and organized data, rather than representing how things are in the world. The values in a relation are strings, not things; the juxtapositions of adjacent cells or column groupings are paratactic, not predicational. Any connection with the world is left to human interpretation. Table-talk (SQL) is about rows, columns, and values, not things (like people), properties (like being German), or relationships (like being the supervisor of an employee). The explicit features of the relational model – tuples, sets, foreign keys, and so on – are natural neither to our common-sense conceptual scheme nor to the domain problems we are attempting to address. Moreover, tables are just one possible logical-level abstraction. Alternative logical-level abstractions record the same information. But what methods or languages do we have to express, in any particular case, what the different but equivalent abstractions of the same data have in common? How can we provide representational continuity when we change from one logical model to another?
Responding to these concerns, models at a higher level of abstraction were proposed: conceptual models. The textbook example is the familiar and widely used entity-relationship model (Chen 1976). The entity-relationship model identifies explicitly the entities, properties, and relationships that are indicated only implicitly by the rows, columns, and values in the relational model. So, just as logical-level models added a layer of abstraction on top of storage models, conceptual models added a layer of abstraction on top of logical-level models, and allowed for system-design decisions independent of decisions about logical models. This greatly facilitated system design, allowing a focus on things and relationships rather than values and relations. In addition, it allowed for wider participation in system design, as only a commonsense informal understanding of the particular needs of users and organizations was required for participation in design meetings – no knowledge of features specific to the relational model (or any logical model) was necessary.
Although typically represented diagrammatically, conceptual models were explicitly defined in first-order predicate logic (plus identity). Entity types indicated monadic predication, relationships dyadic predication, and cardinality was expressed with quantifiers and identity. Such formal constructs could, however, remain behind the scenes as designers operated with their common-sense understandings of things, properties, and relationships. NIAM, IDEF1X, and UML class diagrams are examples of other similar late-20th-century conceptual modeling systems.
As ER diagrams are awkward for XML documents, there is no commonly used conceptual modeling system in the document markup community, although the need for a conceptual model for markup languages was argued for by some markup theorists (Cover 1998), and an academic project, Béchamel, led by Michael Sperberg-McQueen, produced a system that represented an SGML or XML document as a set of logical statements that could be processed according to defined inference rules (Sperberg-McQueen et al. 2000, Renear et al. 2002, Sperberg-McQueen et al. 2002, Dubin et al. 2003). For a survey of proposed conceptual models for SGML/XML logical models, see Nečaský 2006.
Early conceptual models were primarily applied to the development of useful normalized
relational databases, which was the intended purpose. However, as conceptual models
became more wide-ranging in application, they are better described as ontologies with
upper level
categories (such as physical object, event, property, string, number), often including
axioms and executable rules that formalize a detailed understanding of the domain
of interest.
Such ontologies appear to be, finally, describing the world and not simply organizing
data. It was sometimes objected, with some plausibility, that a representation of
the world was sufficiently implicit in logical-level systems all along, and that far
from providing a missing semantics, conceptual models were really just more syntax
. For these critics, the more abstract models seemed unnecessary, and perhaps just
another doomed attempt to externalize, or naturalize
, human thought.
Nevertheless, the data management and software engineering communities, and now many scientific communities as well, generally found much value in higher levels of abstraction. Conceptual models and ontologies do seem to function at a level that approximates human understanding, and to make our assumptions about a domain logically explicit and actionable. It seems that this time we really had, finally, reached peak abstraction.
Paradox I: Ancient Origins
Yet there is a problematic, and perhaps ominous, side to this story.
The relationship between abstraction, paradox, and scientific advance has been a common – even if sometimes suppressed – theme in the history of science, mathematics, and philosophy. Scientific thought as it emerged in antiquity is itself an abstraction agenda, an effort to generalize and deepen our grasp of the world around us by formalizing and making more precise our common-sense understanding of it. But as exemplified by philosophical traditions in multiple historical cultures, this effort often leads to paradox. Early examples include puzzles about motion, space, and time; identity and change; self-referential statements; vague predicates; and various unexpected failures of transitivity. In the West this begins perhaps as early as Parmenides and Heraclitus, and the quickly vivid and notorious in the aporiai of Zeno and Eubulides of Miletus, among others.
For the modern mind, the paradox-generating nature of abstraction is perhaps best
exemplified by Bertrand Russell’s famous antinomy of naïve set theory, in which formalizing
the simple concept of a collection of things that share a property leads quickly to
a contradiction (the set of all sets that are not elements of themselves is both an
element of itself and not an element of itself). Gottlob Frege, while at work building
the logical foundations for computer science, anxiously wrote back to Russell that
this observation undermined mathematics itself (wenn dieses erschüttert ist, so wankt das ganze Gebäude
).
Some of these early paradoxes are now felt to have been resolved in various ways. For instance, the paradoxes of time and motion seem to be resolved by the contemporary theory of limits, and the paradoxes of naïve set theory by axiomatic set theory (ZF) or type hierarchies. But others seem to be with us still. For example, there are paradoxes involving vague predicates (sorites), transitivity of sameness across material change (the Ship of Theseus), elusive subjects of predication (paradoxes of increase and group membership), and intransitive indifference.
However, it is significant that many of these puzzles involve problematic notions that are fully serviceable in their ordinary applications but that generate paradoxes only when formalized and applied to special cases. From that perspective, there is a sense that putative resolutions are not resolutions exactly, but rather workarounds that offer complex alternative concepts that, although paradox-free (so far), are not as useful for the routine work of life and science.
Paradox II: Modern Challenges
Abstraction considerably improves the performance of our digital information systems. To that end, through ontologies and conceptual models, we attempt to make our understanding of the world tractable by reducing or eliminating idiom, metaphor, vagueness, ambiguity, and logical fiction. This is how we can take advantage of formality’s specific affordances, such as semantic compositionality, existential instantiation, and logical inference, all of which are valuable to information management, yet none of which are easily or unproblematically provided by ordinary language.
Local specializing applications of abstraction may seem to improve our systems in the ways described. However, when technical modifications of familiar concepts are being made, it is difficult to formalize notions foundational to these systems in ways that match our common-sense intuitions. Below are some examples. Most of these have been addressed in actual systems and models, but, as described above, those resolutions often appear to be workarounds of some sort, not general resolutions.
-
There is no formal definition of document, text, or file that is consistent with our assumption that such things can change or be modified (Thibodeau 2002, Duranti & Thibodeau 2006, Renear et al. 2008, Renear & Wickett 2009, Renear & Wickett 2010, Yeo 2010). For a clear example of a resulting contradiction in a model: some data models have defined a
file
as a bitstream and specify amodification date
. However, a bitstream is not a mutable object. Recent carefully developed models, Library of Congress’s PREMIS data preservation model, do make files immutable (the associated date becomes a creation date). But this still leaves us without a formalization of our common-sense notion of modifiable files (PREMIS 2015). -
Formal models for digital libaries have sometimes defined collections as mathematical sets. However, there is no commonly accepted formal definition of groups (such as collections) that allows items to be added or removed, which is inconsistent with our formal model of library collections (Renear et al. 2010, Wickett et al. 2011, Galton & Wood 2016).
-
There is no formal account of the distinction between physical objects and their matter that preserves the common belief that they are the same thing in some sense, and yet also different in some other sense. Of course some ontologies address the issue directly (Borgo et al. 2022), but they vary in their approach and often require obscure sui generis concepts such as constitution.
-
There is no formal account of intransitive indifference that matches our intuition about its transitivity under indiscernibility. You may be indifferent betwen a hamburger with 2000 grains of salt and one grain of salt but not ... (etc.) (Luce 1956, Halpern 2008).
-
There is no formal account of vague predicates that matches our use of vague terms (Bittner 2023).
-
There is no formal treatment of propositional attitudes (e.g., RDF reification) that is consistent with the fact that co-referential expressions are not intersubstitutable salve veritate in propositonal attitude contexts (Renear & Choi 2005). For a clear example of how this is evaded rather than resolved: ontologies of information objects fix the identity of a reified proposition by its symbol structure and expressly set interpretation aside, so that referential opacity is bracketed, not addressed (Gangemi & Mika 2003, Doerr et al. 2012).
And so on.
Why It Matters
Francis Fukuyama famously described the apparent success of liberal democracy as the end of history
. It was a happy thought, if a brief one. The imminent success of formal ontologies
was a similar happy thought. But the desired impedance match between such ontologies
and our common-sense conceptual scheme has proved as elusive as the end of history
itself – and the optimistic expectation that it is achievable is, perhaps, similarly
dangerous.
Humans themselves may be abstraction machines, but if so we are selective abstraction
machines. For the most part, when we navigate the world around us, we make liberal
use of idiom and metaphor, tolerate vagueness, leave ambiguities deliberately unresolved,
and so on. Logical fictions abound in our discourse, and if serviceable noun phrases
imply dubious existential instantiations, or common-sense beliefs conceal hidden contradictions,
well, no matter – we simply don’t go there
. OAnd in practice there is typically little evidence of resulting difficulties. No
harm, no foul. We see such ontologically suspicious phrases as merely façons de parler, and any paradoxes as just interesting riddles unrelated to the urgencies of the
current project. We know how to proceed with the work before us, and we get on with
it. Resolution of the problem is a simple matter of programming
. We can keep the model as is and deal with problems in the software.
The underlying issues here are of course well known. Our open textured
concepts resist complete formalization (Waismann 1945). But maybe this is no longer something that can be noted and shrugged off. Ontologies
and conceptual models are used to support reasoning, and increasingly that reasoning
takes place without human oversight, in a world of lights-out
automation.
Sometimes the formalization is outright inconsistent – and as every schoolboy knows,
ex contradictione quodlibet.
But the more common and more insidious case is the workaround that succeeds: a
consistent formalization of part of an open textured concept, applied
automatically to cases it was never adequate to, with no one present to notice
the mismatch between the technical term and the common-sense one term. In the past,
paradoxes were primarily of academic interest; now they are embedded in the systems
that organize and sustain the world we live in (Bowker & Star 1999), atavistically creating new sources of risk and unreliability. And this time, the
consequences may be more than philosophical amusement. H. L. A. Hart held that it
is the responsibility of a judge to resolve the penumbra
of meaning surrounding our concepts (Hart 1961). But in lights-out
automation the judge is a home, sound asleep.
Strangely, the models that seemed to be, and were intended to be, at precisely the same level of abstraction as our common-sense conceptual scheme fail exactly at that. Moreover, they now threaten to recapitulate the logical paradoxes of the last two millennia – as if these paradoxes were not just latent in that scheme, but ineliminable features of it.
But is a paradox-free formalization of our conceptual scheme possible? Or does understanding
the world force us to choose between logical fiction and logical paradox? Is that
yet another tradeoff
? The failure of paradox-free abstraction is, in any case, yet another reason for
keeping a human in the loop. As humans, we know how to intellectually navigate a world
of metaphor, idiom, vagueness, ambiguity, logical fiction, and latent paradox. We
have been doing it, successfully, for a long time.
Postscript: Deep Learning and LLM-based AI
The problems described above have been apparent for a while. But today there is a new development. In just a few years, deep learning has accomplished what Douglas Lenat’s Cyc project could not, after decades of effort and millions in funding: the successful navigation of salad bars and roundabouts.
Deep learning has little use for the familiar logical abstractions of conceptual modeling and ontologies, and so it appears to avoid their attendant vulnerability to paradox. Large language models, specifically, accept ordinary-language prompts and respond in kind. The problems described above are latent in both the prompts and the responses – but only in the same sense in which they are latent as well in our ordinary discourse about the world. Absent formalization, they do not present as problems. More strangely still, our interactions with LLM-based AI are free, on both sides, to enjoy the useful vagueness and meaningful fictions of ordinary language. So is it here, then, that we really do, at last, have an impedance match?
Or have the fundamental issues and dangers simply been hidden, dangerously, behind an unaccountable simulation? Even if much human cognition is connectionist, even LLM-like, humans are accountable in ways that LLMs are not: we can ask for surveyable reasons and actual provenance, and we can provide them.
But that is a story for another time.
Acknowledgements
These thoughts owe much to collaborations with many people, especially Michael Sperberg-McQueen, who gave me a (transitive) identity, and David Dubin, Claus Huitfeldt, Karen Wickett, and Bonnie Mak, as well as many conversations at Balisage and at the School of Information Sciences, University of Illinois Urbana-Champaign. Nevertheless, they must all be held blameless.
References
[ANSI/X3/SPARC 1975] ANSI/X3/SPARC Study Group on Data Base Management Systems. (1975). Interim report. FDT (ACM SIGMOD Bulletin), 7(2).
[Bittner 2023] Bittner, T. (2023). Information, mereology and vagueness
. Applied Ontology, 18(2), 119–167. doi:https://doi.org/10.3233/AO-230277.
[Borgo et al. 2022] Borgo, S., Ferrario, R., Gangemi, A., Guarino, N., Masolo, C., Porello, D., Sanfilippo,
E. M., & Vieu, L. (2022). DOLCE: A descriptive ontology for linguistic and cognitive engineering
. Applied Ontology, 17(1), 45–69. doi:https://doi.org/10.3233/AO-210259.
[Bowker & Star 1999] Bowker, G. C., & Star, S. L. (1999). Sorting Things Out: Classification and Its Consequences. Cambridge, MA: MIT Press.
[Chen 1976] Chen, P. P.-S. (1976). The entity-relationship model – Toward a unified view of data
. ACM Transactions on Database Systems, 1(1), 9–36. doi:https://doi.org/10.1145/320434.32044.
[Codd 1970] Codd, E. F. (1970). A relational model of data for large shared data banks
. Communications of the ACM, 13(6), 377–387. doi:https://doi.org/10.1145/362384.362685.
[Coombs et al. 1987] Coombs, J. H., Renear, A. H., & DeRose, S. J. (1987). Markup systems and the future of scholarly text processing
. Communications of the ACM, 30(11), 933–947. doi:https://doi.org/10.1145/32206.32209.
[DeRose et al. 1990] DeRose, S. J., Durand, D. G., Mylonas, E., & Renear, A. H. (1990). What is text, really?
Journal of Computing in Higher Education, 1, 3–26. doi:https://doi.org/10.1007/BF02941632
[Doerr et al. 2012] Doerr, M., et al. (2012). Information carriers and identification of information objects: An ontological approach
. arXiv:1201.0385. doi:https://doi.org/10.48550/arXiv.1201.0385.
[Dubin et al. 2003] Dubin, D., Renear, A. H., Sperberg-McQueen, C. M., & Huitfeldt, C. (2003). A logic programming environment for document semantics and inference
. Literary and Linguistic Computing, 18(1), 39–47. doi:https://doi.org/10.1093/llc/18.1.39
[Duranti & Thibodeau 2006] Duranti, L., & Thibodeau, K. (2006). The concept of record in interactive, experiential and dynamic environments: the view
of InterPARES
. Archival Science, 6(1), 13–68. doi:https://doi.org/10.1007/s10502-006-9021-7.
[Galton & Wood 2016] Galton, A., & Wood, Z. (2016). Extensional and intensional collectives and the de re/de dicto distinction
. Applied Ontology, 11(3), 205–226. doi:https://doi.org/10.3233/AO-160168.
[Gangemi & Mika 2003] Gangemi, A., & Mika, P. (2003). Understanding the semantic web through descriptions and situations
. In Proceedings of ODBASE 2003. doi:https://doi.org/10.1007/978-3-540-39964-3_44.
[Goldfarb 1981] Goldfarb, C. F. (1981). A generalized approach to document markup
. ACM SIGPLAN Notices, 16(6), 68–73. doi:https://doi.org/10.1145/800209.806456.
[Guarino & Welty 2002] Guarino, N., & Welty, C. (2002). Evaluating ontological decisions with OntoClean
. Communications of the ACM, 45(2), 61–65. doi:https://doi.org/10.1145/503124.503150.
[Halpern 2008] Halpern, J. Y. (2008). Intransitivity and vagueness
. Review of Symbolic Logic, 1(4), 530–547. doi:https://doi.org/10.1017/S1755020308090084.
[Hart 1961] Hart, H. L. A. (1961). The Concept of Law. Oxford: Clarendon Press.
[ISO 8879:1986] International Organization for Standardization. (1986). Information processing – Text and office systems – Standard Generalized Markup Language (SGML) (ISO 8879:1986). Geneva: ISO.
[Kent 1978] Kent, William. (1978). Data and Reality: Basic Assumptions in Data Processing Reconsidered. Amsterdam: North-Holland.
[Luce 1956] Luce, R. Duncan. (1956). Semiorders and a theory of utility discrimination
. Econometrica, 24(2), 178–191. doi:https://doi.org/10.2307/1905751.
[Mak & Renear 2023] Mak, B., & Renear, A. H. (2023). What is information history?
Isis: A Journal of the History of Science Society, 114(4), 747–768. December 2023. doi:https://doi.org/10.1086/727568.
[Nečaský 2006] Nečaský, M. (2006). Conceptual modeling for XML: A survey
. In Proceedings of DATESO 2006 (CEUR-WS Vol. 176).
[PREMIS 2015] PREMIS Editorial Committee. (2015). PREMIS Data Dictionary for Preservation Metadata, Version 3.0. Washington, D.C.: Library of Congress. https://www.loc.gov/standards/premis/v3/premis-3-0-final.pdf.
[Renear & Choi 2005] Renear, A. H., & Choi, Y. (2005). Trouble ahead: Propositional attitudes and metadata
. ASIST Proceedings. doi:https://doi.org/10.1002/meet.14504201275.
[Renear et al. 2002] Renear, Allen, Dubin, David, & Sperberg-McQueen, C. M. (2002). Towards a semantics for XML markup
. In Proceedings of the 2002 ACM Symposium on Document Engineering (DocEng ’02) (pp. 119–126). New York, NY: Association for Computing Machinery. doi:https://doi.org/10.1145/585058.585081.
[Renear et al. 2008] Renear, A. H., Dubin, D., & Wickett, K. (2008). When digital objects change – Exactly what changes?
ASIST Proceedings. doi:https://doi.org/10.1002/meet.2008.14504503143.
[Renear et al. 2010] Renear, A. H., Sacchi, S., & Wickett, K. M. (2010). Definitions of dataset in the scientific and technical literature
. Proceedings of the Association for Information Science and Technology (Pittsburgh). doi:https://doi.org/10.1002/meet.14504701240.
[Renear & Wickett 2009] Renear, A. H., & Wickett, K. M. (2009). Documents cannot be edited
. Proceedings of Balisage: The Markup Conference. doi:https://doi.org/10.4242/BalisageVol3.Renear01.
[Renear & Wickett 2010] Renear, A. H., & Wickett, K. M. (2010). There are no documents
. Proceedings of Balisage: The Markup Conference. doi:https://doi.org/10.4242/BalisageVol5.Renear01.
[Smith 1996] Smith, Brian Cantwell. (1996). On the Origin of Objects. Cambridge, MA: MIT Press.
[Smith 2019] Smith, Brian Cantwell. (2019). The Promise of Artificial Intelligence: Reckoning and Judgment. Cambridge, MA: MIT Press.
[Sperberg-McQueen et al. 2000] Sperberg-McQueen, C. M., Huitfeldt, C., & Renear, A. (2000). Meaning and interpretation of markup
. Markup Languages: Theory & Practice, 2(3), 215–234.
[Sperberg-McQueen et al. 2002] Sperberg-McQueen, C. M., Dubin, D., Huitfeldt, C., & Renear, A. H. (2002). Drawing inferences on the basis of markup
. Extreme Markup Languages®.
[Thibodeau 2002] Thibodeau, K. (2002). Overview of technological approaches to digital preservation and challenges in coming
years
. In The State of Digital Preservation: An International Perspective (CLIR Publication 107). Washington, D.C.: Council on Library and Information Resources.
[Waismann 1945] Waismann, F. (1945). Verifiability
. Proceedings of the Aristotelian Society, Supplementary Volume 19.
[Wickett et al. 2011] Wickett, K. M., Renear, A. H., & Furner, J. (2011). Are collections sets?
ASIST Proceedings. doi:https://doi.org/10.1002/meet.2011.14504801145.
[Yeo 2010] Yeo, G. (2010). Nothing is the same as something else: significant properties and notions of identity
and originality
. Archival Science, 10(2), 85–116. doi:https://doi.org/10.1007/s10502-010-9119-9.
[Cover 1998] Cover, R. (1998). XML and Semantic Transparency. Cover Pages, https://xml.coverpages.org/xmlAndSemantics.html.