When a new document type lands on your desk (a contract, an invoice, a vaccination certificate), the first engineering instinct is almost always the same: design the schema. One data model per document type, rich enough to capture everything important, built once and reused across every workflow that will ever touch that document. It feels like the responsible thing to do. It is also, I have come to think, one of the most expensive mistakes in document AI.
This is the first of a three-part series on how I have learned to think about knowledge in document AI systems. Before the architecture (part two) and the implementation details (part three), I want to start with the stance, because the stance is what makes the rest coherent. That stance is borrowed, almost wholesale, from a hundred-year-old American philosophical tradition called pragmatism. Bear with the philosophy: it pays for itself quickly.
Let me state the claim up front, and state it carefully. It is not that structure never generalizes; for some documents a surprising amount of it does. It is that no single schema captures everything that matters about a document, because what matters is relative to the question you are asking. Knowledge is structured by use, and the categories we extract are tools rather than universal truths. Let’s dive deeper.
The universal-schema instinct
The instinct has a respectable pedigree. It goes back to Aristotle’s Categories, which set out to enumerate the fundamental kinds of things once and for all: substance, quantity, quality, relation, and so on. Every entity falls under a category, categories have structure, and reasoning becomes tractable. Computer science inherited the ambition directly. When Tom Gruber gave the field its canonical definition in 1993, an ontology became “an explicit specification of a conceptualization”: a formal vocabulary of classes, relations, and constraints that a machine can reason over. The word was not borrowed by accident: computer science took over philosophy’s ambition of naming, once and for all, the kinds of things a domain contains.
For part of the problem, this works beautifully. A document, considered as an artifact, really does have stable structure: pages, blocks, tables, reading order, sections, signatures. That layer generalizes across contracts, certificates, and transcripts, because every one of them is, first, a document. Build it once, reuse it everywhere.
The instinct rarely stops at the artifact, though. Having modeled the document’s structure once, it reaches inward for the same universality: one schema for what the document means, rich enough that every future workflow can read what it needs off a single representation. That is the move I want to question. Some of what is inside genuinely generalizes, but the slice of meaning any given task actually needs does not settle into one shared form: how you carve up a document’s meaning depends on what you are trying to do with it.
You might reasonably push back here, and the strongest counterexample is the humble invoice. An invoice is so regular that a single generic schema (issuer and recipient, line items, amounts, tax, dates, a reference number) really does apply to almost all of them. That is a genuine, reusable schema, and I do not want to wave it away: capturing that shared core once, rather than reinventing it per project, is exactly the right move. We will give it a home in the next post, as the middle layer of a three-layer model.
But notice what that reusable core is and is not. It is a stable vocabulary of the entities an invoice tends to contain; it is not a schema that captures everything that matters, because what matters still depends on the question. Ask whether an invoice is a duplicate payment, whether its tax treatment is correct, or whether it is fraudulent, and the decomposition that actually helps diverges each time: which fields you trust, what counts as a match, what “correct” even means. That task-shaped knowledge is the layer the universal instinct cannot pre-model, and the one the next post argues you should build only when a task demands it. A contract makes the point more starkly still: its useful “joints” are completely different when you are checking enforceability, computing tax exposure, or screening for fraud. There is no view-from-nowhere decomposition that captures everything important, because importance is relative to a question. The pragmatist way to put it: decomposition is downstream of use.
Wittgenstein gave us the sharpest tool for seeing why. Concepts like “contract” are not natural kinds with a shared essence; they are family resemblance concepts (Philosophical Investigations, sections 66 to 67), whose members overlap and criss-cross without a single defining feature. Their boundaries shift with jurisdiction and purpose. Trying to write the one definition that captures all contracts is chasing an essence that is not there. Mistaking a meaning problem for a structure problem is the root error, and it is an easy error to make precisely because our tools are so good at the structure part.
Concepts are tools, not mirrors
A category is just a concept put to work: the slice of meaning we decide is worth extracting. So if knowledge is structured by use, what is such a category actually for? Here the pragmatists are blunt. For John Dewey, a concept is not a picture of reality; it is an instrument for resolving a problematic situation. Peirce got there first with his pragmatic maxim of 1878: the meaning of a concept is the sum of the practical consequences of holding it. A category earns its place if it makes some downstream workflow succeed, and the same domain can legitimately warrant different categories for different tasks. A field that helps no inquiry is decoration, however correct it looks on paper.
That is the whole argument, and it points to a single design move: stop trying to model the document, and start trying to equip the inquiry. For any document type there is no one correct ontology, only the ontology that makes a specific downstream question answerable. Expect to hold more than one at a time. Expect to retire them when the question changes. This sounds permissive, but it is the opposite. “Structured by use” is not “anything goes”, it is “justified by use”: every category has to earn its keep against a real workflow.
We already do this, on the evaluation side
Here is the part I find quietly reassuring. Machine learning already practices pragmatism; it just practices it downstream. We do not ask whether a model’s outputs correspond to reality. We ask whether they hold up across distributions, support good decisions, and stay stable under perturbation. That is Peirce’s maxim almost verbatim, meaning cashed out in practical consequences, operationalized as a held-out test set. William James would recognize it at once: truth as what proves itself in practice.
So the field is pragmatist about how it judges models, and rationalist about how it designs schemas, still reaching for the one universal conceptualization that will capture everything. The fix is not more philosophy. It is to bring the stance we already trust for evaluation upstream, to the point where we decide what to extract in the first place.
Where to go from here?
In summary:
- Knowledge is structured by use. A document has no privileged decomposition; its joints depend on the question you are asking.
- Concepts are tools. A category is good if it helps a downstream workflow, so equip the inquiry rather than model the document.
Concretely, the next time you reach for the one universal schema of a document type, stop and ask which task it serves: build the schema for that question, expect to hold more than one, and let each category earn its place against a real workflow. None of this argues against structure. The document-as-artifact layer is real, shared, and worth getting right once. The argument is about where universality stops and use begins, and about being honest that most of the interesting knowledge lives on the “use” side.
In the next part of this series I will make that concrete with a three-layer model that puts each kind of knowledge where it belongs, and a rule that keeps the whole thing diagnosable: never skip a layer. Stay tuned.