Zebrapedia had the right idea. After a decade of work, it also has something more uncomfortable: proof that the right idea, executed with the wrong method, may leave one of the twentieth century's most important philosophical manuscripts permanently unfinished.

The Scale of the Problem

8,000+handwritten pages
~10years of the project
127mentions of 2-3-74 — the central subject — spread across 26+ folders
28+archive folders, each its own workset

Philip K. Dick's Exegesis is the largest body of unpublished work by any major twentieth-century author. From 1974 until his death in 1982, Dick filled notebooks, loose-leaf pages, and typewritten sheets with an eight-year attempt to understand a visionary experience he called "2-3-74" — February and March 1974, when he believed he received direct transmissions from an intelligent cosmic entity he called VALIS. The result is a text simultaneously theological, autobiographical, philosophical, and fictional: a mind working at full throttle in the margins of its own sanity.

A print edition published in 2011 covers roughly a tenth of the full manuscript. Zebrapedia was founded to make the rest accessible — to transcribe every page, index every recurring concept, and build the kind of annotated scholarly apparatus that a text of this density demands. That mission remains the correct one. What went wrong is worth understanding, both for the sake of this project and for every similar effort that follows it.

Before the critique, the credit. Zebrapedia made several architectural decisions that are genuinely excellent — decisions that any successor project should adopt wholesale.

Strength

The subject backlink index. Every entity indexed in the project — a concept, a character, a date, a publication — links back to every manuscript page that mentions it. This is not a small thing. It inverts the reading problem entirely. Instead of confronting 8,000 pages sequentially (impossible), a scholar can ask: "Show me everything Dick wrote about 2-3-74" and receive 127 direct links across 26 folders. This is the right way to navigate a distributed archive.

Strength

The three-level category taxonomy. Zebrapedia groups its subjects into 14 top-level categories — Characters, Notable Dates, Philosophical Concepts, Publication, Places, People, and so on — with up to three levels of nesting. Horselover Fat lives under Characters > VALIS > Fat. This is exactly the right depth for humanistic taxonomy: enough structure to orient a newcomer, not so deep as to become a bureaucratic maze.

Strength

The co-occurrence subject graph. Every subject page displays a network diagram showing which other subjects appear on the same manuscript pages. The closer a node to the centre, the more often the two subjects are mentioned together. This visualisation answers a question a flat list cannot: what does Dick think about at the same time as X? What concepts cluster around 2-3-74? The graph makes implicit structure visible.

Strength

The double-bracket markup system. The [[Canonical|Verbatim]] syntax, borrowed from wiki software, is the simplest possible annotation tool for non-technical scholars. A volunteer does not need to understand XML or TEI to link "2-74" to the canonical entry for "2-3-74." The syntax is learnable in two minutes.

Strength

Multiple export formats. From the same underlying data, Zebrapedia can produce plain text, TEI-XML, HTML, IIIF manifests, and CSV exports of all subject mentions. This is the right model: one source of truth, many derived outputs, each suited to a different downstream use — close reading, archival deposit, network analysis, library interoperability.

These decisions constitute a genuinely thoughtful information architecture. Zebrapedia understood the shape of the problem. What it underestimated was the labour required to fill that architecture with content.

The Central Case: 2-3-74

"2-3-74" is the most frequently mentioned subject in Zebrapedia's index: 127 pages across at least 26 of the archive's folders. It is the name Dick gave to the visionary experience of February–March 1974, the event around which the entire Exegesis spirals. Every problem with Zebrapedia can be illustrated through what happens — and fails to happen — when you navigate to this subject.

The subject page correctly shows a count of 127 mentions and organises them by folder. There is a graph of related subjects. There is a field for a scholarly description of what "2-3-74" means. And yet: the description field is either empty or filled with placeholder text. The co-occurrence graph is a static image with no filters. There is no way to ask whether Dick's language about 2-3-74 changes between 1974 and 1980. There is no semantic search that would let you find pages where he is clearly describing the same experience without using that exact term. And the 127 mentions represent only the pages that a volunteer happened to transcribe and then chose to tag — an unknown fraction of the total.

The subject that defines the whole project is, after a decade, still a stub.

Problem 1 — The Manual Markup Bottleneck

The Exegesis contains thousands of recurring concepts, names, dates, and references. Dick mentions 2-3-74 on at least 127 pages. He mentions VALIS, Horselover Fat, Thomas the Apostle, the Gnostic pleroma, the Black Iron Prison, the Palm Tree Garden, and hundreds of other entities throughout the manuscript. Every single one of those mentions needed to be tagged by a human volunteer, one [[bracket]] at a time.

The problem

Manual markup does not scale to 8,000 pages. Even with Zebrapedia's Autolink feature — which mines previous transcriptions to suggest markup patterns on new pages — the system relies on string matching, not understanding. It misses variant phrasings, conceptual references without the expected words, and anything that hasn't been tagged before. The platform's own documentation acknowledges that Autolink suggestions "are not perfect" and require manual review before saving. After a decade, the index remains partial.

What could be done instead

A large language model can read a transcribed page and identify named entities — people, concepts, dates, places, publications — with a level of contextual understanding no pattern-matcher achieves. An LLM NER (named entity recognition) pipeline run against the fully transcribed corpus could produce a comprehensive first-pass index in hours rather than years. Matches above a confidence threshold are auto-linked; uncertain matches are queued for human review. The human scholar's job shifts from creating the index to curating it — a far better use of scholarly time.

This is not speculative. Projects like Transkribus already use AI to handle handwritten text recognition at scale. The tools exist. What Zebrapedia needed was not more volunteers with brackets — it needed a pipeline.

Problem 2 — The Missing Dimension of Time

Dick wrote the Exegesis over eight years. He returned to the same questions repeatedly: What was 2-3-74? Was it God? VALIS? A hallucination? Psychosis? His answers changed, reversed, and contradicted each other across that span. The temporal evolution of his thinking is one of the most important scholarly questions about the text.

The problem

Zebrapedia has no chronological navigation. There is no date filter, no timeline view, no way to see the 127 mentions of 2-3-74 sorted by when they were written. Folders do not map cleanly to years. Pages within folders are not reliably dated. The archive is navigable by subject and by physical folder, but not by time — the very dimension most relevant to a text that documents a mind changing over eight years.

What could be done instead

Dick's pages frequently contain date headers — "Feb 1974," "2-74," "March, 1975," "1/76." A pattern-matching pass over OCR'd text can extract these with reasonable accuracy, populating a date_tag field for each page. This is deterministic preprocessing requiring no AI at all — just regular expressions and a cleanup pass. The result: a chronological slider across the whole corpus. A scholar could ask "Show me every mention of 2-3-74 between 1975 and 1977" and watch Dick's theory of his own experience evolve in real time.

Problem 3 — A Subject Index With Nothing to Say

The most striking discovery from auditing Zebrapedia's live subject pages is this: the description field for "The Exegesis" — the project's own namesake and central subject — contained advertising joke copy for a fictitious metallic foil product called "TransFoilent." This is not a minor gap. It is a symptom of a systemic failure.

The problem

Zebrapedia's subject index is structurally sound but almost entirely empty of content. The taxonomy tells you that "2-3-74" is a Notable Date and that "Horselover Fat" is a Character in VALIS. It does not tell you what these things mean, why they matter, how Dick's use of the term evolved, or what scholars have said about them. After ten years of volunteer labour, the knowledge base remains a list of names with no knowledge attached.

What could be done instead

For every subject with more than three corpus mentions, an LLM can synthesise a draft description from the top passages that mention it. "Here are the five most informative passages from the Exegesis about 2-3-74. Write a 200-word scholarly summary of how Dick uses this concept, noting its significance and relationship to adjacent ideas." The result is a seeded draft — imperfect, AI-generated, clearly labelled as such — that a human scholar can edit, correct, and curate. A thousand stub articles become a thousand starting points rather than a thousand blank pages.

Problem 4 — A Graph That Cannot Be Questioned

The subject co-occurrence graph is one of Zebrapedia's best ideas and one of its most frustrating implementations.

The problem

The graph is a static image. You cannot filter it by date, folder, or category. You cannot ask "which concepts co-occur with 2-3-74 specifically in the 1974–1975 entries, versus the 1979–1980 entries?" You cannot compare the conceptual neighbourhoods of two subjects side by side. You cannot export the graph data to analyse in Gephi or any other network tool. You can look at it and click on a node to go to that subject page. That is the full extent of the interaction.

A graph that cannot be questioned is closer to decoration than to a research tool.

What could be done instead

The co-occurrence data itself — how many pages mention subject A and subject B together — is straightforward to compute from the entity mentions table and store in a graph edge table. Rendering it as an interactive force-directed graph with D3.js takes perhaps two days of frontend work. Add a date range slider, a category filter, a minimum co-occurrence threshold, and a comparison mode for two subjects, and you have a genuine research instrument. Add GraphML export and it becomes a data source for external network analysis. None of this is technically difficult. It is simply a choice that was not made.

Zebrapedia's subject index is a powerful tool for the scholar who already knows the name of what they are looking for. It is nearly useless for discovery.

The problem

There is no full-text search across the transcribed corpus. There is no semantic or fuzzy search — no way to find pages thematically similar to a query without using exact terminology. A reader unfamiliar with Dick's private vocabulary cannot type "light that reveals reality" and find the passages where Dick describes his anamnetic experience in those terms. They must first know that Dick called it "2-3-74," then navigate to that subject page, then browse the 127 linked pages hoping to find the right one. Zebrapedia rewards prior expertise and punishes genuine curiosity.

What could be done instead

Full-text search over the transcription corpus is a solved problem: SQLite's FTS5 extension provides fast, ranked full-text search over millions of words at no cost. Semantic search — finding passages conceptually similar to a query without relying on exact words — requires embedding the transcription chunks into vector space and retrieving by cosine similarity. This is the technology behind modern AI search systems, and it can run entirely locally using open embedding models. Together, they give the scholar two modes: "find the exact phrase" and "find passages that feel like this."

Problem 6 — The Conversation Happens Elsewhere

The problem

Discussion on Zebrapedia is conducted through an embedded Google Groups forum attached to each subject page and manuscript page. The project's broader community discussion migrated to a Facebook group called "Radio Free Valis Satellite Station." Neither solution integrates with the actual work of annotation. A comment in Google Groups cannot be attached to a specific sentence in a transcription, or linked to the passage in question, or threaded against a particular annotation. The conversation and the text exist in separate siloed systems that do not know about each other.

The migration to Facebook is the telling signal. When a community moves off the platform its project runs on, the platform has failed them.

What could be done instead

Discussion should be threaded, native, and passage-anchored. A scholar should be able to select a span of text in a transcription and attach a comment directly to it — the way a reader marks up a physical book, or the way modern annotation tools like Hypothesis work. Comments should be searchable, versioned, and linked to the subject mentions and annotations they discuss. This is not a difficult feature to build. It simply requires treating discussion as a first-class data type rather than an afterthought bolted onto a third-party forum.

The Social Failure: Why the Swarm Didn't Arrive

The technical problems above are real, but they are only half the story. The other half is organisational.

"Swarm Scholarship" — Zebrapedia's term for its distributed volunteer model — assumes a swarm. The model works when hundreds or thousands of contributors each do a small amount of work. Wikipedia has over 40 million registered editors. Wikisource, which uses a similar model for transcribing public domain texts, has indexed and transcribed hundreds of millions of words with thousands of active contributors.

Zebrapedia had a different experience. Contributing required registering a FromThePage account, agreeing to a terms document hosted on Google Docs, and then emailing the project team directly to be set up as a collaborator. This invitation-only model, designed to maintain quality control, created a friction barrier that filtered out everyone except the most committed — and the most committed are always few.

"To get started, register an account at zebrapedia.psu.edu, agreeing to the project's terms, then send an email to the project team to be set up as a collaborator." — Zebrapedia project description

The choice to require email invitation was understandable. Open contribution to an 8,000-page handwritten archive creates real quality control problems — anyone could add incorrect transcriptions or spurious subject tags. But the solution to that problem is not to restrict the pool of contributors. It is to build a review workflow that makes bad contributions correctible. Wikipedia's model — anyone can edit, vandalism is quickly reverted, quality emerges from volume — only works because reversion is cheap and review is distributed. Zebrapedia made contribution expensive and never made review cheap enough to compensate.

The Facebook group's existence is the clearest symptom. When a community relocates its conversation to a platform the project does not control and cannot index, it is not supplementing the project — it is working around it. The community that cares most about the Exegesis found Zebrapedia insufficient and built a parallel social infrastructure to compensate. That parallel infrastructure generates no transcriptions and no subject tags. The energy is real. The system fails to capture it.

In Defence of the Difficulty

One objection to all of the above deserves a direct answer: is the Exegesis simply too difficult to index? Dick's terminology is idiosyncratic, unstable, and self-referential. He uses "2-3-74" interchangeably with "2-74," "3/74," "February-March 1974," and simply "the experience." He calls the same figure Thomas, the Apostle Thomas, the time-traveller, the Paraclete, and the Holy Spirit at different moments. His categories dissolve and reform across eight years. A conventional controlled vocabulary — the kind used to index, say, a collection of letters — struggles to contain a mind this associative.

This is a genuine observation, not an excuse. The Exegesis is harder to index than most manuscripts. But that difficulty is precisely why the manual-markup approach was always the wrong tool for this job. A human transcriber, confronted with Dick's shifting terminology, faces an intractable normalisation problem: which of the twenty ways Dick writes about his February 1974 experience should be tagged, and how? The result is inconsistent coverage: some phrasings tagged, others not, no systematic way to know which.

A large language model trained on the conventions of philosophical and theological discourse handles this instability better. It can recognise that "the Light" and "the pink beam" and "the Vast Active Living Intelligence System" are Dick's names for related — possibly identical — phenomena, without requiring any of them to have been pre-tagged. The complexity of the source text is an argument for AI assistance, not against the indexing project as a whole.

What Done Looks Like

A decade of partial progress raises an honest question: is completing this project actually achievable? Is "fully transcribed and indexed Exegesis" a realistic goal or a scholarly aspiration that will always remain just out of reach?

The answer depends on the method. Under the manual-markup model, with the current pool of invite-only contributors, the honest answer is probably no — not at any timescale that matters for living scholars. The project as structured does not have a path to completion.

Under an AI-assisted model, the answer is different. The bottleneck is transcription (converting handwritten images to text) and entity extraction (identifying and linking named concepts). Both are now automatable:

"Done" is not a pipe dream. It is a pipeline. The 8,000 pages are a fixed quantity. The compute required to process them is available and affordable. What is needed is not more volunteers willing to type brackets. What is needed is the decision to build the pipeline.

What Other Projects Show Is Possible

Project Corpus Method Lesson for Zebrapedia
Transkribus Millions of handwritten historical pages across hundreds of projects AI handwritten text recognition (HTR) trained on project-specific samples Automated transcription of handwritten text is production-ready. Dick's handwriting is easier than medieval Latin scripts that Transkribus handles routinely.
Wikisource Hundreds of millions of words of public-domain text; thousands of complete works Open contribution (no invite required), layered review, quality emerges from volume The crowdsourcing model can work — but only at Wikipedia scale. Zebrapedia's restricted-access model prevented the volume of contribution that makes distributed quality control viable.
Perseus Digital Library Greek and Latin classical texts, fully annotated Structured systematic annotation with a controlled vocabulary; word-level morphological tagging Systematic annotation at corpus scale is achievable when the annotation schema is clear and the tooling supports it. Perseus shows what a complete, navigable ancient corpus looks like.
Hypothesis Any web document — used across millions of scholarly pages Passage-anchored annotation, open and collaborative, integrated directly into the reading interface Discussion and annotation belong in the same interface as the text. The Google Groups model Zebrapedia uses is a decade behind current annotation practice.

None of these projects proves that a complete, indexed Exegesis is trivial. They prove that each of the individual components — transcription, annotation, discussion, indexing — has been solved at scale by projects with comparable constraints. Zebrapedia's failure is not a failure of ambition. It is a failure to adopt the methods that the field had already developed.

What You Can Do

Zebrapedia, for all its problems, is the only public initiative transcribing the complete Exegesis. The subject index it has built — imperfect, partial, stub-filled — is still the most comprehensive public map of Dick's terminology that exists. That matters. The architecture is sound. The content can be filled in.

If you are a PKD scholar, a digital humanities researcher, a programmer, or simply a reader who cares about this text:

Philip K. Dick spent eight years trying to understand what happened to him in February 1974. His notes on that effort fill 8,000 pages. They deserve better than a decade of partial transcription and a subject index full of blank descriptions. The argument of this page is not that Zebrapedia failed — it is that it has not yet succeeded, and that the path to success is now clear.