A structured analysis of Distributed Scholarship and Philip K. Dick's Exegesis
Zebrapedia is a Penn State–hosted collaborative project for transcribing, annotating, and indexing Philip K. Dick's Exegesis — an 8,000-page handwritten manuscript spanning 1974–1982 that documents Dick's eight-year attempt to understand a visionary experience he called "2-3-74." The project runs on the FromThePage platform and employs what it calls "Swarm Scholarship": distributed volunteer labour guided by a wiki-markup annotation system and hierarchical subject index.
This page explains every major feature of Zebrapedia, how each one works, who can access it, and why it matters for large-scale humanistic manuscript research.
Philip K. Dick's Exegesis is arguably the largest archive of unreleased and unpublished material by any major twentieth-century author. At over 8,000 handwritten pages — daily entries, diagrams, letters, and sketches — it spans from 1974 to 1982 and documents Dick's obsessive philosophical and theological attempt to account for an experience he called "2-3-74": a period in February and March 1974 in which he believed he made contact with an intelligent being he variously identified as VALIS (Vast Active Living Intelligence System), a Gnostic divinity, or a satellite transmitting information directly into human consciousness.
A print edition was published in 2011 (edited by Pamela Jackson and Jonathan Lethem), but it covers only a fraction of the full manuscript. Zebrapedia provides the first public access to the complete manuscript. The project's goal is to transcribe, annotate, and index the whole archive using distributed volunteer labour.
Researchers at Pennsylvania State University developed the portal. The project is formally titled "Zebrapedia: Distributed Scholarship and Philip K. Dick's Exegesis" and currently lives at fromthepage.com/zebrapedia. The original PSU-hosted site (zebrapedia.psu.edu) is no longer accessible.
Zebrapedia's URL structure follows FromThePage's conventions. Every page type maps to a predictable URL pattern that encodes the entity hierarchy:
The project has two active slugs: exegesis (legacy) and exegesis-ii (current). Subject article pages still resolve under exegesis rather than exegesis-ii — a migration artifact. The hierarchy is:
A workset (FromThePage's term) corresponds directly to a physical folder in the PKD archive. Dick's manuscript was physically organised into numbered folders; Zebrapedia mirrors this structure digitally. Each folder gets its own page listing all manuscript documents within it.
Confirmed worksets include 28 or more folders:
The workset overview page shows each document within the folder with thumbnails of the manuscript images, progress indicators (percentage transcribed, percentage reviewed), and status badges. Documents are named by folder and sequence number — e.g., "folder 53 - 023".
Note: Full document listings within worksets appear to require authentication; anonymous users may see only the folder landing page.
All indexed subjects are organised into a three-level category hierarchy accessible at /subjects. This transforms an otherwise flat list of thousands of named entities into a navigable scholarly index. Zebrapedia uses 14 top-level categories:
From the top-level subjects page, a scholar can drill from Characters → VALIS → Fat to reach the specific entry for Horselover Fat, Dick's fictional alter ego, without knowing the exact name to search for. The category tree supports top-down discovery — the mode scholars use when they know the topic but not the terminology.
The "Unknown / Can't read / Spelling" categories function as error bins for illegible or uncertain transcription text — a pragmatic concession to crowdsourced conditions rather than a scholarly category in the strict sense.
Every indexed subject has its own dedicated page at:
Each subject article page contains:
Observed: the description field for "The Exegesis" subject contained placeholder/joke advertising copy ("TransFoilent metallic foil"), suggesting that subject articles in Zebrapedia are largely stubs. The indexing infrastructure is well-built; the knowledge-base content is incomplete.
Every subject article page includes a force-directed network diagram centred on the selected subject. The surrounding nodes are other subjects that appear on the same manuscript pages as the central subject. The closer a node is to the centre, the more frequently the two subjects co-occur.
FromThePage's documentation describes it as: "Everyone who is mentioned in the same context as the subject will surround the subject. The closer they are to the center, the more frequently they are mentioned in the same context."
This graph answers a question that cannot be answered by a list: "What does Dick think about at the same time as X?" It surfaces implicit thematic structure — conceptual constellations — that resists reduction to a flat index. Clicking a surrounding node navigates to that subject's own article page.
Weakness: the graph is static (an image, not interactive). It cannot be filtered by date range, folder, or category. There is no way to compare the neighbourhoods of two subjects side by side, or to export the graph data for analysis in Gephi or similar tools.
The single most important feature of Zebrapedia is the subject backlink index: the ability to ask "where does Dick mention X?" and receive a list of every manuscript page across the entire 8,000-page corpus that contains that subject.
Confirmed examples from live subject pages:
| Subject | Category | Pages Mentioning It | Folders Covered |
|---|---|---|---|
| 2-3-74 | Notable Dates | 127 | 26+ folders |
| The Exegesis | Publication | 17 | 5 folders |
| Horselover Fat | Characters > VALIS > Fat | 3 | 2 folders |
Each mention is listed with the folder name, document title, and page number, and is a clickable link to the manuscript display view. A "Master view" link aggregates all mentions into a single scrollable reading experience (the Read-All-Works view, described below).
This feature solves the fundamental problem of a distributed manuscript: you cannot read 8,000 pages sequentially looking for a concept. The backlink index inverts the reading direction — you navigate thematically rather than sequentially.
Accessible at /display/read_all_works?article_id={id}, this view assembles all transcription passages mentioning a subject into one continuous reading experience — effectively creating a synthetic document: "Everything PKD wrote about 2-3-74."
This is the primary close-reading workflow for a scholar studying a specific concept across the corpus.
Access appears to require authentication. The URL pattern is confirmed but content visibility for anonymous users is unclear.
[[Double-Bracket]] Annotation SyntaxThe mechanism that drives all subject indexing is a simple wiki-markup syntax used inside the transcription editor. Transcribers enclose a word or phrase in double square brackets to link it to a canonical subject:
The pipe-separated form is essential for the Exegesis because Dick uses idiosyncratic shorthand throughout — writing "2-74" or "3/74" when referring to the same experience he elsewhere calls "2-3-74." The verbatim form preserves Dick's exact language; the canonical form normalises it for indexing. Both are stored separately.
Every saved [[link]] creates a subject mention record — a row in the database joining the subject to the manuscript page. The mention count on a subject article page is the total of these records.
This syntax is the foundation of the entire index. Its weakness is that it requires human effort on every page: the same subject must be marked up independently each time it appears. An 8,000-page corpus with thousands of recurring concepts creates enormous markup labour.
The transcription editor includes an Autolink button that analyses the current page's text and suggests [[subject]] markup based on subjects linked in previously transcribed pages. If "Horselover Fat" was manually linked on page 1, pressing Autolink on page 50 will suggest the same link where the text "Fat" appears.
The system's own documentation notes: "the suggestions are not perfect" and require manual review before saving. It performs string and pattern matching only — it cannot infer conceptual references or variant phrasings that a human (or language model) would recognise.
Autolink is nonetheless the platform's most significant gesture toward automation: it makes the project's growing transcription corpus self-reinforcing, where past human effort incrementally reduces future labour.
Autolink is the feature where a modern LLM-assisted pipeline most dramatically outperforms FromThePage's approach. Named entity recognition on a modern language model, combined with embedding-based similarity matching against the existing subject list, can achieve corpus-wide tagging in hours rather than years of crowdsourced effort.
Individual manuscript pages are displayed in a split view: the scanned manuscript image on the left, the transcription text on the right. The URL pattern is:
This is the standard workflow interface for handwritten document transcription. The scholar keeps the image in view while reading or editing the text, enabling continuous comparison between Dick's handwriting and the typed transcription. Page navigation buttons (previous / next) allow sequential reading while maintaining the split-view context.
FromThePage's transcription editor uses the left panel for the image (with zoom controls) and the right panel for a wiki-markup text editor. Subject links are displayed as hyperlinks in the read view and as [[bracket syntax]] in the edit view.
Limitation: there is no synchronisation between image regions and text spans — clicking a word in the transcription does not highlight the corresponding region in the image. Annotation is at the page level, not at the text-span or image-region level.
Every manuscript page has a status that tracks its position in the editorial pipeline. FromThePage's five states are:
Additional boolean flags track specific content states:
hasTranscript — page has typed texthasSubjectTags — at least one [[subject]] link existsneedsReview — flagged by a transcriber or auto-detectionhasTranslation — a parallel translation existsmarkedBlank — no text to transcribeThese states aggregate into the project-level progress metrics: pctComplete, pctTranscribed, pctNeedsReview, pctIndexed, pctMarkedBlank. The workset overview visualises these as progress bars per folder.
The two-button save interface — "Save" (work in progress) vs. "Done" (mark as transcribed) — ensures volunteers explicitly signal completion rather than just closing the editor.
Every edit to a transcription creates a revision record — a full snapshot of the transcription content at that point in time. Subject article pages have a visible Versions tab that shows the history of edits to the subject description, including when [[subject links]] were added or changed.
This follows the wiki model directly: all edits are non-destructive, every change is attributable to a specific user, and any version can in principle be reverted. It is the primary safeguard against vandalism and accidental data loss in a crowdsourced system.
The GitHub repository for FromThePage shows over 10,306 commits on the development branch — the platform itself has a similarly rigorous version history at the code level.
Discussion is attached to both manuscript pages and subject article pages via an embedded Google Groups forum. Collaborators can post questions about difficult handwriting, disputed transcription choices, or uncertain subject identifications within the context of the relevant page or subject.
A broader community forum operates as a Facebook group called "Radio Free Valis Satellite Station" — separate from the transcription interface, used for general project discussion and community building.
Known weakness: the Google Groups embed creates a jarring UX break — discussion is not integrated with the transcription interface, cannot be attached to specific text spans, and is not searchable from within the platform. The migration of community activity to Facebook suggests the Google Groups approach is not well-adopted.
FromThePage provides transcription data in multiple derived formats, all generated from the same canonical wiki-markup source:
| Format | What It Contains | Typical Use |
|---|---|---|
| Verbatim Plaintext | Plain text with line/paragraph/page break conventions. No markup. | Full-text search, further processing |
| Emended Plaintext | Canonical subject names substituted for verbatim forms (e.g. "2-3-74" replacing "2/74") | Programmatic analysis, concordances |
| HTML | Rendered text with hyperlinked subject mentions | Web publication, human reading |
| TEI-XML (P5) | P5-compliant XML encoding unclear text, gaps, contributor metadata, and subject references as <rs> elements |
Scholarly editions, archival deposit |
| Subject CSV (index) | One row per subject mention: page location, work, verbatim text, canonical name | Network analysis (Gephi), quantitative study |
| Subject CSV (details) | One row per subject: name, description, category, coordinates, URI, mention frequency | Knowledge base export, linked data |
| IIIF Manifests | IIIF Presentation API manifests for each work and collection, including page-level status flags and rendering links | Library interoperability, Universal Viewer, Mirador |
| PDF / Word | Formatted document for individual works | Human reading, printing |
The DH conference abstract for the project lists planned future exports including TEI-flavoured encoding and CSV for network analysis — confirming these were intended research outputs from the start.
FromThePage provides a REST API for bulk export of transcription data. An API key is required even for publicly accessible collections — registration is necessary to obtain one.
Available API export endpoints include approximately 15 types:
The API supports per-page download, per-work download, and full-collection bulk download as a zip archive. It also supports date-range filtering via the iiif/contributions endpoint, which returns IIIF collections of pages contributed within a specified time window.
Categories can be GIS-enabled, allowing subjects within that category to carry latitude and longitude coordinates. For the Places category, this means subjects like "Fullerton, California" or "Santa Ana" can be pinned to a map and searched geographically.
For Dick's Exegesis specifically, geographic data is of limited scholarly utility — the manuscript is internally focused, and its "places" tend to be philosophical or cosmological rather than physical. GIS becomes far more valuable in corpora like travel writing, letters, or field notes.
The feature is activated per category via a toggle in the project settings: "Enable GIS for Category." Coordinates appear in the subject CSV export alongside name, description, and mention count.
The statistics tab provides completion metrics at the project level. The same metrics are embedded in IIIF manifests at both the work level and the canvas (page) level, making them machine-readable:
These metrics serve two audiences: volunteers, who need to know where to focus effort ("Folder 53 is only 40% transcribed"), and project administrators, who need to assess overall corpus coverage and allocate resources.
The statistics page appeared to redirect to the landing page during auditing, suggesting it requires authentication or uses a different URL pattern than documented.
Zebrapedia does not allow open public contributions. Participating as a transcriber or annotator requires:
This invitation model reflects a deliberate quality-control decision: unrestricted public contribution would introduce noise and vandalism into a scholarly archive. The trade-off is a smaller, slower contributor pool.
User roles within the platform include: owner, admin, staff, collaborator, and volunteer. Different roles have different levels of access to transcription, review, subject management, and export functions.
[[double bracket]] syntax is simple enough for non-technical scholars to use immediately.[[bracket]] system creates enormous labour overhead that Autolink only partly mitigates.