# Zebrapedia Site Audit
## Humanities Infrastructure Design Perspective
*Audit date: 2026-03-08 | Analyst: Claude Sonnet 4.6*

---

## Executive Summary

Zebrapedia is a Penn State–hosted collaborative transcription and annotation project for Philip K. Dick's *Exegesis*, running on the FromThePage platform. It implements a folder-based workset hierarchy, a wiki-markup subject linking system, a hierarchical category taxonomy, a co-occurrence subject graph, and a multi-state page workflow. Its core scholarly insight — that an 8,000-page handwritten manuscript requires distributed swarm scholarship with automatic subject indexing — maps cleanly onto a local-first AI-assisted pipeline that replaces crowdsourcing with deterministic preprocessing plus LLM tagging. The design patterns most worth borrowing are: subject-to-page reference backlinks, the category drill-down taxonomy, the subject co-occurrence graph, and the workset folder metaphor. The most critical weakness is that FromThePage's subject system requires manual double-bracket markup, which an LLM pipeline can automate far more aggressively and accurately.

---

## What Zebrapedia Is

Zebrapedia is formally titled "Zebrapedia: Distributed Scholarship and Philip K. Dick's Exegesis." It describes itself as "a site devoted to the ongoing transcription, annotation, discussion and editing of Philip K. Dick's Exegesis."

**The source corpus:** Philip K. Dick's *Exegesis* is an 8,000+ page handwritten manuscript (daily entries, diagrams, sketches) spanning approximately 1974–1982. It documents Dick's "eight-year attempt to fathom" a visionary experience he called "2-3-74" (February-March 1974). Only a fraction has been published in print. The complete manuscript is available through the project.

**Institutional context:** Developed by researchers at Pennsylvania State University. Hosted on fromthepage.com under the username `zebrapedia`. Originally at `zebrapedia.psu.edu` (now ECONNREFUSED). Currently operating under two project slugs: `exegesis` (legacy) and `exegesis-ii` (current active project).

**Methodology:** "Swarm Scholarship" — crowdsourced transcription, annotation, and analysis using FromThePage's collaborative wiki-style platform plus community discussion via embedded Google Groups and a Facebook group (Radio Free Valis Satellite Station).

**Platform:** FromThePage (Ruby on Rails, open source, AGPL). Technically a SaaS platform; the project uses it as hosted infrastructure.

---

## Publicly Visible Information Architecture

### Page Type Inventory

| Page Type | Example URL | Contents |
|-----------|------------|----------|
| User/Org profile | `fromthepage.com/zebrapedia` | Project landing, description, featured links, register CTA |
| Project overview | `fromthepage.com/zebrapedia/exegesis-ii` | Workset list (behind auth?), project stats, tabs |
| Workset/Folder | `fromthepage.com/zebrapedia/exegesis-ii/pkd-folder-09` | Document list with thumbnails and status |
| Subjects index | `fromthepage.com/zebrapedia/exegesis-ii/subjects` | Category hierarchy + subject listing |
| Subject/Article | `fromthepage.com/zebrapedia/exegesis/article/32000002` | Subject details, graph, page references, discussion |
| Manuscript display | `fromthepage.com/zebrapedia/exegesis-ii/pkd-folder-01/display/32001919` | Image + transcription split view |
| Statistics | `fromthepage.com/zebrapedia/exegesis-ii/statistics` | Completion metrics |
| Works list | `fromthepage.com/zebrapedia/exegesis-ii/works_list` | Paginated document listing |
| Read all works | `fromthepage.com/display/read_all_works?article_id=32000002` | All transcriptions mentioning a subject |

### URL Structural Patterns

```
/{owner_slug}/{project_slug}/                        — project root
/{owner_slug}/{project_slug}/{workset_slug}          — folder/workset
/{owner_slug}/{project_slug}/{workset_slug}/display/{page_id}  — manuscript page
/{owner_slug}/{project_slug}/subjects                — subject taxonomy
/{owner_slug}/{project_slug}/article/{subject_id}   — subject detail (note: uses exegesis not exegesis-ii)
/{owner_slug}/{project_slug}/statistics              — project metrics
/display/read_all_works?article_id={id}             — cross-workset subject reading view
/iiif/collections                                    — IIIF collection endpoint
/iiif/for/{originating-uri}                         — IIIF content redirect
```

### Navigation Hierarchy

```
Zebrapedia (org profile)
└── Exegesis II (project)
    ├── Overview tab
    │   ├── [Workset grid - likely needs auth for full list]
    │   └── Project description
    ├── Statistics tab
    │   └── Completion metrics (pctComplete, pctTranscribed, pctNeedsReview, etc.)
    ├── Subjects tab
    │   ├── Category tree (14 top-level categories, 3 levels deep)
    │   └── Subject listing per category
    └── Works List tab
        └── Paginated document/work listing

    Worksets (Folders):
    pkd-folder-01, pkd-folder-02, pkd-folder-05, pkd-folder-06,
    pkd-folder-08, pkd-folder-09, pkd-folder-21, pkd-folder-22,
    pkd-folder-24, pkd-folder-27, pkd-folder-28, pkd-folder-30,
    pkd-folder-32, pkd-folder-36, pkd-folder-43, pkd-folder-44,
    pkd-folder-47, pkd-folder-48, pkd-folder-49, pkd-folder-50,
    pkd-folder-53, pkd-folder-55, pkd-folder-62, pkd-folder-63,
    pkd-folder-82, pkd-folder-90,
    "Cosmogony and Cosmology (Part 2)", "Old Folder 50"

    Each Workset:
    └── Works (individual manuscript documents)
        └── Pages (individual scanned images with transcriptions)
```

### Subject Taxonomy (Exegesis II)

```
Characters
  └── VALIS
      └── Fat (Horselover Fat)
Christian References
Condition of Health
  ├── Mental
  └── Physical
Family
Mobius
Notable Dates           [e.g., 2-3-74 — 127 page references]
People
Philosophers
Philosophical Concepts
Places
Politics
Publication             [e.g., The Exegesis — 17 page references]
Unknown
  ├── Can't read
  └── Spelling
Uncategorized
```

---

## Publicly Visible Features

### Confirmed Public (No Login Required)
1. **Landing page** — project description, mission, contact, featured links
2. **Subject detail pages** — view subject description, categories, page references, related subjects graph
3. **Category navigation** — drill from top-level category to specific subject
4. **Subject graph visualization** — co-occurrence network centered on any subject
5. **Cross-workset mention view** — "N pages refer to [Subject]" with links to each mention
6. **Subjects index** — full taxonomy with dynamic loading
7. **Subject description** — article text describing each subject entity

### Likely Public (Inferred from URL Patterns)
1. **Manuscript page display** — split image+transcription view (URL pattern visible)
2. **Read all works view** — aggregated transcription text for a subject
3. **Project statistics** — completion percentages
4. **Works list** — paginated document catalog

### Requires Login (Platform Convention)
1. **Transcription editing** — click "Transcribe" to edit a page's text
2. **Subject markup** — using `[[canonical|verbatim]]` syntax
3. **Autolink button** — AI-assisted subject suggestion
4. **Mark as Done/Needs Review** — page status workflow
5. **Discussion participation** — posting to embedded Google Groups
6. **Account registration** — requires invite/email to [email protected]
7. **Export/API access** — requires API key even for public collections
8. **Volunteer hour tracking** — contributor metrics

### Speculative (Not Confirmed)
1. Admin subject merge/duplicate resolution interface
2. Collaborator-specific dashboards
3. Work ordering/sequencing controls
4. Notification system for edits/reviews

---

## Inferred Entity / Data Model

### Core Entities

#### 1. Collection / Project
- Represents the whole Exegesis transcription effort
- Fields: `slug`, `name`, `description`, `owner_id`, `settings_json`
- Platform-native entity
- Relationships: has many Worksets, has many Subjects, has many Categories

#### 2. Workset / Folder
- Physical folder of the original manuscript archive (Folder 01–90+)
- Also includes thematic groupings ("Cosmogony and Cosmology Part 2")
- Fields: `slug`, `name`, `description`, `collection_id`, `display_order`, `page_count`, `pct_complete`, `pct_transcribed`, `pct_needs_review`
- Platform-native entity
- Relationships: belongs to Collection, has many Works

#### 3. Work / Document
- An individual manuscript item within a folder (e.g., "folder 53 - 023")
- Fields: `title`, `workset_id`, `page_count`, `metadata_json`, `dc_source` (IIIF origin)
- Platform-native entity
- Relationships: belongs to Workset, has many ManuscriptPages

#### 4. Manuscript Page
- A single scanned image within a Work
- Fields: `work_id`, `position`, `image_url`, `width`, `height`, `status` (enum: unedited/transcribed/needs_review/reviewed/blank), `has_subject_tags`, `has_translation`
- Platform-native entity
- Relationships: belongs to Work, has one Transcription, has many SubjectMentions, has many Comments

#### 5. Transcription
- The text content of a manuscript page in multiple formats
- Fields: `page_id`, `content_verbatim` (wiki markup with `[[]]` links), `content_html`, `content_tei`, `content_plaintext_verbatim`, `content_plaintext_emended`, `version`, `created_at`, `updated_at`
- Platform-native entity
- Current transcription = HEAD of revision chain

#### 6. Revision
- A historical version of a transcription
- Fields: `transcription_id`, `page_id`, `user_id`, `content`, `created_at`, `edit_summary`
- Platform-native entity (version history is core to wiki model)

#### 7. Subject
- A named entity/concept identified within the Exegesis
- Fields: `id` (numeric, e.g., 32000002), `name` (canonical), `description` (wiki article text), `category_id`, `uri` (external link, e.g., to FamilySearch, Wikidata), `lat`, `lng` (GIS-enabled categories), `mention_count`
- Hybrid (platform provides infrastructure; project provides content)
- Examples: "The Exegesis" (17 mentions), "2-3-74" (127 mentions), "Horselover Fat" (3 mentions)

#### 8. Subject Alias
- Verbatim text form linked to a canonical Subject
- Fields: `subject_id`, `verbatim_text`, `occurrence_count`
- Created when transcribers use `[[Canonical|Verbatim]]` markup
- Platform-native entity

#### 9. Category
- Hierarchical taxonomy node for Subject classification
- Fields: `id`, `name`, `parent_id`, `project_id`, `depth`, `gis_enabled`
- Project-specific content, platform-provided structure
- Max depth observed: 3 levels (Characters > VALIS > Fat)

#### 10. Subject Mention
- Occurrence of a Subject on a specific Manuscript Page
- Fields: `subject_id`, `page_id`, `work_id`, `workset_id`, `verbatim_text`, `position_in_text`
- Created by transcription markup parsing
- Drives the "N pages refer to [Subject]" display

#### 11. Related Subject Edge
- Co-occurrence relationship between two subjects (appear on the same pages)
- Fields: `subject_a_id`, `subject_b_id`, `co_occurrence_count`, `strength`
- Drives the subject graph visualization
- Computed/derived from SubjectMention data

#### 12. Comment / Note
- Scholarly discussion attached to a page or subject
- Fields: `target_type` (page/subject), `target_id`, `user_id`, `content`, `created_at`
- For pages: appears in Google Groups embed
- For subjects: also in Google Groups embed

#### 13. User / Collaborator
- Fields: `id`, `slug`, `name`, `email`, `role` (admin/staff/collaborator/volunteer), `volunteer_hours`, `joined_at`
- Platform-native; Zebrapedia requires email invitation to become collaborator

#### 14. Source Link
- External URI associated with a Subject (e.g., Wikidata, library catalog)
- Fields: `subject_id`, `uri`, `label`, `uri_type`
- Platform-native URI field on Subject

---

## Research Workflows Supported

### Public Browsing Workflows
1. **Subject-first discovery**: Landing page → Subject article page → "N pages refer to X" → Individual manuscript page
2. **Category drill-down**: Subjects tab → Category (e.g., Characters) → Subcategory (VALIS) → Specific subject (Fat)
3. **Cross-workset subject reading**: Subject page → "Master view" → all transcribed mentions across all folders in one view
4. **Subject graph exploration**: Subject page → Graph visualization → click adjacent node → navigate to related subject

### Likely Logged-In Workflows
5. **Workset-first transcription**: Project overview → Pick folder → Pick document → Open page → Transcribe image
6. **Subject linking during transcription**: Transcribe → type `[[subject]]` or use Autolink → subjects auto-indexed
7. **Review workflow**: Transcribed page → Reviewer reads and edits → marks Done → appears in % reviewed metrics
8. **Subject curation**: Subjects tab → Find duplicate subjects → Merge → Clean taxonomy

---

## UX / IA Strengths

1. **Subject-to-page backlinks**: Every subject page shows every manuscript page that mentions it — this is genuinely powerful for corpus navigation and is the model's killer feature.
2. **Three-level category hierarchy**: Enough depth to distinguish VALIS-related characters from general people without becoming unwieldy.
3. **Co-occurrence subject graph**: Visually communicates conceptual neighborhoods; effective for exploratory browsing.
4. **Category drill-down**: Clean navigation from broad to specific; mirrors how a scholarly index works.
5. **Cross-workset read-all view**: Aggregates scattered mentions into a single reading experience — essential for thematic study of a distributed manuscript.
6. **Multiple export formats**: TEI-XML for scholarly editions, CSV for data analysis, HTML for web display, IIIF for interoperability.
7. **Wiki markup accessibility**: `[[subject]]` is easy for non-technical collaborators to learn and use.
8. **Autolink feature**: Mines previous transcriptions to suggest markup on new pages — reduces repetitive effort.
9. **Page status workflow**: Clear states (unedited → transcribed → needs review → done) support quality control in a distributed team.

---

## UX / IA Weaknesses

1. **Public-facing workset browsing appears auth-gated**: Deeper project pages seem to redirect to the landing page for unauthenticated users, making public browsability unclear.
2. **Subject descriptions are thin**: The "The Exegesis" article shows joke placeholder text (TransFoilent ad copy), suggesting many subject articles are stubs or have been left with test content.
3. **Manual markup required**: Subject linking requires human transcribers to manually type `[[double brackets]]`; Autolink only suggests, doesn't apply. In an 8,000-page corpus this is a massive bottleneck.
4. **Subject graph lacks filtering**: The co-occurrence graph shows everything at once with no way to filter by folder, time period, or concept type.
5. **Discovery is subject-centric but not concept-search**: You can navigate to known subjects but there's no semantic/fuzzy search across the corpus.
6. **Category assignment is manual and inconsistent**: "Unknown / Can't read / Spelling" as subject categories reveals the friction of maintaining a clean taxonomy under crowdsourced conditions.
7. **No timeline/chronological view**: The Exegesis spans 8 years; there's no way to navigate by date or see temporal clustering of subjects.
8. **Platform conventions override project needs**: Google Groups for discussion is clunky; Zebrapedia appears to be constrained by what FromThePage offers rather than what scholars need.
9. **No full-text search visible**: Despite IIIF Voyant integration existing in FTP, it wasn't clearly surfaced.
10. **No entity relationship modeling**: Two subjects can be linked in the graph by co-occurrence, but there's no typed relationship system (e.g., "subject A IS A character IN publication B").

---

## Platform Inference: FromThePage vs. Project-Specific Layers

### FromThePage-Native (Platform Infrastructure)
- User management, roles, permissions
- Page image hosting and IIIF serving
- Transcription editor with wiki markup
- Version history / revision tracking
- Subject linking with `[[]]` markup
- Autolink engine
- Category system (hierarchical)
- Subject co-occurrence graph
- Page status workflow (unedited / transcribed / needs_review / done / blank)
- Progress metrics (pctComplete, pctTranscribed, etc.)
- Export engine (TEI, CSV, HTML, plaintext, IIIF)
- API key access
- Discussion forum embedding (Google Groups integration)

### Zebrapedia Project-Specific (Content Layer)
- Subject taxonomy content (the 14 categories and their children)
- Subject article text (descriptions of each entity)
- Transcription content (all human-produced text)
- Workset/folder naming scheme (pkd-folder-XX mirrors physical archive organization)
- Invitation-only collaboration model
- Community infrastructure (Facebook group, Google Groups)
- Subject IDs in the 32000000+ range (project-allocated IDs)

### Unclear / Speculative
- Whether Zebrapedia has any custom code on top of FTP (unclear; may be vanilla)
- Whether `exegesis-ii` represents a platform migration or a new project phase
- Extent of TEI-XML export use downstream

---

## What to Borrow / Modernize / Avoid

### Borrow Directly
- Subject-to-page reference backlink model (canonical entity → list of source pages)
- Three-level category hierarchy with drill-down navigation
- Co-occurrence subject graph (but add filtering and weighting)
- Cross-workset "read all pages mentioning subject X" aggregated view
- Page status state machine (unedited → processing → needs_review → reviewed)
- Workset/folder as primary organizational unit mirroring physical archive structure
- Multiple export formats from the same underlying data

### Modernize
- Replace manual `[[]]` markup with LLM-assisted entity extraction pipeline
- Replace Autolink (pattern matching) with vector-similarity entity suggestions
- Replace static co-occurrence graph with interactive, filterable, time-aware graph
- Replace manual category assignment with LLM-assisted classification + human confirm
- Replace Google Groups discussion with embedded, threaded, per-passage commentary
- Replace search gap with full-text + semantic vector search across all transcriptions
- Add chronological/temporal navigation layer (the Exegesis has dates)
- Replace thin subject stubs with LLM-generated entity summaries seeded from corpus

### Avoid
- Invitation-only access model (for local-first, you own all the data)
- Platform lock-in to SAAS constraints (build locally with portable formats)
- Over-reliance on human crowdsourcing for markup (bottleneck at scale)
- Static graph visualization without interactivity (D3.js/vis.js instead)
- Category assignments in an "Unknown/Can't read" bucket (pre-clean with OCR/LLM)

---

## Open Questions

1. How many total manuscript pages are in the Exegesis II project? (8,000+ mentioned but exact FTP count unclear)
2. What percentage is currently transcribed? (Statistics page returned landing page)
3. Are worksets publicly browsable without auth, or does public access end at the project landing page?
4. Does `exegesis` vs `exegesis-ii` represent a migration (losing old data) or running in parallel?
5. Is the PSU-hosted site (zebrapedia.psu.edu) defunct, replaced entirely by fromthepage.com/zebrapedia?
6. What is the relationship between the subject IDs in the 32000000 range and the internal FTP database IDs?
7. Are there any custom API endpoints or scrapers running against Zebrapedia to produce downstream research outputs?
8. Has the project produced any TEI-XML exports used in downstream scholarship?
9. What do the "Cosmogony and Cosmology (Part 2)" and thematic worksets contain vs. the folder-numbered worksets?
10. How complete is the subject taxonomy — are the 14 categories all there are, or are there more?

---

## Page Visit Log

| # | URL | Page Type | Contents | Why It Matters |
|---|-----|-----------|----------|----------------|
| 1 | `fromthepage.com/zebrapedia` | Org/User Profile | Project description, featured links, CTA, contact | Entry point; reveals featured subjects and workset |
| 2 | `fromthepage.com/zebrapedia/exegesis` | Project Overview (legacy) | Same as #1 (redirect/landing) | Confirms legacy vs. active project slugs |
| 3 | `fromthepage.com/zebrapedia/exegesis-ii` | Project Overview (active) | Same landing (likely auth-gated) | Active project slug; suggests deeper pages need auth |
| 4 | `fromthepage.com/zebrapedia/exegesis/pkd-folder-09` | Workset Page | Returned landing page | Confirms URL pattern; may need auth for document list |
| 5 | `fromthepage.com/zebrapedia/exegesis/article/32000002` | Subject Article | "The Exegesis" — 17 page refs, graph, categories, discussion embed | Core subject entity model; graph visualization confirmed |
| 6 | `fromthepage.com/zebrapedia/exegesis/article/32001273` | Subject Article | "Horselover Fat" — 3 page refs, Characters>VALIS>Fat taxonomy, graph | Category drill-down confirmed; VALIS sub-taxonomy visible |
| 7 | `fromthepage.com/zebrapedia/exegesis/article/32001295` | Subject Article | "2-3-74" — 127 page refs across 25+ folders, Notable Dates category | Highest-frequency subject; cross-workset distribution confirmed |
| 8 | `fromthepage.com/zebrapedia/exegesis-ii/subjects` | Subjects Index | 14 top-level categories, 3-level hierarchy, dynamic loading | Full taxonomy visible; category structure confirmed |
| 9 | `fromthepage.com/zebrapedia/exegesis-ii/statistics` | Statistics | Returned landing page | Likely auth-gated or different URL |
| 10 | `fromthepage.com/zebrapedia/exegesis-ii/works_list` | Works List | 404 | Correct URL is likely works_list or similar |
| 11 | `fromthepage.com/display/read_all_works?article_id=32000002` | Cross-workset View | Returned landing page | Likely needs auth; URL pattern noted for future use |
| 12 | `zebrapedia.psu.edu/` | Original PSU Site | ECONNREFUSED | Original platform defunct; FTP is canonical home |
| 13 | `content.fromthepage.com/project-owner-documentation/subject-linking/` | FTP Documentation | Complete subject linking docs with markup syntax | Confirmed `[[canonical|verbatim]]` syntax; graph mechanics |
| 14 | `content.fromthepage.com/project-owner-documentation/` | FTP Documentation Index | Full documentation TOC | Complete feature surface of the platform |
| 15 | `github.com/benwbrum/fromthepage` | GitHub Repo | Ruby on Rails, tech stack, 10k+ commits | Platform architecture; open source reference |
| 16 | `github.com/benwbrum/fromthepage/wiki/...IIIF...` | IIIF Documentation | Full IIIF data model, page status flags, export formats | Definitive entity model via IIIF manifest structure |
| 17 | `dh-abstracts.library.virginia.edu/works/2414` | DH Conference Abstract | Technical plan: TEI, OCR-to-TEI, CSV for network analysis | Future roadmap; confirms scholarly intent |
| 18 | `dickiangnosticism.wordpress.com/2018/02/02/...` | Blog Post | Project overview, features described, "Swarm Scholarship" | Secondary source confirming feature set from user perspective |
