Skip to content

Data Model Plan

Draft 3 · 2026-06-30 · Reconciled with Lina's CMS Content Model + IA stand-up refinements

Design rationale — not a spec

The settled schema lives in CMS Architecture (canonical field-level reference) and Database Schema. This document retains the design rationale and reconciliation narrative — read it for the why, not the what.

This is a reading app for C.S. Lewis's body of work — passages, letters, essays, podcasts, and editorial content — where every piece is connected to every other through shared themes. The core promise is a "Wikipedia rabbit hole" feeling: start with one passage and the app surfaces surprising connections across the whole corpus. That mechanic, plus a curated alternative, drives every structural decision below.

Two navigation modes sit on top of the same content.

  • Explore is open-world graph traversal — you follow computed tag-overlap edges from item to item.
  • Journeys are editorially curated paths — a human picks specific pieces, orders them into chapters, and frames them with reflection prompts. On top of both, users accumulate a personal
  • Soul Map (saves, progress, completions) that persists across sessions — which is why accounts are in V1 despite there being no paywall.

Reconciliation note

This draft is reconciled with Lina's CMS Content Model. Where the two models overlapped, hers won on the CMS authoring shape (separate content types, Theme-as-world, standalone Prompt/Author/Portal Image, content blocks) because her field-level reasoning is concrete. This plan keeps the layers her pre-read doesn't cover — the runtime tag-overlap graph, user data / Soul Map, community reflections, daily drop — and folds them onto her content spine. One question stays genuinely open; it's flagged near the end.

Three layers, one content spine

The model splits cleanly into three layers. The Content Library and Editorial Structure come from Lina's CMS model — this is what editors author. The App layer (graph engine, user data, community) is what the backend computes and stores on top, and is not part of the CMS authoring surface. The bridge between them is a polymorphic Content Item — any of Passage / Letter / Essay / Podcast — which is what Journeys reference, what the graph traverses, and what users save.

Legend: 🔵 Content Library · 🩵 Worlds & taxonomy · 🟣 Editorial structure · 🩷 App layer (not in CMS) · 🔮 Content Item bridge

graph TB
  subgraph LIB["Content Library (CMS-authored)"]
    direction LR
    Passage["<b>Passage</b><br/><i>Lewis passage · audio</i>"]
    Letter["<b>Letter</b><br/><i>Lewis, held in full</i>"]
    Essay["<b>Essay</b><br/><i>contributor · Substack</i>"]
    Podcast["<b>Podcast</b><br/><i>external embed</i>"]
    Work["<b>Work</b><br/><i>book proxy + buy link</i>"]
    Author["<b>Author</b>"]
    Prompt["<b>Prompt</b>"]
  end

  subgraph WORLDS["Worlds & Taxonomy"]
    Theme["<b>Theme</b><br/><i>= navigable world / portal</i>"]
    PortalImg["<b>Portal Image</b>"]
    Tags["<b>Tags</b><br/><i>faceted — TBD</i>"]
  end

  subgraph EDIT["Editorial Structure"]
    Journey["<b>Journey</b>"]
    Chapter["<b>Chapter</b>"]
    Blocks["<b>Content Blocks</b><br/><i>Theme page + Home</i>"]
  end

  subgraph APP["App Layer (computed / user — not CMS)"]
    Graph["<b>Tag-Overlap Graph</b><br/><i>computed edges</i>"]
    User["<b>User</b>"]
    Save["<b>Save</b> · Soul Map"]
    Reflection["<b>Reflection</b><br/><i>community, moderated</i>"]
    DailyDrop["<b>DailyDrop</b>"]
  end

  CI(["<b>Content Item</b><br/>Passage | Letter | Essay | Podcast"])

  Passage --> CI
  Letter --> CI
  Essay --> CI
  Podcast --> CI

  Passage -->|source| Work
  Passage -->|source| Letter
  Passage -->|source| Essay
  Essay --> Author
  Passage -.-> Prompt
  Letter -.-> Prompt
  Essay -.-> Prompt
  Podcast -.-> Prompt

  CI -->|tagged| Tags
  CI -->|tagged| Theme
  Theme --> PortalImg
  Theme --> Blocks
  Journey -->|belongs to| Theme
  Journey --> Chapter
  Chapter -->|references| CI

  CI --> Graph
  Save --> CI
  User --> Save
  Reflection --> CI
  DailyDrop --> CI

Content Library

Lina's model keeps the content types separate rather than collapsing them into one unified record, and the reasoning is field-level and concrete. A Passage body is plain text because word-by-word audio sync needs a clean string — rich text would break streaming. A Letter carries a Recipient and an Original Date and is read, not listened to. An Essay is not Lewis — it has a named contributor, a Substack source, and an author-driven card. These differences don't survive a single unified table, so they stay distinct. (My earlier draft unified them; see the open decision at the end for how the graph layer still gets uniform traversal across separate types.)

Everything in the library is created once and referenced many times. Each piece optionally carries a Prompt and is classified by Tags and Themes. The source chain is the key relational bit: a Passage is always an excerpt of something larger, so it links to a source — either a Work (a book we don't hold in full, a proxy that routes to purchase) or directly to a Letter/Essay we do hold in full (the piece is the destination). Never both.

Field-level detail is deliberately omitted below — the settled columns live in CMS Architecture and Database Schema. The tables here name the types and the why, not the what.

Type What it is
Passage Atomic unit. Always Lewis's words, always an excerpt of a larger source. Primary interaction is cloned-voice audio (body is plain text so audio can sync word-by-word).
Letter Lewis's words, held in full. Can be the source of a Passage.
Essay Editorial, by a named contemporary contributor — not Lewis. Sourced from Substack.
Podcast Longer-form external listening — embedded, not CMS-hosted. Distinct from Passage audio. Phase 1: schema stub only, not built.
Work Proxy for books we do not hold in full. Exists to give context and drive hardcopy/ebook sales.
Author Standalone type for essay contributors (Malcolm Guite etc.). Photo drives the essay card.
Prompt A reflection question. Standalone (not a field) so it's reusable and cleanly overrideable per chapter.

Open from Lina's doc + IA stand-up 2026-06-30

The Letter and Essay corpora are not yet in hand — both schemas are provisional and will likely shift after ingest. Raw content was collected on 2026-06-30; the team confirmed it won't map cleanly to Lina's schema — plan is to manually shape a few sample pieces to validate the model first. The Passage audio model is undecided: stream-by-ID from the voice provider (no field needed) vs a stored pre-generated file URL. Work is a pure buy-link proxy with no chapters — Books are source reference only in V1 (cover, synopsis, publication date, future buy link; no e-reader). Letters in-app needs client confirmation (assumed yes, not validated). Podcasts confirmed as a later phase (Phase 1 is a schema stub only).

Worlds & Taxonomy

This is the biggest terminology correction from the reconciliation. The navigable "world" — the portal — is the Theme, not the Journey. Editors create Theme worlds (Grief, Hope, Courage, Love), each with its own landing page, bespoke colour palette, and portal imagery. Themes are a global vocabulary — the same worlds span every content type (Passage, Letter, Essay, Journey), not per-book. Content pieces and Journeys are assigned to Themes via multi-select, and a single piece or Journey can live in several Theme worlds at once. (My earlier draft put "portal" on the Journey; that was wrong — Journey is a path inside the worlds.)

Type What it is
Theme A navigable world that content and Journeys live within. Has its own landing-page template, distinct from Home. Global vocabulary (9 bespoke in V1); may grow pre-launch.
Portal Image CMS-managed reusable image library. Editors pick one per Theme. Grows over time; repetition acceptable.
Tags (deferred) Granular, multi-dimensional taxonomy, distinct from Themes. Likely faceted, not a flat list. Powers fine-grained discovery + the graph.

Tags need their own session

Lina deliberately defers the tag taxonomy — it's faceted and multi-dimensional (subject matter, work types, topical categories). My earlier draft committed to specific axes (motif, life-stage, register, tone-depth); those are candidate facets to bring as input, not locked decisions. The graph engine below depends on whatever facets land here.

Thought provoker = Prompt? (terminology, IA stand-up 2026-06-30)

The stand-up introduced "thought provoker" as the canonical name for a pre-reading prompt/lens shared by every content piece. The data-model plan uses "Prompt" as a standalone type covering reflection questions. These may be the same object (Prompt with a type enum that includes thought-provoker) or two distinct concepts. Confirm terminology — see the glossary for current definitions.

Deeper meaning deprioritized (IA stand-up 2026-06-30)

Client has no editorial team to QA longer AI-assisted interpretive annotations. The shorter shared thought provoker prompt stays; the deeper-meaning block is likely deprioritized for V1.

Editorial Structure

Journeys are the curated counterpart to Explore. A Journey holds no content of its own — it's an ordered sequence of Chapters, and each Chapter references existing Library pieces. A Chapter has one Primary Reading (the only required content field) anchored by optional opening and supporting pieces. Because the Chapter is just a positioned reference, the same piece can appear in many Journeys with different framing — and the Chapter's Prompt Override lets it carry a different reflection question each time.

Type What it is
Journey A curated, sequenced reading experience. Belongs to one or more Themes. Holds no content of its own — an ordered sequence of Chapters.
Chapter One slice of a Journey. Groups pieces around a moment in the arc. Holds no body — references Library content; has one required Primary Reading.
Content Blocks Editor-managed ordered blocks driving Theme landing pages and the Homepage. Not hardcoded layouts.

JourneyChapterContent Item

graph LR
  J["<b>Journey</b><br/>opening + final question<br/>belongs to Theme(s)"]
  C1["<b>Chapter 1</b><br/>opening question<br/>prompt override"]
  C2["<b>Chapter N</b>"]
  OP["opening passage"]
  PR["<b>primary reading</b> (req)"]
  SC["supporting content ×N"]

  J --> C1
  J --> C2
  C1 --> OP
  C1 --> PR
  C1 --> SC

Tag-Overlap Graph (app layer — not in CMS model)

This layer is not in Lina's CMS doc, because it's runtime, not authoring — but it's the load-bearing mechanic for Explore and every "related content" surface. The graph's nodes are Content Items (any Passage / Letter / Essay / Podcast); its edges are computed, not stored, from the shared Tags + Themes the CMS attaches. This is exactly where keeping the content types separate (CMS view) and giving them a uniform traversal interface (app view) coexist — the graph reads through the shared tag surface and doesn't care which underlying collection a node came from.

Overlap scoring produces two flavours of edge: further_in (same thematic thread, deeper) and connects_across (same theme in a different work or format). Hand-curation was rejected — it doesn't scale to Lewis's corpus and kills the serendipity. Tagging is a build-time pipeline: AI drafts at ingest, editors QA before publish, graph recomputes on publish. No runtime AI in the content API path.

IA stand-up 2026-06-30: related-content surfacing deferred

The team called "related content" a gap and deprioritized it for now. This most likely means the related-content surface/UI is deferred within V1, not that the tag-overlap engine is cut — but whether the graph engine itself stays in the V1 build is unresolved. Confirm scope before investing heavily in the engine.

graph LR
  CI(["<b>Content Item</b><br/>Passage | Letter | Essay | Podcast"])
  TH["Themes <i>(M:N)</i>"]
  TG["Tags <i>(faceted, TBD)</i>"]
  OV["<b>Overlap Score</b><br/>pairwise"]
  FI["<b>further_in[]</b><br/>same thread, deeper"]
  CA["<b>connects_across[]</b><br/>same theme, diff. work"]

  CI --> TH
  CI --> TG
  TH --> OV
  TG --> OV
  OV --> FI
  OV --> CA

Candidate tag facets (input to the deferred Tags session): motif (cross-work, powers surprise jumps) · life-stage (youth → grief) · register (bedtime / morning / anytime) · tone-depth (meaning / theological). Edge weighting across facets is TBD and determines the further_in vs connects_across split.

User Data & Soul Map (app layer — not in CMS model)

Also outside Lina's CMS doc, because it's user state, not authored content. The Soul Map is the user's personal record of what they've read and reflected on — and the reason accounts exist in V1. The app is fully usable as a guest; you make an account specifically to keep your Soul Map. No paywall, no content gating in V1.

Soul Map isn't its own table — it's a virtual aggregate of Save (the reflection moment: which item, which prompt was shown, the user's response), JourneyProgress (resume point per active Journey), and JourneyCompletion. Save snapshots the prompt text rather than referencing it by ID, because a Chapter can override a piece's default Prompt — the Save must record what the user actually saw. Reflection is the moderated community layer ("publish, not post") surfaced as ambient presence, not a comment thread. Account deletion is mandatory (App Store policy), done as soft delete.

Type What it is Phase
User Passwordless email-link (magic link) account via Firebase. Soft delete for App Store compliance. MVP
Save Soul Map entry. Snapshots prompt + user response. MVP
JourneyProgress Current chapter + resume item per active Journey. MVP
JourneyCompletion Records finished Journeys. MVP
Reflection Moderated community note. Ambient presence, not a thread. MVP schema · P2 moderation UI
DailyDrop Scheduled "today's piece" pointer. MVP schema · P2 scheduler
DeviceToken APNs/FCM push token. Backlog
graph TB
  U["<b>User</b><br/>magic link · soft delete"]
  S["<b>Saves</b><br/>item + prompt snapshot + response"]
  JP["<b>JourneyProgress</b><br/>current chapter + resume item"]
  JC["<b>JourneyCompletion</b>"]
  U --> S
  U --> JP
  U --> JC

The open decision: separate types vs unified traversal

Decision for the call

Lina's CMS model keeps Passage / Letter / Essay / Podcast as separate collections (right for authoring — their fields genuinely differ). My earlier draft unified them into one record (right for the graph — uniform nodes). These pull in opposite directions only if you force one representation to serve both jobs.

Proposed synthesis

Keep the separate CMS collections as the authoring + storage shape, and expose a thin polymorphic Content Item interface (the shared Tags + Themes + Prompt surface) that the graph, Journeys, Saves, and Reflections all reference. Separate storage, uniform traversal — both views are true at once, no rebuild needed if one side evolves. This plan is drawn that way; the call just needs to bless it (or pick a side).

Resolved (ADR-0006 / ADR-0008)

The call picked neither extreme wholesale: the long-form types (letter / essay / chapter) unified into one Writing model discriminated by kind, while Passage and Book (née Work) stay separate. Cross-type references use two nullable FKs + an XOR constraint (ContentRefMixin), not a runtime polymorphic bridge. The settled shape lives in CMS Architecture and the journeys-and-home design.

Key design decisions

Decision Rationale Source
Separate content types (Passage / Letter / Essay / Podcast) Fields genuinely differ — plain vs rich text for audio sync, Letter's recipient/date, Essay's author/Substack. One table can't hold them cleanly. Lina
Polymorphic Content Item bridge Lets the graph, Journeys, Saves reference any content type uniformly without merging storage. Synthesis
Theme is the world / portal; Journey belongs to Themes Themes have landing pages, palettes, portal imagery. Journeys are paths inside worlds, surfaced in multiple Themes. Lina
Passage source = Work proxy xor Source Piece Out-of-corpus books route to purchase via Work; in-corpus letters/essays are the destination directly. Lina
Standalone Prompt with Type enum Reusable across pieces, cleanly overrideable per Chapter — one record referenced or replaced. Lina
Author + Portal Image as content types Authors contribute multiple essays + drive card visual; portal imagery is a growing reusable library. Lina
Content Blocks for Theme pages + Home Editor-arranged, not hardcoded layouts — same block system both surfaces. Lina
Piece↔Piece edges computed, not stored Rabbit-hole feel needs deterministic tag-overlap; stored edges don't scale + lose multi-facet ranking. Saurabh
Save snapshots prompt text Chapters override prompts; the Save must record what the user actually saw. Saurabh
Accounts in V1, gates out Accounts persist the Soul Map; no paywall/gating until Phase 2. Schema accommodates it without rebuild. Saurabh

Open blockers

Blocker Blocks Owner
Separate-vs-unified content types (bless the synthesis) Whole content spine + graph node shape This call
Tags faceted taxonomy (separate session) CMS multi-selects, AI tagging, graph weighting Editorial + eng
Passage audio model (stream-by-ID vs stored URL) Passage schema, voice-provider integration Eng + voice provider
Work chapter-audio rung — cut or kept? Work schema, source ladder Product
Content block library (block types) Theme landing pages + Homepage Eng + design
Theme Color Palette: CMS reference vs design-token enum Theme schema Design + eng
Letter + Essay corpora not yet in hand Both schemas provisional until ingest Content