Skip to content

About

Turns a PDF into handwriting-ready notes that stay pinned to the exact sentence they came from. Local-first, runs on a local Ollama model.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Marginalia

I remember things I copy out by hand, and nearly everything I read for class is a PDF. Turning a paper into notes I could actually write out used to eat an evening per document. Marginalia is the tool I built to fix that. It reads the PDF, writes structured notes one section at a time, and pins every note to the exact sentence it came from. Notes on the left, source on the right, and the two panes stay linked. The name is the old word for the notes people scrawl in a book's margins.

Everything runs on my laptop. Note generation goes through a local Ollama model, so there's no API key and nothing leaves the machine.

Notes beside the source, split view

Notes stream in as they're written

Drop a PDF on the home screen and hit Generate notes. The backend walks the document section by section and streams each note over SSE the moment the model finishes it. A card gets a definition, the mechanics, an example and a why-it-matters line, so it reads like something worth copying into a notebook instead of a summary blob. A detail dial sets how much each section gets, from a few lines on Concise up to 320 words on Thorough, and there's a range picker when you only care about part of the document.

Notes streaming in during generation

Every note knows where it came from

Each card carries an anchor chip like P6 · EXACT. Click the card and the PDF scrolls to that page and briefly flashes the matched passage. The resolver tries an exact text match first, falls back to a fuzzy match when the wording drifts, and as a last resort anchors to the page with the chip reading PAGE ONLY so you know the difference. Anchors and highlights are stored as coordinates normalized to the page, which is why they survive zoom and window resizes.

Clicking a note jumps the PDF to the matched sentence and flashes it

Highlight the source, hang a note off it

Switch the toolbar to Highlight and drag over the text. Clicking a highlight opens a small editor where a margin note can hang off it, in any of four colors. The → Notes button lifts the highlighted words into the notes column as a new card, anchored back to the passage like any generated note.

A sage highlight with a margin note attached

The outline follows you

The outline rail lists the document's sections and tracks your position as you scroll. The search box covers the source text and the notes together.

Outline rail beside the document

Made to be printed

Export renders the notes as a print layout, wide margin or Cornell. The whole point of the app is paper. I print the Cornell layout and copy it out longhand.

Wide margin export layout Cornell export layout

Ask about the page you're on

A side rail scopes an assistant to the page on screen. Read this page has it write anchored notes straight into your margin, and the ask box answers questions with that page as context. It drives the Claude Code or Codex CLI already on my machine, picked from a dropdown, and the app works fine with the rail collapsed.

The assistant rail scoped to the current page

How it's built

Two strictly separated halves that talk over HTTP and SSE only. The full design lives in SPEC.md.

  • backend/ is FastAPI + SQLModel on SQLite, with PyMuPDF doing the parsing and optional Tesseract OCR for scanned pages. The LLM sits behind one driver interface in services/llm_client.py. Nothing else in the codebase talks to a model, and swapping Ollama for the dormant Anthropic driver is a config flip.
  • frontend/ is React + TypeScript + Vite, with pdf.js rendering the pages. TanStack Query owns server data and a small Zustand store owns view state. The coordinate and anchor math lives in pure modules with unit tests.

The store of record is SQLite plus the PDF on disk. The browser keeps nothing but ephemeral view state. There's also a bulk import that swallows a whole folder of PDFs at once.

Run it

You need Python 3.12 via uv, Node 18+, and Ollama with the notes model pulled:

ollama pull qwen2.5:7b        # about 4.7 GB, serves on localhost:11434

Backend first, then the frontend, in two terminals:

cd backend
uv sync
uv run uvicorn app.main:app --reload --port 8000
cd frontend
npm install
npm run dev                   # Vite on :5173, proxies /api to :8000

Open http://localhost:5173, drop in a PDF and click Generate notes. Scanned PDFs degrade to page-level anchoring unless Tesseract is installed. Configuration is environment driven with local defaults, documented in SPEC.md.

Test

# Backend: services + API against a deterministic MockDriver
cd backend
uv run pytest
uv run ruff check .

# Frontend: unit tests, types, build
cd frontend
npm run typecheck
npm test
npm run build

# End to end
npx playwright install chromium   # one time
npm run e2e

The e2e suite is hermetic. It boots its own MockDriver backend on ports 8001 and 5174 with a scratch data dir and never touches your dev servers or the real ./data.

Layout

backend/app/
  api/        thin HTTP routers (documents, notes, highlights, export, health)
  services/   the logic: pdf_parser, note_generator, anchor_resolver, llm_client, export
  models/     SQLModel tables           schemas/  Pydantic request/response
  prompts/    note-generation templates core/     config + db
frontend/src/
  components/ SplitView, PdfPane, NotesPane, NoteBlock, HighlightLayer, MarginNote, PrintView
  lib/        apiClient, coords, anchors (pure, unit tested)
  hooks/      useNotesStream, useSyncScroll, useHighlights
  state/      Zustand view store         styles/   tokens.css, print.css
frontend/e2e/ Playwright sync + highlight round-trip tests

License

MIT.

About

Turns a PDF into handwriting-ready notes that stay pinned to the exact sentence they came from. Local-first, runs on a local Ollama model.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages