Skip to main content
강홍재/ James
← Work
EMBA2026 · Solo Builder· Started(First Commit date)

데일리 다이제스트

A private reading archive that parses posts from a paid subscription content service day by day and keeps restated summaries instead of the body - the copyright constraint is enforced by the schema and a save gate, not by a document.

  • TypeScript
  • Playwright
  • Local LLM
  • Postgres
  • Self-hosted
  • QA

Setup

Problem

Posts shared in a group chat are buried within days. Text and images drift apart so the context breaks, and anything skipped for lack of time is effectively gone, because scrolling back up is not a real option. What was missing was a place to catch up calmly, day by day. Technically there were two walls. The original opens only through a time-limited access link, and it is an SPA, so fetching it from a serverless function returns a JavaScript shell instead of the body. Without actually rendering the page you get no text at all. The hard problem, though, was not parsing. The source is a paid subscription service's content. The archive had to be worth reading while never copying the body, and that constraint was dangerous not because it is hard to honor but because it is easy to break. There is always the temptation to dump the full body in now and tidy it up later.

Context

So the design doc stated up front that the original body is never stored or re-hosted, and then I worked out what shape that sentence had to take in code for it to actually hold. A rule written in a document gets forgotten, gets broken under deadline, and is invisible to whoever touches the code later. A rule only holds when it lives in the schema. There were three constraints. Serverless cannot render an authenticated SPA, so I needed a machine that stays on. The original body must never leave for an external API. And the site had to stay closed, out of search indexes. That landed on a Playwright worker running on a Mac mini at home, summaries generated only by a local LLM, and a web app blocked from indexing with a full robots.txt disallow plus noindex.

Users

Me, plus a small private reading group. There is no open signup; you only reach it if you have the link. Aggregation tables for likes and views exist, but the only number I can actually check and report is the note count. Everything else is unmeasured.

Hypothesis

Storing no full body at all, only restated summaries and a single pull quote, is enough for a day-by-day catch-up archive - and that constraint only holds if it is pinned into the schema and a save gate rather than a document.

Build

What I did
  • Job queue plus headless rendering pipeline - I paste a source URL, the API enqueues a job, and the Mac mini worker polls it, renders the page for real with Playwright, and extracts the body text along with image order and captions
  • Two-stage summarization - a local LLM (Ollama) produces a short card summary (one pull quote plus three lines) and, separately, 4-8 restated paragraphs for the archive. No external summarization API, so the original body never leaves the machine
  • A separate LLM call generates a comprehension quiz and a "one step further" insight, and the keyword explainers are generated at write time and cached in the database
  • Weekly trends - when a new note lands, the worker gathers the week's other notes and produces a headline, common threads, per-post lessons, and a retrospective alongside it
  • The web app - a hero for the latest note plus a date-ordered timeline home, a detail view, weekly and monthly accordions, and keyword search over summaries and captions. It is a static export, so there is no server
  • Save gate and tone gate - just before writing, runSaveGate splits errors from warnings, normalizeDashes deterministically replaces em and en dashes with hyphens, and lintTone warns when the three-line summary does not have exactly three lines or the register looks wrong
  • Operator tooling - an editor that shows parsing progress as a checklist, a dashboard for cancelling in-flight jobs and deleting notes, and a cron that every 15 minutes requeues jobs held longer than 30 minutes
Product decisions
  • Enforce the copyright constraint at the schema level - the design doc says the original body is never stored or re-hosted, and the notes table has no column for the full body at all. What is stored permanently is the restated summary, my own memo, and a single pull quote. You cannot accidentally write into a column that does not exist
  • Tried to manage images by policy, then deleted the feature - originally images were split three ways (hotlink / user_upload / stored) and the stored case raised a warning at the save gate. But a feature that constantly raises a warning is really a feature that offloads risk onto human attention, so on 2026-06-22 I removed images entirely and went text-only
  • Local model only for summarization - I ripped out the external LLM API and kept Ollama. I give up some output quality and get, in exchange, the property that the original body never reaches an outside service, which points the same direction as the copyright constraint
  • A machine that stays on with a polled job queue instead of serverless - an authenticated SPA returns only a JS shell to fetch, so a real browser was required. The API first ran on a free plan, but idle spin-down made cold starts a problem, so on 2026-07-20 I moved it onto the Mac mini under launchd
  • Removed the read-side login (2026-06-23) - entering a passcode every time was real friction for readers. Reading is open and only writing (the editor) is operator-gated; privacy is held instead by blocking search indexing and by the link being unlisted. That decision is what produced one of the defects I list below
  • Enforce output tone in code at write time, not in the prompt - an LLM is probabilistic, so "do not use em dashes" in the prompt is occasionally ignored. The replacement is done deterministically on the save path, with the verbatim pull quote left untouched as an explicit exception
QA considerations
  • Splitting errors from warnings in the save gate - make everything blocking and the operator routes around the gate; make everything a warning and nobody reads it. Missing source_url, an out-of-range source_tag, and "no archive blocks and no images, so there is nothing transformed to save" block the write. A stored-policy image, or an upload with no file_key, passes as a warning
  • The tone gate rewriting the verbatim quote - normalizeDashes replaces em and en dashes with hyphens deterministically on the save path, but editing the pull quote taken from the original would mean it is no longer a quote, so that field is excluded from replacement. The register check in lintTone is a heuristic on sentence endings and does produce false positives, so it only ever warns, and the code comment says as much
  • The failure mode where the LLM invents links or emits a broken quiz - normalizeQuizInsight drops any item with fewer than two options or an answer index outside the options, and strips keywords that look like URLs. A keyword has to be a concept you can search, not a link that does not exist
  • Parse failures enumerated as a type - ParseErrorCode = LOGIN_OR_EMPTY | NAV_FAILED | NO_CONTENT. A body under 200 characters is treated as an expired access link or a login shell, so it is reported as a failure rather than saved. Half-empty notes accumulating quietly would be the worst outcome. If networkidle never settles, it retries once on domcontentloaded
  • SSRF - the worker opens these URLs in a real browser, so an internal address or a metadata endpoint would simply be fetched. Only https plus the source domain (and its subdomains) passes; everything else is rejected with a 400
  • Two places where the schema and the contract can drift - at compile time, an Exact<> conditional type breaks the build if the inferred type of the Zod WorkerOutputSchema and the WorkerOutput in types.ts stop matching, so with no CI the compiler catches it instead. At runtime, the worker and the API ship separately, and an older worker posting a weekly trend without the newer fields killed the whole job with a 400 until Zod .default([]) accepted it. That second one I fixed after living through it
  • Publish-date timezone - the original publishes at 00:00 KST (15:00Z the previous day), so using the UTC date shifts every publish date back by one. In a date-ordered archive that is the kind of bug that is wrong quietly, so I convert to KST and left the reason in a code comment

Outcome

Metrics
  • 71 commits, 2026-06-22 to 2026-09-11 (measured from git log, KST)
  • 4 pnpm workspace packages (shared / api / worker / web), 6,598 lines of TS and TSX, 28 HTTP routes on the API
  • 117 notes in the production database (100 daily, 17 weekly), note_date spanning 2026-06-15 to 2026-10-10 (measured on 2026-10-10 through the public read API)
  • Since 2026-07-22 there have been only two batches of code changes - four commits on 21-22 August (KST) for the publish-date bug and a ghost worker, and one on 11 September for the false-positive connection overlay. I still have no way to check how many parses failed in between
  • Zero automated tests - no test files and no test-runner dependency in any package.json
  • 3 parse failure codes, a 200-character minimum accepted body, and 2 API boundaries covered by Zod runtime validation
  • Real usage numbers (active readers, total likes and views, parse success rate, summarization time) all unmeasured
Result / Learning

117 notes have accumulated (measured 2026-10-10), and notes have kept landing with only two batches of code changes since 2026-07-22. That is as far as the claim goes, though. There is a path for a dead job to be requeued by the cron and a path for a half-failed parse to be recorded as a failure instead of saved, but I never added the counters, so I do not know how often either one fired or how many parses failed. "I did not have to touch it" and "nothing went wrong" are different statements. The clearest lesson was about where you put a constraint. "We do not store the body" written in a document and a column that was never created hold to very different degrees. Images taught the same lesson. Splitting them into three policies and warning on the dangerous one looked reasonable, but given that the only person reading and judging those warnings is me, it was not a defense - it was a decision deferred. Deleting the feature is what actually removed the judgement call. At the same time the biggest hole in this project is plain. I built the gates and wrote zero tests for them. What the save gate and the tone gate actually block has, to this day, only ever been verified by hand.

Retrospective
  • There are zero automated tests. It is a tool I use alone and the gate logic kept moving early on, so checking by hand felt faster - but that window closed a while ago. runSaveGate in gate.ts and normalizeDashes / normalizeQuizInsight / lintTone in tone.ts are pure functions, input in and output out, which makes them the easiest possible place to test, and they still have none. The next piece of work is vitest on those two files, pinning the errors/warnings split, the pull-quote exception in dash replacement, and the removal of malformed quiz items. I start there rather than at parsing or the LLM because that is the code that actually enforces the copyright constraint.
  • I opened up reading and then applied only half the defense. The jobs list API is operator-only, with a comment saying it is not public because it carries source URLs with access links in them - but the notes list API has no such guard, so the same URLs go out unauthenticated. The policy was applied on one side only, and I found it re-reading my own code. In the same vein there is no rate limiting anywhere, and the device ID is a client-generated UUID, so like and view counts can be forged. Stripping the source URL out of public responses comes first.
  • When I deleted features I did not delete the documentation with them. Login protection, a two-column image gallery, image storage, and entering the editor by clicking the home title were all removed in later commits and are all still described in the README. I enforced constraints through the schema and a gate, and then left documentation drift untouched. Anyone judging this project from the README today gets the wrong picture in four places.
  • Fixed in code and deployed turned out to be different statements. On 21 July (KST) I changed the parser to derive the publish date in KST, and that code did not actually run until 21 August. A worker process that had been alive on the Mac mini since 23 June, started outside launchd, survived the install script's cleanup (which only stops what launchd manages) and kept processing jobs on the old code; racing the new worker for the same queue, it made publish dates come out right or wrong depending on who grabbed the job. The cause of an intermittently reproducing bug was a process, not code. The install script now finds workers launchd does not manage and prints their PID and start time, and the worker logs the commit it is running at startup and warns when it differs from origin/main. The note dates already stored wrong were repaired with a re-parse tool that only inspects by default and only writes with --apply, and the same change fixed its report, which had been counting failed corrections as successes.
  • Whether the code honours what a comment promises cannot be learned by reading the comment. The connection-error overlay kept appearing on phones and, once up, stayed until a manual refresh, while the API was in fact healthy (twenty consecutive 200s on a public route). Two causes. A single rejected fetch raised the overlay immediately, and on mobile a tab going to the background or a switch between Wi-Fi and LTE is enough to produce that rejection. And the comment said an api:ok event would dismiss the overlay, but nothing emitted that event and nothing listened for it, so there was no recovery path at all. The 11 September (KST) fix shows the overlay only after /health confirms the API is really down, re-checks every five seconds while it is up, checks immediately on app return or network recovery, and dismisses itself when the API answers.
Tech stack
  • TypeScript
  • pnpm workspaces (4 packages)
  • Next.js 16 (static export)
  • React 19
  • Tailwind CSS v4
  • Hono 4
  • Neon Postgres (pg)
  • PGlite
  • Zod
  • Playwright (headless Chromium)
  • Ollama local LLM (qwen2.5:32b)
  • node-cron
  • Cloudflare Pages
  • Mac mini + launchd + Tailscale Funnel