SPARK 주간 읽을거리
A weekly article roundup I used to write by hand, moved onto a publishing pipeline with a real validation gate - and when a rule collided with reality, I demoted it to a warning instead of deleting it, so the violation stays on screen.
- Node.js
- SQLite
- Cloudflare Pages
- Local LLM
- Content Pipeline
- QA
Setup
- Problem
Every week I picked articles for the KakaoTalk group chat of SPARK, a startup club. The hard part was never the workload; it was that the system held three separate versions of the truth. The curation rules lived in a prose Markdown file, the week's content lived in a .txt I rewrote by hand, and the actual output lived as plain text pasted into the chat. Three things maintained separately, so when they drifted apart nobody noticed. The symptoms were concrete. The rules document said "summaries end in a noun phrase," but whether I had actually followed it was something I checked by eye right before sending. Same for dead links. Finding what went into a past issue meant scrolling back through the chat, and the moment I added a web archive, the chat text and the web page would start diverging. So I set the goal as "pin each truth to exactly one place" rather than "automate it." Rules in one rules file, data in one database, and every output derived from those. The constraint was zero infrastructure cost.
- Context
In the design doc I compared two serving models. In A, a Mac mini at home answers every request; in B, publishing produces a static snapshot that gets pushed to Cloudflare Pages. I picked B. Weekly content is essentially static once it is set, so there is no reason to hit a personal machine on every request; a static file on the edge keeps the archive up even when the Mac mini is off; and above all, the write path is never exposed to the public internet at all. The zero-cost constraint decided most of the rest. A local Ollama running on the Mac mini instead of a paid LLM API. SQLite as a single file instead of Postgres, and Node's built-in node:sqlite instead of a native module. Runtime dependencies came out to one XML parser. The admin UI is a single HTML file with no framework, and it is never exposed publicly; it opens only over a private network. I also cut automation deliberately in half. Ingestion is automatic; selection, publishing, and deployment are a human pressing a button. The README says that is design intent, not an unfinished edge.
- Users
The readers are SPARK members in the KakaoTalk group chat. I am the only operator doing the curation, and the live site's footer already credits me as the club's president. To be honest about it, I do not know how much those readers actually read. Chat reactions, reader count, and link click-through are all unmeasured. There is a view counter on the site, but it is a plain increment counter that does not separate bots from my own visits, so it cannot be read as traffic.
- Hypothesis
If the curation rules move out of a document people read and into one file the code reads, and get checked automatically right before publishing, the weekly eyeball pass disappears and the chat text and the web archive stop drifting apart.
Build
- What I did
- Poll registered RSS/Atom feeds and load the last ten days of articles as candidates - articles.url carries a UNIQUE constraint and ON CONFLICT DO NOTHING blocks duplicates at the database level
- Send a real HTTP request to every collected link and record a check only for the ones that answer - GET fallback for servers that refuse HEAD, redirect following, concurrency of 6
- Force a JSON schema on the local Ollama to produce category, noun-phrase summary, fit score, language, and overseas flag - a category outside the rule list is dropped to null so a human decides it
- List candidates by fit score in an admin UI with add, reorder, edit-summary, manual-add, and delete (19 admin API routes, a single 374-line HTML file, no build tooling)
- Eight automatic checks against
rules.json right before publishing - count, date range, source diversity, missing description, em dash usage, noun-phrase ending, live links, overseas tag - Render the plain-text chat .txt and the web HTML from the same database, removing the path by which the two outputs could diverge
- One publish button runs validate, publish, static site build, and Cloudflare Pages deploy (killed after a 180-second timeout). View counts, the one thing a static site cannot do, are split out into a separate Worker plus KV
- Product decisions
- Split the truth into three and pin each to one place - rules in
rules.json, data in SQLite, outputs entirely derived. I wrote one sentence into the rules: "never hand-edit the chat .txt." The moment you edit it by hand, the single source is broken again - Split validation rules into publish-blocking and warning - once I was actually running it, the count rule and the source-diversity rule did not match reality, so instead of deleting them I demoted them to
warn_onlyand let the violation keep showing. I have not concluded whether the rule is wrong or the curation is wrong, and until I do I would rather not erase the fact - Compared serving models A and B in the doc and chose snapshot publishing - weekly content barely changes, so there is no reason to hit a personal machine per request; a static file on the edge keeps the archive alive when the machine is off; and the write path is never public
- Cut automation in half - ingestion automatic, selection and publishing and deploy human. Machines are good at gathering candidates; "should this article go to the club this week" is still better answered by a person. The README states this as design intent
- No paid LLM API; it calls a local Ollama instead, for the zero-cost goal. Same reason for SQLite over Postgres and Node's built-in node:sqlite over a native module, which brought native dependencies to zero. Postgres is overkill at weekly scale, and backup is copying a file
- Separated delete from hide - hide only pulls an issue off the live site and keeps the data, so it can be republished at any time. Deleting an issue returns its articles to the candidate pool, and an article already in an issue refuses deletion outright
- Split the truth into three and pin each to one place - rules in
- QA considerations
- Does the gate actually stop, or just log and continue -
publish.jsprints the violations and exits 1 on failure. A check that prints a warning and lets the pipeline flow on is the same as no check - Is the link actually alive, verified over HTTP rather than by eye - GET fallback for servers that refuse HEAD, redirect following, a 10-second HEAD and 15-second GET timeout. A dead link leaves no check record, so the pre-publish gate catches it on its own
- Can an LLM-invented category leak past the rules unnoticed - Ollama gets a JSON schema and temperature 0.2, and any category outside the
rules.json list is dropped to null for a human to decide. Plausible new categories quietly accumulating is the quietest failure mode here - Does one dead feed stop the whole week's ingestion - each feed is isolated in its own try/catch with a 20-second timeout. Normalizing RSS 2.0 and Atom into one shape meant handling title/link arriving as objects rather than strings, HTML tags, and numeric, hex, and named entities, plus a browser User-Agent for servers that answer 403
- Can a retired issue survive in the edge cache - retired issue numbers get a tombstone page that overwrites the Cloudflare edge cache left by a previous deploy, and a 404.html makes unmatched paths return a real 404. I verified on the live site that requesting a non-existent issue number returns 404
- Does reordering break against a UNIQUE constraint -
issue_itemshas UNIQUE(issue_id, position), so reordering runs as a two-pass inside a transaction: push rows to temporary negative positions first, then rewrite 1 through N - The known gaps - there are zero automated tests verifying the validator. The noun-phrase-ending check is a heuristic matching four forbidden endings as strings, so passing it does not guarantee a noun phrase, and the admin API has no CSRF defense and only single-password Basic Auth, so every write route would be wide open the moment the admin UI left the private network. I got as far as moving the rules into code; verifying that code is a step I have not taken
- Does the gate actually stop, or just log and continue -
Outcome
- Metrics
- Nine issues published (
feed.json, measured 2026-10-10) - #1 to #5 weekly (KST 08-10 / 08-16 / 08-23 / 08-30 / 09-06), then a three-week gap, #6 to #8 pushed out together on 09-26, and #9 on 10-03. Commits cluster into a few days, so the evidence of operation is the publish record, not the git history - and that record keeps the three missed weeks visible - Nine live issues carrying 110 articles (#1 8 / #2 10 / #3 18 / #4 17 / #5 13 / #6 11 / #7 14 / #8 5 / #9 14), 9 distinct outlets cited, all 6 categories used, 2 articles tagged overseas, and a link-check record (
link_checked_at) on all 110 - All 8
rules.json checks implemented invalidate.js- 2 of them (count, source diversity) demoted towarn_only - 4 RSS feeds ingested automatically, 19 admin API routes, 9 npm scripts
- About 1,149 lines of backend JavaScript across 11 files plus a 374-line single-file admin UI. One runtime dependency, zero native dependencies, 32 commits (2026-08-08 to 2026-09-06)
- Zero automated tests - no test files, no framework
- There is a view counter, but I do not use it as a metric: it is a plain increment counter that does not separate bots from my own visits, so it cannot be read as traffic. Chat reactions, reader count, and click-through are unmeasured
- Nine issues published (
- Result / Learning
This is not a demo; it actually published. Issues #1 to #5 came out weekly from 10 August to 6 September 2026, then three weeks were skipped, #6 to #8 went out together on 26 September, and #9 on 3 October. The archive at https://spark-weekly.pages.dev still opens with nine issues and a
feed.json. Commits cluster into a few days, so from the git history alone it looks like a weekend hackathon, but the evidence of operation is in the publish record, not the commit log - and that record keeps the three missed weeks visible. The most valuable moment in this project was not code; it was the point where a rule collided with reality.rules.json says an issue holds eight or nine articles, and in practice there were more worth sending that week. In a commit on 10 August 2026, right after publishing started, I demoted the count rule and the source-diversity rule towarn_only. I did not delete the rules, quietly change the numbers, or publish while ignoring them. I lowered the severity and left them visible. Where that call led showed up in the later issues: #3 and #4 carried 18 and 17 articles against a rule that asks for eight or nine, and that fact appears as a warning on every publish. It went the other way too: #8 carried 5 articles, under the minimum of eight, and the same warning showed in the same place. What I took from it is that a gate's value is not only in blocking, but in being able to adjust severity. A validator with nothing but hard failures gets switched off wholesale or ignored the first time it meets reality. Had I deleted the count rule, nobody would know today how far #3 and #4 drifted from the original standard. Keeping the disagreement between rule and reality as data is the most QA-shaped thing this pipeline does.- Retrospective
- I built a validator and wrote zero tests for it. Through nine issues of publishing I had no way to know whether
validate.jswas letting something wrong through; my only defense was rerunning all eight checks by hand whenever I changed a rule. The next task is tests for four cases: date-range violation, em dash, noun-phrase ending, and an unverified link. It is the cheapest fix for the biggest hole, and I postponed it. - This started as a project about a single source of truth, and the human-facing documents drifted apart again. I changed the ingest schedule from Saturdays to daily and never updated the README; the rules file still carries a retired-issue entry that looks like leftover and no longer does anything; and each issue records a
rules_versionwith no logic to reproduce a past issue under the rules of that moment. I held the truth the code reads and lost the one people read. - Ingestion already parses the RSS description and then passes only title and source to the summarization step. The summary is grounded in nothing but a headline, which caps the quality of the LLM output by my own doing. Not forwarding information the pipeline already holds is closer to a design mistake than a shortcut.
- I built a validator and wrote zero tests for it. Through nine issues of publishing I had no way to know whether
- Tech stack
- Node.js >= 22.5
- node:sqlite (내장 모듈)
- SQLite
- node:http (프레임워크 없음)
- fast-xml-parser
- Ollama qwen2.5:32b (로컬 LLM)
- Cloudflare Pages
- Cloudflare Workers + KV
- wrangler
- launchd (macOS)
- Tailscale
- Vanilla HTML/JS 관리 UI