BodyCut Training
A mobile web app that puts body composition, running, and strength work on one screen so you can cut fat without losing muscle - the local LLM writes only the interpretation, and every number on the screen is computed by code.
- TypeScript
- Hono / SQLite
- Local LLM (Ollama)
- Auth / Security
- QA
Setup
- Problem
What actually stalls people mid-cut isn't "the weight won't come off." They want to lose weight but not muscle, they have a body-fat target but no idea how much training is the right amount, and they don't know how to run and lift in the same week. Gym exercise names and machine names are still unfamiliar, and even with equipment at home they don't know what to do with it. Schedules and fatigue break the plan regularly. Then one more thing lands on top. Body composition readings swing hard within a single day. Depending on how much water you drank, when you stepped on the scale, and whether you'd eaten, fat mass can look like it jumped. Hand that number over raw and the user reads it as their own effort failing, when nothing failed except the measurement. So I positioned it as "an app for losing fat while keeping muscle" rather than "an app for losing weight," and wrote a separate principle into the spec: design so that measurement error is never interpreted as failure. That one line is where the confidence rules and the boundary around the LLM both start.
- Context
I built this alone, and got it to a working core loop in a two-day sprint from 8 to 10 August 2026. Three constraints shaped it. Body composition and weight are health data, and I didn't want to ship them to an external LLM API. The two days were all the time the project got. And working solo, I had no reviewer. So the API runs on a single Mac mini at home, exposed only through a Tailscale Funnel, with the frontend as a static SPA on Cloudflare Pages. The LLM is Ollama running qwen2.5:32b on that same machine. Data never leaving the device is both the product's pitch and the reason the deployment looks like this. The cost is just as plain: when the machine is off, the API is off with it. Someone opening the link can get the frontend but no data behind it, and I decided that belongs written down rather than hidden. With no reviewer, I made it a rule to leave the reasoning for each decision in a code comment. Why this status code, why this rate-limit key - written next to the code, it turns whoever reopens the file later, meaning me, into the reviewer. Most of the reasons in this case study came out of those comments.
- Users
The primary persona in the spec is someone cutting who wants to hold on to as much muscle as possible. Someone who wants to run often, can get to a gym two to four times a week, has some basic equipment at home, is a beginner to intermediate lifter still unsure of machine names, checks a smart scale or InBody readings frequently, and plans to enter a 5km or 10km race. Honestly, that persona started as me. But leaving it as a single-user app would have baked one person's target numbers into the code, so I generalized race distances to 5K / 10K / half / full and added signup, terms consent, and an admin member panel, redesigning it to accept multiple users. Being able to accept users structurally is a different claim from having them. Real usage is not measured yet, and no analytics are wired in.
- Hypothesis
If noisy body-composition readings are filtered through a confidence rule first and only the settled numbers are used as the basis for the written interpretation, a single day of measurement error stops reading as failure and the user can keep tracking the real goal, which is preserving muscle while cutting.
Build
- What I did
- A Today dashboard rendered from a single GET /api/today aggregate - today's weight and its delta from yesterday, body fat percent against the 7-day average, a read on skeletal muscle mass, progress toward the goal, today's run and lift, the next long run, race D-day, and a condition check (GREAT/GOOD/TIRED/PAIN), all on one screen
- Body composition entry and history plus automatic confidence grading (lib/
confidence.ts) - eight measurement conditions map to HIGH/MEDIUM/LOW defaults, and a reading is demoted to LOW when fat mass moves 2.0kg absolute or 12% relative, or skeletal muscle mass moves 1.0kg, against the previous HIGH/MEDIUM reading within a 6-hour window. The four thresholds live in one constants object so they can be tuned - Progress trend charts - weight, body fat percent, fat mass, skeletal muscle mass, and running distance over 7/14/30/90 days or all time, with a 14-day muscle-preservation monitor returning GOOD/WARN/NEUTRAL
- An exercise library of 25 movements (searchable by name, category, and muscle) with detail pages, a weekly training plan, today's completion checks, and equipment CRUD including an adjustable-weight stepper
- Running log CRUD with automatic progression - recording a longer long run or race raises your longest distance, and deleting that record recalculates it. Plus race CRUD, D-day, and a race-readiness screen
- Two local-LLM interpretations - GET /api/ai/body-insight for today's body composition and GET /api/ai/muscle-monitor for the 14-day muscle-preservation comment. Every number is computed by code and injected into the prompt as JSON; the model only writes sentences. If the model is unreachable, times out, or returns unparseable output, a rule-based fallback comes back with fallback:true
- Auth and accounts - HS256 JWT and scrypt implemented directly on node:crypto, name/phone/three-way terms consent, account deletion that re-confirms the password and wipes all data in a transaction, an
ADMIN_EMAILS-basedadmin member panel, and three CSV importers (body composition, running, strength) with a downloadable template per type
- Product decisions
- Code computes, AI explains. Trends, averages, confidence, and muscle-preservation state are all computed by deterministic rules and shown as settled values; Ollama receives those numbers as context and writes only the interpretation. The reason is singular: the AI must never invent a number. The boundary is visible in the code too - after the model responds, confidenceNote is overwritten with the value code computed.
- No external LLM API, only local Ollama. Once the pitch is that health data never leaves the user's device, a single convenience call to an outside API makes that pitch a lie. The Ollama client file carries a warning at the top that it is server-side only and must never end up in the browser bundle, so the boundary is stated where someone would cross it.
- Rate limiting keyed on the account (email), not the IP. Behind a tunnel, X-Forwarded-For is client-controllable, so an IP key gets bypassed by rotating a header; and if everyone appears to share one IP, one person can lock out every user. Signup has no account yet, so it's held by a global cap instead (30 in 10 minutes), and I wrote down the trade-off: on an app this small, a legitimate burst of signups shares that cap.
- No hard-coded fallback for
JWT_SECRET- if the value is missing, the server refuses to boot. A fallback means tokens can be forged with a known secret. I chose failing to start over starting in an unsafe state. - Return 403, not 401, when the deletion re-confirmation password is wrong. This client treats every 401 as an expired session and logs the user out, so a 401 here would eject someone for one mistyped password. I picked the server status code while looking at the client's global handling next to it.
- LOW-confidence readings are excluded from averages and min/max aggregates but stay on the chart, dimmed rather than removed. Delete them and the user feels a measurement they took has vanished; leave them in the aggregate and the average is contaminated. It is the "don't hide the error, don't let it read as failure" principle turned into a UI rule.
- QA considerations
- Does the auth gate run before the route handlers, and does login response time leak whether an account exists - the JWT middleware is registered ahead of the routers across all of /api/*, with only signup and login excepted by path comparison. Even when the email doesn't exist, a scrypt verification runs once against a
DUMMY_HASHso timing doesn't differ, and both hash comparison and JWT signature comparison use timingSafeEqual. No test pins any of that down, though. All I have is a manual check against the live API: calling /api/today and /api/ai/health without a token returns 401 for both. - Does the rate limiter lock out real users instead of attackers - login counts only failures and resets the counter on success, so a few typos don't lock anyone out. Ten per email, thirty globally for signup, five per account for a wrong deletion password, each in a 10-minute window, with 429 and a Retry-After header past the limit. The hole on the other side belongs here too: the counters live in a process-memory Map, so restarting the API wipes every window, and repeated restarts effectively remove the limit. Nothing has broken yet because there is one process and no traffic, not because the design prevents it.
- Does the screen go blank when Ollama is down or the network drops - a 1.5s reachability probe fails fast, generation has a 25s timeout, and the response passes a parsing guard that will even try extracting the outermost braces before headline/body presence and tone enum membership are validated. If any of that fails, a rule-based sentence comes back with fallback:true, and a cached value that no longer parses is treated as a cache miss. On the client, network failures are normalized into a status 0 ApiError, and all 20 pages carry an error state. That is the code path, not evidence: I never took Ollama down and recorded the result. There is also no navigator.onLine offline notice, so a user sees that it isn't working without seeing why.
- Can the LLM invent numbers - deltas, averages, and monitor state are computed in code and injected into the prompt as JSON, and the prompt states not to introduce any number beyond the ones provided. confidenceNote, the deterministic signal, is overwritten with the code value after the model responds, so a wrong sentence still can't contaminate a number.
- Does a stale interpretation survive when past readings are deleted or backdated - the AI cache key is not the measurement id but a combination of signal values (7-day deltas for weight, body fat, and muscle, plus monitor state and confidence), so when the average changes, the sentence is regenerated. The key carries a userId prefix so one user's cache can't bleed into another's.
- Can an imported CSV push garbage straight through - the client parser handles commas inside quotes, escaped quotes, CRLF/LF, and a BOM, and the server re-validates weight, body fat percent, and skeletal muscle mass as positive numbers row by row. measuredAt is dropped unless it starts with YYYY-MM-DD, because that column is the basis for both ordering and confidence comparison, so one malformed string would shake the whole grading. Failed rows come back with their row number and reason. Working through this is also how I found that the single POST /api/body, which writes the same data, has no equivalent server-side check. I did not fix it; it went into the retrospective as it stands.
- Does account deletion leave data behind, and can an admin delete themselves - deletion by the user and deletion by an admin share the same cascade, wiping eight user-data tables plus
user_profileplus the userId-prefixed AI cache inside one BEGIN/COMMIT/ROLLBACK transaction. An admin deleting their own account is blocked with a 400 to prevent self-lockout, and admin rights are computed server-side fromADMIN_EMAILS; a role sent by the client is never trusted. What's missing is a test that this transaction really leaves nothing behind. I walked through it by hand once and kept no record, which leaves the one irreversible path in the app as the thinnest-verified one.
- Does the auth gate run before the route handlers, and does login response time leak whether an account exists - the JWT middleware is registered ahead of the routers across all of /api/*, with only signup and login excepted by path comparison. Even when the email doesn't exist, a scrypt verification runs once against a
Outcome
- Metrics
- Frontend live at https://bodycut-training.pages.dev (Cloudflare Pages). The API is a single process on one personal Mac mini, so it goes down when the machine does. I can't promise the link is always up.
- 41 endpoints under /api/* plus one unauthenticated /health, 20 frontend routes (18 behind login plus /login and /signup), a 25-movement exercise catalog, three CSV importers and three templates
- 15,271 lines of source (.ts/.tsx, excluding
node_modules), 19 commits, 14 merged PRs. It got here in a two-day sprint from 2026-08-08 to 08-10, and there have been no commits since. - Zero automated tests. No test files, no test-runner dependency, no test script.
- All 20 pages handle an error state and 16 handle a loading state, with aria-busy in 16 places, aria-live in 12, aria-label in 12, and aria-pressed in 5. There is no automated accessibility check at all, so whether those attributes are actually correct is unverified.
- The GitHub Actions deploy workflow has run 10 times with 0 successes, because I never added the Cloudflare secrets. Pages deploys are done by hand.
- User count, return visits, retention, and LLM generation time are all not measured yet. No analytics are wired in, and the warm/cold generation times written in the docs have no benchmark log behind them, so I don't cite them.
- Result / Learning
The core loop went live in two days. Enter a body composition reading and it gets a confidence grade automatically, noisy readings drop out of the aggregates, and the Today screen shows the settled numbers alongside a Korean sentence explaining them. Running and strength records sit on the same screen, so what to do today is visible in one look. A temporary edge password gate I had put up before launch got removed once real login existed, because it turned into a double prompt. That said, the goal simulator the spec listed inside MVP scope is still on the TODO, so a working core loop and a finished MVP by the spec's own definition are not the same thing. The clearest thing I learned is that where you draw the boundary around an AI feature is itself a quality decision. "The AI interprets your body fat percentage" quietly contains two jobs, computing numbers and writing sentences, and if you don't separate them there is no way to catch the model stating a wrong average. Separating them created somewhere to verify: the numbers can be checked against a rule you can read, and a wrong sentence can't contaminate a number. Designing the failure path before the happy path was also the right order for a local LLM, and I'd do it again. But it's worth being exact about how much of that is confirmed. In the code path, the screen shows something whether Ollama is up or not; there is no test that actually cuts the model off and holds that behavior in place. And from outside you cannot tell whether the sentence on the screen was written by the model or by the rule-based fallback. The AI health check sits behind auth too, so only someone logged in can check.
- Retrospective
- I shipped it with zero automated tests. That is the sorest point in this project, and a two-day sprint is not an excuse. The confidence grading, the JWT code, and the rate limiter are already deterministic rules with their thresholds pulled out as constants, which made them the cheapest possible place to pin boundary tests. I skipped the cheapest thing to lock down. If I open this again, I start with the 2.0kg / 12% / 1.0kg / 6-hour boundaries in
confidence.ts. - Validation is asymmetric across paths. The CSV importer validates positive numbers row by row, while the single POST /api/body that writes the same data inserts the request JSON with no server-side check. SQLite is loose about types, so a value like a negative weight can land if the client validation is bypassed. I defended the path I built later and left the earlier one as it was, which is a textbook ordering bias. The CORS origin has the same shape of problem: its default is open, and not making it refuse to boot the way
JWT_SECRETdoes is inconsistent. - I built the deploy workflow before wiring the secrets. It has run ten times and failed ten times, and main has carried a red badge ever since. I did write a comment up front saying it just fails fast until it's configured and breaks nothing else, but a permanently red CI kills the signal itself. Either finish wiring it or delete the workflow. The rest of the operational side is just as thin: there is no SQLite backup procedure in the deployment doc, and schema changes go through a homegrown check-then-ALTER-TABLE-ADD-COLUMN step, so rollback isn't a concept that exists here.
- I shipped it with zero automated tests. That is the sorest point in this project, and a two-day sprint is not an excuse. The confidence grading, the JWT code, and the rate limiter are already deterministic rules with their thresholds pulled out as constants, which made them the cheapest possible place to pin boundary tests. I skipped the cheapest thing to lock down. If I open this again, I start with the 2.0kg / 12% / 1.0kg / 6-hour boundaries in
- Tech stack
- TypeScript 5.6
- React 18 + React Router 6
- Vite 5
- TailwindCSS 3
- Hono 4 + @hono/node-server
- Node 24 (node:sqlite 내장 드라이버)
- SQLite (WAL, foreign_keys ON)
- pnpm workspaces 모노레포 (web / api / shared)
- Ollama qwen2.5:32b (로컬 LLM)
- 자체 구현 HS256 JWT + scrypt (node:crypto)
- Cloudflare Pages
- Tailscale Funnel
- macOS launchd LaunchAgent
- GitHub Actions (설정 미완, 실행 전부 실패 상태)