Skip to main content
강홍재/ James
← Work
EMBA2026 · Solo Builder· Started(First Commit date)

스파크 AI 창업 점집

An unstaffed booth web app that runs on each visitor's own phone from a QR code, built on one GPU and a fixed event window - so the hard part was reliability, not AI.

  • Next.js 15
  • Self-hosted LLM
  • Fallback Design
  • Live Event
  • QA

Setup

Problem

The booth had to introduce a club and collect signups at a school orientation event, but there was no room for devices, no printer, and nobody who could stand at the table. What was left was the phone every visitor already had in hand, and one Mac mini at home running the model. Under those conditions the real question was never "what should the AI do." The event window was fixed and could not slip. I set thirty seconds in front of the booth as the target for the whole thing, and the moment those thirty seconds stall, the next person just walks past. That makes it an availability problem, not a model-quality problem. So before writing the app I listed what could break: several requests landing on one GPU at once, the model getting evicted from memory so the first visitor pays the cold start, the tunnel dropping, a shared campus IP making one rate limit cover the entire booth, and the case where I cannot open a laptop at all.

Context

I skipped commercial AI APIs and ran everything on a self-hosted Ollama with gemma3:27b on a Mac mini. The reason was API cost. In exchange I took on two different problems, cold start and concurrency, and paid them back with a warmup path and a serial queue. I first picked Cloudflare Workers for deployment and then walked that back to Vercel: the worker subdomain carries my personal account name, and changing the address would have broken the addresses of other workers too. Print was its own constraint. A QR code cannot be corrected once it is printed, so I forced the order: fix the deployed address first, burn the QR against that address, then print. After that the infrastructure behind it can move and only an environment variable changes, while the printed material stays valid. The whole thing took five days and forty commits. With the event date fixed there was no room to grow the scope.

Users

Visitors at a school orientation event and the club members running the booth. Visitors scan a QR code and go through it on their own phone; the organisers watch the booth through a /stats dashboard, a /draw raffle screen, and /api/health. It did run on the day. Participation, though, was never instrumented as a metric: logging only records successful AI calls, so anyone who fell through to question mode never appears in the counts at all. That means this project has no participation number worth putting forward, and I am leaving it that way.

Hypothesis

Even with one GPU and nobody staffing the table, designing the fallback, the queueing, and the recovery procedure before the code keeps the booth experience running for the whole event window.

Build

What I did
  • A QR deep-link entry that branches by mode - /?m=face for a selfie reading, /?m=idea for an idea appraisal, /?m=about for the intro. /poster renders five print-ready posters with their QR codes generated at runtime against the deployed address
  • Selfies resized to a 768px JPEG in the browser, then classified by the self-hosted gemma3:27b vision model into one of twelve types, with three reading points, four stats, and a lucky item. The photo is never stored or logged; it is discarded within the request scope
  • A one-line idea appraised on three axes (market, originality, feasibility) with an S-to-D grade, two strengths, two risks, and one first step that fits inside this week
  • Automatic switch to a five-question mode when the AI fails - it runs client-side with no server and no network, and decides against the same twelve-type system deterministically from a seed
  • A FIFO serialisation queue (one request in flight) inside a dependency-free Node gateway of 315 lines. Clients that leave while waiting are dropped from the queue and their abort is propagated upstream, so an abandoned request never keeps holding the GPU
  • Results drawn to a share card image with the Canvas API and handed to the Web Share API, with a download fallback on browsers that do not support it
  • Three operator tools - /stats (auto-refreshing every 60 seconds, type distribution), /draw (weighted raffle with cohort filtering, phone-number dedup, masked numbers, and automatic exclusion of previous winners), and /api/health (three status values plus a model warmup)
Product decisions
  • Self-hosted Ollama on a Mac mini instead of a commercial AI API - the reason was API cost. In exchange I took on cold start and concurrency and committed to paying them back with warmup and a serial queue. It was less about saving money than about choosing which problem I wanted
  • Walked the deployment back from Cloudflare Workers to Vercel - the worker subdomain carries my personal account name, and changing the address would have broken other worker addresses with it
  • Cut the palm-reading mode even though it had already passed verification, narrowing to two modes - the remaining time went into the fallback path and the operating procedure for two modes rather than into a third one. The only rationale I wrote down was a note pointing at the git history so it can be brought back
  • Removed every external link (form, group chat, social) and made signup finish inside the app with four fields. Any moment that sends someone out of the app during those thirty seconds at the booth is where they drop
  • Kept the deployed address off the posters so the only way in is the QR code, because it is not a proper domain. And I forced the order of fixing the address, generating the QR, then printing, so the printed material does not move when the infrastructure behind it does
  • Face-reading results keep only the type and the one-line comment, anonymously, and the photo is never collected under any circumstance. That was the decision that makes the "discarded immediately, never stored" notice on the first screen a property of the code rather than a sentence
QA considerations
  • Do concurrent requests on a single GPU all slow down and then time out together? The record in my notes has one of four concurrent requests succeeding, and the original logs were not kept, so I could not reproduce it. I put a FIFO serialisation queue (one in flight) in the gateway, and clients that leave while waiting get dropped from the queue with their abort propagated upstream, which cuts the self-amplifying congestion where abandoned requests keep occupying the GPU
  • If the timeouts on each layer are not aligned you cannot tell which layer cut first. I lined them up: client 85s, deployment function maxDuration 90s, model call 70s plus a 30s retry, warmup 50s, gateway proxy 90s. To catch the case where the tunnel forwards headers and then stalls on the body, I wrote a fetchJsonWithTimeout that counts reading the response body inside the timeout
  • If the model gets evicted from memory the first visitor eats the cold start. /api/health keeps it resident with a keep_alive:-1 warmup request, and the morning-of check is documented as reading three values: ollama_reachable, model_loaded, join_ready
  • Behind a shared campus IP (NAT), rate limiting on IP alone makes the entire booth share one quota. The key is IP plus user agent plus a device id (a localStorage UUID), with 8 face readings, 6 idea appraisals, and 5 signups per minute. Fixing the sweep so it only clears expired keys also fixed a bug where a normal user's counter reset whenever the user agent rotated
  • What breaks on screen if you trust LLM output as-is? The server does type checking, string truncation, score clamping, a grade whitelist (S through D), padding empty arrays with fallbacks, and JSON extraction after stripping code fences. It retries once on a parse failure and never on a connection failure, because the fallback must not be delayed when the AI is down. The prompt carries its own rules: no mocking anyone's appearance, no inventing competitor names that do not exist (the anti-hallucination rule went in as its own commit), and do not refuse a photo just because there is no face in it. Meaningless input (under five characters, one repeated character) is rejected with a 400 before it ever reaches the GPU
  • How far does personal data leak? Photos are discarded without storage or logging and the notice sits on the first screen; gateway logs mask phone numbers; the roster CSV is key-protected; the raffle screen masks numbers on the assumption it will be on a projector; CSV formula injection (=, +, -, @) is escaped; and a corrupted JSONL line is skipped so the rest of the roster survives
  • If the gateway process dies, the AI and signups stop at the same time. It is kept alive by an auto-restart loop and an uncaughtException handler, and the roughly five-minute recovery procedure for a dead tunnel, along with a phone-only contingency for when no laptop is available, lives as a runbook in the same repository as the code

Outcome

Metrics
  • 40 commits, 2026-08-18 to 2026-08-22 (KST), sole author
  • About 3,156 lines across the main sources, 6 API routes, 12 types, a 5-question fallback mode, 5 print poster variants
  • Timeout chain - client 85s / deployment function 90s / model call 70s plus 30s retry / warmup 50s / gateway proxy 90s
  • Rate limits - 8 face readings, 6 idea appraisals, and 5 signups per minute, keyed on IP plus user agent plus device id
  • Zero automated tests. I never added a test framework. Instead I walked the whole flow by hand with the AI switched off, through an AI_PROVIDER=off path and a mock gateway
  • Two documented code reviews - 8 fixes in the first, then an adversarial pass with 7 fixes and 3 findings rejected
  • It ran on the day, but participation was never instrumented as a metric (logging only records successful AI calls, so question-mode fallbacks are not counted): not measured. The load figures in my notes are also unverified, since the original logs were not kept
Result / Learning

Five days to build, running on the day of the event. The clearest outcome is not a feature but the point where a measurement changed the architecture. I found out first that concurrent requests against a 27B vision model all slow down and then mostly time out - by the record in my notes, one of four got through - so I redesigned the gateway from "a server that handles things quickly" into "a line that handles exactly one thing at a time." At an unstaffed booth, a predictable wait beats an unpredictable failure. That is the whole of that decision. The second thing I took away is that a fallback is a contract, not a feature. Question mode runs client-side with no server and no network and decides against the same twelve-type system the AI uses, so the visitor still gets a result card even if the whole backend is down. In exchange I wrote the limit into the README plainly: signups and idea collection only work while the tunnel is alive. If you are going to claim what does not die, you have to name what does. And there is a lesson pointing the other way. Writing a good pre-event risk list and actually executing that list to the end are two different jobs. Everything before the event got done. The items after it did not.

Retrospective
  • I never executed the post-event teardown in my own runbook: shut the tunnel down, clean up the collected data. When I checked two weeks later the gateway was still up. I had written three layers of failure response and then skipped the simplest closing procedure. The gateway itself has no auth and no rate limiting, only a path whitelist blocking the admin API, and that design leaned on the assumption that it is only switched on during the event. Not tearing down broke that assumption. From now on, anything with an end time gets managed as a task with a reminder attached, not as a checklist line.
  • Zero automated tests. With the date fixed I spent the time a test framework would have cost on the fallback path and the runbook, and for a five-day one-off booth I still think that priority was right. But the fallback classifier (twelve types from a seed) and the LLM response normaliser are pure functions with fixed inputs and outputs, so testing them would have cost almost nothing. Not going that far was laziness, not judgement.
  • The scale is small. One booth, one day. The log numbers from it cannot be used as a result and I am not using them. What this project carries is not a number but a procedure: count the failure modes first, then change the structure to match them.
Tech stack
  • Next.js 15 (App Router)
  • React 19
  • TypeScript 5
  • Vercel (route별 maxDuration)
  • Node.js 무의존성 게이트웨이
  • 자체 호스팅 Ollama + gemma3:27b (비전)
  • Cloudflare Tunnel
  • Canvas API + Web Share API
  • qrcode (런타임 QR 생성)
  • JSONL 파일 저장