Case studyPersonal · AI feature05 / 05 in Lab
Ask-about-me chat
It only knows what this site says.
The AI chat on this site. It answers only from my real content, for under a cent a question.
- Role
- Solo · design, front-end, back-end, AI
- Timeline
- 2026
- Stack
- Next.js · TypeScript · Anthropic API · Upstash Redis · Zod
- Status
- Live

Problem
A visitor has one specific question and little time, and a generic chatbot would happily invent the answer. This chat answers only from the site's content, passes its whole eval set, and each question costs under a cent.
A recruiter has 30 to 90 seconds and one specific question: has he used Stripe, can he start soon, does he know Python? Scanning a whole portfolio for that is slow, and a generic chatbot would happily invent the answer.
What I built
- 01
One source of truth
The knowledge base is built from the same Markdown files that render the site, so the chat can't drift from the pages.
- 02
Answers with sources
Every answer ends with the files it used; the server turns them into source chips and a link to the case study.
- 03
Guardrails
It declines salary, personal topics, NDA details and prompt injection politely, and answers in the visitor's language.
- 04
Hard spend limits
A per-IP rate limit, a daily token budget and a request timeout, with a friendly fallback to my email.
Architecture
Ask-about-me chat
Deterministic code builds the knowledge base and enforces every limit. The model sees one cached prompt and the visitor's last few messages, and only writes the answer.
- Input
- Code
- LLM
- Output
01 · Input
Visitor asks a question
02 · Code
Zod validation and length limits
Rate limit and daily budget in Redis
03 · LLM
The model answers from the cached prompt
04 · Code
Sources line stripped from the stream
Spend recorded per day
05 · Output
Streamed answer, source chips
Build time · Code
Project, experience and about-me files → knowledge base, drafts skipped
Decisions & trade-offs
Chose the whole knowledge base in a cached prompt over embeddings and retrieval
because it is about 19k tokens, so everything fits. Prompt caching makes repeat reads cost a tenth, and there is no retrieval step that could miss the right fact.
Chose building the knowledge from site content over a separate hand-written chat document
because one definition of every fact. When I edit a project page, the chat knows it on the next build.
Chose a budget in input-token equivalents over counting requests
because output, cache writes and cache reads cost different amounts. One weighted number caps the real daily spend, whichever model runs: about $3 a day on Sonnet 5.5.
AI specifics
- Model
- The one I pick through CHAT_MODEL, no code change; now Claude Sonnet 5.5
- Grounding
- Knowledge base only; unknown answers point to my email
- Tool use
- None: one streamed call per question
- Memory
- Last 10 messages of the current page; nothing stored on the server
- Cost controls
- Prompt caching (~16 fresh input tokens per question), 20 requests per 10 min, daily budget
- Evals
- 29 golden questions: facts, refusals, NDA, injection, language, client work; all pass on Sonnet 5.5
What I'd do differently
I'd rely less on rules in the prompt. Told not to use Markdown, the model still formatted answers now and then, so the UI now renders paragraphs and bold instead of fighting it. Asked to refuse a prompt injection, it quoted its own rules back; a neutral refusal fixed that.
Results
Works end to end with streaming, source chips, a floating drawer on every page and the section on the home page. The eval set passes 29 of 29, and the prompt cache holds: after the first request, each question reads about 19k cached tokens and costs well under a cent.
Next
Log anonymous questions to see what recruiters actually ask. Move to embeddings only if the knowledge base outgrows the prompt.