Skip to content

Case studyPersonal · AI feature

Ask-about-me chat

It only knows what this site says.

Role
Solo · design, front-end, back-end, AI
Timeline
2026
Stack
Next.js · TypeScript · Anthropic API · Upstash Redis · Zod
Status
Live
italik.dev/#ask
Ask-about-me chat: screenshot

Problem

A visitor has one specific question and little time, and a generic chatbot would happily invent the answer. This chat answers only from the site's content, passes its whole eval set, and each question costs under a cent.

A recruiter has 30 to 90 seconds and one specific question: has he used Stripe, can he start soon, does he know Python? Scanning a whole portfolio for that is slow, and a generic chatbot would happily invent the answer.

What I built

  1. 01

    One source of truth

    The knowledge base is built from the same Markdown files that render the site, so the chat can't drift from the pages.

  2. 02

    Answers with sources

    Every answer ends with the files it used; the server turns them into source chips and a link to the case study.

  3. 03

    Guardrails

    It declines salary, personal topics, NDA details and prompt injection politely, and answers in the visitor's language.

  4. 04

    Hard spend limits

    A per-IP rate limit, a daily token budget and a request timeout, with a friendly fallback to my email.

Architecture

Deterministic code builds the knowledge base and enforces every limit. The model sees one cached prompt and the visitor's last few messages, and only writes the answer.

  • Input
  • Code
  • LLM
  • Output
  1. 01 · Input

    Visitor asks a question

  2. 02 · Code

    Zod validation and length limits

    Rate limit and daily budget in Redis

  3. 03 · LLM

    The model answers from the cached prompt

  4. 04 · Code

    Sources line stripped from the stream

    Spend recorded per day

  5. 05 · Output

    Streamed answer, source chips

Build time · Code

Project, experience and about-me files → knowledge base, drafts skipped

Decisions & trade-offs

  • Chose the whole knowledge base in a cached prompt over embeddings and retrieval

    because it is about 19k tokens, so everything fits. Prompt caching makes repeat reads cost a tenth, and there is no retrieval step that could miss the right fact.

  • Chose building the knowledge from site content over a separate hand-written chat document

    because one definition of every fact. When I edit a project page, the chat knows it on the next build.

  • Chose a budget in input-token equivalents over counting requests

    because output, cache writes and cache reads cost different amounts. One weighted number caps the real daily spend, whichever model runs: about $3 a day on Sonnet 5.5.

AI specifics

Model
The one I pick through CHAT_MODEL, no code change; now Claude Sonnet 5.5
Grounding
Knowledge base only; unknown answers point to my email
Tool use
None: one streamed call per question
Memory
Last 10 messages of the current page; nothing stored on the server
Cost controls
Prompt caching (~16 fresh input tokens per question), 20 requests per 10 min, daily budget
Evals
29 golden questions: facts, refusals, NDA, injection, language, client work; all pass on Sonnet 5.5

What I'd do differently

I'd rely less on rules in the prompt. Told not to use Markdown, the model still formatted answers now and then, so the UI now renders paragraphs and bold instead of fighting it. Asked to refuse a prompt injection, it quoted its own rules back; a neutral refusal fixed that.

Results

Works end to end with streaming, source chips, a floating drawer on every page and the section on the home page. The eval set passes 29 of 29, and the prompt cache holds: after the first request, each question reads about 19k cached tokens and costs well under a cent.

Next

Log anonymous questions to see what recruiters actually ask. Move to embeddings only if the knowledge base outgrows the prompt.