Case studyPersonal · AI bot03 / 05 in Lab
AI Telegram Assistant
Code decides when the model runs.
A group-chat AI bot where code decides when the model runs: an estimated $3-6 a month instead of $75-150.
- Role
- Solo · design, back-end, AI, ops
- Timeline
- 2026 · ~2 months
- Stack
- TypeScript · Cloudflare Workers · grammY · D1 · xAI Grok · Workers AI · Vectorize · R2
- Status
- v2 in progress

Problem
Calling the model on every message in a busy group chat would cost an estimated $75-150 a month. Cheap code checks run first (is the bot addressed, cooldown, chance), so the model runs only when it's needed: about $3-6 a month.
Our friends' group chat produces hundreds of messages a day, and anyone who steps away can't catch up. Existing bots are generic: they forget context and can't take the chat's tone. I wanted one that logs quietly, remembers who is who and answers in the chat's own voice.
What I built
- 01
Summaries on demand
/summary 2h turns a stretch of chat into a retelling built from the transcript, daily digests and stored facts.
- 02
Semantic memory
Facts are tied to real chat members, embedded in Vectorize and recalled by similarity when they matter.
- 03
Plain-language actions
Reminders, todos, stats, translation, search and drawing all work by simply addressing the bot.
- 04
Learns from reactions
Reactions and replies to the bot's messages are distilled nightly into a tone note that steers later replies.
Architecture
AI Telegram Assistant
The webhook returns 200 at once and the work runs in the background. Deterministic code filters, logs and routes; Grok is called only to classify, reply or distil.
- Input
- Code
- LLM
- Check
- Output
01 · Input
Telegram webhook, secret check
Voice, photo, video, docs to text
02 · Code
Log the message in D1
03 · Check
Gates: addressed, cooldown, chance
04 · LLM
Grok classifies intent, argument
05 · Code
Dispatch to a handler
Recall facts, moods, vibe
06 · LLM
Grok replies in persona
07 · Output
Reply sent to the chat
Log reply, reaction scores
Nightly · LLM
Digests, facts, style
Decisions & trade-offs
Chose an AI intent classifier with keyword hints over keyword routing and regex alone
because regex argument extraction was brittle. The model picks the action and a clean argument, and routing falls back to keywords if the AI call fails.
Chose cheap gates before every model call over calling Grok on every message
because my estimate was $75–150 a month ungated versus about $3–6 with gates and prompt caching. Address checks, cooldowns and chance run first.
Chose D1 plus a per-minute cron for reminders over a Durable Object per chat
because it is simpler and needs no paid tier. Repeating reminders just move their due time forward instead of being deleted.
AI specifics
- Model
- grok-4-fast via Cloudflare AI Gateway; separate env vars for chat, reply and profile models
- Grounding
- Recall uses only stored facts above a 0.45 similarity score; chat text is treated as untrusted
- Tool use
- None yet: intent classifier plus a hardcoded dispatch; function calling is next
- Memory
- Facts per member in D1, aliases, profiles, Vectorize (bge-m3), a nightly vibe note
- Cost controls
- Deterministic gates, static-first prompts for caching, classifier capped at 200 tokens
What I'd do differently
It grew to about 6,400 lines before it got structure: two files took on too much and the persona lived in three prompts. v2 starts from tests and a config module. An early prompt-injection loop taught me to treat chat history as untrusted input, not as instructions.
Results
Deployed and running in a live group chat: 14 D1 migrations, about 35 modules, and most media handled by Workers AI (Whisper, captions, embeddings) so Grok is kept for the work where quality matters.
Next
Split the two largest modules and add lint and tests. Replace the classifier and switch with Grok function calling. Add reranking and hybrid search to memory recall.