Skip to content

Case studyPersonal · AI bot

AI Telegram Assistant

Code decides when the model runs.

Role
Solo · design, back-end, AI, ops
Timeline
2026 · ~2 months
Stack
TypeScript · Cloudflare Workers · grammY · D1 · xAI Grok · Workers AI · Vectorize · R2
Status
v2 in progress
t.me · group chat
AI Telegram Assistant: screenshot

Problem

Calling the model on every message in a busy group chat would cost an estimated $75-150 a month. Cheap code checks run first (is the bot addressed, cooldown, chance), so the model runs only when it's needed: about $3-6 a month.

Our friends' group chat produces hundreds of messages a day, and anyone who steps away can't catch up. Existing bots are generic: they forget context and can't take the chat's tone. I wanted one that logs quietly, remembers who is who and answers in the chat's own voice.

What I built

  1. 01

    Summaries on demand

    /summary 2h turns a stretch of chat into a retelling built from the transcript, daily digests and stored facts.

  2. 02

    Semantic memory

    Facts are tied to real chat members, embedded in Vectorize and recalled by similarity when they matter.

  3. 03

    Plain-language actions

    Reminders, todos, stats, translation, search and drawing all work by simply addressing the bot.

  4. 04

    Learns from reactions

    Reactions and replies to the bot's messages are distilled nightly into a tone note that steers later replies.

Architecture

The webhook returns 200 at once and the work runs in the background. Deterministic code filters, logs and routes; Grok is called only to classify, reply or distil.

  • Input
  • Code
  • LLM
  • Check
  • Output
  1. 01 · Input

    Telegram webhook, secret check

    Voice, photo, video, docs to text

  2. 02 · Code

    Log the message in D1

  3. 03 · Check

    Gates: addressed, cooldown, chance

  4. 04 · LLM

    Grok classifies intent, argument

  5. 05 · Code

    Dispatch to a handler

    Recall facts, moods, vibe

  6. 06 · LLM

    Grok replies in persona

  7. 07 · Output

    Reply sent to the chat

    Log reply, reaction scores

Nightly · LLM

Digests, facts, style

Decisions & trade-offs

  • Chose an AI intent classifier with keyword hints over keyword routing and regex alone

    because regex argument extraction was brittle. The model picks the action and a clean argument, and routing falls back to keywords if the AI call fails.

  • Chose cheap gates before every model call over calling Grok on every message

    because my estimate was $75–150 a month ungated versus about $3–6 with gates and prompt caching. Address checks, cooldowns and chance run first.

  • Chose D1 plus a per-minute cron for reminders over a Durable Object per chat

    because it is simpler and needs no paid tier. Repeating reminders just move their due time forward instead of being deleted.

AI specifics

Model
grok-4-fast via Cloudflare AI Gateway; separate env vars for chat, reply and profile models
Grounding
Recall uses only stored facts above a 0.45 similarity score; chat text is treated as untrusted
Tool use
None yet: intent classifier plus a hardcoded dispatch; function calling is next
Memory
Facts per member in D1, aliases, profiles, Vectorize (bge-m3), a nightly vibe note
Cost controls
Deterministic gates, static-first prompts for caching, classifier capped at 200 tokens

What I'd do differently

It grew to about 6,400 lines before it got structure: two files took on too much and the persona lived in three prompts. v2 starts from tests and a config module. An early prompt-injection loop taught me to treat chat history as untrusted input, not as instructions.

Results

Deployed and running in a live group chat: 14 D1 migrations, about 35 modules, and most media handled by Workers AI (Whisper, captions, embeddings) so Grok is kept for the work where quality matters.

Next

Split the two largest modules and add lint and tests. Replace the classifier and switch with Grok function calling. Add reranking and hybrid search to memory recall.