Skip to content

Case studyPersonal · AI product

Money Track

The model never computes a number.

Role
Solo · design, front-end, back-end, AI
Timeline
2026 · ~3 months
Stack
TypeScript · React · RTK Query · Cloudflare Workers · Durable Objects (SQLite) · D1 · Hono · Anthropic API · TypeSafe Jev
Status
Live · open source
money.italik.dev/demo
Money Track: screenshot

Problem

The AI advisor invented numbers on real bank data, and better prompts didn't stop it. In a finance app, a confidently wrong number is worse than a missing feature. So I moved every calculation into SQL, added a check that blocks any figure the database didn't produce, and let rules categorise spending before a model does: ~97% accuracy at ~80% lower cost.

I wanted one place for my own money: Monobank, cash and subscriptions, with an assistant I could ask "can I afford this?" Bank apps show transactions but don't explain them, and general chatbots will state a plausible total for anything.

What I built

  1. 01

    Grounded AI advisor

    Chat, advice and reports read the same snapshot the screens use; a check drops any figure or date the data doesn't contain.

  2. 02

    Deterministic-first categorisation

    Aliases, subscriptions, merchant consensus and MCC rules run first; a model only sees rows those rules could not settle.

  3. 03

    Your ledger over MCP

    A read-only MCP server with its own OAuth 2.1 server, so Claude connects with just a URL and a consent screen.

  4. 04

    Throwaway public demo

    /demo creates a private sandbox with six months of seeded data that resets after 24 hours.

money.italik.dev/demo
Money Track: AI chat
Money Track on mobile: AI chat
Money Track on mobile: AI overview
money.italik.dev/demo
Money Track: AI financial overview
Money Track on mobile: Dashboard
Money Track on mobile: Subscriptions
money.italik.dev/demo
Money Track: Monthly trends
money.italik.dev/demo
Money Track: Subscriptions

Architecture

The model never touches the data.

Every number is computed by deterministic code inside the user's own Durable Object. Models only classify rows the rules can't settle and phrase answers, and every answer passes a grounding check before it renders.

  • Code
  • LLM
  • Check
  • Output
  • External

System map. Clients: Web app, Telegram bot, MCP client. Cloudflare edge · deterministic: Worker, OAuth 2.1 server, Bank normaliser, User's Durable Object, Canonical SQL, Grounding check. External: Monobank API. LLM zone · read-only, no writes: Jev → Haiku, Claude. Connections: Web app to Worker; Telegram bot to Worker; MCP client to Worker; Monobank API to Bank normaliser (webhook); Bank normaliser to User's Durable Object (upsert); Worker to User's Durable Object (routes); Worker to OAuth 2.1 server (MCP auth); User's Durable Object to Jev → Haiku (unknown rows); Jev → Haiku to User's Durable Object (category); User's Durable Object to Canonical SQL (read); Canonical SQL to Claude (snapshot); Canonical SQL to Claude (+ tools); Claude to Grounding check (draft answer); Grounding check to MCP client (checked answer).

One question, end to end

“Can I afford a $400 laptop this month?”

  1. 01Code

    Worker checks the session and routes to the user's own Durable Object

  2. 02Code

    Canonical SQL builds the snapshot the screens use: balances, burn, budgets

  3. 03LLM

    Claude gets the snapshot plus read-only tools: query_spend, find_transactions

  4. 04LLM

    It drafts the answer and explains it. It never computes a total itself

  5. 05Check

    Grounding check drops any figure or date the snapshot doesn't contain

  6. 06Code

    The checked answer streams to the web app, Telegram or an MCP client

  • Every figure comes from SQL.

    Checked before it renders

  • AI changes are logged and revertible.

    No silent writes

  • 35 / 36 on held-out merchants.

    npm run eval · $0.03 per run

  • Public demo capped at $1 a day.

    Cost is a feature

Decisions & trade-offs

  • Chose one Durable Object per user over user_id filters on shared tables

    because with dozens of canonical queries, one missing filter would leak another user's data. Physical isolation keeps the queries unchanged, and forwarding the whole request into the object is one network hop.

  • Chose my own OAuth 2.1 authorization server over delegating MCP auth to Google

    because Google knows who the user is, not what access this ledger grants. Tokens are bound to this server's audience, as the MCP spec requires, with PKCE S256 and rotating refresh tokens.

  • Chose a Jev → Haiku cascade over Claude categorising every row

    because on cases written before the code, Jev files a row only at confidence ≥ 0.8 and passes the rest to Haiku. That kept Haiku's accuracy at about a fifth of the cost.

AI specifics

Models
claude-haiku-4-5, claude-sonnet-5, claude-opus-4-8, jev-latest; routed per task
Grounding
Figures or dates not in the canonical snapshot are dropped before rendering
Tool use
query_spend, find_transactions, list_categories, remember_fact; read-only MCP server
Memory
Server-side chat history, user-confirmed facts, advice history to avoid repeats
Cost controls
Per-task model routing, 1h prompt cache on bulk enrich, demo capped at $1/day
Evals
npm run eval: 168 real bank descriptions plus 36 held-out merchants

What I'd do differently

I'd write the eval set before the first prompt. A judgment model tuned on my dataset scored 98.9%, but only 72% on 36 merchants written afterwards. The cause was a hand-condensed copy of the category guide: a second definition of the same thing, a pattern I've since removed across the codebase.

Results

Live in production with open Google sign-up and a public demo. `npm run check` runs lint and the tests, including golden analytics snapshots. Latest held-out eval: Haiku gets 35 of 36 root categories right, for $0.03 per run.

Next

Connect a second MCP client (ChatGPT) to the live server. Resolve sole-trader tax edge cases and verify the PrivatBank live API. Re-measure search reranking before wiring it in.