Case studyPersonal · AI tool04 / 05 in Lab
Job Radar
The model extracts facts. Code makes every decision.
A job-search tool where code makes every decision and the model only reads new text: cost down from ~$15 to an estimated $1-2 a month.
- Role
- Solo · design, front-end, back-end, AI
- Timeline
- 2026 · ~1 month
- Stack
- TypeScript · Hono · Drizzle ORM · Cloudflare Workers · D1 · Anthropic API · Workers AI · React · Mantine · Zod
- Status
- Live

Problem
Model cost was about $15 a month. Now free code filters run first, the model only extracts facts from text it hasn't seen, and the score is plain code: an estimated $1-2 a month.
Job boards optimise for volume, and a list of 40 positions leads to no letters at all. I wanted a short daily list of relevant vacancies and companies, with a memory of who I had already contacted or turned down.
What I built
- 01
Ten-card daily queue
A hard cap of 10 cards a day, and every decision (interesting, contacted, blocked, snoozed) removes the company from later queues.
- 02
Free filters first
Stop words, role, geography and experience checks run before any model call. Rejected vacancies are kept for statistics.
- 03
Catalogues through my browser
A Chrome extension parses catalogues behind Cloudflare in my own session, paginating at a human pace with page limits.
- 04
Faster, personal letters
I choose who to write to and press send myself; drafts start from my own templates, and replies are tracked so nobody gets chased twice.



Architecture
Job Radar
Fetchers and normalised page diffs feed a deterministic filter chain. The model only extracts facts from new text as strict JSON; scoring, dedupe and sending rules are plain code.
- Input
- Code
- LLM
- Check
- Output
01 · Input
ATS APIs, RSS, job boards
Chrome extension scrapes catalogues
02 · Code
Normalise pages, block-level diff
Vacancies and snapshots in D1; stop words, role, geo filters
03 · LLM
Haiku extracts strict JSON
04 · Check
Zod check, score, dedupe
05 · LLM
Sonnet drafts the first paragraph
06 · Check
Validate, else use my template
07 · Output
Daily queue, 10 cards
I review, edit and send by hand
- Reply tracking, Telegram digest
Decisions & trade-offs
Chose deterministic scoring in code over letting the model score vacancies
because the same input always gives the same score, so I can explain any card. The model's relevance is only one small input.
Chose a block-level diff on normalised text over classifying whole pages on every crawl
because stripping dates, counters and tokens keeps hashes stable, so a vacancy is classified once. That is what cut the model spend.
Chose an AI intro validated by code, with a template fallback over sending model text directly
because a made-up fact in a letter to a real person costs more than a generic paragraph. Any failed check or low confidence falls back to my template.
AI specifics
- Models
- claude-haiku-4-5 for extraction, claude-sonnet-5 for drafts; Workers AI Llama as an alternative
- Grounding
- Strict JSON, missing facts must be null, Zod parse, one retry, then manual review
- Tool use
- None: single-shot calls with structured JSON output
- Memory
- Company state and outreach history in the DB; LLM cache keyed by model and prompt version
- Cost controls
- Free filters first, daily call limit, text truncation, cache, AI Gateway
- Tests
- Unit tests with the model call stubbed; no eval set yet
What I'd do differently
At first I called the model before the free filters, and it burned the 500-call daily limit on 498 vacancies, 6 of them relevant. Now I set the filter order and a spend budget before the first model call, not after the first bill.
Results
It runs daily on Cloudflare Workers with D1. Every source adapter has a fixture test, and the model call is stubbed in the unit tests. Model cost fell from about $15 to an estimated $1–2 a month.
Next
Batch API for scheduled classification. Classify only what reaches the queue. A budget ceiling in money, not in call count.