Case studyPersonal · agent
Quiz Dock
The model writes questions; code grades them.
Study any document with an agent that quizzes you on it. The model writes the questions, code grades the answers.
- Role
- Solo · back-end, AI
- Timeline
- 2026 · since September
- Stack
- TypeScript · Node.js · Anthropic SDK · Zod · tsx
- Status
- Building

Problem
Generic chatbots answer from training data, so for new or niche libraries they confidently mix in facts the docs never state. Quiz Dock stays inside one document: the model reads it through a tool and writes the questions, and code grades the answers.
Reading documentation doesn't mean understanding it.
I wanted a study tool that stays inside one document: it explains what I ask, tests me on it and explains my mistakes from the exact section they come from.
What I built
- 01
Hand-written agent loop
Tool registry, message history, step limit and error-to-model feedback, written from scratch and shared by quiz and chat modes.
- 02
Grounded document chat
The model sees only section titles and reads sections through a tool before it answers.
- 03
Interactive quiz tool
The model writes questions as schema-validated tool input; the user answers by number and code grades them.
- 04
Event-driven output
The core emits typed events instead of printing, so the same loop can drive a CLI or a web UI.
Architecture
Quiz Dock
Code splits a Markdown document into sections. The model decides which sections to read and when to quiz; parsing, validation and grading are deterministic code.
- Input
- Code
- LLM
- Check
01 · Input
Markdown file via CLI flags
User message in the terminal
02 · Code
Split into sections by headings
03 · LLM
Model plans from the section index
readSection returns section text; it answers from that text
04 · LLM
Model writes a quiz as tool input
05 · Check
Zod validates the questions
06 · Code
Code grades the answers
User picks options by number
07 · LLM
Model explains the mistakes
Events rendered in the terminal
Decisions & trade-offs
Chose my own agent loop on the raw SDK over an agent framework
because I wanted direct control over message history, tool results, errors and step limits. Every tool call gets a result; failures go back to the model with is_error.
Chose a section index plus a read tool over the whole document in every request
because history holds only the sections the model opened, so long chats send fewer input tokens and every answer traces to a specific section.
Chose grading answers in code over letting the model judge answers
because comparing the chosen option with the correct index is exact, free and deterministic. The model only explains the mistakes.
AI specifics
- Model
- claude-sonnet-4-6, the default in the LLM call wrapper
- Grounding
- Section index in the system prompt; sections must be read before answering
- Tool use
- readSection, startQuiz, askHuman; parallel tool calls off for quizzes
- Memory
- Full message history per session, in process memory
- Cost controls
- Step limit, max_tokens 1024, token usage summed per run
What I'd do differently
I'd separate the generic loop from the quiz from day one: my first core had the quiz prompt baked in, and I refactored it before adding chat. I'd also design tools that pause the loop earlier; the readline human-in-the-loop works in a CLI but not over HTTP.
Results
The CLI works end to end in two modes: grounded chat about a document, and quizzes the model starts when asked. Malformed tool input is rejected by Zod and the model corrects it on the next step.
Next
In progress: a web version with a Next.js front end and a Hono server, sessions and history in SQLite (later D1), and tools that pause the loop until the user answers in the UI. Then an ingestion agent that reads multi-page docs with subagents.