Skip to content

Case studyPersonal · agent

Quiz Dock

The model writes questions; code grades them.

Role
Solo · back-end, AI
Timeline
2026 · since September
Stack
TypeScript · Node.js · Anthropic SDK · Zod · tsx
Status
Building
Quiz Dock: screenshot

Problem

Generic chatbots answer from training data, so for new or niche libraries they confidently mix in facts the docs never state. Quiz Dock stays inside one document: the model reads it through a tool and writes the questions, and code grades the answers.

Reading documentation doesn't mean understanding it.

I wanted a study tool that stays inside one document: it explains what I ask, tests me on it and explains my mistakes from the exact section they come from.

What I built

  1. 01

    Hand-written agent loop

    Tool registry, message history, step limit and error-to-model feedback, written from scratch and shared by quiz and chat modes.

  2. 02

    Grounded document chat

    The model sees only section titles and reads sections through a tool before it answers.

  3. 03

    Interactive quiz tool

    The model writes questions as schema-validated tool input; the user answers by number and code grades them.

  4. 04

    Event-driven output

    The core emits typed events instead of printing, so the same loop can drive a CLI or a web UI.

Architecture

Code splits a Markdown document into sections. The model decides which sections to read and when to quiz; parsing, validation and grading are deterministic code.

  • Input
  • Code
  • LLM
  • Check
  1. 01 · Input

    Markdown file via CLI flags

    User message in the terminal

  2. 02 · Code

    Split into sections by headings

  3. 03 · LLM

    Model plans from the section index

    readSection returns section text; it answers from that text

  4. 04 · LLM

    Model writes a quiz as tool input

  5. 05 · Check

    Zod validates the questions

  6. 06 · Code

    Code grades the answers

    User picks options by number

  7. 07 · LLM

    Model explains the mistakes

    Events rendered in the terminal

Decisions & trade-offs

  • Chose my own agent loop on the raw SDK over an agent framework

    because I wanted direct control over message history, tool results, errors and step limits. Every tool call gets a result; failures go back to the model with is_error.

  • Chose a section index plus a read tool over the whole document in every request

    because history holds only the sections the model opened, so long chats send fewer input tokens and every answer traces to a specific section.

  • Chose grading answers in code over letting the model judge answers

    because comparing the chosen option with the correct index is exact, free and deterministic. The model only explains the mistakes.

AI specifics

Model
claude-sonnet-4-6, the default in the LLM call wrapper
Grounding
Section index in the system prompt; sections must be read before answering
Tool use
readSection, startQuiz, askHuman; parallel tool calls off for quizzes
Memory
Full message history per session, in process memory
Cost controls
Step limit, max_tokens 1024, token usage summed per run

What I'd do differently

I'd separate the generic loop from the quiz from day one: my first core had the quiz prompt baked in, and I refactored it before adding chat. I'd also design tools that pause the loop earlier; the readline human-in-the-loop works in a CLI but not over HTTP.

Results

The CLI works end to end in two modes: grounded chat about a document, and quizzes the model starts when asked. Malformed tool input is rejected by Zod and the model corrects it on the next step.

Next

In progress: a web version with a Next.js front end and a Hono server, sessions and history in SQLite (later D1), and tools that pause the loop until the user answers in the UI. Then an ingestion agent that reads multi-page docs with subagents.