An interactive guide
Inside the agent loop
Every coding agent, Claude Code included, runs on one simple idea: a program asks an AI model what to do next, does it, and asks again. This page builds that up one picture at a time. Everything with a button is meant to be clicked.
I made this to learn how agents really work, while going through Geoffrey Huntley’s free workshop, how to build a coding agent.
The one analogy to remember
Picture a brilliant expert on the phone who forgets everything the moment you hang up. They can’t see your computer. So every time you call, you read them the whole conversation so far. They tell you one thing to do: “open math.js and read it to me.” You do it, then call back with the result. The expert is the model. You, keeping the notes and doing the legwork, are the harness.
You
Type what you want done.
Harness
The program. Keeps the notes, runs the tools, loops. Claude Code is a harness.
Model
The brain. Reads the notes and decides the next step. Can’t touch anything itself.
Tools
The hands. Read a file, edit a file, run a command, search the code.
Part 1 · The loop
It’s a loop that keeps asking “what next?”
Watch an agent fix a real bug, one step at a time. Keep an eye on the notes on the right. They’re the agent’s only memory, and the whole list gets sent to the model on every call.
Fix a failing test
click Next, or use ← and →Step 1
notes = [instructions, tool_menu]notes.add(your_request)loop: reply = model(notes) # sends ALL the notes notes.add(reply) if reply has no tool request: stop result = run_tool(reply) # the harness does the work notes.add(result)show(reply)
The notes
size in tokensThe box under the diagram is the whole agent, in plain pseudo-code. Huntley’s workshop turns it into about 300 lines of real code, and most of those lines are the tools themselves. The model does the thinking. The harness just keeps the list and runs what it’s asked to run.
Part 2 · Tokens
The model reads tokens, not words
Before the model sees anything, the text gets chopped into chunks called tokens, and each chunk becomes a number. Short common words are usually one token. Long or unusual words get split into pieces. Type below, or pick an example.
Chop it into tokens
· marks a space that’s part of the tokenThis is a simplified splitter for learning. Real tokenizers learn their chunks from huge amounts of text, so the exact splits and numbers differ by model. The idea is the same.
It’s the unit for everything
The size limit, the price and the speed are all counted in tokens. Rough rule: one token is about four characters of English.
Code costs more than prose
Symbols, odd names and spacing tend to split into more pieces, so a line of code often uses more tokens than a line of English.
It explains odd mistakes
Ask how many r’s are in “strawberry.” The model sees a few chunks, not ten separate letters, so counting letters is surprisingly hard.
Part 3 · How the model writes
It writes one token at a time
A model is autocomplete at enormous scale. It reads every token so far, scores every possible next token, picks one, adds it to the end, and repeats. That process is called inference. Watch it decide how to reply to “fix the failing test.”
Watch the model decide
top 5 guesses shownLow temperature makes the top guess dominate. High temperature flattens the bars, so picks get more random.
A few steps are merged so the demo stays short. Probabilities here are made up to show the shape.
Reading is fast. Writing is slow.
The model takes in your whole prompt in one parallel pass. That’s the short pause before the first word appears. Writing is different: each new token needs its own pass, so the reply comes out one token at a time. That’s the typing speed you see, and it’s part of why output tokens cost more than input tokens.
Reuse what you’ve already read
While reading, the model keeps scratch notes on every token (the KV cache). If your next request starts with exactly the same tokens, the provider can reuse that work. That’s prompt caching: faster and much cheaper. It only works up to the first difference, which is why agents try hard to only add to the end of their notes.
Part 4 · The context window
Everything has to fit in one box
The context window is the most the model can take in on a single call, counted in tokens. Every note has to fit: the instructions, the tool menu, your messages, file contents, command output. Think of it as a desk. The more you pile on, the easier it is to miss something. Past the edge, nothing fits at all.
Pile more on
Keep it under control
Try “Print a giant log” a few times, then fix it with the buttons above.
What’s in the box
Sizes are typical ballparks, and window size varies by model. The “getting crowded” warnings are a rule of thumb, not a hard line.
Part 5 · Toy vs. the real thing
The loop is easy. Making it good is the job.
Huntley’s point is that the core of an agent like Claude Code is small, and that’s true. What separates a weekend project from a real product is how it handles everything that goes wrong. Most of these follow straight from Parts 2 to 4.
The box fills up
Long sessions bury the model in old output, and it starts missing things.
Compaction. Summarize the old notes into a short recap and keep going with a mostly empty box.
Side quests clutter the notes
Searching 40 files leaves 40 files of leftovers in the context.
Subagents. A helper with its own fresh box does the digging and hands back only the answer.
One command prints a novel
A single noisy test run can eat a big slice of the box.
Output caps. Trim long output before it goes in the notes, keeping the parts that matter, like the errors at the end.
Re-sending everything adds up
In Part 1, five calls re-read about 25,000 tokens to fix a one-line bug.
Prompt caching. Keep the start of the notes identical and only add to the end, so the provider reuses its work.
Edits land in the wrong place
Line numbers shift as you edit. Rewriting whole files drops code by accident.
Find-and-replace edits. The exact old text must appear once, or the edit is refused and the model tries again.
Finding the right code
The model can’t see your codebase until something finds the right files.
Plain text search. Claude Code skips building an index and searches with ripgrep. Code is full of exact names, so simple search works and never goes stale. Some tools, like Cursor, add a semantic index on top.
Every app needs its own tools
Writing a custom tool menu for each agent doesn’t scale.
MCP. A standard plug for tools, so one tool server works in Claude, ChatGPT and other agents. Gonkbot and the drone lab on this site are built this way.
The model runs whatever it decides to
Deleting a folder is just another tool request.
Permissions. Ask before risky commands, allow safe ones, and sandbox the rest.
Small prompt changes break things
The model’s answers vary from run to run.
Evals. Test prompts and tool descriptions like code, on a set of real tasks, before shipping a change.
Part 6 · In five sentences
The whole thing, from the bottom up
If you can explain these five things in order, you understand how a coding agent works.
Tokens. Text gets chopped into numbered chunks. The limit, the price and the speed are all counted in tokens.
Inference. The model predicts the next token, adds it, and repeats. Reading the prompt is one fast parallel pass. Writing is one token at a time.
The context window. Everything the model sees on a call has to fit in a fixed token budget, and it gets less reliable as the box fills up.
Tools. A menu of tools is written into the prompt. The model writes a request in a set format. The harness runs it and pastes the result back. MCP is a standard way to plug in more.
The loop. Repeat until the model answers in plain text. The model remembers nothing, so the harness re-sends everything each time. That’s why caching, compaction and subagents exist.
Check yourself
Built from Geoffrey Huntley’s free workshop, how to build a coding agent, and his post ghuntley.com/replaced.
Token counts, context sizes and probabilities on this page are illustrative ballparks chosen to show the shape of things, not measurements from a specific model.