8 original lessons3 linked videos22 recall cardsPython 3.12+ · your laptop
This is a deeper coding-agent project alongside our Microsoft and OpenAI paths. The upstream implementation uses Pydantic AI with Gemini, OpenRouter or Modal. The engineering ideas transfer; this is not an OpenAI or Microsoft SDK tutorial.
Articles and videos stay with Decoding AI. Recall & Build adds original practice tasks, recall cards and five self-reported project checkpoints. The evidence ZIP contains worksheets and source links; clone the upstream repository for runnable code.
BEFORE YOUR NEXT FEATURE / SPEC-DRIVEN DEVELOPMENT
Specify before you build.
Pair the harness course with Paul Everitt’s Spec-Driven Development with Coding Agents, made with DeepLearning.AI and JetBrains. Use a specification to make the expected outcome reviewable before a coding agent changes your project.
The original AgentClinic snapshots use Node/TypeScript and include Hono examples. Our practice below applies the workflow to your existing Python project. This is an original companion exercise, not a reproduction of the provider’s quiz or an official course completion.
Work through the five-part practice exercise
1. Write the project rules
Use your current support-agent project. Record its purpose, permitted data, non-goals, chosen stack and next small feature. Separate observed behaviour from intended behaviour in an existing codebase.
Evidence: A short constitution that a teammate can challenge before implementation.
Choose a narrow change: prevent duplicate ticket creation after a retry. Define caller identity, operation IDs, expected results, conflict behaviour and exclusions. Resolve open questions before asking an agent to implement it.
Evidence: requirements.md plus a small implementation plan, with stable acceptance-criterion IDs.
Connect each criterion to a test. Check repeated calls, changed payloads, missing identity and another tenant. Review the diff and comments against the specification; a green test suite can still miss a requirement.
Evidence: A requirement-to-test table, actual test output and one human-reviewed diff.
Introduce a conflicting requirement, such as operation-ID expiry. Document the decision and update affected specs and tests. Do not weaken a criterion just to accept the current code.
Evidence: A decision record describing what changed, why, and which earlier evidence is now stale.
After completing the loop, write a small workflow for a future feature: inputs, outputs, evidence checks and stopping conditions. Try it in a fresh coding-agent session. Keep tool-specific commands in an adapter or clearly labeled note.
Evidence: A reviewed workflow file plus a handoff attempt showing which assumptions needed clarification.
Use SPEC-WORKSHEET.md in the evidence kit. Record the resulting test evidence under the existing Test checkpoint and the decision record under Evidence. These are self-reported checkpoints; watching the course does not complete them.
Checked 2026-10-05 at revision 4a9c2ac8cdac. We inspected the public materials; we did not run AgentClinic or complete the provider’s graded assessment. Original post.
Start here: workspace, requirements and costs
Use a fresh, disposable workspace. You need Python 3.12+, uv and git; Windows users follow the upstream WSL2 guidance. Docker is needed for the local sandbox and Kitaru replay work.
TERMINAL
git clone https://github.com/decodingai-magazine/building-a-coding-agent-from-scratch-course.git
cd building-a-coding-agent-from-scratch-course
git checkout 3f8d219961498f939b29c506d90d0bf36e5f91d2
uv sync --locked
uv run decode --help
This pins the inspected revision. uv sync --locked installs from its lockfile; use the upstream install guide if you also want its Git hooks and global CLI.
Before running an agent: choose and configure one supported provider in a private environment file. The repository defaults to host execution (SANDBOX_MODE=none). Read the sandbox guide before executing agent-written commands; its Docker setup is described as protection against accidents, not hostile code. Keep repository-write credentials out of the first exercise.
Live model calls, cloud runs and evaluations may cost money. Free tiers and credits depend on your account. The author estimates 4–8 hours for a first pass; allow extra time for setup, experiments and revision.
We inspected the articles, code and commands at the pinned revision. We have not run the upstream test suite or a live Decode session. Record your actual results rather than treating the reference as verified execution.
Inspect local and remote entrypoints. First test the headless contract locally. Optionally follow the Modal guide with your own account and a bounded task.
Keep this evidenceLocal test results; for an actual cloud run, add its run ID, status, cost and cleanup record. A design-only review is not deployment evidence.
Explain it before you move on
What changes when an agent runs without a terminal UI?
Then challenge yourself: Does a successful deployment prove the agent completed a task?
Inspect one benchmark verifier and its oracle sanity check. Introduce a broken solution that should fail. Keep infrastructure errors separate from agent failures.
Inspect recording failure handling, then follow the Kitaru guide on a small disposable task. Change one variable and inspect the replay tool policy before executing it.
Keep this evidenceBaseline and replay identifiers, changed setting, tool policy, result and a regression case. Label inspection-only work separately.
Explain it before you move on
Why is a replay not automatically a fresh end-to-end test?
Then challenge yourself: What happens when a replay needs an unrecorded tool result?
Sources you can inspect.
Course by Paul Iusztin / Decoding AI, introduced in this LinkedIn post. Upstream code is Apache-2.0; the articles and videos remain on their original sites. This companion is independently created and is not an endorsement by the course author.
Source snapshot: 3f8d21996149 · checked 2026-10-05. The outline links three distinct videos and marks the later videos coming soon, despite broader headline video counts. Availability can change.
Remote deployment is optional. A worksheet, mocked test, live model run and cloud deployment are different evidence. State which you completed.