recall&build
Browse freely. Sign in to save project checkpoints and card reviews.Sign in

DEEPER BUILD / HARNESS ENGINEERING

Build the runtime
around the model.

Follow Paul Iusztin’s Decode course. Trace the agent loop, change the code, test the boundaries and explain your evidence.

Evidence kit
8 original lessons3 linked videos22 recall cardsPython 3.12+ · your laptop

This is a deeper coding-agent project alongside our Microsoft and OpenAI paths. The upstream implementation uses Pydantic AI with Gemini, OpenRouter or Modal. The engineering ideas transfer; this is not an OpenAI or Microsoft SDK tutorial.

Articles and videos stay with Decoding AI. Recall & Build adds original practice tasks, recall cards and five self-reported project checkpoints. The evidence ZIP contains worksheets and source links; clone the upstream repository for runnable code.

Open the upstream course Practice the harness cards →

BEFORE YOUR NEXT FEATURE / SPEC-DRIVEN DEVELOPMENT

Specify before you build.

Pair the harness course with Paul Everitt’s Spec-Driven Development with Coding Agents, made with DeepLearning.AI and JetBrains. Use a specification to make the expected outcome reviewable before a coding agent changes your project.

The original AgentClinic snapshots use Node/TypeScript and include Hono examples. Our practice below applies the workflow to your existing Python project. This is an original companion exercise, not a reproduction of the provider’s quiz or an official course completion.

Work through the five-part practice exercise

1. Write the project rules

Use your current support-agent project. Record its purpose, permitted data, non-goals, chosen stack and next small feature. Separate observed behaviour from intended behaviour in an existing codebase.

Evidence: A short constitution that a teammate can challenge before implementation.

Inspect the original example ↗

2. Specify one feature

Choose a narrow change: prevent duplicate ticket creation after a retry. Define caller identity, operation IDs, expected results, conflict behaviour and exclusions. Resolve open questions before asking an agent to implement it.

Evidence: requirements.md plus a small implementation plan, with stable acceptance-criterion IDs.

Inspect the original example ↗

3. Implement and verify independently

Connect each criterion to a test. Check repeated calls, changed payloads, missing identity and another tenant. Review the diff and comments against the specification; a green test suite can still miss a requirement.

Evidence: A requirement-to-test table, actual test output and one human-reviewed diff.

Inspect the original example ↗

4. Replan when evidence changes

Introduce a conflicting requirement, such as operation-ID expiry. Document the decision and update affected specs and tests. Do not weaken a criterion just to accept the current code.

Evidence: A decision record describing what changed, why, and which earlier evidence is now stale.

Inspect the original example ↗

5. Package the repeatable workflow

After completing the loop, write a small workflow for a future feature: inputs, outputs, evidence checks and stopping conditions. Try it in a fresh coding-agent session. Keep tool-specific commands in an adapter or clearly labeled note.

Evidence: A reviewed workflow file plus a handoff attempt showing which assumptions needed clarification.

Inspect the original example ↗

Use SPEC-WORKSHEET.md in the evidence kit. Record the resulting test evidence under the existing Test checkpoint and the decision record under Evidence. These are self-reported checkpoints; watching the course does not complete them.

Open the ticket lab to apply it →

Checked 2026-10-05 at revision 4a9c2ac8cdac. We inspected the public materials; we did not run AgentClinic or complete the provider’s graded assessment. Original post.

Start here: workspace, requirements and costs

Use a fresh, disposable workspace. You need Python 3.12+, uv and git; Windows users follow the upstream WSL2 guidance. Docker is needed for the local sandbox and Kitaru replay work.

TERMINAL
git clone https://github.com/decodingai-magazine/building-a-coding-agent-from-scratch-course.git
cd building-a-coding-agent-from-scratch-course
git checkout 3f8d219961498f939b29c506d90d0bf36e5f91d2
uv sync --locked
uv run decode --help

This pins the inspected revision. uv sync --locked installs from its lockfile; use the upstream install guide if you also want its Git hooks and global CLI.

Before running an agent: choose and configure one supported provider in a private environment file. The repository defaults to host execution (SANDBOX_MODE=none). Read the sandbox guide before executing agent-written commands; its Docker setup is described as protection against accidents, not hostile code. Keep repository-write credentials out of the first exercise.

Live model calls, cloud runs and evaluations may cost money. Free tiers and credits depend on your account. The author estimates 4–8 hours for a first pass; allow extra time for setup, experiments and revision.

We inspected the articles, code and commands at the pinned revision. We have not run the upstream test suite or a live Decode session. Record your actual results rather than treating the reference as verified execution.

Upstream setup guide · Sandbox configuration

LESSON 01

Map the harness

Separate the model loop from the runtime that controls it.

Lessons 1 and 2 share the first video.

Your build task

Trace one request from the terminal interface through the loop, permission gate and tool result. Draw where state and decisions live.

Inspect the code and run a focused checksrc/decode/agent/loop.py
TERMINAL
uv run decode --help
Follow the upstream run guide

Keep this evidenceA one-page diagram and one trace annotated with the responsible modules.

Explain it before you move on

What belongs to the harness rather than the model?

Then challenge yourself: A stronger model passes more examples. Can you remove runtime checks?

LESSON 02

Follow the model–tool loop

Understand tool requests, observations and termination.

Lessons 1 and 2 share the first video.

Your build task

Inspect the loop and its unit tests. Add a case where the model requests another tool after a tool error; state how your version must stop.

Inspect the code and run a focused checksrc/decode/agent/loop.py
TERMINAL
uv run pytest tests/unit/decode/agent/test_loop.py
Follow the upstream run guide

Keep this evidenceYour added test, its failure before the change and its result afterward.

Explain it before you move on

Why must a tool result return to the agent loop?

Then challenge yourself: What should stop an agent that keeps requesting failing tools?

LESSON 03

Enforce permissions and isolation

Distinguish authorization from the place a command runs.

Your build task

Inspect allow/ask/deny precedence. Add a denied-mutation case and inspect the Docker workspace mapping before any agent-written commands execute.

Inspect the code and run a focused checksrc/decode/permissions/gate.py
TERMINAL
uv run pytest tests/unit/decode/permissions/test_gate.py tests/unit/decode/sandbox/test_workspace.py
Follow the upstream run guide

Keep this evidenceA passing boundary test and a diagram of mounted files, credentials and network access.

Explain it before you move on

Does using a sandbox make every tool call authorized?

Then challenge yourself: Is the course Docker setup a proven boundary for hostile code?

LESSON 04

Manage context, memory and skills

Keep useful evidence while controlling context growth.

Your build task

Inspect compaction, memory loading and skill discovery. Add a test that a tool call/result pair stays intact when older context is reduced.

Inspect the code and run a focused checksrc/decode/context/compaction.py
TERMINAL
uv run pytest tests/unit/decode/context/test_compaction.py
Follow the upstream run guide

Keep this evidenceThe test plus a before/after context example identifying retained and lost information.

Explain it before you move on

What can go wrong when you summarize old context?

Then challenge yourself: Does fitting into the token window prove the summary is adequate?

LESSON 05

Scope a subagent’s work

Make delegation useful without widening permissions.

Your build task

Inspect the exploration tool. Compare one focused child task with an unnecessarily broad task; check report size, tool scope and failure reporting.

Inspect the code and run a focused checksrc/decode/tools/agent.py
TERMINAL
uv run pytest tests/unit/decode/tools/test_agent.py
Follow the upstream run guide

Keep this evidenceA delegated-task contract, a failed-child case and a reason to prefer either one agent or delegation.

Explain it before you move on

What should a parent specify when delegating a task?

Then challenge yourself: Why can parallel children make a task worse?

LESSON 06

Run the harness headlessly

Separate execution from the terminal interface.

Your build task

Inspect local and remote entrypoints. First test the headless contract locally. Optionally follow the Modal guide with your own account and a bounded task.

Inspect the code and run a focused checksrc/decode/remote/headless.py
TERMINAL
uv run pytest tests/unit/decode/remote/test_headless.py
Follow the upstream run guide

Keep this evidenceLocal test results; for an actual cloud run, add its run ID, status, cost and cleanup record. A design-only review is not deployment evidence.

Explain it before you move on

What changes when an agent runs without a terminal UI?

Then challenge yourself: Does a successful deployment prove the agent completed a task?

LESSON 07

Build an evaluation gate

Measure task outcomes and regressions separately.

Your build task

Inspect one benchmark verifier and its oracle sanity check. Introduce a broken solution that should fail. Keep infrastructure errors separate from agent failures.

Inspect the code and run a focused checktests/unit/evals/benchmark/test_oracle_sanity.py
TERMINAL
uv run pytest tests/unit/evals/benchmark/test_oracle_sanity.py
Follow the upstream run guide

Keep this evidenceA passing and failing artifact, grading criteria, sample size and an explicit release decision.

Explain it before you move on

How do you test the evaluator itself?

Then challenge yourself: What if infrastructure errors are silently dropped from the report?

LESSON 08

Turn a recorded failure into a regression

Know what a replay controls and what it does not.

Your build task

Inspect recording failure handling, then follow the Kitaru guide on a small disposable task. Change one variable and inspect the replay tool policy before executing it.

Inspect the code and run a focused checksrc/decode/runtime/recording.py
TERMINAL
uv run pytest tests/unit/decode/runtime/test_recording.py
Follow the upstream run guide

Keep this evidenceBaseline and replay identifiers, changed setting, tool policy, result and a regression case. Label inspection-only work separately.

Explain it before you move on

Why is a replay not automatically a fresh end-to-end test?

Then challenge yourself: What happens when a replay needs an unrecorded tool result?

Sources you can inspect.

Course by Paul Iusztin / Decoding AI, introduced in this LinkedIn post. Upstream code is Apache-2.0; the articles and videos remain on their original sites. This companion is independently created and is not an endorsement by the course author.

Source snapshot: 3f8d21996149 · checked 2026-10-05. The outline links three distinct videos and marks the later videos coming soon, despite broader headline video counts. Availability can change.

Remote deployment is optional. A worksheet, mocked test, live model run and cloud deployment are different evidence. State which you completed.