Week 5 of 6
Writing and testing code with an agent
Zoom · 2.00–4.30pm ET
20 September 2026
Derrick Schultz

Week 1 · Aug 16
Introduction and Hermes Agent setup
Week 2 · Aug 23
Models, text generation, and everyday tasks
Week 3 · Aug 30
Skills, MCP, and image generation
Week 4 · Sep 13
Memory, scheduled jobs, and ongoing projects
Week 5 · Sep 20
Writing and testing code with an agent
Week 6 · Sep 27
Computer use and show and tell
Today: describe a small tool, have the agent build it, and test the result.
2.00
Review your Week 4 scheduled job
2.20
Code, creative processes, and agentic development
3.10
Break
3.15
Lab: build Frame Stack, one Compound Engineering step at a time
4.05
Let the agent use the tool you built
4.15
Homework, presentation prep, and questions
Useful output
Show one result you’d keep. How does it meet your brief?
Unexpected output
Show one surprise. Did it reveal a problem or suggest a useful change?
One correction
What did you change? Compare a result before and after.
Repeated behavior
What worked consistently? What failed more than once?
Tool idea
Name one task a small tool could help you do.
Post your tool idea in the chat. You’ll use it in the lab.
today’s question
Include deciding what to make, evaluating results, and changing direction.
This week, you’ll learn how to write better code with agents. The larger question is how to describe and implement a creative process: its actions, context, judgments, and revisions.
How much can we make explicit? What can an agent discover through doing? What still depends on the artist?
01
2.20–3.10 · 50 minutes
Marc Andreessen · 20 August 2011
Fifteen years ago, Andreessen argued that software was becoming the basis of businesses across the economy. Entire industries were reorganizing around software and online services.
He also identified a shortage of people with the skills to build these systems. Coding agents let us revisit who can build software and what processes it can carry out.
from description to execution
Specify the inputs, actions, decisions, and expected result. Ask the agent to implement the process, then test its behavior.
Intention
You can begin with a question, a constraint, or a rough direction.
Experiment
Write or change a process, then inspect what it produces.
Selection
Choose what to keep and what deserves more work.
Revision
The results can change both the code and your initial intention.
A coding agent can help translate each new direction into an implementation. You can begin before you know the final result.
Code
Defines when a decision happens, what information is available, and what happens next.
Model
Interprets that information and proposes a choice or revision.
Feedback
Records what happened so the result can inform later decisions.
You can implement this process without writing an explicit rule for every possible judgment.
As an artist, you make choices without explaining every decision out loud. Your experience, taste, references, and intentions inform those choices.
An agent needs access to that context. Explain what you want, show examples, and describe why you keep one result or reject another.
a proposed artistic system
01
Choose a question or intention to pursue.
02
Produce alternatives and decide what to try next.
03
Decide which results are worth developing and why.
04
Change the method or goal, or decide the work is finished.
These decisions can be shared between a person, a model, and fixed rules. Revising the goal can start the process again.
Your choices
Goals, constraints, examples, and feedback describe what matters to you.
The model
Patterns learned during training affect what it proposes and how it evaluates the results.
The system’s history
Saved results and evaluations can provide context for subsequent choices.
Saved feedback can change later choices without retraining the model. Decide whether the system can also revise its evaluation criteria.
Richard Sutton · 2019
Sutton argues that general methods using search and learning have repeatedly surpassed approaches built around expert knowledge as more computing power became available.
For an artistic system, this raises a design question: which decisions should be prescribed, and which could develop through exploration and feedback?
The essay concerns progress in AI research. It does not establish that creativity can be fully automated. · Read Sutton’s essay (PDF)
example 01 · computer vision
Task-specific systems
Engineers often designed features, collected labeled data, or trained models around a particular problem.
General pretraining
CLIP (2021) learned relationships between images and text from a large, varied dataset.
A new task
For a new classification task, describe the categories in words. The same model scores how well each description matches an image.
Example
Change the categories from “portrait / landscape / still life” to “photograph / painting / sketch.” Reuse the model without training a new classifier.
This is zero-shot classification. Check the model on your own images; performance varies across tasks. · CLIP: research and limitations
before · HOG person detector · 2005


New task: train a different classifier; the features may also need to change.
after · CLIP · 2021

Text encoder ↓
New task: change the label descriptions and reuse the pretrained model.
Simplified workflows, not an accuracy comparison. · Dalal & Triggs, 2005 · Figures 1 and 6 · Radford et al., 2021 · Figure 1
example 02 · agent tools
Custom tools
Each tool exposes actions its author has implemented. A new operation may require additional tool code.
Shell access
Bash is a command-line shell. An agent can combine installed programs, write scripts, run tests, and use errors to revise its approach.
A reported result
Vercel’s d0 turns questions into database queries. Replacing most custom tools with shell access to documented files improved success from 4/5 to 5/5 in its five-query test.
Vercel kept a separate tool for running database queries. Clear goals and verification still matter. · Vercel’s case study · mini-SWE-agent: coding with Bash alone
before · specialized tools
Custom tools, grouped by purpose
ExecuteSQL→Database→AnswerDevelopers maintain the tools and the rules for retrieving context.
after · Bash and files
ExecuteCommandls · find · grep · cat
ExecuteSQL→Database→AnswerThe agent chooses commands to explore well-documented source files.
Simplified from Vercel’s d0 case study. The database query tool remains separate from Bash. · Vercel · 22 December 2025
last week’s Three.js bot
probably not yet.


why coding helps
Use code to generate variations, inspect results, record judgments, and revise the next attempt. Check how each step works and where it still needs your input.
from task to script
01
Pick a repeated process with clear inputs and outputs, such as resizing a folder of images.
02
List the steps, the expected result, and what should happen with missing or invalid input.
03
Ask the agent to write a script. Run it on a small sample and compare the output with your description.
04
Save the working script and its run command. Schedule it if the task needs to repeat automatically.
A script needs a running computer or hosted environment to run on a schedule.
the development process
These steps still apply when an agent writes the code.
SDLCSoftware development life cycle
ADLCAgentic development life cycle
Plan
People turn requirements into a design and a task list.
You provide goals, constraints, and context. The agent proposes a plan for you to review.
Build
Developers write code and carry out the implementation.
The agent writes code and uses tools. You review its choices and redirect the work.
Verify
People and automated tests check the software. Developers fix failures.
The agent runs checks and uses failures to revise its work. You inspect the evidence and test the result.
Document
People record decisions and fixes for the team to reuse.
Save decisions and fixes as project instructions, skills, and tests the agent can use on later tasks.
Here, ADLC means applying agents throughout software development. People still own requirements, release decisions, and maintenance. · SDLC · Agentic coding practices
four steps
01
Define the inputs, outputs, and tests. Review the agent’s proposed approach before it builds.
02
Ask for one working version. Check that the agent stays within the task you described.
03
Have the agent run tests. Try the tool yourself with normal, empty, and invalid input.
04
Save the setup instructions, problems, and fixes in the project. Ask the agent to read them next time.
The agent can help with all four steps. You decide whether the result meets your needs.
verification in code
Expected behavior
Define the result you expect, including what should happen with missing or invalid inputs.
Automated tests
Unit tests check individual pieces of code. Other tests check how those pieces work together.
Direct observation
Run the actual program. Use its interface and inspect its outputs, screenshots, or recordings.
Revision
Compare the evidence with the expected behavior. Use mismatches to guide changes, then check again.
Passing a check supports a specific claim about behavior. It does not prove the whole program is correct. · Lauren Tan · The Complete Guide to pstack, Pt. 1
verification in a creative process
Observation
Inspect the actual output: use vision for images, listening for audio, or reading for text.
Criteria
Compare it with the brief, references, and desired qualities such as composition, pacing, or tone.
Judgment
A model or a person can compare versions and explain a preference. Aesthetic evaluations depend on the criteria.
Next decision
Keep, revise, or discard the result. An unexpected outcome may also justify changing the goal or criteria.
Compound Engineering
Every’s collection includes planning, implementation, review, and documentation skills. /ce-compound records a solved problem for future reference. Repository
pstack
Lauren Tan’s workflow collection. poteto-mode selects a process for the task, coordinates the work, and checks results. Community port for Claude Code and Codex
Matt Pocock’s skills
grilling questions a plan, tdd guides development through tests, and domain-modeling defines the project’s concepts and terms. Repository
All three are plain markdown you can read. Compound Engineering is the one for today because each step is a separate command you run yourself.
LaB
10–15 minutes
quick aside · TypeSafe · introduced 15 September 2026
The model
TypeSafe calls Jev a “System One” model, built for fast, repeated judgments inside software.
Input
Provide text or structured context. Define the question and its answer choices or scoring criteria.
Output
Choose an option, score something against criteria, or estimate whether a statement is true. Your code uses the result to decide what happens next.
Creative example
Use a written critique to choose between revising, generating another version, or asking the artist.
Jev returns values and probabilities. It does not generate prose or code. Test its judgments on your own examples. · TypeSafe documentation · Launch announcement
quick aside · access
1. Join the waitlist
Visit typesafe.ai, choose Join Waitlist, and complete the form.
2. Sign in
Once you have access, open console.typesafe.ai. Sign in with Google or email.
3. Try the Playground
Open the Playground, paste some text, and add a question. Inspect the answer before using it in a project.
4. Get an API key
Create a key in the console’s API keys page when you are ready to call Jev from code.
Early access may require waiting. Installing the skill does not grant API access. · Official quick start
quick aside · TypeSafe skill
02
3.15–4.05 · 18 steps · 32 build
setup
1. Run in your terminal
2. Start a new Hermes session
Review the installer’s summary and confirm. This installs the full skill collection into ~/.hermes/skills/. · Skills CLI · Hermes skills
the lab project
We’ll build one part of a creative process: a tool for exploring video through different arrangements and viewpoints. You’ll make the creative judgments while the agent helps implement and test the tool.
Tier 1. Load a video, stack the frames, orbit and zoom.
Tier 2. Sliders for depth spacing and trail opacity.
Tier 3. Save the current view as a PNG. From a side angle that image is a slit-scan, assembled from a hundred instants that were never a single frame.
Use an MP4 from your Week 4 job. No video in your practice? Bring a folder of stills and stack those instead. Each Compound Engineering step gets two slides: what it does, then what it looks like on this project. They stay up during the lab.
/ce-brainstorm to define what you want to build.What it does
Asks questions about your idea, compares approaches, and helps you choose a scope. It does not write code.
What you give it
Describe the problem, who will use the tool, and any constraints you already know.
What comes back
Agreed requirements in chat or a saved plan, ready for /ce-plan.
What you check
Confirm the input, output, and first version’s features. Correct assumptions and remove additions you did not agree to.
Start here when the requirements are unclear. If you already have an agreed brief, go to /ce-plan.
step 00 · /ce-brainstorm · example
Agree on the tool’s behavior before choosing libraries or writing code.
Read the requirements back. Correct anything the agent inferred without asking.
Continue with /ce-plan once the requirements match what you intend to build.
/ce-brainstorm I want to make 3D images from the videos my Week 4 job produced. Help me decide what to build.
You could stack frames along a depth axis or use brightness to set the height of a surface. Which would you like to try?
Stack the frames. I want to rotate the view and save a PNG. Keep the first version to one video at a time, with no editing or audio.
Requirements: load one video, stack its frames in 3D, rotate the view, and export it as a PNG. Process one video at a time. Leave out editing and audio. Is this the scope you want?
Yes. Use those requirements for /ce-plan.
/ce-plan turns your brief into an approach. No code yet.What it does
Reads the brief and writes down which files it will create, which libraries it will use, in what order, and what it will leave out.
What you give it
The brief in one message, with every constraint you already know. A constraint you leave out of the brief is not in the plan.
What comes back
A plan file under plans/. It usually asks one or two questions first. Answer the ones that change the output.
What you check
Confirm both constraints made it into the plan, then read the rest of it for additions you did not request.
Anthropic’s rule of thumb is to skip the plan when you could describe the change in one sentence. This build has two constraints that are easy to lose, so plan it.
step 01 · /ce-plan · example
Without a frame cap and a reused texture pool, the tool dies after about forty seconds. State both in the prompt, because the planner works from what you wrote and will not add constraints on its own.
Read the plan for two things. Did it keep your constraints, and is it building more than you asked for. Send it back if the texture pool is missing.
A working version is at artificial-images.com/frame-stack/. Open it after the lab rather than before, and read the comments at the top of the file.
/ce-plan Work in ~/tools/framestack. One HTML page: load a video, draw each frame on a plane stacked back along Z, orbit with the mouse. Keep at most 120 history frames at 160x90 and reuse a fixed pool of textures, never one per frame. Size the planes from the video’s real dimensions. Empty or unreadable file: say so on the page and stop. List the files you will create, then wait.
One index.html, three.js from a CDN via importmap, no build step. I will preallocate 120 canvases and 120 CanvasTextures once, overwrite them in a ring, and set needsUpdate. Planes sized from videoWidth and videoHeight on loadedmetadata. One question: sample every presented video frame, or every second frame?
Every presented frame. Use requestVideoFrameCallback.
Plan written to plans/framestack.md. Ready for /ce-work.
/ce-work builds from the plan, file by file.What it does
Follows the plan in order, writes the files, and runs whatever it can run. It stops when every item in the plan is complete. Whether the tool works is a separate question, answered in the verify step.
What you give it
/ce-work and nothing else. A new request here changes the build without changing the plan, and next week you will not know why the two disagree.
What comes back
The files, a summary of what it did, and the output of anything it ran. For this project, one index.html.
What you check
The file list. Expect one page. If a src/ folder with many files appears, the build has drifted from the plan. Stop it.
Serve the page over http, not by double-clicking the file. A file:// page cannot load three.js from the CDN and shows a blank screen with no error.
step 02 · /ce-work · example
Build
Serve and open
Commit as soon as the stack appears. Everything after this point is a change you can undo to here.
/ce-debug finds the cause before it changes anything.What it does
Reads the error or the wrong behaviour, forms a hypothesis about the cause, tests it, and edits only after the test confirms it.
What you give it
The exact error text, or exactly what you saw and what you expected instead. Paste it rather than retyping it, so the detail it needs survives.
What comes back
A diagnosis, then a fix. If a fix arrives without a diagnosis, ask for the cause before you accept it.
What you check
/rollback diff shows what it touched. The fix should touch only the code that was broken.
After three failed fixes for the same bug, stop. Roll back and re-plan, because by then the problem is usually in the plan.
step 03 · /ce-debug · example
What you type
What the cause turns out to be
All three happened while building the reference version. The third one was found by a portrait video, not by reading the code.
Checks the agent runs
/ce-code-review reads the diff for bugs and for drift from the plan. It returns findings. It cannot drop a portrait video into the page.
Checks only you can run
Open the tool. Use your own files. Watch it for a minute. The four checks on the next slide are this kind.
What comes back
From the review, a list of findings. From your checks, screenshots and numbers. Ask for the command it ran and what it returned.
What you check
That what came back is a measurement or a screenshot. A sentence saying it works is not evidence.
A reviewer asked to find problems will find some even when the work is sound. Act on the findings that change whether the tool works and leave the rest.
It plays
Drop in a video and orbit to the side. You should see separate planes receding, not one flat rectangle.
It survives a minute
Let it run for 60 seconds. Ask the agent for the frame rate at the start and at the end. If the second number is much lower, the texture pool is not being reused.
It respects the shape
Load a portrait video, or one shot on a phone. Tall video, tall planes. This is the check people skip.
It fails politely
Drop in a PDF, or a video with no frames. It should say so on the page rather than going blank.
Ask for the command and what it returned, not a summary. Week 5’s rule, on your own project.
/ce-compound writes down a problem you solved so the next project starts from the answer.What it does
Records the problem, the cause, and why the fix was right, in a file inside the project that the agent reads the next time it works there.
What you give it
/ce-compound after a real problem. Not after every commit. A routine fix whose code explains itself does not need an entry.
What comes back
A short entry under the project’s docs. Read it. If you would not understand it in a month, ask it to rewrite.
What you check
That it recorded the cause, not only the fix. The entry should explain that a texture allocated every frame leaks GPU memory, rather than only saying textures are now reused.
Without this entry, the texture leak comes back in your next project and costs the same time again.
Your Turn
Save: the plan file, a commit of the working version, and one PNG you exported.
Check: the four checks from the verify slide. Write down what each one did.
If you finish early: add the spacing slider, then run /ce-compound on whatever cost you the most time.
Post your exported PNG in Discord with one sentence about the clip it came from. Raise a hand in Zoom if you are stuck for more than five minutes.
step 05 · /ce-compound · example
The entry on the right records the symptom, the underlying cause, and the rule that prevents it next time.
Run it on whichever of the three failures cost you the most time today.
Next week, when you extend this tool, the agent reads this before it writes a line.
/ce-compound The tab crashed after about a minute of playback. The cause was a new CanvasTexture created every frame instead of reusing a pool.
write_file · docs/solutions/texture-pool.md
Recorded. Symptom: playback stutters, then the tab dies at roughly forty seconds. Cause: a texture allocated per frame is thirty GPU allocations a second and nothing frees them. Rule: allocate a fixed pool once, overwrite in a ring, set needsUpdate. Applies to any per-frame canvas or video texture in this project.
Add that plane sizes must come from videoWidth and videoHeight, not a fixed ratio.
Added as a second entry. Both load next time anyone works in ~/tools/framestack.
03
4.05–4.15 · 10 minutes
a possible extension to Frame Stack
Actions
Let the agent load a video, adjust parameters, render a preview, and export an image through callable tools.
Observation
Give it access to each rendered preview so it can inspect the result.
Criteria
Describe the qualities to look for and provide references. Test whether its judgments match yours.
Feedback
Record results, evaluations, and your corrections so they can inform the next attempt.
This would extend the lab tool. Decide which actions the agent can take on its own and which require your review.
example · Scene Ripper
Scene Ripper divides video into shots. Its agent can use the app’s tools to find and return shots matching a request.
Your request defines what to look for. The agent calls the available actions, and you evaluate the returned shots.

return to the creative process
Name one decision you could describe for the agent and one you want to keep. Consider source material, aesthetic choices, and deciding when to stop.
04
4.15–4.30 · 15 minutes
Show a result
Share output from your scheduled job, demonstrate your tool, or show both. You don’t need slides.
Explain a correction
Show a change to your brief, SOUL.md, README, or memory. Compare the results before and after.
Describe a boundary
Name a decision or action you kept for yourself. Explain why.
Answer two questions
Be ready to explain what you tested, what worked, and what remains unresolved.
Sign up in Discord by Wednesday. The class agent presents last.
Use your tool twice
Apply it to real work. Record problems and possible improvements before adding features.
Review the scheduled job
Continue through the second week if it’s running as intended. Decide whether to keep, pause, or revise it.
Review saved instructions
Read MEMORY.md and SOUL.md. Correct outdated entries and save a copy before major changes.
Prepare your presentation
Choose one result, one correction, and one boundary. Practice explaining them in six minutes.
Coding agents
Agent access to apps
Related course materials
NYU course slides: additional lessons on building and deploying apps
Week 5 slides
before next week
Week 6 · final class
Computer use and show and tell
27 September · 2.00pm ET
Questions: Discord or derrick@titles.xyz · artificial-images.com