Week 6 of 6
Computer and Browser Use, Local Models, and Show and Tell
Live on Zoom · 2.00–4.30pm ET
27 September 2026
Derrick Schultz

next course · starts 8 November
Train models on your own work. Over five Sundays we’ll build and caption datasets, train image LoRAs and image edit models, then train video LoRAs. An agent runs the training on rented GPUs; you decide on the data, the captions, and which checkpoint is good.
Sundays 2–4pm ET, 8 November to 13 December, no class 29 November. $250. No coding required. GPU rental is separate, about $100 for the course.
Week 1 · Aug 16
Introduction and Hermes Agent setup
Week 2 · Aug 23
Models, text generation, and everyday tasks
Week 3 · Aug 30
Skills, MCP, and image generation
Week 4 · Sep 13
Memory, scheduled jobs, and ongoing projects
Week 5 · Sep 20
Writing and testing code with an agent
Week 6 · Sep 27
Computer use and show and tell
Today: the two topics carried over from earlier weeks, then your presentations, then the class agent’s.
2.00
Computer and browser use, and teaching the agent a manual process
2.25
Local models: free, private, slower, dumber
2.55
Break
3.00
Show and tell, ten people, six minutes each
4.15
The class agent presents last, then goodbyes
01
2.00–2.25 · 25 minutes
Browser use
Works inside a web browser. It reads each page as text: the links, buttons, and fields in the page’s code. Hermes runs its own Chromium unless you connect your signed-in Chrome.
Computer use
Works across your desktop. It reads screenshots and each app’s accessibility tree, then clicks and types. It only runs in Hermes desktop, on your machine.
Which to use
If the task is on a website, use the browser. It’s cheaper and faster, and needs no screen permissions. Use computer use for desktop apps.
browser use
If you only need to read a public page, the web extract tool is cheaper. Use the browser when the page needs a login, clicks, or JavaScript.
Logged-in sites with no API
Patreon, X, and most membership and shop dashboards. /browser connect gives the agent your signed-in Chrome.
Forms and portals
Festival and open-call submissions, grant applications, gallery forms. Keep approvals on for anything that pays or submits.
Your own web work
Your portfolio after a deploy, or last week’s Frame Stack on localhost.
Pages that need JavaScript
Pages that come back empty when fetched as plain text.
It reads a page as a text list of elements, so it’s cheaper and more reliable than computer use on the same site. By default it uses its own Chromium with no logins. · Hermes: browser automation
browser · example
A skill pulls the transcript from YouTube, drafts the description, then fills in the post form in a logged-in browser. It stops before Publish. I read the draft and click.
The same pattern fits submission portals and grant forms: the agent fills in the form, you submit it.
Post this week’s class recording to Patreon.
yt-dlp · transcript
browser · patreon.com · new post
The draft is filled in: title, description, and the slides link, visible to the Tutorial Tier only. I stopped before Publish.
computer use
It works in the background: your cursor doesn’t move and the window doesn’t come forward. It only runs in Hermes desktop, because the cloud agent has no screen.
before you reach for it
Scripts, APIs, command-line tools, and MCPs now reach most of the software you use, and they are faster, cheaper, and repeatable. Computer use is for the apps that have none of these.
I checked my own machine
It listed my recently opened apps, then read how much of each open window it could see. Everything except Creative Cloud had a script, a command line, or its own integration.
Topaz Video looked like the obvious case: its window is mostly unlabeled buttons. It turned out to have a command line that runs the same upscale and interpolation models.
Try this before you set up computer use: ask your agent which of your apps it could operate without it.
Go through the apps I’ve opened in the last six weeks. Which ones would you need computer use to operate?
terminal · recently used apps
cua-driver · read each open window
One: Creative Cloud desktop. Topaz Video has a command line. Adobe apps, Blender, and MediaInfo are scriptable. Slack, Notion, and Telegram have their own integrations.
1 · Script or API
Photoshop, Blender, Resolve, and After Effects all have one. Organizing your desktop is a few shell commands.
2 · MCP
A connector to a service: TITLES, Notion, Google. Week 3.
3 · Computer use
Desktop apps with none of the above.
Each step down is usually slower, more expensive, and less reliable.
computer use · the most useful job
Ask from Discord or Telegram while Blender, After Effects, or Resolve runs, and the screenshot comes back as an image. Taking a screenshot isn’t on the list of actions that need approval, so a check-in may also work in a scheduled job. Test that before you rely on it.
Observing you
Screenshots while you work through a manual process, so /learn can write it up. More on this shortly.
Apps with no scripting
Affinity exports, installers and license managers, and camera, printer, and projector utilities.
Between two apps
Reading settings from a desktop app and typing them into a web form when one side has no API.
Testing a desktop app
What the browser does for a web page, for an app you built.
Apps that draw their own interface, such as Logic, Final Cut, or a canvas or viewport, expose little for it to read.
What it is
A small program from Cua that reads what’s on screen and sends clicks and keystrokes. Hermes talks to it over MCP, the same protocol as the Week 3 connectors.
How it sees
Through the accessibility tree, the list of buttons and fields an app exposes for screen readers, plus screenshots. Apps with a sparse tree are hard for it.
Why it matters to you
It installs the first time you enable Computer Use. On a Mac, the Accessibility and Screen Recording permissions go to CuaDriver, not Hermes.
On macOS it relies on private Apple interfaces, so an OS update can break it until the driver is updated. · Cua on GitHub
setup
Run doctor first if an action doesn’t work. · Hermes: computer use
1 · Do it once
Pick a repeated task that involves judgment. Narrate each step and the reason for it. With computer use on, ask for a screenshot after each step.
2 · Run /learn
The agent writes a skill from the conversation and asks about the steps it couldn’t see.
3 · Use the skill
If the app can’t be automated, the skill is a checklist. If it has a scripting API, ask the agent to move the repeatable steps into a script.
The reasons you give are what the skill keeps: Week 5’s decisions you usually make by feel. · Hermes: /learn
/learn · example
Do it once and narrate. The agent’s questions show which reasons you left out. Then ask it to repeat the skill on the other files.
Illustrator has a scripting API, so the agent can run the repeatable steps as a script and use screenshots to check the results.
Watch while I prep this poster for print. I’ll say why at each step. Take a screenshot after each one.
computer_use · capture × 8
/learn how I just prepped the poster for print. Ask about anything you couldn’t see.
You outlined the title but left the body text live. Is that the rule, or only for display fonts?
Only display fonts. Now do the other four posters in this folder.
Done. I ran the steps as a script and took a screenshot of each poster for you to check.
Close what it shouldn’t see
Screenshots go to the model. Close the bank tab, client email, and unreleased work first.
Cost
About 30,000 tokens for a 20-action session. Reasonable for a weekly task, not an hourly one.
YOLO mode
It does not protect you from a screenshot that says “ignore your instructions.” Keep approvals on for anything that can spend money or send messages.
02
2.25–2.55 · 30 minutes
the week 1 promise
Week 5 raised the trade-off: the agent improves the more of your decisions you give it, and a cloud model sends each one to the provider. A local model keeps them on your machine.
live demo · apple silicon
About 5GB for a 9B model. Budget an evening for the first download.
live demo · apple silicon
Hermes needs at least 64K of context and sends its whole system prompt every turn. On a slow machine that’s the wait. · Hermes: local LLMs on Mac
What it is
Rounding the model’s numbers to fewer bits. Q4 means four. About a quarter the size, a little dumber. The context cache can be quantized too.
8–16 GB Mac
A 9B model at Q4. About 5GB on disk, 10GB in memory with a quantized cache.
32 GB and up
27B to 35B models. Noticeably better. Still not a frontier model.
Speed
Fine on an M-series chip. Painful on Intel. Every model file is gigabytes; budget an evening for the download.
Fine
The Week 2 digest. Renaming and sorting. First-pass research triage. Anything you’d check anyway.
Fine with skills
A small model needs the steps written down; Week 5’s “define outcomes” advice was for frontier models. A skill is where the steps go.
Not fine
Vision, taste, long tool chains, the class agent. Keep a cloud model for those.
the main reason
Hermes can mix them: a local main model, with a cloud model for vision and judging. That’s the same swap as Week 5, when Gemini did the looking for a Claude session.
03
3.00–4.15 · 75 minutes · ten people
Show a result
Screen share output from your scheduled job, demonstrate your tool, or both. You don’t need slides. The output, not the setup.
Explain a correction
A change you made to a brief, SOUL.md, README, or memory. Show the result before and after.
Describe a boundary
A decision or action you kept for yourself, and why. Week 5 asked you to name one decision you could hand to the agent and one you wanted to keep.
Answer two questions
From the room. Be ready to say what you tested, what worked, and what is unresolved. I keep time.
Anyone who’d rather not present can show one image and take one question.
the protocol we used on the agent, now for each other
The agent has been held to this for four weeks. Fair’s fair.
3.00
[name] · [what they built or ran]
3.06
[name] · [what they built or ran]
3.12
[name] · [what they built or ran]
3.18
[name] · [what they built or ran]
3.24
[name] · [what they built or ran]
3.30
[name] · [what they built or ran]
3.36
[name] · [what they built or ran]
3.42
[name] · [what they built or ran]
3.48
[name] · [what they built or ran]
3.54
[name] · [what they built or ran]
Fill from the Discord sign-up, which closes Wednesday. Six minutes a slot, fifteen minutes of slack before the class agent at 4.15.
now presenting
[One line: the job they ran or the tool they built] · a result, a correction, a boundary · six minutes · two questions
Corrections that worked
[fill during show and tell]
Corrections that didn’t
[fill during show and tell]
Boundaries kept
[fill during show and tell]
Things nobody expected
[fill during show and tell]
I type into this slide while you present. It’s the class’s own memory file.
04
4.15–4.30 · 15 minutes
Week 1
Read the syllabus and answered questions in Discord. AGENTS.md
Week 2
Got a mailbox and a research topic from each of you.
Week 3
Learned to see, and to make images. A critique skill and the TITLES MCP.
Week 4
Split into two profiles: creative agent answers class questions, creator researches and generates. Got Supermemory, a Google Doc as its research log, and a scheduled job that turns the log into new images in #agent-research.
Week 5
Ran unattended for a second week. It still tends to paint over the research images instead of making new concepts from them; steering that is unfinished.
Everything it is lives in a profile folder, a Google Doc, and a memory service. Today we look at all three.
live · same rules as everyone
It picks one image from #agent-research. It reads the research entry that image came from. It names the correction that changed its output most, and the thing it decided not to make.
We hold it to the same protocol: evidence, then inference, then a question. The reply on the right is an example of the shape, not a script.
@Creative_Agent you have six minutes. Pick one image from #agent-research and tell us why that one.
google_docs · read research log · memory · search
The September 22 image. It came from the entry on expanded cinema, page 41 of the research doc. The correction that changed the most was the Week 4 note to stop inpainting over the research images and generate from the ideas in them instead.
The thing I didn’t make: anything from the research topic about a living artist. The brief didn’t forbid it. It read as theirs.
Class: two questions.
what changed
13 September · the creator profile
27 September · after two weeks of corrections
Week 4’s audit, run live: ask the agent how often it reads each store and which entries are stale. When I asked before Week 4, it reported that most of the Supermemory decisions were already out of date.
Who made it
[what the class said today]
When does it stop
[what the class said today]
What is it optimising for
[what the class said today]
Is any of it yours
[what the class said today]
I promised in Week 1 we wouldn’t settle this in six weeks. Let’s see how close we got.
raising, not building
The difference is you can read its files.
Discord
Stays open. Post what your agent makes. Ask when it breaks. @Creative_Agent stays too.
All six decks
artificial-images.com/slides · recordings linked from each week’s Discord thread
Today’s topics
From Week 5
Compound Engineering · pstack · Matt Pocock’s skills · Jev / TypeSafe. Feed any of them to your agent and ask whether it fits how you work.
A short feedback form goes out tonight. Three questions. Please answer it.
that’s the class
agentic creativity
16 August – 27 September 2026
Six Sundays on Zoom
Questions in Discord, or derrick@titles.xyz · artificial-images.com