Working with a Coding Agent
In the Agent Lab, robot arms assemble code faster than anyone can watch. Their foreman, Unit-7, is cheerful, tireless and sometimes confidently wrong: it will report "All done, all tests pass!" about work it never checked. A coding agent is different from a chat assistant because it acts. It opens files, edits several of them, runs commands and reports back. That makes it far more useful and far easier to lose control of. These habits keep you in charge.
Scope the task tightly
"Improve the app" invites the agent to touch everything. A good task names one outcome, where the change should happen, and what it must not touch.
Add input validation to POST /api/signup in routes/auth.js.
Reject an email without "@" and a password under 8 characters
with a 400 and a JSON error message. Do not change other routes,
the database schema or package.json. Add tests in auth.test.js.
A task this size produces a diff you can read in a few minutes. If you notice you are writing "and also...", split it into a second task.
Work on a branch, keep commits small
Before the agent starts, create a branch (git switch -c add-signup-validation) and make sure your working tree is clean. Now everything the agent does is separate from your working code, and throwing it all away is one command.
Commit after each piece that works. Small commits mean that when step four breaks something, you can go back to step three instead of untangling a morning of mixed changes. An agent that makes one giant change across twenty files leaves you only two choices: accept all of it or none of it.
Read the diff before you accept
The agent's summary is a claim, not evidence. Unit-7 might say "added validation" while the diff also shows it deleted a failing test, changed a config value, or reformatted three unrelated files. Always look at the actual changes, with git diff or your editor's review view, and ask for each file: did I expect this file to change? Does this change do what the task asked, and only that?
Watch in particular for deleted or skipped tests, loosened checks, new dependencies, and edits to configuration, environment or lock files you didn't ask about.
Run the tests yourself
"All tests pass" from an agent can mean it ran them, ran only some, or ran nothing and predicted the result. Run the test suite yourself on the branch, then try the feature in the running app. It takes a minute and turns a claim into something you have seen.
If the agent made tests pass by changing the tests rather than the code, that is not a fix. A test that was rewritten to expect the wrong answer is worse than no test at all, because it now tells you something false.
Stop it when it wanders
Agents can go in circles: fixing one error creates another, which it fixes by changing something else, and twenty minutes later half the project is different. Signs it's wandering include edits to files unrelated to the task, repeated attempts at the same fix, new libraries appearing to work around a problem, and disabled checks.
When you see that, stop it. Don't let it dig deeper. Reset the branch to the last good commit, work out what it misunderstood, and restart with a smaller, clearer task, often with the error message and the relevant file included.
You own the merge
The agent does the typing. You decide what gets merged. Nothing leaves the branch until you have read the diff, run the tests and understood the change well enough to explain it to someone else. If you can't, the task was too big or the change isn't ready.
Resources
Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.