Verifying with Tests and Evidence
The Verification Core sits at the top of the campus, and nothing built in Universe 2099 counts until it passes through. The question it asks is not "does this look right?" but "how do you know?". AI-written code nearly always looks right. It is neat, well named, and comes with a confident summary. None of that is evidence. Evidence is something you ran and saw: a test that passed and would have failed if the code were wrong, or a feature that behaved correctly in the running app.
Looking right is not evidence
It is easy to read generated code, nod, and move on. But reading can't tell you whether an off-by-one error is hiding in a loop, whether a date is in the wrong time zone, or whether an empty list crashes the page. Your brain fills in what the code probably does. Running it shows what it actually does.
Treat every claim, from the AI or from yourself, as a hypothesis until something confirms it. "The function handles empty input" becomes true when you have called it with empty input and seen the result.
Tests that would fail if the code were wrong
A test is only useful if it can fail. Compare these two:
import { test } from 'node:test';
import assert from 'node:assert/strict';
import { parsePrice } from './price.js';
// Weak: passes for almost any wrong answer
test('parsePrice returns something', () => {
assert.ok(parsePrice('$1,299.50'));
});
// Strong: checks the exact value, including the edge cases
test('parsePrice converts dollars to cents', () => {
assert.equal(parsePrice('$1,299.50'), 129950);
assert.equal(parsePrice('7'), 700);
assert.equal(parsePrice(''), null);
assert.equal(parsePrice('abc'), null);
});
The weak test passes if parsePrice returns 1299.5, 12, or the string "oops". The strong one fails on each of them.
A quick way to check a test is to break the code on purpose, for example by deleting the line that removes commas, and run the test again. If it still passes, it isn't testing what you think.
Asking the AI for tests
Assistants can write tests quickly, and that is a good use of them, with care. If the same model writes both the code and the tests, it can bake the same misunderstanding into both, and the tests will happily confirm the bug. Tests can also end up copying the implementation's logic instead of checking results.
So decide the expected values yourself, from the requirements, not from what the code happens to return. Give the assistant the cases and answers you care about ("empty string returns null, '$1,299.50' returns 129950") and let it write the test code around them. Then read the tests as carefully as the code.
Run them, and test the cases you care about
A test that was written but never run is not evidence. Run the suite yourself and read the output: how many tests ran, how many passed, and whether any were skipped.
Focus on the cases that matter for your feature, not just the one in the example: empty and missing input, the largest input you expect, invalid data, and the boundary values, such as exactly 8 characters for an 8-character minimum password. Those are where generated code most often breaks.
Check the running app
Tests check the pieces. Users see the whole. After the tests pass, start the app and use the feature the way a user would: submit the form, refresh the page, try it logged out, try it on a narrow phone-sized window. Watch the browser console and the network tab for errors. Many bugs, such as a route that isn't connected, a missing environment variable or a button that calls the wrong endpoint, only show up here.
A draft you are responsible for
Treat AI output as a draft from a fast collaborator who never checks their own work. The draft can be very good. But when it goes live, it's your name on the commit and your users who hit the bug. Before you ship, you should be able to say what you ran, what you saw, and why you believe it works. If the only answer is "the AI said so", it isn't finished.
Resources
Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.