Before asking an AI agent to fix a bug, define what would prove the fix works.
“Fix this bug” gives the agent a goal. It leaves the acceptance criteria open to interpretation.
Verify the failure you need to prevent
For example, suppose a checkout form submits the same order twice when someone clicks the button repeatedly.
Before implementation, define how to reproduce the duplicate submission, what should happen when the user clicks again, and what should happen if the first request fails.
Then verify the change against those expectations.
A disabled button may prevent repeated clicks. It doesn’t, by itself, prove that retries cannot create duplicate orders.
The evidence needs to match the failure you care about.
A useful check should expose the original problem, pass after the fix, and cover relevant behavior that the change could break.
Passing tests still depend on what those tests check. If they only confirm that the button becomes disabled, duplicate orders may remain possible elsewhere in the flow.
Match the evidence to the consequences
Scale verification to the consequences of being wrong. A spacing adjustment and a payment flow need different levels of evidence.
Before delegating, answer:
“What evidence would make me accept or reject the result?”
Make verification part of the task you assign.