Paste a function, ask for tests, get a green suite that proves nothing. The fix is to hand over the boundaries before the code.
2026-10-06 · 约 3 分钟读完
作者:周野(CPCX 撰稿人 · 编程与数据)
The coupon bug in the opening is real, and the original green suite really did declare that function safe for weeks before a boundary test caught it. The review-the-assertions habit came straight out of that incident.
I once asked an AI to write tests for a discount function. Thirty seconds later I had twelve tests, all passing, all testing roughly the same thing: the example from my own docstring. Coverage looked great. Two weeks later a customer paid full price on a coupon that should have stacked — a boundary no test ever touched.
The AI did what I asked. The problem was what I asked. “Write unit tests for this function” gets you tests of the happy path, because the happy path is the most statistically likely reading of the word “test”.

You know your function's real boundaries; the AI doesn't. It can guess generic ones, but the valuable boundary list always comes from the person who has seen the weird support tickets.
Here's the subtle trap: paste the implementation and ask for tests, and the AI reads the code and asserts what the code does. Broken code produces confident tests of broken behavior — the suite goes green and nothing improves.
What works better: describe what the function should do — the contract — including each boundary and its expected outcome, and hold the code back. Assert on the contract. Then when a test fails against the implementation, that failure is information, not noise. Sometimes it catches a bug that has been shipping for weeks.
When the tests come back, don't check the coverage percentage first. Read the assertions. A test that asserts “result is truthy” will never fail when it should. I cross out roughly a third of first-draft assertions for being too weak and ask for the specific expected value each time.
Weak assertions are how a suite looks healthy while testing nothing. The coverage number counts lines executed, not lines verified — the two get confused constantly, and only one of them is worth anything.
So the order is: boundaries from you, contract from you, tests from the AI, assertions reviewed by you. The AI is fast at the mechanical part — writing ten test cases — and unreliable at the judgment part — deciding what deserves a test. Splitting it this way costs maybe ten extra minutes and turns a decorative suite into one that has actually caught things.
I know it has caught things, because one of them was mine. The coupon bug at the start of this article was found by the first test written this way — weeks after the original green suite had declared the function safe.