Five hundred passing tests, and it missed 99% of the real thing
I vendored a well-tested AI-writing detector instead of building my own. Its tests all passed. Then I ran it against 678 real samples from two published datasets and it caught one in 270 — not after someone tried to fool it, on ordinary unmodified output. What five hundred passing tests never checked, and the two signals that actually held up.