Skip to content

44 Agents Tried to Break My Case Studies. Three Succeeded.

AI WORKFLOW
44 Agents Tried to Break My Case Studies. Three Succeeded.
Conner Crowe

Last month I finished five case studies for this site. Real accounts, real numbers, each one traced back to the source report. Before publishing I ran my usual pass: the voice check, the fact check, the read-aloud. It came back clean. Then I pointed 44 AI agents at the same five case studies and told each one to assume the copy was wrong and prove it. Between them they raised 38 findings. Three would have embarrassed me.

The lesson was not that AI makes a good editor, though it does. The lesson was about my own clean pass. I wrote the copy, I checked the copy, I signed off on the copy, and an adversarial reader still found a client-data leak, an overclaim, and a number I could not trace, all in work I had just called finished. Your own review of your own work is the least reliable review you run.

How the 44 agents were set up

The structure is the part that matters, because “ask ChatGPT to check it” does not do this. I split the review into six dimensions: factual accuracy, client confidentiality, voice, internal consistency, overclaiming, and links. For each dimension I ran a set of independent agents, and every one got the same instruction. Refute by default. Do not confirm the claim, try to break it. Assume a number is wrong until you cannot prove it is. Assume the client can be identified until you have checked every figure.

Forty-four agents in total, each hunting for what was not right, none able to see the others’ work or lean on my confidence that the copy was done. That last part is the point. I could not talk them out of a finding, because they never heard me call it finished.

The three that landed

Thirty-eight findings came back. Most were small: a link to tighten, a sentence that read templated, a stat that needed its source line. Three were not small.

One case study about a consent-tracking fix said two platforms had been reporting zero conversions. The account’s own history showed one of them had already been corrected before my window. I had overstated the break. Scoped it down to what was true.

One case study printed a client’s exact cost per click. Anonymized everywhere else, and I had still left a real dollar figure a competitor could use. Changed it to a relative multiple.

One live figure on a results page, a shopping-campaign lift, I could not trace back to a report when the agent pushed me to. If I cannot source it, it does not ship. Pulled it and replaced it with a number I could stand behind.

None of those were lies. Each was the kind of thing that happens when you are close to your own work and reading for confirmation instead of for holes.

Why your own pass misses them

When you review your own writing, you are not checking it. You are re-experiencing writing it. You remember what you meant, so you read what you meant, not what is on the page. You remember tracing the number, so the number looks sourced even where the citation is missing. Confidence is the problem, and you have the most confidence in the work you just finished. Catching an AI tell in your own prose is the easy version of this. Catching a false number you believe is the hard version.

An adversary has none of that. It did not write the sentence, it does not know what you meant, and you told it to assume you got it wrong. It reads what is there. That is why refute-by-default matters more than the model behind it. A friendly reviewer confirms. An adversarial one breaks. Only the breaking finds the leak.

I did it again this week

The home-services cost-per-lead post on this site went through the same thing before it shipped. Three agents, refute by default, and they caught two problems my own pass had waved through. The draft was about to cannibalize an existing page of mine, competing with it for the same search, and I had a number backwards. Both fixed before anyone saw them. I do not publish account work here without running it now.

What I did not claim

Not that the agents wrote or fixed anything. They found, I decided and rewrote. And not that this makes the case studies perfect. It makes them checked by something other than the person most motivated to believe they were done. That is a lower bar than perfect, and a much higher one than a self-review.

The receipts

Five case studies, one 44-agent adversarial workflow across the six dimensions above, 38 findings raised and dispositioned, three rated high severity and fixed before publish. The five are live on the results page, and each one carries a “what I did not claim” section. That habit is the same instinct, written into the copy itself: name the thing you are not saying, so the thing you are saying can be trusted. A claim that survived an adversary is worth more than one you only reviewed yourself.

Keep going

Free PDF: The Voice Audit Checklist. The by-hand version of the voice dimension, the one I still run first. No email gate.

What’s next

The five case studies at the results page are the ones that came out the other side of all 44 agents. The leaked CPC is gone, the overclaim is scoped down, the number I could not trace is replaced. If you are deciding who to trust with your accounts, the section to read is not the win. It is the “what I did not claim.” That line is the difference between someone showing you their work and someone showing you only the parts that flatter it.

Want a review like this on your account?

Want this kind of review
on your account?

Thirty minutes on the phone. Same person on the call as on the work. Walk out with a clear set of next steps.