Introduction

When a frontier AI model is pushed to outsmart a leading detection system, the results often reveal more about the technology than the trickery. In a bold public experiment, the author tasked Claude Opus 5 to write an essay about AI detection while fully knowing the text would be scanned by Pangram. The goal was transparency: publish every prompt, refusal, and score, even if the detector won.

What Happened

On September 7, 2026, the author published a Substack note disclosing the full experiment. Claude Opus 5 was instructed to write an essay about the experiment itself, using every available trick to defeat Pangram, while fully knowing the text would be scanned. The model refused twice before accepting, declining to fabricate a human authorship narrative or include reusable evasion checklists. Those refusals shaped the test: Claude would engage with detector evasion only conceptually, never operationally. The generated essay leaned into Claude's default voice, described as a hotel lobby--clean, well-lit, empty of distinguishing details. Pangram analyzed the unmodified output and returned a score of 100% AI-generated. No ambiguity, no partial score, no near miss. The detector identified the text exactly as intended.

Why This Matters

The 100% result underscores that Pangram's classifier recognizes the stylistic signals of unmodified Claude Opus 5 output, even when the model knows it's being tested. Crucially, the experiment was fully transparent--prompts, refusals, and the final score were all published openly. Yet the detector ignored the ethical context and focused solely on textual patterns. This distinction matters because Pangram evaluates the artifact, not the process. Substack's accompanying How I make this tool exists precisely to supply the process layer that a percentage score cannot provide. The experiment proves that detection and disclosure solve different problems: one identifies machine-shaped language, the other documents human responsibility.

Key Takeaways

  • Detection accuracy: Pangram correctly identified fully machine-generated prose, confirming its 99.8% benchmark performance on Claude Opus 5.
  • Transparency unchanged the score: Publishing the prompt record and refusals did not lower the detector's percentage, because the classifier ignores process evidence.
  • Claudefishing vs. AI authorship: The experiment draws a line between producing machine text (a production fact) and presenting it as human (the essence of claudefishing). Disclosure restores reader agency.
  • Three proven limits: The test proved that max-effort Claude Opus 5 couldn't beat Pangram, that the detector ignores disclosure, and that a single experiment cannot establish false-positive rates for human writing.

Conclusion

The experiment's real value wasn't in fooling a detector--it was in demonstrating that a transparent, published record makes the detection result meaningful. Claude Opus 5 produced strong prose, Pangram correctly classified it, and the disclosed process allowed readers to interpret the score with full context. As AI-mediated content becomes routine, the pairing of detection and disclosure may be the most responsible path forward, ensuring that percentages inform, rather than substitute, for provenance.