Stylometric Summer? Pangram Is Having a Moment
A flag is not a verdict.

Yesterday Granta published Jamir Nazir’s “The Serpent in the Grove” (archived) as the Caribbean regional winner of the 2026 Commonwealth Short Story Prize. By dinner time, Nabeel Qureshi had posted a Pangram detection report on the published text and the screenshot was everywhere I look. The Pangram report (archived) had it pegged at 100% AI-generated, high confidence. The verdict: We believe that this document is fully AI-generated. Twelve consecutive 400-word chunks, every one flagged at high confidence, no hedging.
I spent an hour with the story last night, running my own tools to see what I’d find.
If you’ve been paying attention this spring, we’ve all watched some version of this happen four or five times now. In April, Kelsey Piper at The Argument announced that Claude could identify her from 125 words of unpublished writing. Megan McArdle replicated the test in the Washington Post a week later. Two days ago Richard Hanania ran a competitive version on his own Substack readers, and found they could barely distinguish his real op-eds from Claude imitations of his style. Taylor Lorenz published a Pangram-powered survey at her newsletter User Mag at the end of April, claiming about a third of Substack newsletters in certain bestseller categories test as AI-generated. ArXiv announced on May 16 that it would impose one-year submission bans on authors whose papers contain unchecked LLM use. Yesterday: Granta.
Pangram is having a moment, and that moment is starting to spiral out of control.
What the day looks like
By this morning the case has become a Rorschach. On X, especially in literary and tech-policy circles, it is the story of the week, with Bluesky more muted. AI partisans are celebrating proof that ChatGPT can produce literary beauty at prize-winning quality. AI skeptics are taking the opposite lesson: the AI-character of the prose was (supposedly) obvious to anyone paying close attention, more evidence for the Russell paper showing that heavy LLM users detect AI writing more accurately. A nastier strand is treating the case as evidence against postcolonial literary prestige, MFA-style fiction, or inclusive prize culture generally.

My take, which is hot enough to be wrong: the submitter planned (or at least entertained) a Sokal-style reveal. Get an AI-generated story past the gatekeepers, accept the prize, then publish a self-exposing essay about how the institution failed to spot it. The point would be to embarrass Granta and the Commonwealth Foundation along with the prize machinery’s confidence in its own editorial judgment. Qureshi’s Pangram screenshot scooped that plan and pre-empted the reveal. The submitter now gets to watch the chaos, but has to decide whether to reveal themselves.
What is Pangram?
Pangram is, pretty clearly, the best commercial AI-detection tool on the market. That’s a technical claim, not a marketing one, but there’s millions of dollars at stake. The Russell, Karpinska, and Iyyer paper at the University of Maryland tested it head-to-head against several other detectors plus a panel of expert human annotators on 300 articles, and Pangram was the only commercial tool that matched human expert accuracy. Pangram’s own published technical report describes an innovative training procedure that iterates on binocular detection: for each human document in the training set, the team generates an AI-mirror that matches topic and length and style as closely as possible, then trains a transformer classifier to discriminate the human document from its AI-mirror. Because the mirror controls for nearly everything except the generation process, the classifier ends up learning whatever fingerprint LLMs produce that is intrinsic to LLM generation, regardless of register. This is why Pangram works across many text types where other detectors break; it’s also why Pangram cannot really tell you why it flagged a given text (the notes in the detector are put in later by another model), because the features its transformer learned to discriminate on are not features any human would name.
That is the whole problem: it’s a black box.
I cannot replicate Pangram with the computing resources I have (donations welcome!) though I can run a single-shot version of the underlying intuition. What you're about to see looked damning to me on first read. That is exactly why it was dangerous. Take the first 500 words of Nazir's story. Open a fresh GPT-5.x session. Paste those 500 words. Ask the model to write the next 150 words of continuation in the same voice. The model never sees what Nazir actually wrote next.
Here is what GPT-5 produced, blind:
...crouched low in the corner of the road like a crab with one claw, its boards soft from years of spill and sweat. Inside, shadows hung thick, the air sweet with cane and kerosene. Men lined along a counter scarred by bottles, trading coins meant for flour or lamp oil for measures of rum that burned like memory. One drink loosened the tongue, two made promises come easy, three drowned the small voice saying better sense. Zoongie moved between them with a sway that made stools shift. Some said she was Wilfred’s niece, some said she came out of the bush on a wet night, carrying her own thunder.
Here is Nazir’s actual next 150 words:
...leaned into the road like a rotten tooth. Inside, boards blackened by smoke and sweat, the air sweet with cane and forgetting. Coins meant for rice or kerosene slid across the counter and came back white rum hot as apology. One drink opened the chest, two turned fear into courage’s cheap cousin, three steadied the hand enough to write the future in invisible ink. She moved through that shop like heat through dry bush. They called her Zoongie. Maybe it was a name; maybe rain took a shape and decided to keep it.
Read them side by side. Both passages:
personify the rum shop with a simile in the opening clause (crab claw, rotten tooth)
describe boards weathered by sweat
produce the phrase the air sweet with cane almost verbatim
run a three-drink crescendo with the same one-two-three structure, and
introduce Zoongie with the same mysterious-origin framing.
The prose is different but the sequence of moves is nearly identical.
Feels like a closed case. The vocabulary overlap between GPT’s blind prediction and Nazir’s actual continuation came in higher than my matched control (a literary story GPT itself had written from scratch.) Nazir’s story looked more LLM-shaped than the LLM’s own output, at least according to a test designed by an LLM.
The temptation here is to post that finding and walk away.
The First Result Looked Damning
I didn’t. I’m trying to build something a bit more useful than another AI detector. Serious empirical work on text classification has a rule about single-shot results like this one: they are noisy. Maybe the rum-shop scene has so few alternative things to describe that any LLM continuation will converge on similar moves. Maybe after the phrase Wilfred’s rum-shop there aren’t many directions the next 150 words can go. So my rule is to run at least three more windows at different points in the story before treating any single-window finding as evidence of generation source.
So let’s look at: the scene where Sita lifts the planks over the well, the scene of vigil and recovery after the rescue, the years-later closing. At each, we can compare GPT’s blind continuation to Nazir’s actual using the same metrics. The vocabulary-overlap signal that looked so strong at the rum-shop window dropped substantially at the other three. The four-window aggregate placed Nazir between the known-AI control and a known-human control (a 2026 Granta memoir essay I used as the negative baseline), no longer above the AI control. The single-window finding had fallen inside the noise range the procedure exists to catch.
The methodology corrected my over-claim, which is, I think, what serious empirical work looks like when it works. But we don’t have the same insight into Pangram.
Against conclusions
Where I land is unsatisfying, because I’m the kind of philosopher who trusts uncertainty more than clean answers. I’m not going to tell you whether “The Serpent in the Grove” is AI-generated, even though the evidence looks damning, because the evidence I have direct access to won’t let me.
Pangram says 100% AI at high confidence. The distributional measurements I ran didn’t show the kind of smoothing we usually associate with LLM-assisted prose, including no obvious collapse in lexical variety, sentence-shape variation, or syntactic texture. For Apodictic, I built a craft-pattern audit, which found one elevated signal, but it looked less like generic machine prose than the story’s central thematic device repeating under pressure. The fixed-window mirror tests landed in ambiguous middle ground. The expanding-context mirror tests recovered discrimination in a direction consistent with GPT-family generation, but the only human control I ran was a memoir essay rather than another literary-fiction baseline that would have controlled for genre register issues.
So what I have is an evidence pack, not a verdict: several measurements pointing in several directions, each with its own caveats… but crucially, it seems to me, each replicable by anyone willing to run the same scripts. On prose alone, the case is more genuinely ambiguous than Pangram and the foofaraw let on, and the right next moves are institutional rather than computational. Did Granta or the Commonwealth Foundation run any AI-detection screening as part of the prize selection process? What documentation of drafting history does Nazir have? Is there an online footprint for the author that predates the prize? Since the headshot accompanying the submission appears itself to be AI-generated, the case is likely going to be resolved quickly. There’s a guy on the other end, or there isn’t.
Still, these are the questions we’d want a responsible institutional review to ask, and Pangram’s verdict ends the inquiry a bit earlier than I’d like, but just in time for a social media pile-on.
This matters beyond this one prize
The Nazir case is an instance of a pattern: it’s got a prestigious adjudication of merit and honor, a commercial verdict ripe for screenshots, trial by Twitter in real time, an author with thin online footprint, and a lightly-smoking gun.
Look… Pangram is very good. In Pangram’s own published numbers, the overall false positive rate is around 1 in 10,000, dropping to 1 in 25,000 on academic essays and roughly 1 in 11,000 on creative writing and short fiction (the domain Nazir’s story sits in). The Russell paper at UMD measured something closer to 2 to 3 percent on their evaluation set, which is higher… but not so high you’d want to lean on it. Those numbers don’t port cleanly to a case like Nazir’s, of course. A benchmark false-positive rate isn’t the same as a live accusation rate in a strange, adversarial, highly selected case. But they give us the scale of the problem. There’s a gap, but even the higher number is state of the art for commercial detection by a wide margin. As far as I can tell nothing else on the commercial market comes close!
That gap, though: the range (somewhere between 1 in 11,000 by Pangram’s own measurement and 2 to 3 percent by an independent academic one) is fine for some uses and disastrous for others. Pangram catches my own AI copy-editing on blog drafts, for what it’s worth, and I’m not ashamed of that. (Hire me a copy editor cheaper than Claude!) For student work at scale, though, even Pangram’s own creative-writing rate of 1 in 11,000 means roughly one false-positive accusation per large university per year if the tool is treated as decisive, and the Russell number means dozens to hundreds. That’s too high to run experimentally on millions of students, and faculty should not ask their universities to adopt Pangram as an automated adjudication system over student work, nor run it on every paper. Limited investigative use, though? If you see a specific signal and want a second opinion? That is a different question.
Suffice it to say, for cases like Granta, when something hits the public discourse, Pangram is a red flag.
But shouldn’t a flag start an investigation, rather than finish it? What I haven’t seen happen this spring on any of our contested cases is a Pangram verdict followed by the investigation that would tell us whether the verdict was right. The Nazir case is the closest we’ve come, and the hour I spent on it last night doesn’t settle the question either way. That’s the whole point of what I’m trying to build: we all have to get up to speed on this, fast!
Obviously, this is going to keep happening, but I don’t think a verdict is the right shape for the work it’s being asked to do. 100% AI at high confidence is a sentence that implies a certitude the underlying, auditable evidence can’t support. It can’t be probed for where the confidence comes from, and it can’t self-correct because there’s no reasoning trail, just a result from an older corpus of texts. When this kind of verdict is wrong, the accused writer has no recourse.
So I say: provide an evidence pack.
Distributional measurements, each one spelling out what it actually shows and what you can’t conclude from it alone. Sentence-length variance can tell you the prose is unusually flat. It can’t tell you whether the flatness is from an LLM or from a human revising too hard. Bad editors can produce flatter prose than any machine.
Craft-pattern audits that flag rhetorical tics LLMs lean on, then go passage by passage to ask whether each flagged passage is doing the story’s work or is just generic machine dressing. (Perhaps this isn’t an endorsement, but when SETEC flagged image conjunctions in Nazir’s story at 1.35x the literary baseline, the passage-by-passage review found every flagged instance was the story’s central metaphor doing exactly what the story needed. That’s just how it works.)
Mirror-discrimination tests where you can see the prompts used, the model versions named, and, again, a safeguard that warns that single-window results can’t resolve the question alone. (These are massive autocorrect machines, they’re good at predicting text.) You watched that one work in real time above: the rum-shop K=1 result was striking; the K=4 aggregate caught and corrected it.
Caveats spelled out instead of buried. The whole thing is replicable by anyone holding the same subscriptions.
That kind of output is an artifact a reader can argue with, agree with, or learn from. Using something like this, writers can respond to accusations, and institutions can defend their reasoning. They can also, of course, pre-run the detector and hide their tracks: that’s part of the reason the detectors have settled into black boxes. (I’m sure the money helps, too.)
Unlike Pangram, I don’t have a solution to sell. I do have one for free, though. It’s called SETEC, it’s open-source, and it runs on any laptop with Python (full disclosure: I built it, and it’s a Claude plugin). I call it “glass-box” because every step it takes is exposed and inspectable. You can see which measurements it runs, identify baseline texts and corpora against, and wrangle with what the numbers let you say, what all the math means. It’s slower than an AI-percentage, and less definitive.
Right now it provides three kinds of measurement. Distributional measurements look for the statistical smoothness that LLMs tend to leave behind: flat in sentence-length, flat in vocabulary, flat in syntactic templates. Craft-pattern assessments hunt the rhetorical tics LLMs overproduce: I’ve named most of them after their bad habits (“the kicker,” “the triplet,” “the manifesto cadence,” the disguised “not X, but Y” correctio I now notice myself using every. single. time). Voice-attribution scores compare the prose to a corpus of the writer’s own prior work, or to other people’s work, looking for drift and differences. Each measurement reports what it found and the tool then refuses to claim more than the underlying signal entitles. The mirror-discrimination test I walked you through above is the newest piece, still in active development. Everything I ran on Nazir’s story can be reproduced by anyone willing to install SETEC.
Whatever happens, the disciplines and practices of how this case gets adjudicated will outlast the verdict. For the next case, we should demand for our institutions to have access to evidence, instead of verdict-shaped outputs from instruments that can’t be inspected. That’s why I won’t let Pangram have its moment without butting in.
If you want to see the framework, it’s at github.com/anotherpanacea-eng/setec-voiceprint. Let me know if you’re suspicious of something but can’t run it, and I’ll try to run it for you!


