Pangram, a startup whose model scores any text for how much of it a machine wrote, has become the instrument publishing reaches for when it suspects an author. Its chief executive posted on X that 78% of the manuscript of Shy Girl was machine-generated; Hachette, which had planned to publish the self-released novel, cancelled it. Pangram scores have since been attached to an instalment of the New York Times Modern Love column (100%), a Commonwealth Short Story Prize winner (100%), the novel Daggermouth (60%) and a thriller sold for $2.4 million (97%). Substack is now building the detector into its platform so readers can check a post themselves. The company has raised $9 million and shipped a new model, Pangram 4.
A yellow ribbon over two glowing hands holding a third hand writing on a manuscript with a quill
Source: wired.com
The product returns a single estimated percentage, produced, inevitably, by AI. Substack has not disclosed the terms of the partnership or said whether it pays Pangram per check. A spokesperson said the platform is not taking an anti-AI position, only that readers should understand what they are reading. Some writers are wary of the integration; one said a single press of a button can destroy an entire career. Jane Friedman, a writer and publishing expert, says many people regard AI detectors with disgust and anger, and consider them at times as dangerous as the companies building AI, or more so.
Max Spero, 30, co-founded Pangram with Bradley Emi, whom he met at Stanford. He grew up in La Crescenta, a Los Angeles suburb, wrote code and was on his school's robotics team. At Google he worked on FLoC, the scheme meant to replace third-party tracking by sorting Chrome users into interest groups; Google shut it down after privacy complaints. He moved to the autonomous-vehicle company Nuro; Emi passed through Tesla and Absci, a biotech firm that uses AI. After ChatGPT appeared, the two saw a business in preparing for a world full of machine-written material. They founded Checkfor.ai and later renamed it Pangram.
They were late. At least twelve companies were already selling AI text detection, among them Originality.ai, GPTZero and Turnitin. Pangram performed well in several early independent tests and pushed into the lead. Six roles are open on its site; filling them would grow the company by 25%. The team is small and includes former teachers and one person with publishing experience, and a company representative noted that one researcher holds a degree in English literature.
The detection method has two parts. In what the company calls synthetic mirroring, Pangram takes human writing and asks an LLM to produce a close copy, so the model learns the texture of machine prose. The second part is hard negative mining: the company hunts through datasets for human texts that get mistaken for AI, mirrors those synthetically, and feeds them back into training. Spero says all of Pangram's datasets are used lawfully, partly because the service needs far less data than ChatGPT or Claude. Pangram sells into education, law, recruiting and creative writing, and creative writing is the largest share of its training material.
The numbers the company publishes about itself are where the story gets complicated. Pangram says its new model falsely flags human writing in 0.0041% of cases, down from a previously reported 0.01%. Spero says the service is deliberately cautious: when a text sits on the border, Pangram calls it human. That choice suppresses false positives and lets more machine writing pass unnoticed. By Pangram's own estimate, essays written by a person and then heavily reworked by AI still read as human 41.37% of the time.
Independent testing pulls in the other direction. A Notre Dame working paper titled "Why AI detection doesn't work for academic integrity" found that Pangram 3.2 classified lightly AI-edited academic abstracts as fully machine-written in 64–80% of cases. Run AI text through a humanizer — a tool that makes machine prose look more human — and Pangram caught it less than 4% of the time. Tuhin Chakrabarty, an assistant professor of computer science at Stony Brook University who works closely with the company, acknowledges the detector can be fooled; some writers have edited their prose deliberately, and others have written in an AI register on purpose to trigger a false positive.
The company also publishes its limits. Pangram works worse on short passages, especially under 100 words. Its free tier hands out a limited number of credits a day, roughly 2,000 words, and the median uploaded text is 350 words. And Pangram concedes that the same passage can score differently depending on whether it is checked alone or inside a longer document — which is exactly what happens when someone copies an excerpt out of a novel and uploads it on its own.
Put those together and the instrument looks well designed for one job and badly suited to another. A classifier tuned to say human whenever it is unsure is the right instrument for a teacher deciding whether to open a conversation with a student. It is a poor basis for a public percentage posted to X beside an author's name, because everything that makes it safe — the caution at the boundary, the sensitivity to passage length, the collapse under light editing in either direction — is invisible in the output. The number arrives looking like a measurement of a manuscript. It is a model's verdict on whatever fragment was pasted in, and the company says as much in its own documentation.
The Shy Girl episode is also less tidy than the company's account of it. Spero says he was tagged in a Reddit thread, was sent the manuscript file, uploaded it, published the result, and was later asked to comment by the New York Times; he thinks Pangram's role in the affair is overstated. The chain ran differently. A Pangram employee responsible for customer work discussed the story with a publishing industry analyst, who passed it to the Times. Critics including the investigative project The Drey Dossier established that the copy of the manuscript Spero scored had been downloaded from a pirate site. Asked about it, he said he had not examined the file's provenance and had not read the document, but uploaded it straight into Pangram. When the Commonwealth Short Story Prize winner scored high, Pangram ran every past winner of the prize and flagged three more possible cases. Spero maintains that, as far as he knows, a Pangram score has never been the sole reason a book deal collapsed — one input in a larger picture.
The network around those scores is small and mutually reinforcing. Chakrabarty first heard about Pangram from a post by Spero; after he signed up, Spero gave him API access, they became close friends, and the company has kept supplying him with credits for research. He used Pangram to scan 14,419 self-published novels and found that nearly 20% scored high for AI; the working paper was widely cited and its data became the raw material for the accusations against Daggermouth. He posts Pangram results publicly, defends the company regularly, and has met publishers to discuss detection — while arguing that Pangram should not be the only basis for a decision, and that a human read combined with a Pangram score is more useful. Chakrabarty is in a relationship with Todd Shuster, co-director of Aevitas, a large New York literary agency, and suggested Shuster meet the company. Shuster says the agency almost immediately had to start putting results to its authors, with Pangram returning 50 to 95% machine text; he now runs manuscripts and book proposals through it, advises Pangram, and introduces the company to publishers. Some authors get defensive or appear to hide the truth, he says; others describe plainly how they used AI, and in some cases he has asked for a rewrite in the writer's own voice and ideas.
That is not a chain of independent checks, and it should not be read as one. The researcher whose paper supplies the evidence gets his access from the vendor; the agent who deploys the tool on manuscripts also sells the vendor into publishing houses. Each link is defensible on its own terms and the whole thing still runs as a closed loop, with Pangram's output entering the industry at several points that all trace back to the same two people.
The rest of publishing is more guarded. Simon & Schuster and HarperCollins declined to comment on their use of detection tools; Hachette and Macmillan did not respond. A Penguin Random House representative confirmed editors may use approved detection tools as an additional signal, explicitly not decisive and only one part of a broader editorial process. Friedman notes that the cancelled deals that become public are a fraction of what happens: agents usually say nothing.
The bias objection has not been answered. Sam Illingworth, a professor of critical AI literacy at Edinburgh Napier University, argues detection tools broadly do not work as they should and can be biased against particular people. One study found such systems more often mistake writing by non-native English speakers for machine text, and neurodivergent writers have said their habits trigger detectors disproportionately. The study predates Pangram, and Spero points to internal data he says contradicts it. The three books at the centre of publishing's biggest detection scandals — Shy Girl, Daggermouth and Call Me, I'll Hide the Body — were all written by authors from racial minorities. Two of those authors did not respond to questions and one declined over legal risk. Regina Brooks, who is African American and president of the Association of American Literary Agents, says the industry needs to look harder at inequity: who gets tested and who gets publicly accused is decided by people, and carries their bias.
What the accused actually have is a bounty. In some cases where writers said they had been wrongly flagged, Spero offered money to anyone who could prove they wrote the text themselves — an offer he says he made when he was nearly certain the person was lying, while adding that if the author was telling the truth the company needed to know about the error. Nobody has claimed it. The absent piece of the story is the appeal: there is no process for a writer who has been scored, only a public number, a publisher already reading it, and an invitation to prove a negative. Defamation suits have been hinted at; Spero says that as far as he knows none has been filed.
The company's posture has already shifted once. Rod Breslau, a journalist and former Pangram contractor, offered himself to Spero as an attacking player, tired of watching machine text fill social media. He advised on social strategy, helped arrange interviews, browsed the web with the Pangram extension and publicly accused suspected AI users, wanting to push harder because he thought people were getting off too easily. Pangram ended the arrangement, in part because the strategy changed: the company now expects AI use to become more acceptable. Spero no longer appears to want to chase every user, though he says Pangram will speak up when authors mislead readers. The roadmap is to catch AI even under light editing, add granularity, explain results better, and give end users more information — with Pangram positioned as the arbiter of originality for publishing, education and beyond. Because the rest of the market is a set of black boxes, Spero says, the company intends to win trust by publishing technical reports on its models and training and by talking openly online.
Transparency is a reasonable strategy for a company that wants infrastructural status, and Pangram is more forthcoming about its failure modes than its competitors. It is also the wrong unit of remedy. The reports sit on the company's site and describe a classifier with a 41.37% miss rate on heavily reworked human essays and a documented collapse against humanizers. What leaves the building is a percentage, attached to a name, with a decimal point that implies a precision the underlying method does not have. Substack's button ships either way, and the number will travel a great deal further than the caveats behind it.