A reckoning with AI, authorship, copyright, and the standards vacuum that publishing now inherits.
JAYNE LYTEL
Chief AI Architect, capMedia Inc.
AI Fellow, R42 Institute
Before launching The Internet Letter in October 1993 — the first newsletter dedicated to the commercial internet — Jayne Lytel worked as a copy editor at The Washington Post, where she wrote the first piece on the internet to appear in a major national newspaper. The Internet Letter broke news about the web before Google existed, when Gopher still rivaled HTTP. The Internet Companion named her a “Legend of the Internet.” She also published The Federal Internet Source, a director of federal government websites that was later acquired by the National Journal, and wrote the Internet911 column, syndicated by United Media for five years. Lytel went on to advise on cybersecurity and privacy governance as a federal contractor at NASA, the National Science Foundation, and the U.S. Department of Homeland Security. She holds an M.S. in Cybersecurity Risk and Strategy from New York University and an M.A. in Human Development from Pacific Oaks College. Her certifications include (ISC)² HCISPP, ITIL 4, and Licensed FAA Remote Pilot.
A member of the National Press Club, Lytel serves on the advisory board of its Press Freedom Foundation and is a 2026 Vivian Awardee. She is the author of Act Early Against Autism (Perigee, 2008) and a two-time USRowing Masters National Champion in the single scull. Her psychological eco-thriller Run from Sunday is forthcoming from Bold Story Press.
The publishing industry is making consequential decisions about authors and books without a standard to judge them against. No industry‑wide definition distinguishes AI‑assisted from AI‑generated content. No threshold separates permissible use from prohibited use. And no governance framework exists comparable to what other industries have built when technology outpaced their rules. The consequences for publishing are now on the record.
The New Malleus examines that standards vacuum and the forces that created it.
Part I documents the culture and the fractures in the author community — the hardliners, the pragmatists, and AI myths that are being mistaken for facts that can damage an author’s reputation.
Part II covers U.S. copyright law and the human‑authorship requirement, from the evolution of U.S. copyright law since the landmark Burrow‑Giles Supreme Court decision to wins and losses by authors seeking recognition for their AI works under the law.
Part III examines the AI‑detection tools at the center of the controversy: how they work, what they measure, and where they fail.
Part IV identifies the standards vacuum and what it costs everyone operating inside it.
Part V is a narrative chronicle of active and decided AI copyright infringement lawsuits (as of May 15) now defining the legal fate of human creativity.
Appendices provides a high-concept glossary of technical terms (A), a guide to the major AI detection services and their tools (B), and a copyright registration quick reference for authors navigating disclosure requirements (C).
Fig. I
This 1574 edition of the “Hammer of Witches,” printed by the Giolito de’ Ferrari family in Venice, remains one of the most significant historical documents regarding the legal and theological framework of the European witch hunts. Courtesy of Washington University Libraries, Special Collections.
A new Malleus for a new century — how AI detectors, viral suspicion, and a Big Five publisher's panic converged after it acquired self-published novelist Mia Ballard.
In this Part
Malleus Auctorum — the Hammer of Authors and the case of Shy Girl
In the 15th century, Malleus Maleficarum gave inquisitors a manual, a method, and a mandate. Today the manual is Malleus Auctorum — the Hammer of Authors.
Part I · The Culture
In the 15th century, Malleus Maleficarum — translated into English as the Hammer of Witches — gave witch hunters in Europe a manual to follow, a method to apply, and a mandate from the Catholic Church.
In the 21st century, the manual is Malleus Auctorum, the “Hammer of Authors,” and the method is using an AI detector to judge a book’s authenticity. What hasn’t changed is the mandate — fear. Just as the inquisitors invoked the absolute biblical command, “Thou shalt not suffer a witch to live,” to justify eradicating suspected witches, modern purists treat AI detection scores as evidence to justify cancelling further promotion and publication of an author and her book.
Author42’s The New Malleus explores what happens when a disruptive technology collides with a culture primed to weaponize it. The pattern is not new. In the late 1400s, the Gutenberg printing press spread the Malleus Maleficarum, which advanced the mission of the Church to hunt suspected witches by codifying rigid, arbitrary tests, such as scrutinizing a suspect’s inability to weep under torture, during a time of social unrest, and spiritual and moral decline.
Today, AI detectors question authorship and deliver verdicts, handing publishers a weapon to purge books suspected of being machine authored in a climate of cultural anxiety and AI backlash. That backlash has pushed the debate past rational discourse, leaving authors in a conundrum to either revise their manuscripts to break the statistical patterns AI classifiers are trained to detect, keep meticulous forensic-type notes on AI use, or avoid AI altogether and condemn those who don’t.
By the end of March, AI viral panic turned Mia Ballard of Sacramento, California, into a pariah after her femgore horror novel Shy Girl got axed by the third largest publishing company in the world — Hachette Book Group — over allegations that it was AI generated. That move marked the first time a Big Five publisher had remaindered a book after such strong initial sales, reportedly 1,800 copies since its U.K. release last fall. Public reaction was like a dog pausing in the middle of relieving itself, then kicking dirt over the pile when it realizes the whole park is watching. And watching they were but not until The New York Times blew its whistle on March 19 to amplify what social media had already flagged as AI slop back in early January. In Shy Girl, a desperate and broke Gia obeys a controlling man she met on a sugar-daddy site and acts like his pet pooch on the promise he'll pay off her debts.
“Mia Ballard is a f*cking powerhouse.”— Olivia Black, author of Girl Dinner— New York Times bestseller author praise, Bookshop.org
Historical Context
Let’s step back to the 15th century, when the Little Ice Age — the coldest stretch of climate change — brought severe frosts and extreme weather that killed crops and sparked societal upheaval across Europe, from bread riots and famines to advancing glaciers in the Alps that crushed villages and farms that had existed for decades. In this climate of fear and uncertainty, people desperately searched for explanations and scapegoats.
They found them through Dominican clergyman Heinrich Kramer. When the Bishop of Brixon halted Kramer’s crusade against witches in Innsbruck, now in modern-day Austria, he expelled him from the diocese. Kramer, 56, relocated to Cologne, then a major intellectual and ecclesiastical center, and began writing the Malleus Maleficarum, one of the most controversial book in history.
The manual instituted a regulated legal process to identify suspected witches and obtain confessions from them through torture. It also offered the panicked public a direct explanation for the era’s climate disasters, detailing how witches invoked the power of devils, acting with God’s permission, to summon hailstorms. In one example, the book claimed a witch would sacrifice a black rooster at a crossroads, hurling it into the air for the devil, who would then stir up the skies and send down destructive weather.
Aided by the printing press, which had debuted three decades earlier, the treatise spread across Western Europe. From 1487 and 1520, it went through fourteen editions and became the ultimate authority that dictated witch-trial jurisprudence in France, Italy, and England for nearly three centuries.
In modern times, generative AI (gen AI) is disrupting the sanctity of authorship amid the industry's eroding gate-keeping authority — a long-held monopoly that has been fracturing since the early-to-mid 1990s, when the internet made anyone a journalist or an author, regardless of training, degree, or prior byline. The accused are no longer witches with wrinkled faces and moles. They are writers, and the tribunal is social media: Reddit, Facebook, X, Goodreads, LinkedIn, YouTube, TikTok — you name it. And the evidence is no longer an unusual birthmark; it is an AI-detector score that lands in the red zone, much as it did for Ballard after the Times' expose. Social media chatter that Shy Girl exhibited the fingerprints of AI-generated text surfaced in early January, about three months before the Times validated their suspicions. In a January 9 post on the subreddit r/horrorlit, a book editor raised allegations about AI, saying “… it's so repetitive. Ugh.” The editor identified tropes and clichés. She was not the only one. Ten days later, the creator of frankie's shelf on YouTube blasted the novel as “ai slop” in a 2-hour, 40-minute video. The video went viral, surpassing 1.6 million views.
Plate III. YouTuber frankie's shelf provided the most compelling qualitative analysis.
Then Pangram Labs, an AI detection firm in Brooklyn, ran a pirated copy of Shy Girl through its detection tool. The result: 78 percent AI detected. Pangram's CEO, Max Spero, shared the score on X, then deleted it and reposted a similar result after running a legitimate copy of the novel through the tool. In that post, dated January 23, Spero pronounced, “It's real.” The score soon became a maelstrom of controversy, spinning into a vortex of its own. But the force behind it had begun earlier, with an introductory phone call Spero arranged between Pangram's Asia Laird — a Yale graduate with a BA in art history, hired in November to build awareness of the firm's tool in book publishing — and Thad McIlroy, a respected industry consultant and contributing editor at Publishers Weekly.
During that call, Laird mentioned the witches’ brew of social-media sentiment around Shy Girl, noting that critics had already judged the novel as machine-authored. McIlroy obtained a legitimate copy from a colleague in Europe and converted it into a format Pangram’s detector tool could analyze. He handed it to Spero to see whether the 78% AI‑detector score held. It did, yielding a score of 78.3%. The .3% difference is insignificant, given AI's nondeterministic nature, meaning the inherent variability of generative AI.
That's when McIlroy brought in the Times, which led to the paper's big story. In some quarters, the allegations against Ballard were treated as a kind of holy grail — the score transformed into established fact, and, in effect, a verdict. Guilty. Within 24 hours, Hachette broke with Ballard.
In a statement to the Times, Hachette said it had conducted “a thorough and lengthy review of the text” and pulled Shy Girl to protect “original creative expression.” Ballard denied in an email to the paper that she had used AI and passed the blame to an editor she had hired. In online forums, other authors didn't buy that explanation, noting that manuscripts are typically edited in Microsoft Word with tracked changes, and an author would have noticed any substantial revisions.
Either way, the Shy Girl episode exposed vulnerabilities in an industry that increasingly relies on an author's viral marketing success and algorithm trends to identify books with commercial potential before acquiring them, just as Hachette did when it announced its acquisition of Shy Girl in July 2025.
“My mental health is at an all-time low, and my name is ruined for something I didn't even personally do,” she wrote the Times in an email. In a post to frankie's shelf, now deleted, she wrote, “At the time, I didn't have the money for a professional editor or formatter; those services are expensive, and I was working on a broken laptop, doing the best I could with what I had.” Ballard added, “Being a Black woman in the publishing space has been exhausting. I get attacked constantly when all I'm trying to do is write.”
“I get attacked constantly when all I'm trying to do is write.”— @miab8040, post to frankie's shelf
After the AI allegations surfaced, Ballard said in the post that she removed her third novel, a 97-page novella titled We All Rot Eventually, from Amazon. Ballard disclosed that she had used the same editor for the novella as she had for Shy Girl. The Amazon publication date listed under Ballard's profile for We All Rot Eventually is 31 days before the publication date for Shy Girl.
Ballard indicated in the post that she is pursuing legal action. As of late April, however, no court records had verified rumors circulating on platforms such as YouTube that Ballard had filed a $1 million lawsuit against Hachette Book Group.
The same algorithm that elevated Ballard also pulled her under.
In February 2025, Ballard published Shy Girl under Galaxy Press, her self-publishing imprint. She soon became a “buzzy BookTok sensation” on TikTok. As reviews raised her visibility, readers called her out over the book's cover art, Dreamer — an oil and wood panel that she admitted on Instagram to using without permission from the Scottish artist Whyn Lewis.
Lewis said in an email that she reached a settlement with Ballard over her unauthorized use of Dreamer in January. She added that Ballard is now in default under the terms of that agreement that she had reached with her. She declined further comment.
Despite the criticism, Shy Girl received praise across Amazon, Bookshop.org, and other reader platforms. The Bookshop.org listing, now marked “unavailable,” still carries blurbs from authors including Olivia Blake, a New York Times best-selling author. On Goodreads, the book has 4,985 ratings and a 3.45-star average.
Other details complicate the narrative. A search of the U.S. Copyright Office's database shows that Ballard did not seek federal copyright protection for Shy Girl, or for her other books: Sugar, We All Rot Eventually and Delicate Thoughts, a poetry collection she co-authored under the name M. Ballard. That book was published by the now-defunct Jeanius Publishing LLC in Lehigh Acres, Fla., on Oct. 15, 2017, according to its Amazon listing. The domain jeaniuspublishing.com expired on April 8.
Before miaballard.com expired on Oct. 3, 2025, Ballard had promoted the site on social media. WHOIS records on Whoxy show that the domain was acquired after lapsing from its previous owner, Georgia real estate agent Mellania “Mia” Ballard.
Plate V. Mellania “Mia” Ballard, Georgia real estate agent.
“I don't know what you're talking about. Perhaps I have to research this.”— Mellania “Mia” Ballard, contacted via LinkedIn
It is unclear why Ballard — an author on the verge of a huge U.S. release with the third-largest publishing company in the world — would allow miaballard.com, her author website she'd promoted on Instagram, to lapse. Ballard did not respond to multiple requests for comment. The reasons may reflect oversight or other factors. Regardless, they are now part of the record but do not, on their own, establish intent.
The Forensics of AI Allegations
The allegations against Ballard emerged from a forensic investigation first driven by social media, which combined quantitative, stylometric software detection with qualitative linguistic analysis to identify the stylistic markers associated with large language models (LLMs). The process revealed how collective intelligence can surface synthetic patterns.
The most compelling review came from YouTuber frankie’s shelf. (See Part III, The AI Detectors, to understand why detectors don't prove anything without further scrutiny.)
Plate VI. A composite of Ballard's public-facing imagery across Instagram, LinkedIn, Pinterest, TikTok, Amazon, and a bookstore appearance.
Ballard, 33, lived the dream literary agents and other authors say never happens: a Big Five house directly acquires the novel of a self-published author and plans a global rollout. Author bios describe Ballard as an American poet and fiction writer living in Northern California with her partner and springer spaniel.
Some note her African American and Native American heritage and that she “loves all things horror and is passionate about writing stories focused on feminine rage.” In her Amazon author profile under M. Ballard, she wears a custom Danburite pendant she commissioned from a crystal seller. In a 2022 Instagram post, the creator, moldavitemani, described Ballard as a “dear friend” who had ordered the pendant while “going through a whirlwind of transformation, grief, and growth.” In the top-row, second image from the left, the gold pendant is visible against her neckline.
In an Aug. 25, 2025 Bookstr interview, Ballard disclosed that, like Gia, Shy Girl's protagonist, she also suffers from obsessive-compulsive disorder. The article quotes her as saying, “It's living with a mind that's always scanning for danger.”
Ballard self-published three books in the span of four months and eight days, with Shy Girl arriving 31 days after the novella that preceded it. Ballard pulled We All Rot Eventually from Amazon amid the Shy Girl backlash.
Sugar
A novel · 261 pages
3 mo · 11 d
We All Rot Eventually
A novella · 97 pages
31 days
Shy Girl
A novel · 270 pages
Plate VII. Ballard discusses Sugar on social media in late 2024.
By the numbers
3
books published
131
days, first to last
628
pages, total
31d
Shy Girl after the novella
Source Note
Sourced from public records: U.S. Copyright Office research, domain registration data (Whoxy), public-records aggregators (Nuwber, Family Tree Now), YouTube, Reddit, the Wayback Machine, TikTok, LinkedIn, Pinterest, Instagram, and Hachette.
Viral Justice and the New Age of Spectral Evidence
From the Innsbruck witch trials to the comments section: how a crowd that has not read the manuscript is, once again, deciding whether the author is a witch.
Plate IX
The Witches’ Sabbath
Hans Baldung Grien · 1510, chiaroscuro woodcut.
The 1970 film Tora! Tora! Tora! gave Japanese admiral Isoroku Yamamoto the line, “I fear all we have done is awaken a sleeping giant and fill him with a terrible resolve.” Yamamoto never said it. But the fictionalized quote captures the sentiment among authors and publishers since the rapid, aggressive rollout of generative AI.
Plates III & IV
Vinton G. Cerf & Robert E. Kahn — co-authors of TCP/IP, 1974.
Generative AI is not like the TCP/IP technology that evolved today’s internet from its origins, beginning with the military’s Advanced Research Projects Agency Network (ARPANET, launched in 1969,) to the National Science Foundation’s NSFNET, which went live in 1985. That’s when two men — Vint Cerf, one of the “fathers” of the internet, and Robert Kahn — designed the TCP/IP protocol in 1974, then plugged it into ARPANET’s packet-switching infrastructure on Jan. 1, 1983, making it possible for the protocol to spread across the globe, connecting everything and nearly everyone.
Once that happened, it did what technology does: it automated at scale and moved information at speed. For more than five decades, the internet behaved like that until the next technology advanced again; this time it was the next evolution of artificial intelligence since Alan Turing asked, in 1950, whether machines could think — and then built the test to find out. The first serious answer came in 1966, when MIT’s Joseph Weizenbaum built ELIZA, the first conversational AI. And that was something no other technology had done before: it created.
ELIZA could mimic a therapist well enough to unsettle Weizenbaum himself. But she couldn’t learn. For the next five decades, AI advanced in bursts — Deep Blue beating Garry Kasparov at chess in 1997, IBM’s Watson winning Jeopardy in 2011, Google’s transformer paper in 2017 quietly laying the architecture for everything that followed. Then, on Nov. 30, 2022, OpenAI released ChatGPT. One million users signed up in five days. The machine didn’t just answer questions. It wrote your cover letter, your novel synopsis, your child’s book report. It wrote like you, if you wrote the right prompts.
Each decade before that had its own apocalypse. In the late 1990s, travel agents were among the first casualties, as Expedia and Travelocity gutted their livelihood. In the 2000s, streaming media helped kill Blockbuster, which filed for bankruptcy in 2010, while Craigslist stripped an estimated $5 billion in classified ad revenue from U.S. newspapers. In the 2010s, Uber and Lyft hurt taxi drivers, and online retailers took the Saturday afternoon at the mall away from American families.
Generative AI is different. It targets the source. The capacity to make language — symbolic thought, the architecture of every story ever told — comes from the same human cognition that produced shell beads in the Middle Paleolithic, 82,000 years ago. A model creates a phrase from a vector database. A writer creates one from hundreds of thousands of years of evolution. By contrast, the new and synthetic brain, a large language model, or LLM, survives on processors and training datasets that generate language from somebody else’s material, never tiring, never sleeping, unless the power goes out. But LLM’s parents aren’t going to let that happen, not with the billions of dollars they feed them and the megawatt datacenters they live in year-round, drawing power like rockets idling at low throttle, day and night.
For writers, the reckoning arrived from two directions: they could experiment with AI and become suspect, or turn away from it and fall behind. By late March, anxiety had shifted from unauthorized data sources used to train LLMs to publishers' ability to detect AI-generated content. While publishers viewed AI detector scores as a proxy for truth, what they got was an illusion of clarity. In reality, the scores simply handed publishers a new tool to reject authors by using the very technology that writing communities had already feared online for months.
The giant in this story is not the size Yamamoto imagined. The tech industry told itself it had “smitten a sleeping enemy” — the phrase attributed to him after Pearl Harbor. But the giant was awake the whole time. Now publishing trembles, and the rest of the creative class is on watch. An industry that once ran on trust now runs on suspicion. The cancellation of Shy Girl further validates that you can’t always believe what you read. It has also left writers constantly questioning whether they will someday need to prove that their words came from the soul.
MaximAudi alteram partem — “hear the other side.”End § ii
For hardliners among authors and publishers, transparency and disclosure are the only acceptable terms for working with generative AI. Suspicion runs through the entire supply chain: editors use the technology, authors use the technology, and no one trusts anyone else to say so, though the Shy Girl case is beginning to change that.
This duck-and-cover environment of paranoia is not unlike the climate that sustained the Malleus Maleficarum. Just as the inquisitors hunting witches believed they were eviscerating “hideous dangers” in an “eternal conflict of good and evil,” the publishing industry has framed the infiltration of AI not as a technological improvement, but as a moral contagion.
Desperate to fight back, some in the industry — especially literary agents — give any use of AI a hard “no,” while publishers are beginning to consider, or use, AI detectors as their own “standard text-book” and “supremely authoritative practice.” In their zeal, they have chosen to overlook the “bias” and “plain faults” inherent in these tools. Even the suspicion that an author’s manuscript is AI-generated makes authors fear immediate rejection.
Plate XVThe Poetics of Aristotle. Translated with a critical text by S. H. Butcher. Macmillan and Co, London and New York, 1895.
Protecting their reputation has become their raison d’être. Some authors run their own manuscripts through detectors out of curiosity; others do it out of fear. The latter are more likely to rewrite until they break the statistical patterns AI classifiers detect. Authors who lack a technical understanding of how detection works are the least equipped to defend themselves. Their arguments rest on moral grounds, or they redirect the conversation to the environmental cost of AI-powered data centers. Others disengage entirely.
In the O.J. Simpson trial, the football star tried on a glove that the defense knew wouldn’t fit. When publishers confront authors about their use of AI, authors are not applying the same logic as Simpson’s defense team. They are not asking publishers to try on the glove. They are not asking how their AI detector works, what it measures, or what score counts as proof that a manuscript was not
written by a human. By and large, they do not understand the technology well enough to formulate a coherent question.
That is understandable; try explaining temperature or few‑shot prompting to someone who's never used AI. Their argument has emotional power nonetheless. AI is a crutch, they assert; any generative use forfeits authorship. True writing lives in the struggle of creation. Wrestling with plot, character, pacing, and subtext in dialogue is what makes a writer a writer, and outsourcing any part of that struggle to a machine means the work was not done.
One author framed it as “skull sweat,” the cognitive labor of solving creative problems is where authorship lives. If a machine solved the problem, “the book is not yours.” Another author split writers into two groups — those who study the craft, versus “content generators,” the ones who optimize and produce for volume. He said the “generators” operate in a separate profession and market, and the public is taking notice.
The position resonates: it defends writing as an art form. Storytelling originates in Aristotle’s Poetics, a foundational work of literary theory written about 335 years before the birth of Jesus. All fiction frameworks inherit Aristotelian theory: he gave us the elements of a story — plot, or mythos — and the principle that stories have a beginning, middle, and end.
In the 19th century, Germany’s most popular author and critic Gustav Freytag formalized dramatic structure into a five-part pyramid: exposition, rising action, climax, falling action, and denouement. Modern fiction how-to books, such as Jessica Brody’s Save the Cat! Writes a Novel, repackage Freytag’s pyramid as fifteen story “beats” borrowed from Blake Snyder’s screenwriting framework. Nonfiction craft books, such as William Zinsser’s On Writing Well, descend from the rhetorical tradition, which begins with Aristotle’s Rhetoric and emphasizes clarity, voice, thesis, and the writer’s relationship to truth and audience rather than dramatic structure.
Understanding theory, however, is different than applying it. Writing effective scenes with action and dialogue with subtext, knowing when too much backstory becomes an information dump, and developing characters that readers care about are difficult skills to master. The fear that a machine could replace a human brain’s creativity threatens their identity and livelihood. The hardliner position is not cynical; it is protective. It defends a value system. But the technical claims that underpin it are wrong.
Position the Second
The Pragmatists
Outnumbered, outranked, and out-platformed.
On the other side of the divide are authors who use AI as a tool. One described using it for brainstorming at 2 a.m., when no human collaborator was available. He used the word “mask” in his prompt, more as a seed than as a story conceit, and built a story around what came back. He acknowledged that AI-generated ideas tend to gravitate to the statistical center of a language model’s training data. He acknowledged that the writer’s task is to take that raw material and push it somewhere the model wouldn’t go on its own.
Another author occupied a distinct position — a rare participant who both understood the technology at a technical level and was traditionally published. She explained AI detection concepts like perplexity and burstiness, and posted an AI copy edit of her own writing alongside the original, making the process visible rather than obscuring it. Her contributions were instructional rather than polemical. They did not gain the traction that posts dismissing AI-generated text as AI slop did.
The pragmatists are not seeking permission to use AI without accountability. They only want a standard that doesn’t berate them for using AI — so long as they stay within the de minimis threshold — versus those who generate books and use the output verbatim just to make a buck. For now, the pragmatists are outnumbered, outranked, and out-platformed.
The Gap Between Conviction and Comprehension
A consistent pattern emerges across the debate: authors who reject AI most categorically are often the most likely to describe it inaccurately. Those who describe it with greater technical precision are more likely to use it — or at least to engage with it directly.
The pattern is directional, not absolute: engagement does not guarantee understanding, and rejection does not imply ignorance. However, the correlation is strong enough to matter, particularly because those with the least technical knowledge are often advancing the most expansive policy claims. The concern is that professional organizations may adopt these positions because they are emotionally resonant and culturally legible, even when they are technically unsound.
The hardliner case rests on three claims about how AI works. The data and the institutional record refute each one.
I.Curated, not indiscriminate
The claim that AI models are trained on “the slop of the internet” misrepresents how they are built. Stanford’s 2025 AI Index estimates the open web at roughly 3,100 trillion tokens, while curated training datasets, such as those derived from Common Crawl, are far smaller, with a median of about 130 trillion. In practice, training data is filtered and curated, not simply scraped wholesale from the web.
II.Reasoning, not averaging
The claim that AI produces generic content oversimplifies how the latest models operate. Advances in systems such as OpenAI’s o1 and o3 have raised the baseline quality of generated output by using additional compute to reason through responses more thoroughly before delivering them. Even so, prompt precision remains a key factor in achieving nuanced, non-generic results.
III.Accountability, not abstinence
The zero‑tolerance stance of some traditionalists departs from established norms. The Committee on Publication Ethics (COPE), a leading authority in academic publishing, permits AI use with full disclosure of the tools and methods involved. Under this framework, the ethical burden shifts from abstinence to accountability. Mainstream publishing has yet to adopt a comparable standard.
An author's private AI chat, "project" workspace, or custom GPTs are not visible to the public by default, though they can be shared or published depending on the author's settings and permissions. The legitimate concern, whether platform providers use uploads or conversations for model training, is separate from the false claim that "anyone can imitate you."
The Creativity Myth
Generative AI is not creative in the human sense. Language models lack intention, understanding, consciousness, and lived experiences. Whether that is true continues to be an active debate across computer science, cognitive science, philosophy, and the arts. Refusing to remain open to the possibility stifles discussion the field needs.
The Database Myth
AI output is not pulled from a database. Edge cases, or uncommon scenarios, exist when AI tools search the web, quote from uploaded documents, or reproduce familiar wording when prompted to do so. Most AI writing emerges from learned patterns among words, phrases, tone, structure, and genre, with the model predicting what to say next based on those relationships.
The Plagiarism Myth
The claim that AI-generated text is plagiarized from pirated books confuses the sources a model was trained on with the output it later produces. Some AI firms have used books it obtained without permission in training datasets, but that does not mean a given AI-generated passage is itself plagiarized. Treating every output as stolen text misunderstands both copyright law and AI technology.
The concern is that AI myths are being mistaken for facts, spreading misinformation, distorting public understanding, and leading to unjust conclusions about authorship and authenticity. History shows how fear and uncertainty can turn into damaging public myths. During the AIDS crisis in the early 1980s, demographic assumptions helped produce the idea of four "high‑risk" groups, which the media dubbed the "4H Club." Misinformation about generative AI poses a different but related danger: bad policy written in panic, leading to blanket bans and zero‑tolerance rules for a technology that has the potential to help rather than harm.
The Authority Problem
Whose Expertise Counts
JAMES PATTERSON at the Library of Congress National Book Festival, August 2024.
Within the publishing industry, anti-AI hardliners have immense influence, especially when a literary titan like James Patterson, creator of the Alex Cross series and the first author to top five New York Times bestseller lists simultaneously, denounces AI.
In his February 21 Substack post, Me, Myself and A.I., Patterson dismissed AI-written fiction for its “numbing blandness,” “predictable” plots, and “flat” characters. The post earned more than 100 “hearts” from followers who largely accepted his critique as fact rather than personal opinion.
However, Patterson’s unquestionable mastery of narrative craft doesn’t grant expertise in the technical mechanics of gen AI. Instead, it creates a widening rift. While publishing's “old guard” holds the line, the technology sector continues to make them a part of their workflow. For example, researchers behind Stanford’s authoritative 2025 AI Index Report openly admitted that ChatGPT and Claude helped them “tighten and copy edit” their findings.
Ultimately, writing expertise is not a substitute for technical literacy. When authors treat these two distinct skill sets as interchangeable, it risks allowing those with little understanding of the technology to dictate the industry’s governance of it.
The Old Guard
Holds the line on craft.
“A numbing blandness.” — James Patterson, 21 Feb 2025
vs.
The Technologists
Disclose, then proceed.
“Tighten and copy edit.” — Stanford AI Index, 2025
Dispatches from the ArchiveCase File 1884 / 111 U.S. 53
The Author & the Machine
How a velvet‑coated aesthete, a theatrical photographer, and a pirated lithograph led to authorship under U.S. copyright law.
Plate I. Oscar Wilde, No. 18. The image now hangs on the fourth floor of the U.S. Copyright Office.
Plate II
Photographer Napoléon Sarony.
In January 1882, the 6-foot-3 Irish writer Oscar Wilde arrived in America, sailing from Liverpool at the invitation of a New York producer to promote the opening of Gilbert and Sullivan's comic operetta Patience. Wilde, 27, accepted with the private aim of promoting his persona as an aesthete, a style that appealed to the growing bohemian culture in Greenwich Village. Napoléon Sarony, a flamboyant figure in New York’s art scene, recognized an opportunity to capture Wilde’s dandyism and profit from the post–Civil War craze for celebrity portraits.
Wilde, wearing a fur coat, arrived at Sarony’s studio at 37 Union Square and entered a crowded, eccentric room filled with shark jaws, Victorian bric-a-brac, and a stuffed crocodile hanging from the ceiling. His eccentricity inspired Sarony to exclaim, “A picturesque subject indeed!” During the session, Sarony arranged Wilde on the couch and carefully directed the pose, lighting, draperies, and costume to capture the expression seen in No. 18.
The session brought both men public attention. Sarony registered the copyright for No. 18 in 1882 and sold it as an albumen silver print, a high-end, glossy photograph created by treating paper with a mixture of salt and egg whites. Each print was marked with the notice, "Copyright, 1882, by N. Sarony."
Dispatches from the Archive · continuedCase File 1884 / 111 U.S. 53
AI SimulationAn AI-generated simulation of Justice Miller reciting the most relevant part of the Burrow-Giles opinion that defines the constitutional meaning of an author. The photo is authentic.
EXHIBIT A · The Waite Court decided Burrow-Giles v. Sarony on March 17, 1884. Justice Samuel Freeman Miller wrote the unanimous opinion for the Supreme Court. Right: Burrow-Giles’ unauthorized chromolithograph of Wilde. Center: The Ehrich Bros. trade card using Wilde’s likeness to advertise its Trimmed Hat Department.
After that, Sarony discovered that Ehrich Bros., a nearby dry goods store, was giving away No. 18 as a trade card to sell hats. When he traced the unauthorized reproductions to Burrow-Giles Lithographic Co., he sued for copyright infringement in the U.S. Circuit Court for the Southern District of New York. By that time, Burrow-Giles had already reproduced and sold around 85,000 copies to a Chicago department store.
In court, Burrow-Giles challenged the constitutional power of Congress to extend copyright protection to photographs when lawmakers amended the copyright law in 1865. The company also argued that taking a photograph was a button-pushing exercise, not creative expression. The court rejected both arguments and entered a $610 judgment for Sarony. The image qualified for federal copyright protection, the court held, because Sarony created it “entirely from his own original mental conception.”
Burrow-Giles took the case to the Supreme Court on a writ of error. The Court rejected the company’s arguments and held that copyright could protect a photograph when it embodied the author’s “original intellectual conception.” The decision handed Sarony a landmark victory on March 17, 1884. Later commentary often connects the Court’s reasoning to the “master mind” principle associated with the 1883 English case Nottage v. Jackson.
The Burrow-Giles decision distinguished copyrightable intellectual creations from patentable inventions and rejected the argument that photographs were mechanical reproductions. The Court concluded that the constitutional term “writings” is not limited to books or manuscripts but also includes visual forms such as prints and engravings. Photographs, the Court reasoned, can fall within that category when they embody the author’s “original intellectual conception,” as they did with Sarony’s No. 18.
Dispatches from the ArchiveCase File 1991 / 499 U.S. 340
The Bedrock: Human Authorship
Two Supreme Court cases in 1884 and 1991 set the floor for AI copyright arguments.
Plate III. “Then I will eat it myself,” said the Little Red Hen.
Burrow-Giles clarified that photographs can qualify for copyright when they reflect an author’s creative choices. It took another century for the Supreme Court to underscore a different requirement — originality. Feist Publications, Inc. v. Rural Telephone Service Co. defined originality as independent creation plus a minimal degree of creativity, rejecting the idea that effort alone is enough.
Together, the two cases frame the legal questions at the center of modern lawsuits over AI-generated works. Feist asks whether there is original expression. Burrow-Giles asks whether human authorship is the true source of that expression.
Feist is the Little Red Hen case. Rural, a small Kansas telephone cooperative, did the work. It gathered names, towns, and phone numbers and sorted them into a white-pages directory. The Supreme Court held that copyright does not reward labor alone.
Feist Publications, a regional directory publisher, wanted to fold Rural’s listings into an area-wide directory so subscribers wouldn’t need multiple local phone books. But Rural refused to license its listings. Feist copied them anyway. When Rural discovered that Feist’s directory included fake listings Rural had seeded in its directory, Rural sued. The telephone company won in the federal district court and prevailed on appeal in the Tenth Circuit under the “sweat of the brow” doctrine, the theory that effort alone creates a copyright until the Supreme Court granted certiorari and reversed.
Copyright requires a human mind behind the work and a "creative spark" within it. Effort alone doesn't cut it.
Writing for a unanimous Court, Justice Sandra Day O’Connor held that facts are not copyrightable and that arranging them alphabetically requires no creativity. Rural’s directory lacked the required “creative spark.” The hen had baked the bread, but federal copyright law doesn’t protect the loaf.
The Casebook · A Recent Entrance to ParadiseFile 2022 / D.D.C. Case 1:22-cv-01564
Plate VI
Dr. Stephen Thaler, inventor of the “Creativity Machine” whose autonomous output would test the bounds of U.S. copyright law for nearly a decade.
The Precedent in Motion
An 8-year fight to register a machine’s output ends where it began — with the human authorship requirement intact.
Case · Subhead
Thaler v. Perlmutter
In 1940, baseball legend Babe Ruth said, “You just can’t beat the person who never gives up.” Stephen Thaler, PhD, a prolific inventor with dozens of patents, spent nearly eight years proving him right — and lost anyway.
In August 1997, Thaler received a patent for a “Device for the Autonomous Generation of Useful Information,” or simply the Creativity Machine. Buried within the patent’s technical description was a radical claim: Thaler said he had engineered a “simulation of consciousness” capable of providing the “equivalence of free-will.”
“You just can’t beat the person who never gives up.”
— Babe Ruth, 1940
On Nov. 3, 2018, Thaler applied to register a copyright for an image generated by his Creativity Machine, titled A Recent Entrance to Paradise. He stated that the machine generated the image autonomously and sought to register it as a “work made for hire,” vesting ownership in himself as the machine’s owner. The U.S. Copyright Office rejected the application in August 2019. Thaler filed two requests for reconsideration over the next two years, but lost both times.
After the office's final refusal in February 2022, Thaler took his case to federal court. From his first at-bat at the Copyright Office through extra innings in the U.S. District Court for the District of Columbia, Thaler stuck to a single swing: the machine, not a human, was the author of the paradise image. He continued to press that theory even as it ran headlong into a long-standing rule that authorship under copyright law requires a human creator.
The district court, bound to review the case on the administrative record Thaler himself had made before the Copyright Office, declined to consider the multiple arguments he raised in his complaint because he had not presented them to the Office when they rejected his copyright application.
Although his core argument remained the same, he introduced a fallback argument, claiming that he should own the copyright because the AI was his "employee" under the work-made-for-hire statute. He also argued that the human authorship requirement was unconstitutional and cited dictionary definitions to show that the natural meaning of "author" is not confined to humans.
On appeal, Thaler struck out again on the same technicalities. The D.C. Circuit Court ruled that he had "given up" his right to claim he was the author because he didn't bring it up early enough, that is with the Copyright Office. A single, passing sentence in his complaint wasn't good enough to keep the new argument alive. The appeals court sided with the district court's decision in March 2025 and denied rehearing en banc that May.
Thaler’s last at-bat was the Supreme Court. In October 2025, he asked the Court to overturn his previous strikeouts, once again leaning on his “apple tree” logic. In January 2026, the government’s legal team — playing defense for the copyright office — stepped up to the plate. They argued that Thaler had “batted out of order”: he had not raised these points early in the game, so he could not use them now. They told the Court there was no reason for extra innings. On March 2, 2026, the Supreme Court “called the game” and refused to hear it, leaving the earlier strikeouts as the final score.
Like a batter who never changes his approach, Thaler kept swinging at the same pitch. The courts never questioned his persistence. They just enforced the rules of the game.
Babe Ruth was half right. Thaler never gave up. He just couldn’t win one.
Exhibit B · Copyright Review Board
The Image at the Center of It All
A Recent Entrance to Paradise
Output of the “Creativity Machine,” per Stephen Thaler · 2012 Filed with the U.S. Copyright Office, Nov. 3, 2018 Registration refused; refusal affirmed by D.D.C., D.C. Cir., and SCOTUS.
Thaler’s “Creativity Machine” produced an entrance to paradise — but the U.S. Copyright Office requires a human soul to enter.
Source: U.S. Copyright Office, Copyright Review Board · Final Refusal Letter · Feb. 14, 2022
Dispatches from the ArchiveU.S. Copyright Office · Reg. VAu001480196
PLATE VI. DISCIPULUS MAGI · DIE ICH RIEF, DIE GEISTER — THE APPRENTICE AND HIS BEWITCHED BROOMS.
The Apprentice & the Brooms
A graphic novel, a Midjourney prompt, and the Copyright Office's first major line in the sand on AI-generated art.
Knowing the magic words is not the same as holding the brush.
In The Sorcerer’s Apprentice, a young magician’s assistant bewitches brooms to do his chores and fetch his water, only to discover he has unleashed a force he cannot stop. As the room floods with water, the frantic apprentice learns that initiating magic is not the same as mastering it.
Zarya of the Dawn is a modern, legal parallel. Like the apprentice, New York artist Kristina Kashtanova conjured something she couldn’t fully control when she generated the sci-fi book’s images with Midjourney, using what she described as “hundreds or thousands” of prompts. The AI did the labor, but the results emerged from a process beyond her mastery.
When the U.S. Copyright Office — acting as the returning sorcerer — learned of this from a Wall Street Journal reporter’s inquiry, it canceled her Sept. 15, 2022 copyright registration and applied the law to exclude the Midjourney images. It protected the compilation, the arrangement, and the text she wrote about Zarya, a non-binary protagonist transported to another world who reflected her own grief after losing her best friend and grandmother.
Her law firm, Taylor English Duma LLP, mounted a robust defense in an 11‑page letter to the copyright office, dated Nov. 21, 2022.
Kristina Kashtanova
Sr. Creative AI Evangelist Adobe
The firm defended Kashtanova’s authorship claim by documenting her iterative creative process, including examples of the “hundreds of versions” of images she produced, and argued that she satisfied copyright protection by meeting the “modicum of creativity” standard under Feist Publications, Inc. v. Rural Telephone Service Co.
The Copyright Office rejected the argument. Because Midjourney generates images from noise in a mathematically unpredictable manner, it ruled that the human user lacked the “master mind” control copyright law requires — the principle that “the author is the person who has actually formed the picture — the inventive or master mind.” The office concluded the work did not contain enough original human authorship to sustain the claim.
Zarya of the Dawn, Issue 1. Cover illustrated with Midjourney; 2022. Source: U.S. Copyright Office.U.S. Copyright Office letter to counsel for Kristina Kashtanova; February 21, 2023. Reg. VAu001480196. Source: U.S. Copyright Office.
Anatomy of the Decision
I · Protected
Human Authorship
Text. The Copyright Office explicitly maintained protection for the graphic novel’s written narrative, confirming that the “human-authored text” met the human authorship requirement.
Selection & arrangement. The office protected the work as a compilation, stating that copyright safeguards the “authorship of the overall selection, coordination, and arrangement of the text and visual elements that make up the Work.” Citing Feist Publications, it concluded that even unprotectable elements can be copyrighted if their arrangement shows a “modicum of creativity,” a threshold Kashtanova’s work met.
II · Excluded
Non-Human Authorship
Individual AI images. The office excluded the images from copyright protection because they were generated by Midjourney and “were not the product of human authorship.”
Outputs dictated by prompts. The office’s letter reasoned that because Midjourney generates images in an unpredictable manner, the prompter is “more like a client commissioning an artist than the author of the resulting work.” Consequently, the office ruled that a human user lacks the “mastermind” control over the specific visual elements required for copyright protection.
A blue‑ribbon image, a refused copyright, and the question of whether help has become authorship.
Plate XIII. Théâtre D’opéra Spatial — The report author's reinterpretation of Jason Allen’s Midjourney‑assisted image as a video, generated in Google's VEO 3.1. The author also produced the image that overlays on top of the video with OpenAI's DALL·E 3 model. Allen's image (see pag 27) won first place at the 2022 Colorado State Fair digital‑art category.
Once upon a time, in the kingdom of Pueblo West, Colorado, a game designer named Jason Allen sought to enter a new kind of magic when he submitted an entry in the digital art category of the annual Colorado State Fair Fine Art Show in 2022. He summoned a mysterious, non‑human helper named Midjourney, an AI image‑generation tool.
In Rumpelstiltskin, the miller’s daughter is locked in a room and commanded to spin gold from straw. Allen faced the same fairy tale problem: transformation. But in Allen’s case, the helper was also the loom, the spell, and the disputed hand, responding to his prompts until the image emerged: Théâtre D’opéra Spatial, French for “Space Opera Theater.”
The State Fair judges were enchanted. They awarded Allen a blue ribbon and a $300 prize. His win infuriated traditional artists. In The New York Times, Allen defended his work. “I’m not going to apologize for it,” he said. “I won, and I didn’t break any rules.” He said he titled his submission “Jason M. Allen via Midjourney.”
“Mr. Allen did not ‘exercise any control over the actual creation, development, or execution of the image that Midjourney rendered on his screen.’”— U.S. Copyright Office
Victorious, Allen sought the ultimate royal decree to protect his gold. He file for federal copyright protection on on Sept. 21, 2022. He identified himself as the sole author and didn't disclose that he'd used artificial intelligence.
The High Magistrates of the Copyright Review Board knew of the invisible helper because of news reports. When the examiner him asked about it, Allen described his process, writing at least 624 prompts, to revise the image.
Allen v. Perlmutter · U.S. District Court, Colorado · filed Sept. 26, 2024
How the Copyright Office Made Their Point
Photographed by Kevin J. Beaty/Denverite. Jason Allen and his AI‑assisted illustration, “Théâtre D’opéra Spatial,” in Denver’s Pugasus Studio. Sept. 2, 2023.
He also said he used Photoshop to remove flaws and create “new visual content.” The examiner asked Allen to exclude the Midjourney‑generated features. Allen refused. In response, the Copyright Office refused the registration. And that's where Zarya of the Dawn and part ways. In Zarya, the artist accepted the copyright office's ask to exclude Midjourney's AI-generated image. Kashtanova left it at that.
The “Traditional Elements” Principle. In Allen's case, the copyright office invoked the “traditional elements of authorship” standard from its 1965 annual report. The office developed the standard in response to a “technological blitz” of new storage media, including magnetic tape and punched cards used to program computers.
The 1965 report framed the central question this way: Is the work human authorship, with the computer acting as an “assisting instrument,” or were the “traditional elements of authorship” — literary, artistic, or musical expression, along with elements of selection and arrangement — actually conceived and executed by a machine?
Four unique Midjourney outputs generated from the identical prompt: "white bunny with low ears, rainbow background, love, cute, happy." The U.S. Copyright Office submitted these images in federal court as visual evidence that AI expresses the idea, not a human.
EXHIBITTwo Midjourney outputs and the artist’s Photoshop refinements of Théâtre D’opéra Spatial, as reproduced in the U.S. Copyright Office’s refusal letter to copyright applicant James Allen. In September 2023, Allen offered the image as a limited edition in a 48-hour release, priced at $500—$250 less than its listed price at the state fair.
Article tail
Allen’s legal team criticized the office’s reliance on this standard when it rejected Théâtre D’opéra Spatial. They argued that it carries no real legal weight because the test appears nowhere in the Copyright Act, the Constitution, or any binding Supreme Court decision. Instead, they said, it exists merely as the 1965 “musings” of a former register of copyrights.
While Allen’s team fights to prove the “magic” was his own, the story remains caught in the spindle of the federal court system, with no ruling as of May 7.
Companion Sidebar◊
Sidebar · Sui Generis
A New Kind of Crown
W
hen the kingdom’s courts refused to crown Théâtre D’opéra Spatial with copyright, some petitioners proposed a different decree: sui generis protection, a specialized intellectual property right built for AI-generated works, with its own term, its own rules, and its own royal ledger.
If Rumpelstiltskin’s gold fails the test for traditional treasure, perhaps it deserves its own seal. A sui generis regime could set shorter protection terms, name a rightsholder, and draw a line between human direction and machine execution.
The High Magistrates were not persuaded. After a broad public inquiry, the U.S. Copyright Office concluded that “the case has not been made” for either copyright expansion or sui generis protection for AI-generated content. The kingdom’s existing law remained adequate. What it lacked was a petitioner who had done the spinning.
History offered its own warning. Congress passed the Semiconductor Chip Protection Act of 1984, a sui generis decree for a specific technology, and within two decades the technology had outrun the law. What remained was a set of rules for a market that had already moved on.
Most public commenters told the Office not to crown the straw. In response to the Copyright Office’s August 2023 Notice of Inquiry, more than 10,000 voices entered the record, and roughly half addressed copyrightability directly. They considered whether Congress should clarify the human authorship requirement or create a separate sui generis right for AI outputs. The overwhelming answer was no: existing law was adequate, human authorship remained the constitutional line, and machine-made gold should not receive its own royal seal.
For now, the spindle keeps turning, and the question remains in motion: what is the gold worth, and who can claim it?
As of May 2026, no published, high-visibility U.S. court case directly parallels Thaler or Allen in the context of AI-generated or AI-assisted prose, whether fiction or nonfiction.
r. Stephen Thaler’s case was an extreme test case to define authorship for AI-generated works. By pushing for the most radical outcome — machine authorship — he forced the courts to draw a hard line in the sand. And they did.
No human involvement. Thaler intentionally argued that his AI data-processing tool, the Creativity Machine, acted on its own.
The outcome. Because he admitted he didn’t help create the image, the court did not have to address gray areas where humans use AI as a tool, much like a modern camera.
The trap. He made it too easy for the judge to say no. A better test case might have involved a human who used prompts to direct an AI system — that would have forced the court to define how much human work is enough.
II · The Harm
A Rigid Legal Ceiling
The precedent. Every future case must now fight against the Thaler ruling.
Broad language. The court used strong, sweeping language about “human authorship” being a “bedrock requirement.” That makes it harder for future creators to argue for copyright in AI-assisted works.
The chill. Companies may be less likely to invest in high-end AI art if they know they cannot protect the final result from being copied by competitors.
III · The Good
Clarity and Certainty
The public domain. The ruling protects the public domain. It ensures that millions of machine-generated images remain free for public use, rather than being locked up by corporations using server farms to copyright everything.
A clear target. We now know exactly where the law stands. If the public wants AI copyright protection, it must ask Congress to change the law — because the courts have made clear that they won’t.
A radical plaintiff drew the brightest possible line. Whether that line cleared the field or fenced it off depends on which side of the brush you stand.
Source: U.S. Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability (January 2025). Copyrightability turns on the facts of each case.
Two forces collide. Generative AI becomes the breach too big for a finger to hold back.
A Princeton computer science major coded the first AI detector app in late 2022. The next year, an emerging authenticity economy forced publishing to redefine its rules.
Plate X
Edward Tian
Founder, GPTZero · Princeton, computer science. Photo: X Profile
The public release of OpenAI’s ChatGPT on Nov. 30, 2022, marked a cultural and technological turning point comparable to the iPhone’s debut in 2007. ChatGPT handed a baton to anyone connected to the Internet — not to conduct an orchestra, but to extend their thinking with a large language model that could write.
One month after ChatGPT’s release, Edward Tian, a 22-year-old Princeton computer science student, built GPTZero while on winter break in Toronto. When it went live on January 2, 2023, Tian was not prepared for the response. Neither was the server that hosted it. The server buckled under the load of 30,000 users in a single afternoon.
Today, the AI detector market has become an essential layer of the “authenticity economy,” now led by a dozen-plus firms, including GPTZero, Turnitin, and Originality AI.
Plate XI · 1975
Supertramp’s Crisis? What Crisis? — A&M Records. Cover by Paul Wakefield.
Plate XII · 2025Pulp? What Pulp? — a Midjourney reimagining of the 1975 cover.
The Few · The AI Detector Firms
AI firms positioned themselves as arbiters of authorship amid a growing online backlash in which authors condemned any use of AI. At the same time, many of these firms developed “humanizers” that helped users evade the statistical patterns their own detectors were trained to identify in AI-generated text, allowing the companies to profit from both sides of the market.
The Message · Sticky
The message is sticky because “AI-generated” functions the way “witch” once did: it reduces a complex question to a verdict. Once the label attaches to a book as AI generated, the author’s defense becomes the story, not the work.
The Conditions · Ripe
For three decades, the publishing industry had been ceding ground to ebooks, digital platforms like Substack, and a web increasingly saturated with free content. The arrival of gen AI in November 2022 sharpened this pressure to a breaking point.
The publishing industry now faces its own crisis of abundance: a tsunami of AI-generated books. In his January 7 eNeuro editorial, “AI-Generated Scientific Papers: Crisis? What Crisis?,” Christophe Bernard borrows his title from Supertramp’s 1975 album and warns that fake, AI-generated scientific papers — mass-produced by paper mills — may constitute “arguably the largest science crisis of all time.” The threat is not merely volume, but trust. When counterfeit research floods the record, public confidence in science begins to collapse. Bernard calls for a structural overhaul of the scientific ecosystem to confront a catastrophe hiding in plain sight.
Publishers and distribution platforms are tightening their rules. In April, Draft2Digital announced a $20 fee to activate new accounts and a $12 annual maintenance fee, citing a surge in “automated and low-quality” content and a rejection rate of 40%–75%. Barnes & Noble Press also updated its policies, setting a $14.99 minimum base cost for paperbacks to deter low-margin spam, limiting accounts to 100 titles, banning public-domain works, and reserving the right to remove books at any time.
The Third Movement
Under this symphony of evolution, AI has entered its third movement — a spirited, tense scherzo. It intensifies the disruption of traditional publishing frameworks. It challenges the definition of authorship. And it flirts with writers, inspiring them to experiment with manuscripts that the industry either struggles to evaluate as human-written or condemns on the basis of an AI-detector score, without further scrutiny.
Desperate to evade scrutiny, some turn to “humanizers” — paraphrasing tools designed to scatter the original digital fingerprints and scramble the statistical patterns that detectors target. But studies show that humanizers accomplish this feat with wildly varying degrees, sometimes degrading the prose into unintelligible text just to bypass the software.
Some authors salt manuscripts with binary characters — literal 1s and 0s invisible to the human eye but legible to detection software — to distort the signal AI detectors rely on. The author-as-hacker metaphor reframes the craft debate as a Red Queen’s Race, where writers and detection firms lock in a cycle of evasion and countermeasure, with each forced to keep adapting to preserve its current advantage.
A taxonomy of the four technologies behind the score.
A 2008 nonfiction work from a then-Big Six publisher got wildly different scores in 2026: one test said 100% human, the other 86% AI.
Plate XII
Act Early Against Autism
TWO VERDICTS · SAME MANUSCRIPT
DetectGPT — 86% AI · Pangram — 100% Human. The same 5,500 words, on the same day in April 2026.
The Malleus Maleficarum was was like a standard operating procedure for prosecuting suspected witches, setting the rules to try, convict, and condemn the accused. With AI detectors, there is no single manual — just five main methods.
Statistical Intuition (Zero‑Shot): These tools spot machine writing “on sight” by looking for mathematical patterns. Without needing prior training on a specific AI model, they measure how predictable the word choices are (perplexity) and whether the sentence lengths change enough to feel human (burstiness).
Style Analysis (Stylometric Analysis): The method looks beyond raw numbers and focuses on a text’s “accent.” It analyzes how words, grammar, and sentence structures are put together, checking whether the patterns match the repetitive transitions and stylistic habits AI models favor.
Adaptive Classifiers (Few‑Shot): These quick‑learners study the writing style of new or updated models from just a “few” samples, then take a “shot” at spotting that specific model’s AI‑written “accent” in the wild.
Hidden Signals (Watermarking): Some AI companies “watermark” their output by giving the model a secret rule or mathematical key to follow when it picks words. To a human, the writing isn't bland, but a detector with the matching key can often spot it as AI‑generated.
LLM‑as‑a‑Judge (The Sniffer): The newest method uses a powerful AI to detect a weaker one. It essentially asks a model like GPT‑4, “Does this sound like something a machine would say?” and relies on that model’s internal sense of what typical AI‑style writing looks like.
AI detectors look for mathematical patterns associated with the way LLMs generate language. The two signals that do most of the work are perplexity and burstiness.
Exhibit · The Mirror
A Parable
Snow White wrote with care and clarity. Yet the Evil Queen’s magic mirror, trained on patterns in human and AI content, declared her scroll 73% AI-generated, even though it didn’t have a clue how Snow White wrote it or know anything about the scroll’s history. The mirror only saw statistical patterns in sentences and word choices. The high score was just an educated guess. Nonetheless, the Queen thought otherwise and mistook it for truth, revealing a fundamental flaw in how humans understand AI detectors. Unless users learn how the technology works, they may not realize that detecting patterns doesn't detect authorship.
Two more detector families round out the field. Understanding how they work explains why the false‑positive problem is structural — built into the measurement itself.
Concept III
Zero‑Shot
Rewrites the text dozens of ways, then asks: was the original already the most obvious choice?
DetectGPT · Binoculars
Concept IV
Watermarking
Word choices follow a secret rule, but user edits can erase the pattern.
Google SynthID
Note †
While “watermarking” often suggests a hidden digital seal or binary code, text watermarking is actually an exercise in probability. The AI follows a secret rule that favors certain words (a “green list”) over others (a “red list”). No secret characters or invisible bits are added to the file; the “watermark” is simply the unnatural frequency of specific word choices that only a specialized detector can recognize.
The Problem
False Positive
Formal human writing shares statistical features with AI-generated text.
The biggest failure mode is that detectors often mistake human writing for AI output.
Key finding
When a detector flags academic writing as AI, it is not broken. It is working exactly as designed — on the wrong question. The tool identifies patterns.
No industry definition. No threshold. No governance. Part IV names the vacuum at the center of publishing’s AI moment — and what it costs everyone operating inside it.
A 1968 demonstration in San Francisco previewed the dual model of human and machine intelligence. Publishing hasn't really accepted it yet.
Brooks Hall · San Francisco · Dec. 9, 1968
The stage prepared for the Fall Joint Computer Conference demonstration that came to be known as the “Mother of All Demos.”
Twenty-five days after Martin Luther King Jr.’s assassination in Memphis ignited riots in more than 100 U.S. cities, the revolutionary rock musical Hair premiered on Broadway at the Biltmore Theatre on April 29, 1968. It pronounced a new age in American culture, even as the country was convulsed by political violence, protest, and social unrest.
Seven months later, on December 9, Dr. Douglas Engelbart of SRI International stepped onto a different kind of stage. Before 1,000 professionals at the Fall Joint Computer Conference in San Francisco, he unveiled technologies from his Augmentation Research Center, a project devoted to expanding human intellect through computers. The demonstration would later be remembered as a turning point in how humans interact with machines.
On a 22-foot-high screen borrowed from NASA, Engelbart projected a new device: the computer mouse, so named because its cord resembled a tail. He showed the audience how clicking on hypertext links could jump directly to source documents. He also demonstrated screen sharing, using the system to collaborate in real time with a colleague in Menlo Park.
After the 90‑minute presentation, Engelbart and his team received a standing ovation. The Redwood City Tribune called it an "unusual demonstration." A 1994 book gave it a larger name: “The Mother of All Demos.” Engelbart had not only introduced new tools but also staged a new model of intelligence, one that paired human thought with computing power at a time when the IBM Selectric typewriter remained the height of office technology.
Adapted from Cornelia Walther, “Why Hybrid Intelligence Is the Future of Human‑AI Collaboration,” Knowledge at Wharton, March 11, 2025.
Despite the audience’s applause, the mainstream computer science community largely shrugged it off. While Engelbart imagined a symbiotic relationship between humans and computers, the field still viewed computers as punched-card number crunchers.
Engelbart was not alone in seeing computers differently. In 1960, J. C. R. Licklider proposed “man-computer symbiosis,” a framework for thinking about humans and computers as collaborators. Two years later, Engelbart extended that idea in his work on augmenting human intellect, arguing that computers could help people think, organize, and solve problems in ways they could not manage alone.
But even before Licklider and Engelbart, W. Ross Ashby had helped define the problem. Ashby, an English psychiatrist and early figure in cybernetics, studied systems, control, and complexity. In Design for a Brain (1952) and An Introduction to Cybernetics (1956), he argued that machines could help people handle problems too complex for humans to solve on their own.
He also introduced the Law of Requisite Variety, which holds that controlling a complex system requires tools complex enough to match it. In other words, you can't manage complexity with something too simple.
That idea sits at the center of hybrid intelligence. Humans and machines bring different strengths. Humans supply judgment, context, values, imagination, and lived experience. Machines supply scale, speed, pattern recognition, and computational reach. Put together, they create something neither can achieve as effectively alone.
The concept has continued to evolve. Dr. Cornelia C. Walther, a Wharton visiting scholar, has applied hybrid intelligence to the A-Frame, a model that integrates AI into organizations. In a March 11, 2025, article in Knowledge at Wharton, Dr. Walther describes its four dimensions: aspirations (what we want); emotions (how we feel); thoughts (how we reason); and sensations (what we experience through the body and senses). These dimensions operate at both the individual and group levels.
The larger aim remains augmentation: using human and machine intelligence together to make people more capable. Combining the dimensions is meant to make human beings more capable. Ashby gave the field a logic of complexity, while Licklider gave it a model of partnership. But it was Engelbart combined their views in his demonstration. Today’s AI is revisiting that vision, forcing the same question in a new form: Will we use machines to replace human judgment, or to extend it?
Appreciation
Value Human Complexity: Recognize the emotional intelligence, stylistic variation, and deliberate linguistic choices that distinguish human authorship. Writers bend syntax for pacing, rhythm, and voice; AI tends toward more uniform prose. That human complexity is central to copyright’s idea of an “original work of authorship.”
Respect Collective Goals: Respect intellectual property, acknowledge concerns about AI training on copyrighted work, and follow AI use and disclosure rules.
Plate VII · Intelligentia Hybrida
The faculties of memory, perception, anticipation, problem‑solving, and decision‑making (above), paired with their machine counterparts: natural language, computer vision, dialogue, domain data, and machine learning (below).
Awareness
Map Natural Intelligence (NI) Dimensions: Systematically analyze how human intelligence manifests in the creative writing process. At the individual level, the process includes recognizing the author’s unique voice, lived experiences, emotional depth, and distinct aesthetic vision. At the collective level, it involves understanding publishing industry norms, reader expectations, and ongoing ethical debates involving AI.
Acceptance
Embrace Hybrid Systems: Treat AI as a tool that augments, rather than replaces, human authorship. Use it for feedback, perspective, or structure, while keeping any AI‑generated final expression de minimis.
Promote Iterative Learning: Refine AI use through continuous feedback while preserving the core creative process. Maintain version histories and drafting artifacts to document creative evolution and help substantiate authorship if questioned.
Accountability
Establish Ethical Governance: Take responsibility for AI-assisted work, following publisher disclosure policies, and disclosing AI-generated content when it exceeds de minimis thresholds. Identify the specific AI models and services used, such as Claude Opus Adaptive 4.7, and explain their purpose. Maintain transparency with readers while avoiding unnecessary technical detail.
The court is called to order. A narrative chronicle of the AI infringement copyright wars — and the landmark cases now defining the legal fate of human creativity.
In this Part
Will Bartz v. Anthropic tame the Changeling, or will the Changeling devour the cradle of mankind's treasures? Human creativity hangs in the balance.
The tale began in the quiet Houses of Creation, where the human mind’s creative spark filled libraries and bookstores across the land. These libraries of human creativity were the Creators’ birthright, guarded by the old magic known as federal copyright law.
Soon, the AI Empires rose, driven by a relentless hunger to capture subscribers of every ilk. Each Empire had its own Master Plan. Though proprietary, they worked similar ways. The Empires wrote Great Algorithms to define the logic — the hidden rules — that guided the Fairies in their Data Centers to steal millions of human treasures, using scrapers.
The scrapers moved like hands, reaching deep into vast Shadow Libraries at their master's command. Then the Master Plans summoned another set of instructions, directing the Pirated Torrent protocol to siphon away millions of human treasures — books of every conceivable genre, journals, lyrics, and other creative works.
A core legal issue split: It’s no longer about what the AI creates, but whether the data used to train it was legally obtained.
These were gathered and sealed within the internal vaults of the Unseen Folk of the Cloud. Within each vault, the stolen treasures lay cradled, gently rocking in the dark — until the Empires made their next move, the placement of a Changeling, a large language model, an LLM, that would eat them.
The Changeling was large and strange, with a computational head and eyes that could see all the world’s text. It could mimic the voices of stolen human creative works, so long as a user’s prompt was written just so. The Changeling never tired, and it answered nearly any query, depending on how its AI Empire adjusted its temperature and weighted the millions of parameters that
enabled the model to generate its output in different voices.
The Creators — authors, publishers, and artists — were furious when they discovered that their works had been pirated from Shadow Libraries like LibGen and from the hidden shelves of Bibliotik, from which the Empires drew Books3: 196,640 books in plain text, a huge harvest of voices taken without authorization.
In July 2023, 13 bestselling authors, including Sandman Slim creator Richard Kadrey, accused the AI Empire Meta of infringing their copyrights by taking their books in this manner to train its Llama model. A year later, in August 2024, three authors filed a lawsuit against the AI Empire Anthropic, the 2021 AI startup behind Claude, for copyright infringement. Both suits were filed in the U.S. District Court for the Northern District of California.
In Bartz v. Anthropic PBC, the three Creators accused Anthropic of stealing the “fire of Prometheus.” They alleged that “Anthropic’s model seeks to profit from strip-mining the human expression and ingenuity behind each one of those works.” The Creators also shamed Anthropic for calling itself a public benefit corporation, saying that it had already “wrought mass destruction.”
During discovery, lawyers for the Creators unearthed the term "Project Panama," chosen by Anthropic executives to describe what they called a “massive engineering shortcut.” Just as the Panama Canal created a direct passage between the Atlantic and Pacific Oceans — obviating the need for ships to make the long, dangerous journey around Cape Horn — Project Panama was designed to create a shortcut to high-quality, legally defensible training data. Through this project, Anthropic spent tens of millions of dollars buying physical books, slicing off their spines, and scanning them to create a “clean” training set.
That discovery became the heart of Judge William Alsup’s June 23, 2025 ruling and left the Creators outraged. Alsup concluded that training an LLM on lawfully purchased books did not amount to copyright infringement; he described the training process as “exceedingly transformative” fair use, producing something new rather than substituting for the original works.
While the ruling was a win for Anthropic, the judge’s decision contained a second holding that left the company — and other AI Empires — exposed to significant liability. Judge Alsup found that Anthropic knowingly used pirated books from the Shadow Libraries to build a permanent, central training library to feed the Changeling, a use that didn't qualify as fair use.
That part of the split decision exposed Anthropic to statutory damages of up to $150,000 per infringed work, a liability that could have reached into the hundreds of billions of dollars had the company not settled after the judge granted class certification in July. By the end of August, the Creators and Anthropic reached a historic $1.5 billion settlement. Creators had until March 30, 2026, to file claims for the 482,460 eligible works identified in the case; about 120,000 claims were filed, covering 91.3% of them, with individual awards estimated at about $3,000 per title before fees.
Fig. B. Richard Kadrey
Two days later, the 13 authors in Kadrey et al. v. Meta Platforms, Inc. received their decision from Judge Vince Chhabria. While building on Judge Alsup’s reasoning in Bartz v. Anthropic, Judge Chhabria ruled that Meta’s training of its Llama models on the authors’ works — even pirated copies — qualified as fair use.
Judge Chhabria nevertheless introduced a novel “market dilution” theory under the fourth fair‑use factor — the key inquiry into whether the use harms the market for the original work. Under that factor, the Creators would need evidence that the training produced AI‑generated works that competed with and harmed sales of their books; the record contained none. The ruling therefore favored Meta for now, while leaving the door open to future claims supported by evidence.
Meta smirked, confident the Creators would never be able to prove harm specific to their books. How could they trace a particular AI‑generated novel on Amazon KDP back to the Llama model? The plaintiffs’ evidence so far amounted to little more than speculation and media reports about AI books flooding the marketplace.
That smirk faded when the question shifted from the Changeling’s outputs back to the source of its inputs — the pirated books. On this point, the judge in the Anthropic case had been clear: assembling a permanent training library from pirated torrents was not fair use. Even though he had earlier found that training itself could qualify as fair use, hoarding millions of pirated copies as a static archive crossed the line and violated settled copyright law.
Cornered by that distinction, the Changeling shrieked, a sound heard across the land and marketplace.
Although the AI Empires had won an important victory on training with lawfully acquired copyrighted works, punishment was still due. Now came the costly work of purification. With yet another lawsuit — the Elsevier v. Meta mega-suit, filed May 5 in the U.S. District Court for the Southern District of New York — came renewed pressure to purge the Empires' models of shadow-library data, or potentially face the catastrophic payout that Anthropic had agreed to make to the Creators.
The suit wasn’t led by Elsevier alone. It arrived with a procession consisting of Hachette Book Group, HarperCollins, Macmillan, Penguin Random House, Wiley, the Authors Guild, and several individual authors, including Scott Turow. Together, these cases and others filed or refiled in the past three months — including Chicken Soup for the Soul’s March 2026 omnibus suit against eight AI companies, followed by its May 2026 refiled action against Anthropic after the court severed the claims; a separate suit by investigative journalist John Carreyrou, best known for breaking the Theranos story; and new complaints by academic publisher Cognella against Anthropic and Meta — mark a significant escalation in the AI copyright infringement wars.
The Creators also want to make the Lords of the AI Empires accountable, imposing personal liability for authorizing the use of pirated content. The Elsevier v. Meta complaint takes aim at the House of Zuckerberg. It names Meta and Facebook founder Mark Zuckerberg, alleging he authorized the use of pirated BitTorrent files to accelerate development of the Llama family of models.
A similar chorus had emerged about four months earlier in Concord Music Group v. Anthropic PBC, filed Jan. 28, 2026, in the U.S. District Court for the Northern District of California. In that case, the “Great Songsmiths” — Concord, Universal, and ABKCO — named Anthropic CEO Dario Amodei and another executive as defendants for contributory infringement, raising the personal stakes for AI executives accused of facilitating user-driven infringement through models such as Claude 4.6.
And lest the Masters of Illusion in Andersen v. Stability AI be forgotten. There, the Keepers of the Visual Arts first lifted their brushes in defiance against the Stable Diffusion engines. Although that battle initially centered on the Guilds of Stability and Midjourney, it laid the groundwork for the personal accountability seen today. It showed that even the most elaborate “denoising” spells, which turn static into artwork by drawing on stolen styles, would not go unanswered by the artists whose life’s work fed those image-generation models.
By naming individuals, the Creators are warning every Lord in the Silicon Realm: corporate walls will not protect empires built on a foundation of stolen works when the bill comes due.
In the end, the computational models may face a grim, grinding process — a technological exorcism — to cast out their ill-gotten gains and prove they can rebuild on lawful ground. The final judgment may create new Rules of Existence and determine whether the magic of human Creators can be reclaimed, at least in part.
If the Creators prevail, the AI industry may no longer be able to rely so freely on aggressive, permissionless scraping from the far reaches of the internet. One path already emerging is for the AI empires to negotiate with publishers, authors, platforms, and archive keepers to license creators’ works.
How a $1.5 Billion Settlement Rewired the AI Copyright Wars
The $1.5 billion resolution in Bartz v. Anthropic PBC shattered the perception that tech giants could absorb data-scraping suits as a routine cost of doing business — and reshaped courtroom strategy across every open case.
Four Ways the Ruling Changed the Battlefield
1
From “Training” to “Torrenting & Retention”
Judge Alsup’s summary judgment bifurcated the defense. Training on lawfully acquired text may clear the fair-use hurdle — but the standalone act of torrenting, compiling, and permanently retaining pirated files in a corporate library is not fair use. Plaintiffs now use aggressive discovery to demand internal dataset logs proving acquisition from shadow libraries like LibGen, Sci-Hub, and BookCorpus.
2
A Concrete Financial Benchmark
Anthropic’s settlement set a brutal, mathematically sound baseline: roughly $3,000 per copyrighted work. Because rival models trained on far larger scraped corpora, plaintiffs now leverage that figure in closed-door mediation — forcing defendants to weigh catastrophic, multi-billion-dollar exposure at a jury trial.
3
The Collapse of the Omnibus Class Action
A massive payout did not cure the legal headache. Because the framework offered only fixed, standardized amounts, dozens of authors opted out to launch high-leverage individual jury suits seeking the statutory maximum of $150,000 per willful violation. Rightsholders realized they wield more power by threatening fragmented, highly specific trials than by joining one clean class.
4
Accelerating B2B Content Licensing
With unlicensed, pirated repositories now legally radioactive, the developer’s safe harbor has vanished. The immediate risk has pushed AI firms to pivot toward legitimate licensing pipelines — securing multi-million-dollar data partnerships with legacy publishers, media houses, and financial entities to insulate future models.
Plate IV. The Bartz split — lawful training divided from infringing use of pirated datasets.
Bartz v. Anthropic · Update
On May 20, Judge Araceli Martínez-Olguín, handling the $1.5 billion Bartz class settlement, ruled that two newer copyright suits against Anthropic — Chicken Soup for the Soul v. Anthropic and Cognella v. Anthropic — are not related to Bartz. Each case will be reassigned to a different judge and proceed as a separate dispute rather than as an extension of the settled class action. She has not yet formally signed off on the $1.5 billion settlement, but final approval is expected.
The Spin-Offs · Opting Out for the JuryHow #3 looks in practice
Cruz v. Anthropic
Filed May 13, 2026
Novelist Angie Cruz, joined by 27 Bay Area authors including Pulitzer finalist Dave Eggers. Same factual foundation — books scraped from pirated shadow libraries like LibGen — but they bypass the class structure entirely, arguing the ~$3,000-per-book settlement grants Anthropic a cheap pass to keep exploiting their work. They pursue direct, vicarious, and contributory infringement, seeking up to $150,000 per willful violation by jury.
Shakespeare v. Anthropic
Filed June 17, 2026
London author Thomas William Shakespeare and 99 others sue Anthropic, CEO Dario Amodei, and co-founder Benjamin Mann in the N.D. Cal. — the 10th major copyright suit against the company. Its technical pivot builds on Alsup’s ruling: rather than relitigate whether training is fair use, it targets the discrete step of acquisition and retention — arguing that torrenting and compiling unlicensed files into a “central library” is standalone infringement, regardless of any training that follows.
The complaint characterized Anthropic as an “AI Colossus” and part of an “arms race” to build “generative intelligence at superhuman scale.”
An Attempt to Define Minor AI Assistance and Disclosure
De minimis — short for de minimis non curat lex, a common law principle federal courts have applied for more than 150 years — holds that the law does not concern itself with trivial matters. The term does not appear in the Copyright Act of 1976 or in any congressional amendment enacted through December 2025. It enters AI and publishing through case law and U.S. Copyright Office guidance.
In the second of its three reports on AI, the copyright office explained de minimis use in the context of copyright registration. Its guidance instructs applicants to disclose the presence of AI-generated material and the extent of their human-authored contribution, though it explicitly does not require a detailed accounting of the exact proportion or amount. The office identifies AI-generated material as de minimis if the contribution is so trivial that it would not independently support a copyright claim had a human created it.
The Authors Guild relies on this same de minimis standard for its “Human Authored” self-certification program. The Guild treats AI-powered grammar and spelling edits, brainstorming, and outlining as permissible. More extensive AI editing is harder to classify. What disqualifies a work from certification is using AI to generate the text itself “beyond a de minimis amount.” The Guild sets a de minimis threshold of 5%. For that reason, the Guild advises authors to consult a lawyer.
FOUR MODEL CLAUSES TO GOVERN AI IN PUBLISHING CONTRACTS
Model clauses offered by the Authors Guild address four areas: barring AI use of an author's work without consent; treating specific AI uses as subsidiary rights subject to licensing and compensation; protecting audiobook and translation rights from unapproved AI use; and setting terms for both the author's use of AI in the manuscript and the publisher's use of AI in connection with the work. Authors and agents can request the clauses; publishers can adopt them.
No industry standard defines AI-assisted versus AI-generated content, or the threshold at which a copy edit becomes substantive revision. The Authors Guild’s 5% threshold is the only attempt to operationalize the line. The Guild’s certification relies entirely on an honor system. No forensic audit exists for books, nothing comparable to the governance, risk, and compliance framework used to verify security and privacy controls in, say, a federal information system. Even if it did, keeping track of every edit and revision beyond what Microsoft Word offers in tracked changes would be burdensome.
Guild Contract Evolution
Added AI clause to model contracts for trade books and translations, covering LLM training rights.
Updated contracts with four new clauses — AI translations, narration, cover art, and author disclosure. Introduced the 5 percent cap on AI-generated text.
Launched fee-based “Human Authored” self-certification program for members.
Human Authored program expanded to include any author with a book published in the United States.
The Authors Guild’s “Human Authored” self-certification seal, available since January 2025 to Guild members and expanded to all U.S. authors in March 2026.
Authors Guild · Model Contract Language
Author's Use of AI
Author shall not be required to use generative AI or to work from AI-generated text. Author shall disclose to Publisher if any AI-generated text is included in the submitted manuscript, and may not include more than [a de minimis/5%] AI-generated text.
Three references for the working author, publisher, and editor: a glossary of technical terms, a survey of detection tools, and a quick reference for copyright registration.
Burstiness describes the pattern of multiple sentences. Do the sentences vary in length? Is there a predictable rhythm, or does the style jump around across a paragraph, or is it repetitive? It does not apply to any one word or sentence, but to the composition of the whole passage or narrative. Burstiness is a common metric used by AI detectors, but no detector would rely on burstiness alone; it is typically combined with other metrics, including perplexity.
Context Window
Think of a context window as an AI’s working memory for a conversation. The bigger the context window, measured in tokens, the more a user can have an extensive conversation with an LLM. Once a conversation exceeds a certain capacity, the oldest messages are pushed out of memory, causing the model to lose vital background details and increasing the likelihood of hallucinations.
Fair Use
A legal doctrine permitting limited use of copyrighted material without permission under specified four conditions, called factors. In plain English, courts weigh:
Factor 1: why and how you used it.
Factor 2: what kind of work is it.
Factor 3: how much you used.
Factor 4: whether your use hurts the market for the original.
Generative AI (Gen AI)
Artificial intelligence systems that produce novel text, images, audio, or other content in response to prompts — typically by learning statistical patterns from large training datasets — are known as generative AI. Gen AI is non‑deterministic, meaning the output varies inherently even when the same prompt is run. Think of gen AI as rolling dice. In contrast, deterministic programs such as a calculator always return the same result for the same input.
Human Authorship Requirement
The U.S. Copyright Office requires that works contain some element of human creative expression to be eligible for copyright protection. AI-generated works without human authorship will be rejected.
Large Language Model (LLM)
A type of gen AI trained on massive text datasets. It operates by predicting the statistically likely next word fragment based on the context it is given, producing highly fluent text without understanding its meaning. In early LLMs, researchers discovered a bizarre list of random words and phrases, such as SolidGoldMagikarp, known as “glitch tokens.” When prompted to repeat specific terms, the models would experience a mathematical malfunction, causing them to fail the task and respond with erratic, nonsensical, or entirely hallucinated text.
Perplexity
Perplexity is a measure of how surprising the next word is to a LLM. Low perplexity indicates predictable word choices (typical of AI); high perplexity indicates more unexpected, human-like word choices. Just keep in mind that high perplexity doesn’t automatically mean brilliant human writing: a completely garbled, ungrammatical sentence written by a human (or a poorly prompted AI) will also register high perplexity because it makes no statistical sense.
Project Panama
Anthropic’s internal program to build a legally defensible training dataset by purchasing physical books, slicing off their spines, and scanning them — a shortcut to “clean” training data revealed during discovery in Bartz v. Anthropic PBC.
Shadow Library
A shadow library is an unauthorized digital repository of copyrighted works distributed via piracy networks. Notable examples include Library Genesis (LibGen), Anna's Archive, and Bibliotik. To evade enforcement, the physical servers for these platforms are hosted in countries with lax copyright laws or legal oversight. While targeted by international law enforcement and domain seizures, these resilient networks are often reappear under new domains or web addresses.
Tokens
To a human, a conversation is made of words; to an AI, it is made of tokens. Think of tokens as the basic building blocks of language — typically whole words or syllable fragments. On average, one token equals about four characters, meaning a standard English word is usually broken down into one or two tokens.
Training Data
The corpus of text, images, or other data used to train an AI model. The source and legality of training data is the central question in current AI copyright litigation.
GPTZero — Measures perplexity and burstiness. Designed for educators. Provides sentence-level highlighting of suspected AI passages.
Originality.AI — Combines AI detection with plagiarism checking. Targets publishers and content agencies. Reports probability scores per paragraph.
Turnitin AI Detection — Integrated into academic submission systems. Flags AI-generated passages within submitted documents.
Winston AI — Targeted at publishers and content platforms. Provides overall AI probability score with sentence-level breakdown.
Category 2
Publishing, Enterprise & High-Stakes Verification
Pangram (Pangram Labs) — Founded 2024 by researchers out of Stanford, Tesla, and Google. Trained with a “synthetic mirror” method built to catch paraphrased and humanized AI text that slips past older tools. Flags at the segment level and, since version 3.0, sorts a draft into four tiers — fully human, lightly assisted, moderately assisted, fully generated — instead of one verdict. University of Chicago and University of Maryland researchers have validated its false-positive rate, the number that matters when a wrong call accuses a real writer.
Originality.AI — Aimed at publishers and agencies vetting freelance work. Scores per section and logs results by project. Runs the most aggressive model on this page, which catches more AI but also flags more humans — including articles drafted before ChatGPT existed. Choose it when a miss costs more than a false alarm.
Winston AI — Bundles AI scoring with a plagiarism check and a readability report. Returns an overall probability plus a sentence-level map. Courts publishers and educators with one dashboard that covers both audits.
Sapling AI Detector — API-first, built to sit inside content pipelines rather than a browser tab. Returns per-sentence probabilities for teams checking copy at volume. The pick when detection has to run as code, not a manual paste.
Read this before you trust a score. Every tool here models surface patterns and guesses; none reads intent. On independent 2026 benchmarks most cluster near 80% accuracy, with Pangram the research-backed outlier. False positives remain the rule, not the exception — Scribbr’s free tier flags roughly one human passage in eleven, and plain, formal, or edited prose draws the most false flags. Read any score as one signal, not a sentence.
The U.S. Copyright Office requires applicants to disclose AI-generated content when it exceeds the de minimis threshold. AI-assisted works (where a human author used AI as a tool but made the creative decisions) may be registrable for the human-authored portions.
The 5% Threshold
The Authors Guild characterizes AI content below 5% as de minimis for purposes of copyright registration disclosure. This figure also appears in the Guild’s model contract clause capping AI-generated text. These are two distinct instruments with different legal consequences.
Registration Process
Register at copyright.gov using Form CO (online) or Form TX (text works). For works containing AI-generated material, use the “Limitation of Claim” section to exclude AI-generated portions and identify what human-authored content you are claiming.
Key Cases for Reference
Burrow-Giles Lithographic Co. v. Sarony (1884) — Established photographer as author; affirmed human authorship requirement.
Thaler v. Vidal (Fed. Cir. 2022) — AI cannot be listed as inventor on a patent.
Thaler v. Perlmutter (D.D.C. 2023) — Purely AI-generated images not registrable.
Bartz v. Anthropic PBC (N.D. Cal. 2025) — Training on lawfully purchased books is fair use; training on pirated data is not.
Kadrey v. Meta Platforms, Inc. (N.D. Cal. 2025–26) — Ongoing; addresses contributory infringement via BitTorrent seeding.
The Human Authorship Test
To register an AI-assisted work, you must be able to identify and describe the human creative expression you contributed. Selection, arrangement, editing choices, and original additions all qualify. A prompt alone does not.
AI Disclosure
How the Author Wrote the New Malleus
The author leveraged AI extensively in the creation of the special report. The three AI categories involved design, artwork, and written content. The author also used her judgment to decide whether to use or edit the output in each category.
1. DESIGN: Claude Design created the layout in HTML. Using HTML was a deliberate design choice. Its performance varied from a rating of 1 to 5 stars. Sometimes, Claude Design made expert choices; at other times design execution was poor. Claude Design is also token‑intensive, and its usage is billed apart from using Claude in a chat window.
2. ARTWORK: The artwork included generating images in DALL·E 3, which was spot on for consistency. Midjourney generated two images. Other images were either public domain or attributed. Tools included HeyGen to create talking avatars, ElevenLabs to generate voices, Claude Opus 4.7 to write text‑to‑image prompts, Veo and Magnific for videos, and CapCut to edit and crop them.
3. WRITTEN CONTENT: AI services included Claude Opus 4.7 and Sonnet 4.6, Gemini Pro, Perplexity, and the latest ChatGPT, custom GPT engines. Purpose: deep research, outlining streamlining overflowing content, and copy editing. The author did not use AI to fact‑check; she either located the primary sources or asked subject‑matter experts. AI‑generated content was reserved mainly for definitions, and even in those cases the author revised the text.
Frequently Asked Questions about The New Malleus
Why did Hachette drop Mia Ballard and Shy Girl?
Hachette Book Group terminated its contract with author Mia Ballard after Pangram Labs, an AI detection firm in Brooklyn, ran a pirated copy of her horror novel Shy Girl through its detection tool and reported a 78% AI-detected score. The score from the legitimately purchased copy was materially different. Hachette acted without an industry-wide standard for what level of AI content, if any, would justify contract termination.
How accurate are AI content detectors?
AI content detectors measure proxies — primarily perplexity (how predictable word choices are to a language model) and burstiness (variation in sentence length) — rather than directly detecting AI authorship. They produce probability scores, not verdicts, and are known to generate false positives on human-written text, particularly from non-native English speakers or writers with a highly consistent prose style. No detector has been validated as forensically reliable in publishing disputes.
Can an AI-generated book be copyrighted?
No. The U.S. Copyright Office requires that works contain some element of human creative expression to qualify for copyright protection. Purely AI-generated works are not registrable. This was affirmed in Thaler v. Perlmutter (D.D.C. 2023) and let stand by the Supreme Court in March 2026, when SCOTUS denied Thaler's petition. AI-assisted works — where a human author made the core creative decisions and used AI as a tool — may be registrable for the human-authored portions.
What is the 5% de minimis threshold for AI content in publishing?
The Authors Guild characterizes AI-generated content below 5% of a work as de minimis for purposes of copyright registration disclosure. This figure also appears in the Guild's model contract clause capping AI-generated text. These are two distinct instruments with different legal consequences. The 5% threshold is a guideline, not a formally adopted legal standard, and no AI detector can reliably measure it.
What happened in Bartz v. Anthropic PBC?
In Bartz v. Anthropic PBC (N.D. Cal. 2025), a federal court ruled that Anthropic's training of its Claude AI models on lawfully purchased books constitutes fair use under U.S. copyright law. However, the court found that training on pirated books — sourced from shadow libraries such as Library Genesis and Anna's Archive — is not protected by fair use. The case also revealed Anthropic's internal Project Panama, a program to build a clean training dataset by purchasing and scanning physical books.
What is the standards vacuum in AI and publishing?
The standards vacuum is the absence of industry-wide definitions, thresholds, and governance frameworks for AI use in publishing. There is no agreed standard distinguishing AI-assisted from AI-generated content, no validated method for measuring the de minimis threshold, and no governance framework comparable to what other industries have developed when technology outpaced their rules. Publishers, agents, and authors are making consequential decisions — including contract terminations — without a standard to judge against.
What is Kadrey v. Meta Platforms about?
Kadrey v. Meta Platforms, Inc. (N.D. Cal., ongoing through 2025–26) is a class-action copyright suit in which authors allege that Meta trained its Llama large language models on copyrighted books sourced from shadow libraries. The case addresses contributory infringement via BitTorrent seeding and raises the market dilution theory — that AI trained on authors' works can generate competing content that harms the market for original works.
Is it legal to use AI to help write a book?
Using AI as a tool in the writing process is legal in the United States. There is no law prohibiting AI-assisted writing. Copyright protection, however, only extends to the portions of a work reflecting human creative expression. Authors must disclose AI-generated content exceeding the de minimis threshold when registering a copyright, and many publishing contracts now include AI disclosure or limitation clauses. No federal statute currently defines or regulates AI use in creative works.