Meet the Guy Who Destroys Antique Books for AI

"I'm the guy who destroys antique books after we scan them into our company's AI" is a real job description, not an exaggerated Reddit confession. Court records unsealed in 2026, according to Euronews's review of the filings, describe exactly this role at Anthropic: buy used and antique books by the truckload, slice off the bindings with a hydraulic machine, feed the loose pages through an industrial scanner, and send whatever paper is left to a recycling plant.

Short answer: Yes — this job is real. Anthropic's confidential "Project Panama" bought tens of thousands of used and antique books through resellers like Better World Books, cut off the spines to scan them faster, and shipped the destroyed pages off to become toilet paper and cardboard. If you own a book you'd hate to lose, scan it yourself before it ever reaches a bin like that one.

Claude homepage — screenshot of claude.ai
Claude homepage — screenshot of claude.ai

Last updated: August 30, 2026 · By Vishal Swami, Founder & Lead AI Reviewer, AISagely

What actually happens to the books

According to Euronews’ reporting on the unsealed court filings, Anthropic went looking for books that predated the internet, on the theory that pre-digital writing would teach Claude better prose than "low quality internet speak." Buying that volume of print meant working through secondhand dealers, mainly Better World Books and the UK-based World of Books, tens of thousands of titles at a time. The company reportedly wanted a vendor able to destructively scan somewhere between 500,000 and 2 million books in six months, per Euronews's account of the filings.

The physical process is blunt: a vendor sliced the spine and individual pages off each book with hydraulic cutting equipment, ran the loose sheets through high-speed imaging machines, and sent the shredded remains to a recycler that turns waste paper into things like toilet paper and cardboard. A federal judge ruled in June 2025 that doing this to a book Anthropic legitimately owned was fair use, since one physical copy simply became one digital copy. Anthropic's separate habit of downloading pirated books from shadow libraries wasn't covered by that same reasoning, and it's what led to a $1.5 billion settlement approved in July 2026, covering roughly 500,000 works — per Euronews, the largest known copyright recovery on record.

None of that is illegal for a secondhand bookstore or a bulk-buyback service to feed into, either. A box of paperbacks sold to a reseller for a few dollars can legally end up exactly where those antique books ended up: sliced apart in a warehouse. If a book matters to you, in my testing the fastest way to protect it from that fate is a phone, not a plan to keep it forever.

What you'll need

You don't need anything close to industrial scanning gear. A smartphone with a working camera handles almost every book. Add a free scanning app, a flat surface with even light (a window during the day works better than an overhead bulb), and two or three paperbacks to gently weigh pages open without cracking old glue. Set aside real time — a 250-page book takes roughly 25–35 minutes to photograph carefully, more if the paper is brittle or the print is small. For a genuinely rare or fragile volume, budget extra time to research its value first, since that changes whether you scan it, insure it, or just leave it alone.

Step-by-step: scanning an antique book without destroying it

1. Decide what the book is actually worth

Before you touch the spine, find out if the edition is common or scarce. I ran a first-edition guess through how to use Perplexity for research and had a sourced answer, with citations to sold listings, in under two minutes — far faster than digging through forums by hand.

2. Prop the book open, don't flatten it

Rest the spine at roughly 100–120 degrees on a table, not pressed flat against a scanner glass. Old binding glue cracks under flat pressure long before it cracks from a gentle angle.

3. Photograph each spread with a scanning app's edge detection on

In my testing, Adobe Scan's auto-capture and Genius Scan's edge-crop both did a solid job squaring up a curved page automatically, which matters more than camera resolution once you're shooting one spread at a time.

4. Run OCR so the scan becomes searchable text

A photo of a page is not yet useful text. I compared free and paid OCR options for exactly this kind of bulk job in my PixelRead AI OCR review; when I needed to process a couple hundred pages at once rather than one at a time, Mistral OCR chewed through the batch far faster than any phone app I tested.

5. Have an AI assistant flag pages, not rewrite them

OCR mangles old typefaces and yellowed paper more than people expect — misreads pile up around footnotes and page numbers especially. Paste a chapter into an AI assistant and ask it to flag suspect lines rather than "fix" the prose. See the exact prompt below.

6. Archive it somewhere you can actually search

Once the text is clean, dumping it into a folder buries it. I load finished chapters into how to use NotebookLM, which turns a pile of digitized files into something you can ask direct questions of later.

Example prompts you can copy

For catching pages the OCR likely scrambled, without touching the wording: > "Here is raw OCR text from a printed book chapter. Do not paraphrase or reword anything. List only the line numbers where you suspect a scanning error — garbled words, a missing page break, or a repeated paragraph — and briefly say why for each: [paste text]"

For checking whether the page order survived the scan: > "These are OCR'd page headers and the first line of each page, in order. Flag any place the sequence looks out of order, skips a number, or repeats a number: [paste list]"

For building a quick index once a chapter is clean: > "Pull every proper name, date, and place mentioned in this chapter into a short list I can use as an index, without summarizing or altering the surrounding text: [paste text]"

Common mistakes to avoid

In my testing, the costliest mistake wasn't a scanning error — it was skipping the research step and assuming a "we'll pay cash for your old books" pickup means the books get resold. Some of them do. Plenty get pulped or, per the court record above, get bought precisely because a buyer wants to destroy them for training data. Beyond that:

  • Don't hand off a box of books before checking what's inside it. A quick pass with an AI research tool takes minutes and can save a first edition from a giveaway pile.
  • Don't scan under a single overhead light. Shadows near the gutter confuse edge-detection software and leave you re-shooting half the book.
  • Don't let an AI cleanup pass rewrite anything. Tell it explicitly to flag errors, not improve sentences — otherwise you've quietly swapped the author's words for a paraphrase.
  • Don't trust a scanning app's star rating without reading what the free tier actually includes. I break down how to read those ratings honestly in AI tool ratings: how to read them without getting fooled.
  • Don't keep the only digital copy on the phone you scanned it with. One drop in a sink undoes the entire project.

Tools that make this easier

I checked each price against the vendor's own page rather than a roundup, since app pricing shifts fast.

Tool Best for Price (checked August 2026) Cuts the spine?
Adobe Scan Quick phone scans with built-in OCR Free app; Acrobat Pro from $19.99/mo (annual, billed monthly) for advanced editing No
Genius Scan Fast batch scanning, clean edge-crop Free (fully functional); Genius Scan Ultra $39.99/year for cloud backup + extra OCR, per the vendor’s pricing page No
Mistral OCR Bulk OCR on hundreds of pages at once See my Mistral OCR review for current API pricing No
A hydraulic guillotine + industrial scanner Destructively digitizing a warehouse of used books for an AI training set Not for sale to consumers — this is what Project Panama's vendor used Yes

That last row isn't a joke entry. It's there to make the contrast obvious: everything a home user needs to protect a book costs under $40 a year, and none of it involves a blade anywhere near the binding.

Frequently Asked Questions

Is "I'm the guy who destroys antique books" a real job?

Yes. Unsealed court filings describe workers at a vendor for Anthropic's Project Panama doing exactly this — slicing spines off purchased books so they could be scanned faster, then discarding what was left.

Which AI companies have destroyed physical books to train their models?

Anthropic is the documented case, revealed through 2026 court filings in its copyright litigation. The company bought used and antique books through resellers like Better World Books and World of Books specifically for destructive scanning.

Could donating my books to a thrift store put them in an AI training pipeline?

It's possible, though not guaranteed. Anthropic bought books through mainstream secondhand resellers rather than directly from donors, so a box that moves through a thrift store or a bulk book-buyback service could plausibly end up resold to a bulk buyer somewhere down the chain.

What's the safest free way to digitize a book without damaging it?

A phone scanning app like Adobe Scan or Genius Scan, used spread by spread with the book propped open rather than pressed flat. Both have functional free tiers, and neither requires cutting anything.

Is it legal for AI companies to destroy books after scanning them?

For books they legally purchased, a federal judge ruled in June 2025 that it counts as fair use. Anthropic's separate use of pirated books was treated differently and led to a $1.5 billion settlement approved in July 2026, according to Euronews's coverage of the case.