Last updated: August 18, 2026 · By Vishal Swami, Founder & Lead AI Reviewer, AISagely
The headline sounds like a stretch until you read the investigation behind it: AirTag reveals Amazon is trashing rare books to train AI, and the method is exactly what it sounds like. A journalist hid an Apple AirTag inside a bulk order of out-of-print books and watched it travel straight into an Amazon warehouse, where workers slice off the spines, scan every page, and throw away what's left. 404 Media published the investigation on August 17, 2026, confirming what booksellers had suspected for months: a wave of odd, high-volume orders for obscure titles wasn't collectors — it was Amazon buying raw material for its AI models.
Short answer: An AirTag hidden inside a 1,000-book order, tracked by 404 Media, followed a shipment bought through the marketplace Biblio to Amazon's VGT3 team at a Las Vegas warehouse, where workers cut book spines to scan pages for AI training and discard the originals. Amazon confirmed it buys books "through commercial channels." A near-identical Anthropic program was ruled fair use in June 2025.
I spend most of my week testing the AI products these companies ship, not tracking shipping containers. But in my testing of chatbots from Amazon, Anthropic, and OpenAI, I've never gotten a straight answer about where the training data actually comes from — the models either dodge the question or give a vague answer about "publicly available and licensed sources." This investigation is one of the first times anyone has actually followed the paper trail instead of asking the vendor to explain itself.
What the AirTag actually found
Rare-book dealers had noticed something strange for months: buyers placing large, oddly specific orders for obscure, out-of-print titles through the secondhand marketplace Biblio, paying without haggling and asking no questions. To find out who was buying, 404 Media placed an AirTag inside a book included in a roughly 1,000-book order and tracked it as it moved from California through Milwaukee, Wisconsin, and Grand Junction, Colorado, before landing at an Amazon fulfillment complex (warehouse code LAS8) in Las Vegas.
Inside, the shipment was handled by a team called VGT3, whose internal logo shows a Tyrannosaurus rex clutching a book. Workers there described a simple, brutal process: books arrive by the pallet, get their spines cut off with a blade so the loose pages feed through a scanner faster, and the stripped pages are discarded once digitized. One worker put it plainly to 404 Media: "I work at VGT3 here in Vegas, and all we do is scan books." Every book's ISBN is logged during the process, which is why booksellers suspect the goal is to systematically capture every edition by catalog number, not just popular titles.
Print books that predate 2022 are unusually valuable for this purpose. They're less likely to already be sitting in a training set scraped from the open web, and they're free of the AI-generated text now flooding online book listings — a problem I've covered separately in my look at how generative AI is flooding and diluting the market for books. A source in the rare-book trade summed up the mismatch in values to 404 Media: "There's monetary value, obviously, but there are a lot of other types of value… All of those the AI companies don't care about. They just want the content as a bunch of words strung together." Amazon's response to the investigation was short: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use."
Amazon vs. Anthropic: the same playbook
Amazon isn't the first company caught running this exact process. Court filings from Bartz v. Anthropic describe a nearly identical operation, code-named Project Panama, in which Anthropic bought millions of physical books, cut the bindings, scanned the pages, and destroyed the originals. The two cases aren't identical legally, and Amazon's version hasn't been tested in court yet — but the physical process is the same one now showing up at VGT3.
| Amazon (VGT3, Las Vegas) | Anthropic (Project Panama) | |
|---|---|---|
| Reported by | 404 Media, Aug 17, 2026 | Court filings in Bartz v. Anthropic |
| Scale confirmed | ~1,000 books in one tracked shipment | Millions of books, per litigation |
| Method | Cut spines, scan pages, discard originals | Cut spines, scan pages, discard originals |
| Court ruling on the practice | Untested for Amazon specifically | Destroying print originals after scanning ruled fair use, June 23, 2025 |
| Separate legal exposure | None reported yet | $1.5B settlement (Sept. 2025) over unrelated pirated ebook copies |
Judge William Alsup’s June 2025 order in Bartz v. Anthropic found that training an AI model on books is "spectacularly" transformative, and that scanning a legally purchased print book and destroying the original counts as fair use because the digital copy simply replaces the paper one. That ruling covers the scan-and-destroy method itself, not the separate question of where a company gets the books in the first place. Anthropic's $1.5 billion settlement, finalized months later, was for a different problem: millions of ebooks the company had downloaded from pirate sites rather than purchased. Amazon's operation, as reported, appears to be buying its books legitimately through Biblio — which is closer to the part of Anthropic's practice a court already blessed.
Step-by-step: how to check whether your book (or your content) is training an AI model
You can't stop a company from buying a book you already sold on the open market — once it's purchased legally, current case law treats scanning and destroying it as fair use. But you can find out whether it's happening to your own work, and decide what to do about it.
1. Check for the buying pattern, not just one sale
A single sale means nothing. The tell booksellers spotted was a pattern: large orders, unusual ISBN spread, no negotiation, no questions about condition. If you sell books, watch your own order history for that shape of purchase.
2. Search your own titles and content for reuse
Ask a chatbot to summarize your out-of-print book or a long-form article you wrote and see how much it can reproduce. If it can quote specific passages verbatim, that's a stronger signal your text is in a training set than a vague summary would be. I cover how to think about this from the other direction — using your own writing to train a model on purpose — in my guide to using your outputs to train an AI model.
3. Read the marketplace's bulk-buyer policy
Biblio, AbeBooks, and similar marketplaces don't currently flag or restrict AI-training buyers. Before listing rare stock in bulk, check whether the platform discloses anything about who's buying, since right now that disclosure mostly comes from journalism, not policy.
4. Know where the legal line actually sits
Buying a book and destroying it after scanning is currently legal, per the Alsup ruling. Distributing pirated copies to build a training library is not, which is the distinction that cost Anthropic $1.5 billion. If you suspect a company used unauthorized digital copies of your work rather than a purchased physical copy, that's the stronger legal claim.
5. Track a shipment yourself if you're testing a theory
404 Media's method wasn't exotic: a $29 AirTag, a book, and patience. In my testing of Bluetooth trackers for other projects, AirTags reliably report location every few minutes near populated areas, which is precise enough to identify a destination warehouse, not just a city.
Example prompts you can copy
Use these to dig into a claim like this one before you repeat it:
- "Summarize the investigation at [paste URL]. List every factual claim separately from any opinion or speculation in the piece."
- "Explain the legal difference between training an AI model on a purchased physical book versus a pirated digital copy, based on U.S. case law as of 2026."
- "I sell [type of book] online. What buying patterns would suggest a bulk buyer is purchasing for AI training rather than resale or collecting?"
These work because they force a model to separate sourced fact from spin, instead of handing you a confident-sounding summary that blurs the two.
Common mistakes to avoid
The biggest mistake is assuming this practice is illegal because it feels invasive — right now, buying a book and destroying it after scanning is exactly the kind of activity a federal court has already called fair use, so outrage alone won't get a book pulled out of a training set. The second is assuming this is only an Amazon problem; Anthropic ran a near-identical operation at a larger scale, and there's no reason to think other AI labs buying bulk print inventory are doing anything different. Third, don't conflate the scan-and-destroy practice with book piracy — they're legally distinct, and mixing them up weakens your argument if you're trying to raise this with a platform, a publisher, or a lawyer.
Tools that make this easier
If you're trying to figure out whether AI training claims about a company or a tool hold up, the fastest way is to test the product yourself instead of taking a press statement at face value. My AI tool reviews page explains the process I use to test AI products hands-on, and my AI tool ratings guide covers how to read a review without getting fooled by a cherry-picked demo. If you're deciding which AI model to actually use, Best AI Models (2026) ranks the ones I currently recommend based on testing, not marketing copy. And if this story has you wondering whether the used-book market is quietly being reshaped by AI demand, my piece on why secondhand book sales are booming digs into that exact question.
Frequently Asked Questions
Is it legal for Amazon to buy and destroy books to train AI?
Based on the closest precedent — Judge William Alsup's June 2025 ruling in Bartz v. Anthropic — buying a book legally and destroying the print copy after scanning it is fair use, because the digital copy is treated as a replacement for the paper one. Amazon's own practice hasn't been tested in court as of this writing.
How does 404 Media's AirTag investigation differ from the Anthropic book lawsuit?
The AirTag investigation is journalism that documented Amazon's process in real time by physically tracking a shipment. The Anthropic case was litigation, where court filings about a program called Project Panama described a similar scan-and-destroy process at a much larger scale, plus a separate, unrelated claim over pirated ebooks that led to a $1.5 billion settlement.
What AI model is this book data used to train?
Neither 404 Media's investigation nor Amazon's statement to reporters names a specific model. Amazon said only that it purchases books "through commercial channels to help develop and improve the products and services our customers use."
Can I find out if my own book was scanned this way?
Not directly — neither Amazon nor Anthropic publishes a list of scanned titles. The most reliable check is testing whether a chatbot can reproduce specific passages from your book, which suggests the text made it into a training set, though it won't tell you which company's warehouse it passed through.