Is your AI vendor's training data a legal risk to you?

Published September 20, 2026 · Updated September 20, 2026My Business AI Audit · Tag: AI training data
AI training data vendor risk indemnity
The short answer

Is your AI vendor's training data a legal risk?

Dated analysis — filing and reporting dates: September 17–18, 2026. Every quote below comes from the September 17–18, 2026 reporting; the filing itself is dated September 17, 2026.

TechCrunch reported on September 17, 2026 that newly unredacted filings quote Microsoft's own director of Applied Science calling AI web scraping "the largest theft of labor in human history." No court has ruled on the copyright question, so you are not a defendant — but training-data provenance and indemnity are now diligence items to put in writing.

What was reported on September 17–18, 2026

TechCrunch and Engadget reported on newly unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft — the case filed in 2023 over training generative AI models on the Times' own content. According to TechCrunch's Rebecca Bellan, a January 2023 internal Microsoft memo written by Brent Hecht, Microsoft's director of Applied Science, called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." A separate internal Microsoft presentation, written by Hecht in January 2024, described a "doom loop" that would "hurt the performance of our models and the entire web at the same time." And OpenAI's head of ChatGPT, Nick Turley, "wrote in internal communication that publishers face an" "existential threat" from chatbot products that are "largely substitutive" and "will get more and more substitutive as they get better."

Three sentences of caution belong here, not in a footnote. TechCrunch notes that "much of the new information comes from The Times’ own brief, not the underlying exhibits, which remain sealed." Its own caution adds that "The quotes below are presented without their original context." " TechCrunch also reports that "The question of whether AI firms can legally use copyrighted material to train AI has no clear answer," while judges "have been largely favorable to AI companies' arguments that training constitutes" fair use. Microsoft's position is on the record: "Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism," Microsoft spokesperson Alex Haurek told The New York Times in a statement. And, per TechCrunch, "OpenAI and Microsoft did not return requests for comment."

Why a vendor's internal words matter more than a critic's

Every AI vendor sells the same reassurance: the model is trained on public data, the output is yours, the terms of service handle it. That reassurance has always been a claim, not a warranty. Now the vendors' own documents, quoted in a plaintiff's brief, describe the training-data supply chain as a threat to the publishers who supply it — one internal Microsoft line says that "an end-product threatens the economic foundations of its essential suppliers".

For a small business, that is the point — and it has nothing to do with who wins. The gap between what a vendor markets and what a vendor will sign is now documented in public filings rather than suspected. If the people who built these systems describe their own content supply chain as contested, a buyer who never asked about provenance, indemnity and model-training opt-outs was never told there was anything to ask.

What it does and does not change for a business already using AI

It does not change today's liability picture. No ruling is reported, the exhibits remain sealed, and no court has reached a conclusion on the merits. TechCrunch reports that several of the new admissions "run counter to OpenAI’s fair use defense" — a statement about a defense in a case, not about your business. Training a model on someone else's work is not the same as using a tool whose model was trained that way, and nothing in the reporting supports reading it that way.

It does change what is defensible inside your own operation. Three things are now harder to wave through:

Questions to ask your AI vendor

  1. Provenance. Which datasets and which model version sit behind the product you are paying for, and will the vendor describe them in writing — model card, dataset documentation, or a contract exhibit? See what a watermark actually proves (and what it misses).
  2. Indemnity. Does the agreement indemnify you for third-party intellectual-property claims arising from model output? What is carved out, and who pays defence costs while a claim is open?
  3. Your data. Is your input or business data used for training, can you opt out in writing, and what is the retention and deletion window for both?
  4. Notice. Does the vendor commit to telling you when litigation, a regulator or a rightsholder raises a claim touching the model you depend on, and within what period?
  5. Sub-processors. Which upstream model providers, data vendors and hosted services sit behind the product, and are they named in your agreement?
  6. Evidence for your own file. What can the vendor hand you — attestation, audit report, or a documented provenance statement — that you could show a client, an insurer or a regulator asking how you vetted the tool?

Bottom line

A vendor's internal memo is not a ruling against you, but it does settle whether training-data provenance is a real diligence item: the vendors' own documents, quoted in a brief filed September 17, 2026, describe the supply chain in exactly those terms. The takeaway is one sentence — if nobody can tell you where the training data came from, or who indemnifies you if that answer is wrong, that gap is closable on paper.

Run the questions above against the AI tools you actually pay for. If the answers come back incomplete, that is the audit.

---

Run the free AI audit tool

Sources

Accuracy note: Quotes are attributed to the outlet that reported them, on the dates shown; filing language is quoted as it appears in the September 17–18, 2026 reporting. The material is one side's brief and sealed exhibits described in press reports — it is litigation risk disclosure, not a finding of liability, and no ruling is reported. This page is informational only and not legal advice; indemnity and liability language needs your own counsel.