Back to Blog
AI

AI Training Data Copyright: What Belongs in Your Contracts and Pipelines

Disputes over generative AI training data start in contracts long before legislation catches up. Using Korea's 2026 fair use guide and the new KOGL AI license type as a baseline, we map the risks across four data sources, the clauses your contracts need, and the evidence your pipeline must retain.

POLYGLOTSOFT Tech Team2026-08-318 min read0
AI CopyrightTraining DataFair UseData ContractsAI Governance

Contracts Become a Problem Before Regulation Does

When companies adopt generative AI, their first concern is usually "how will the law change?" In practice, however, disputes erupt much earlier — in the contracts used to acquire data. With no settled legislation on training data use, agreements between parties effectively serve as the operating rules.

On February 26, 2026, Korea's Ministry of Culture, Sports and Tourism and the Korea Copyright Commission published the Guide to Fair Use for Generative AI Training, an attempt to fill that gap. The guide walks through the four statutory factors under Article 35-5 of the Copyright Act — the purpose and character of the use, the type and purpose of the work, the amount and significance of the portion used, and the effect on the current or potential market for the work — and how each applies to generative AI training. It also notes that a commercial purpose or the use of web crawling does not by itself disqualify training from fair use. The key point is that training is neither automatically permitted nor automatically prohibited. The specific facts of each case determine the outcome — and the party holding evidence of those facts holds the advantage.

Procurement options changed as well. Under the government plan to expand public-works use for AI training, announced on January 28, 2026, the Korea Open Government License added Type 0 (unrestricted use) and an AI type, letting organizations verify AI training eligibility at the license level. Companies that only ever checked Types 1 through 4 need to update that habit.

The principle governing outputs is also becoming clear. Material generated from a simple prompt alone does not receive copyright protection. Only portions reflecting human creative contribution are protected. Practically, an AI-generated image or copy dropped straight into a product gives you no basis to stop a competitor from using something identical. If an asset needs to be exclusive, preserve evidence of human editing and composition.

The Four Kinds of Data You Actually Handle

Training data in real operations falls into roughly four buckets, each carrying a different kind of risk.

  • Internally held data: It looks safest, but there is a catch. Images uploaded by customers, contributions from external writers, and works employees created outside their job scope may not belong to the company. If existing terms of service permit use only "for service provision," training use falls outside that scope.
  • Purchased third-party data: Confirm that the contract explicitly names training use and that the provider secured sublicensing rights from the original rights holder. The phrase "data provision" alone does not implicitly include training.
  • Web-scraped data: You must consider robots.txt, site terms of service, and database producers' rights together. Even when the underlying facts are not copyrightable, substantially reproducing a systematically organized database raises a separate issue.
  • Vendor deliverables: Clients often have no idea what data a development firm used for fine-tuning. When an infringement claim arrives, the client operating the service is targeted first.
  • A Contract Clause Checklist

    These are the clauses most often missing in reviews we conduct.

  • Stage-by-stage scope of use: Pre-training, fine-tuning, evaluation, and output redistribution are distinct acts. A single catch-all phrase like "use in AI development" invites each side to read it differently once a dispute begins.
  • Warranties and indemnification: The provider should warrant non-infringement of third-party rights, and the contract should name who leads the response, how costs are shared, and what caps damages. Whether that cap sits at 100% of contract value or allows overage is usually the crux of negotiation.
  • Deletion and retraining obligations: Decide in advance whether withdrawal of consent requires only deleting the data or retraining the model that absorbed it. Retraining can cost tens of thousands of dollars, which makes after-the-fact agreement extremely difficult.
  • Ownership of outputs: Specify who holds usage rights to outputs, and define the notification and remediation procedure when a result closely resembles a third party's work.
  • The Evidence Your Pipeline Should Retain

    Contracts alone are not enough, because a dispute ultimately turns on proving what data you used at the time.

  • A data provenance register: Record, per dataset, the source URL or contract number, license type, acquisition date, permitted scope, and expiration. Starting in a spreadsheet is fine; what matters is that it stays current with no gaps.
  • Records of technical measures: Deduplication, similarity filtering on outputs, and blocking prompts that name specific artists can count as good-faith effort in a fair use assessment. Logging that you took the measure is not enough — record when it was applied and how many items it blocked.
  • Model cards and training history: Retain the dataset list, training timestamp, and hyperparameters for each model version. Given copyright limitation periods and contractual warranty windows, we recommend keeping these for at least five years.
  • A Realistic Sequence by Company Size

    You do not need every control at once. Work through it in stages.

    Companies early in AI adoption resolve much of their risk simply by cleaning up procurement contracts. If you are only calling an external API, you hold no training data yourself, so the priority is checking vendor terms for whether your inputs are reused for training and whether opt-out is available.

    Companies running their own fine-tuning need the register, the filters, and the audit logs. If data and legal responsibilities sit in separate teams, making the license field mandatory at dataset registration alone eliminates most omissions.

    POLYGLOTSOFT's AI adoption consulting covers data procurement contract review, provenance register design, and automated evidence capture in training pipelines. We also support remediation for organizations already running fine-tuning in production. If you want to design your training data governance properly from the start, please reach out through our [contact page](https://polyglotsoft.dev/en/support/contact).

    Need Technical Consultation?

    Our expert consultants in smart factory, AI, and logistics automation will analyze your requirements.

    Request Free Consultation