Contracts Become a Problem Before Regulation Does
When companies adopt generative AI, their first concern is usually "how will the law change?" In practice, however, disputes erupt much earlier — in the contracts used to acquire data. With no settled legislation on training data use, agreements between parties effectively serve as the operating rules.
On February 26, 2026, Korea's Ministry of Culture, Sports and Tourism and the Korea Copyright Commission published the Guide to Fair Use for Generative AI Training, an attempt to fill that gap. The guide walks through the four statutory factors under Article 35-5 of the Copyright Act — the purpose and character of the use, the type and purpose of the work, the amount and significance of the portion used, and the effect on the current or potential market for the work — and how each applies to generative AI training. It also notes that a commercial purpose or the use of web crawling does not by itself disqualify training from fair use. The key point is that training is neither automatically permitted nor automatically prohibited. The specific facts of each case determine the outcome — and the party holding evidence of those facts holds the advantage.
Procurement options changed as well. Under the government plan to expand public-works use for AI training, announced on January 28, 2026, the Korea Open Government License added Type 0 (unrestricted use) and an AI type, letting organizations verify AI training eligibility at the license level. Companies that only ever checked Types 1 through 4 need to update that habit.
The principle governing outputs is also becoming clear. Material generated from a simple prompt alone does not receive copyright protection. Only portions reflecting human creative contribution are protected. Practically, an AI-generated image or copy dropped straight into a product gives you no basis to stop a competitor from using something identical. If an asset needs to be exclusive, preserve evidence of human editing and composition.
The Four Kinds of Data You Actually Handle
Training data in real operations falls into roughly four buckets, each carrying a different kind of risk.
A Contract Clause Checklist
These are the clauses most often missing in reviews we conduct.
The Evidence Your Pipeline Should Retain
Contracts alone are not enough, because a dispute ultimately turns on proving what data you used at the time.
A Realistic Sequence by Company Size
You do not need every control at once. Work through it in stages.
Companies early in AI adoption resolve much of their risk simply by cleaning up procurement contracts. If you are only calling an external API, you hold no training data yourself, so the priority is checking vendor terms for whether your inputs are reused for training and whether opt-out is available.
Companies running their own fine-tuning need the register, the filters, and the audit logs. If data and legal responsibilities sit in separate teams, making the license field mandatory at dataset registration alone eliminates most omissions.
POLYGLOTSOFT's AI adoption consulting covers data procurement contract review, provenance register design, and automated evidence capture in training pipelines. We also support remediation for organizations already running fine-tuning in production. If you want to design your training data governance properly from the start, please reach out through our [contact page](https://polyglotsoft.dev/en/support/contact).
