Research story · AI, law & institutional design

One dataset.
Two kinds of rights.

The same training document can be an author’s work and information about other people. Published research examines how AI governance can address both sets of interests throughout the model’s life.

A developer licenses an archive for model training. The agreement addresses the publisher’s rights in the expression. The archive also contains names, quotations, photographs and biographical details about people who are not parties to that agreement. One transaction has answered a copyright question while leaving a separate data-protection analysis to be completed.

My article, Training on the tightrope: AI copyright and data privacy as colliding regulatory regimes, published in Computer Law & Security Review, examines that overlap. It asks how institutions should govern training data when the legal obligations attached to the same material follow different objects, rights holders and remedies.

Follow the data through the system

The question changes as data moves through an AI pipeline. At collection, an organisation needs to establish what material it holds and on what terms it may use it. During training, the material contributes to a model whose parameters do not reproduce the organisation of the source archive. At deployment, outputs create another point at which protected expression or personal information may appear.

A rule directed at the source document is therefore not automatically a complete rule for the trained model or its outputs. Copyright analysis concerns protected expression and its use. Data-protection analysis concerns information relating to people and the conditions under which it is processed. The article’s central argument is that governance needs to examine their interaction across those stages.

The study uses comparative doctrinal analysis, centred on the United States and European Union. It maps elementary obligations, identifies where they attach to the same artefact and distinguishes compatibility from conflict. The result is a taxonomy of five problem areas, used to assess existing approaches and develop a proposed coordinating framework.

Permissions need to match the interests involved

The consent–licence gap illustrates the issue at the point of acquisition. A copyright permission and a lawful basis for processing personal data answer different questions. The person entitled to license a work need not be the person described in it, and an individual’s privacy authorisation need not confer rights in someone else’s expression.

Consent is one possible privacy basis, rather than a universal requirement for every processing operation. The practical research point is the need for a separate, applicable justification. A licence should not become a shortcut that silently treats all other interests in the material as resolved.

The article examines private licensing and intermediary arrangements as ways to reduce transaction costs, while asking which interests they can actually represent. That shifts attention from whether a dataset has a contract attached to whether the contract and the processing basis cover the relevant use and parties.

Retention and remedies operate on different levels

The retention problem appears when evidence-preservation obligations meet requests to remove personal data. The legal-claims exception to erasure is part of that analysis; the question is what retention is necessary for the particular claim and which material remains within its scope. Treating an entire corpus as one undifferentiated object can obscure those distinctions.

The article also separates remedies aimed at source data, particular outputs and a model as a whole. These interventions have different effects. Removing a record from a dataset does not by itself specify what has changed in parameters already trained on it. Restricting an output is different again.

Its remedial-mismatch argument asks how decision-makers should assess those operations together. A remedy chosen within one proceeding can affect another protected interest or another enforcement process. The proposed solution requires institutions to make that cross-domain consequence part of the decision.

Technical assurance must describe what it proves

Machine unlearning gives this institutional problem a technical core. The article distinguishes retraining without targeted material, approximate modifications intended to remove its influence, and filters that restrict what a deployed system returns. Each changes a different part of the system and supports a different assurance claim.

Retraining can impose substantial economic costs at foundation-model scale. Approximate methods raise questions about whether information remains recoverable under adversarial testing. Output filtering can prevent particular disclosures without establishing that the underlying model has forgotten the material.

The research therefore separates cost from verification. An expensive operation is not the same problem as an operation whose claimed effect cannot be demonstrated to the required standard. Its governance proposal asks for explicit tests, stated limits and procedures that can evolve with the technical evidence, rather than a generic declaration that the model has been “cleaned.”

Coordinate the institutions, not just the paperwork

The comparative analysis examines how different regulatory and enforcement structures divide responsibility. Copyright, data protection and AI-specific oversight can involve different institutions, procedures and timing. The article argues that a developer’s compliance file cannot resolve every coordination problem by treating each requirement as an isolated checklist.

It evaluates risk-tiering, sector-led governance, incident reporting and model-specific obligations as partial responses. Each performs useful work: classifying risk, allocating responsibility, recording failures or imposing model-level duties. The proposed framework combines those functions around training data’s dual character.

The proposal has five connected elements: a shared regulatory category, tiers reflecting the interests present in the data, independent model audits, conditional statutory safe harbours for the training phase and a coordination mechanism. These are legislative and institutional proposals in the article, rather than a description of protections already available simply by adopting them.

Make the proposal testable

The audit component would examine memorisation and personal-data leakage with adversarial probes as well as standard tests. Incident reporting would supply evidence between scheduled audits and help refine the thresholds. Safe-harbour eligibility would depend on specified classification, provenance, testing and response practices.

Those choices raise substantive questions about acceptable residual risk, auditor competence, incentives and the rights of affected people. The paper presents a design for evaluating such questions; it does not report an empirical trial proving that its framework has solved them. Thresholds and institutional arrangements would need scrutiny and calibration.

The interdisciplinary contribution joins legal doctrine, machine-learning assurance and institutional economics. It follows the same material from acquisition through training, deployment and remedy, asking which interests attach at each stage and who can coordinate the response. That perspective makes AI data governance a problem of accountable system design, with the interaction between rights visible from the beginning.