A filing that should not have been necessary
On 29 June 2026 a judge in the Southern District of New York told Sony Music it could not add more than 30,000 sound recordings to the case it was already running against Udio. Three weeks later Sony filed them anyway, as a separate lawsuit. The complaint landed on 20 July and asserts 30,117 recordings.
Procedurally this is unremarkable. A court declining to expand an existing case does not extinguish the claims; it just makes the plaintiff open a second front. What makes it worth an operator's attention is the arithmetic that follows. The original action put 333 recordings in dispute. Statutory damages in that case topped out around 50 million dollars, roughly 46 million euros. At 30,117 recordings the theoretical ceiling moves to about 4.5 billion dollars, around 4.2 billion euros.
Sony is also now the odd one out. Universal and Warner both reached licensing arrangements with Udio. Sony did not, and the absence of a settlement is what keeps the discovery record in play rather than sealed behind a deal.
How 333 became 30,117
The interesting part is not the number. It is where the number came from. Sony did not sample the internet and infer what a music model must have been trained on. Through discovery in the first case it obtained access to Udio's training data, then ran audio fingerprinting against it and identified hundreds of thousands of its own recordings. The 30,117 asserted in the new complaint are a subset of what the matching turned up.
A training corpus, it turns out, is an inventory. That sentence would have been contested two years ago. Model builders routinely described their datasets as too large, too transformed and too intermingled to be enumerated, and the argument had some force while nobody had looked. Somebody has now looked, with a commodity technology that the music industry has used for content identification for two decades, and produced a list with five significant figures.
Nothing about this is specific to music. Fingerprinting works on audio because audio has robust perceptual hashes, but the general capability, matching a known catalogue against a corpus you have been ordered to hand over, transfers to images, video, text and code with different tooling and the same legal effect.
The claim that is not about copyright
The complaint brings three claims. Two are conventional infringement: post-1972 recordings, and pre-1972 recordings protected under the Music Modernization Act. The third is different in kind. It alleges circumvention of technological protection measures under the Digital Millennium Copyright Act.
Circumvention is an acquisition claim. It does not ask whether training on a work is transformative, or whether the output competes with the original, or whether any of the fair-use factors point one way. It asks how the defendant got past the lock. That distinction is doing quiet structural work in this case, because a defendant can win the argument about what training is and still lose on how the material was collected.
European readers should not file this as an American curiosity. The mining exception in the 2019 copyright directive, the provision most EU firms rely on when they scrape, is conditioned on lawful access and on the rightsholder not having reserved its rights in a machine-readable way. Both conditions attach to acquisition. The doctrinal vocabulary differs on either side of the Atlantic and the pressure point is identical: how you obtained the data is a separate question from what you did with it, and it is answered first.
Why not knowing what you trained on stopped being safe
For three years the practical answer to a provenance question from a customer or a regulator has been a shrug dressed up as engineering. The corpus is enormous, the pipeline ran years ago, the intermediate artefacts were not retained. That answer used to function as a defence because the counterparty had no way to check.
The Udio discovery record ends that. Once a rightsholder can obtain the corpus and match it, the party who cannot describe its own training data is not protected by the gap, it is simply the party with less information about its own exposure. Uncertainty stopped being a shield and became an unpriced liability, and it is now sitting on the balance sheet of every firm that fine-tuned on material it did not document.
The regulatory direction reinforces the same point rather than causing it. The European AI Act requires providers of general-purpose models to publish a sufficiently detailed summary of the content used for training, and the Commission's enforcement powers over those obligations become exercisable on 2 August 2026. A firm that cannot produce that summary internally will not be able to produce it for a supervisor either. The litigation and the regulation are asking the same question from opposite directions.
Put provenance in the contract, not in the assumptions
Three concrete changes, none of which require a legal opinion to start. First, in your next model or fine-tuning agreement, ask for a written provenance warranty that names the principal datasets and states what rights-reservation checks were performed, with indemnity attached. A vendor that will describe its data verbally and not contractually has told you something useful.
Second, keep a manifest of anything you fine-tune on yourself: source, date acquired, licence or exception relied on, and whether a machine-readable reservation was present at the time of collection. That last field is the one nobody records and the one the mining exception turns on. Reconstructing it later is materially harder than logging it now.
Third, separate the two questions in your own risk register. Acquisition lawfulness and training lawfulness fail independently, and a model can be clean on the second and exposed on the first. Most internal AI risk assessments written before this month treat them as one line item. They are two, and the first one is the one a plaintiff reaches with a fingerprint.
Read next: Germany's Record Seed Bets Robots Need Data | Suno's Training Data Is Now an Itemised List



