Method
Vocal Versus Backing: Why Hybrids Break Detectors
A real singer over a generated bed is the case that gets past almost everything. Scoring the mix averages the answer away.
The hardest real-world case in AI music detection is not a fully generated track. It is a hybrid: a genuine vocal performance over a generated instrumental, or a human arrangement with one generated element buried in it.
Why the mix score fails
Imagine a track where the vocal is unmistakably human and the backing is unmistakably generated. Scored as one piece of audio, the signals pull in opposite directions and the detector lands somewhere in the middle. That middle number is not a measurement of anything. It is an artefact of averaging two different answers.
Worse, it is usually below whatever threshold triggers action, so the track passes. This is the single most reliable way to get generated music through a detection gate, and it is widely known.
What separation changes
Split the same track and score the parts independently and you get two numbers that each mean something. A low vocal score alongside a high backing score is a specific, legible finding: a real singer, a generated bed. That is an entirely different commercial and licensing situation from a fully generated track, and it deserves a different label.
The taxonomy problem
This is where the whole category gets vague. "AI-generated" is treated as binary when the reality is a spectrum: fully generated, generated then substantially reworked by a person, human composition with generated elements, human performance with AI-assisted mixing, and so on.
The TISMIR authors put it directly: the question of how much AI must be present before the label applies deserves careful attention, and binary classification is reductive. Only one commercial vendor currently distinguishes fully-AI from AI-assisted at all.
Why this matters commercially
- Platform policies differ by category. Some restrict fully-AI tracks while permitting AI-assisted ones, so a binary flag cannot support either decision.
- Licensing consequences differ by generator, and two of the three major labels have signed deals with some of them.
- A disclosure obligation is about what was used, not whether anything was. A yes or no cannot answer it.
Which is why the stem split is not a premium feature bolted on top. It is the part that makes the finding mean anything.
Common questions
Because a mix containing one human element and one generated element averages to an ambiguous middle. The average is the least informative number available and it is the one most detectors report.
Somewhat, and that has to be accounted for in the model rather than ignored. A separation artefact must not be read as a generation artefact, which is one reason separation and scoring have to be designed together.
When the tool opens
Three checks a day, free, no account.
We will write once when the benchmark is published and once when the detector opens. Nothing else.
Keep reading
The basics
How to Tell If a Song Is AI-Generated
The signals that actually separate a generated record from a played one, and the ones that only look like they do.
If it happened to you
Flagged as AI on a Track You Made Yourself
What a false positive costs you, why there is no appeals process, and how to assemble evidence that a distributor will read.
Method
Why Detectors Flag Human Music
Lo-fi, quantised electronic, and anything made entirely in the box. The failure is systematic and it is documented.