Method
Why Detectors Flag Human Music
Lo-fi, quantised electronic, and anything made entirely in the box. The failure is systematic and it is documented.
False positives in this category are not random noise. They cluster, and they cluster on identifiable kinds of music, which means some producers are far more exposed than others through no fault of their own.
The mechanism
Most detection signals are proxies for "was this performed in a physical space by a person". Room tone, breath, handling noise, micro-timing drift. Music made entirely inside a computer has none of those, not because it is fake but because there was no room and no microphone.
So the detector is not really answering "was this generated". It is answering something closer to "does this have the acoustic fingerprint of a recorded performance". For a large amount of legitimate modern music, the honest answer to that second question is no.
The deeper problem: pipelines, not authorship
The peer-reviewed TISMIR audit found something worse than a tuning issue. Generated output from the major tools tends to arrive at particular sample rates and bitrates. Much of the human reference material in these test sets arrives at a different one. A rule based on sample rate alone reached 83% precision on their data.
Their conclusion, stated plainly, was that detectors may be identifying the production pipeline rather than whether a machine composed the music. They called it shortcut learning.
Cros Vila, Sturm, Casini & Dalmazzo, TISMIR 8(1), DOI 10.5334/tismir.254What that means if you make music
- Exporting at an unusual sample rate or bitrate can change your result, which tells you the result was never purely about the music.
- A detector that has not been tested against re-encoding has not been tested at all.
- A single blended accuracy number hides all of this. The per-genre and per-codec breakdown is where the truth is.
What an honest detector owes you
It should tell you the false-positive rate for music like yours, not music in general. It should be tested under resampling and re-encoding and publish those results. It should have an inconclusive band and use it rather than forcing a decisive-looking answer. And it should separate the stems, because a full-mix score on a hybrid is close to meaningless.
None of that is exotic. It is simply not what the category currently does.
Common questions
Lo-fi, quantised electronic and techno, trap, in-the-box pop production, sample-based hip-hop, and instrumental or orchestral music built from virtual instruments. The common factor is the absence of an acoustic recording environment.
Not reliably. Detection improves on the models it has seen and degrades on the ones it has not. One study found a detector trained largely on one generator missing three quarters of another generator\u2019s output.
When the tool opens
Three checks a day, free, no account.
We will write once when the benchmark is published and once when the detector opens. Nothing else.
Keep reading
The basics
How to Tell If a Song Is AI-Generated
The signals that actually separate a generated record from a played one, and the ones that only look like they do.
If it happened to you
Flagged as AI on a Track You Made Yourself
What a false positive costs you, why there is no appeals process, and how to assemble evidence that a distributor will read.
Method
What a Detection Score Actually Means
A number is not a verdict. The threshold that matters is the one your distributor uses, and it is invisible to you.