Skip to content
Join the waitlist

Method

Why Detectors Flag Human Music

Lo-fi, quantised electronic, and anything made entirely in the box. The failure is systematic and it is documented.

False positives in this category are not random noise. They cluster, and they cluster on identifiable kinds of music, which means some producers are far more exposed than others through no fault of their own.

The mechanism

Most detection signals are proxies for "was this performed in a physical space by a person". Room tone, breath, handling noise, micro-timing drift. Music made entirely inside a computer has none of those, not because it is fake but because there was no room and no microphone.

So the detector is not really answering "was this generated". It is answering something closer to "does this have the acoustic fingerprint of a recorded performance". For a large amount of legitimate modern music, the honest answer to that second question is no.

The deeper problem: pipelines, not authorship

The peer-reviewed TISMIR audit found something worse than a tuning issue. Generated output from the major tools tends to arrive at particular sample rates and bitrates. Much of the human reference material in these test sets arrives at a different one. A rule based on sample rate alone reached 83% precision on their data.

Their conclusion, stated plainly, was that detectors may be identifying the production pipeline rather than whether a machine composed the music. They called it shortcut learning.

Cros Vila, Sturm, Casini & Dalmazzo, TISMIR 8(1), DOI 10.5334/tismir.254

What that means if you make music

  • Exporting at an unusual sample rate or bitrate can change your result, which tells you the result was never purely about the music.
  • A detector that has not been tested against re-encoding has not been tested at all.
  • A single blended accuracy number hides all of this. The per-genre and per-codec breakdown is where the truth is.

What an honest detector owes you

It should tell you the false-positive rate for music like yours, not music in general. It should be tested under resampling and re-encoding and publish those results. It should have an inconclusive band and use it rather than forcing a decisive-looking answer. And it should separate the stems, because a full-mix score on a hybrid is close to meaningless.

None of that is exotic. It is simply not what the category currently does.

Common questions

Lo-fi, quantised electronic and techno, trap, in-the-box pop production, sample-based hip-hop, and instrumental or orchestral music built from virtual instruments. The common factor is the absence of an acoustic recording environment.

Not reliably. Detection improves on the models it has seen and degrades on the ones it has not. One study found a detector trained largely on one generator missing three quarters of another generator\u2019s output.

When the tool opens

Three checks a day, free, no account.

We will write once when the benchmark is published and once when the detector opens. Nothing else.