Method
What a Detection Score Actually Means
A number is not a verdict. The threshold that matters is the one your distributor uses, and it is invisible to you.
A detection score looks like a probability. It usually is not one. It is a model output mapped onto a 0 to 100 range, and what it means depends entirely on where the thresholds sit and what the error rate is around them.
The trap nobody explains
If you get 72 out of 100, the intuitive reading is "28% chance I am fine". That is not how it works in practice, because you are not the one making the decision. Your distributor or platform has an internal action threshold, and it is invisible to you. If theirs sits at 60, your 72 is a rejection regardless of how you read the number.
The score is not the decision. The threshold is the decision, and it belongs to whoever is acting on the finding. Always ask what threshold was applied and what the error rate is in that band.
Why an inconclusive band has to exist
Any honest detector has a range where it does not know. Removing that band makes the output look more decisive and makes the product worse, because the cases that fall in it are exactly the cases that matter: hybrids, heavily processed audio, short clips, instrumental-only material.
A tool that always returns a confident answer is not more accurate. It is less honest.
Reading a split score
Where the vocal and the backing are scored separately, the interesting information is in the gap between them. A high backing score with a low vocal score is the classic signature of a real singer over a generated bed. A single mix score of around 44 for that same track tells you nothing at all.
Questions worth asking any vendor
- What is the false-positive rate in the band my track landed in?
- Is that rate measured on data the model has never seen, or on a held-out slice of its own training corpus?
- Has it been tested under resampling and re-encoding, and where are those results?
- What is the rate for my genre specifically?
- Can I see the composition of the test set?
If a vendor cannot answer the first question, the score is not evidence and should not be treated as any.
Common questions
Not necessarily, and treating it that way is a mistake. A score is a model output, not a calibrated probability, unless the vendor has explicitly calibrated it and published how. Ask what the false-positive rate is in that band.
Not on a score alone, and not at all without knowing the error rate in that band. A published band-level false-positive rate is the minimum standard for acting on one.
When the tool opens
Three checks a day, free, no account.
We will write once when the benchmark is published and once when the detector opens. Nothing else.
Keep reading
The basics
How to Tell If a Song Is AI-Generated
The signals that actually separate a generated record from a played one, and the ones that only look like they do.
If it happened to you
Flagged as AI on a Track You Made Yourself
What a false positive costs you, why there is no appeals process, and how to assemble evidence that a distributor will read.
Method
Why Detectors Flag Human Music
Lo-fi, quantised electronic, and anything made entirely in the box. The failure is systematic and it is documented.