Skip to content
Join the waitlist

Accuracy

We publish where we are weak.

There is no accuracy figure on this page, because we have not earned one yet. Here is exactly what we will publish, and the standard we are holding ourselves to.

Current status

The benchmark set is in design. No detection accuracy has been measured, and no score has been produced by this site. Any number you see elsewhere in this category that is not accompanied by a released test set is an assertion, including ours when it arrives.

The state of the field

What happened when somebody finally checked.

In 2025 a group at KTH published a peer-reviewed audit of commercial and academic AI-music detectors in TISMIR. They tested against 10,000 Suno tracks, 10,000 Udio tracks and 10,000 human recordings.

The one commercial detector that publishes a false-positive bound advertises under 1%. Measured, it was 4.7%.

Reducing the sample rate to 22.05kHz caused it to misclassify every Suno sample tested. High-pass filtering made it call everything machine-made. Low-pass filtering made it miss everything. Confidence on those wrong answers ran from 50% to 97%.

The leading academic model in the same test missed 75.3% of Udio tracks, having been trained largely on Suno.

The authors' conclusion is the uncomfortable one: the detectors may be identifying production pipelines rather than authorship.

Independent findings

  • 4.7% measured false-positive rate against an advertised <1%
    TISMIR 8(1), DOI 10.5334/tismir.254
  • 69.26% false-positive rate measured for one research model on unseen data
    ArtifactNet, arXiv:2604.16254
  • A re-encode moved one model's output probability by 0.95
    Same
  • The market leader's own research arm published a formal caveat against "a flourishing market of artificial content checkers"
    Deezer Research, arXiv:2501.10111

Our commitment

The shape of the number, before the number.

Separate rates

False positives and false negatives reported apart, never blended into one accuracy figure. The two errors have completely different costs: a missed fake is a rounding error, a wrongly flagged musician is somebody losing a release.

Broken out

By generator, by genre, and by codec and sample rate. A single headline number hides the cell that matters to you. If you ship production music, the instrumental figure is the only one you should care about.

Adversarially tested

Resample, pitch shift, re-encode at low bitrate, high-pass and low-pass. Published results under each. This is precisely where the audited commercial detector broke, and precisely what nobody publishes.

Fixed and dated

The benchmark set frozen, dated and described, so a figure can be reproduced rather than taken on trust.

Versioned findings

Every finding carries the index version that produced it, so a verdict from March can be re-litigated against the corpus that existed in March. That is the only way a finding survives a dispute.

Weak spots printed

The worst class shown next to the best. A detector that hides its floor is not publishing accuracy, it is publishing marketing.

Questions about accuracy

Because we have not measured one, and the category is full of figures nobody can reproduce. When we have run the benchmark we will publish the set composition alongside the result so the number can be checked rather than believed.

It averages the rate of correctly identifying generated tracks with the rate of correctly identifying human ones. A detector can score 95% overall while being catastrophic on human music if the test set is mostly generated. Separate false-positive and false-negative rates are the honest presentation.

Instrumental-only output, on every published result in the field. No vocal means no breath, no handling noise and no micro-timing, which removes most of the human-performance signal at once. Any detector quoting one blended number is hiding this.