Skip to content
Join the waitlist

The basics

How to Tell If a Song Is AI-Generated

The signals that actually separate a generated record from a played one, and the ones that only look like they do.

There is a short version of this answer and a long one. The short version is that by ear, in 2026, you probably cannot, and the research agrees with you: in a survey of 9,000 listeners, 97% could not reliably distinguish AI-generated music from human music.

Deezer / Ipsos survey, 9,000 respondents

What actually carries signal

These are the things a detector looks at. Every one of them also occurs innocently, which is exactly why no single item is a verdict.

Room tone and the noise floor

A microphone in a room picks up the room. Air conditioning, traffic, the chair, the performer breathing between phrases. Generated audio often has an unnaturally clean floor, because there was never a room. The catch: a track made entirely inside a DAW with virtual instruments also has no room, and that is most electronic music ever released.

Breath and handling noise

Singers inhale. Guitarists' fingers squeak on wound strings. Drummers hit slightly off-centre. Generated vocals frequently have breath that is either absent or placed too regularly, because it was modelled rather than taken.

Micro-timing

Human performance drifts against the grid in ways that are not random: players rush into fills and lay back on the two and four. Generated output can be too even, or unevenly even in a way that does not correspond to how anybody plays. The catch, again: heavily quantised electronic music is deliberately grid-locked, and gets flagged for it constantly.

Spectral residue from the vocoder

This is the one that is genuinely technical rather than aesthetic. The neural vocoders inside generative music models use upsampling layers, and the mathematics of those layers requires them to leave regularly spaced peaks in the frequency spectrum. It is a property of the architecture, not the training data. It is published, peer-reviewed and reproducible with an ordinary spectrogram.

It is also removable. Pitch shifting moves the peaks. Re-encoding at low bitrate can smear them. Future models can use different upsampling and erase them entirely. Which is why it can be a strong signal and never the only one.

Afchar, Meseguer-Brocal & Hennequin, ISMIR 2025 best paper, arXiv:2506.19108

What does not work

  • "It sounds soulless." Not measurable, not reproducible, and demonstrably wrong at the rate listeners actually perform.
  • Lyrical banality. Plenty of human songs are lyrically banal, and generators have improved fast on exactly this axis.
  • Checking the metadata. Stripped by ordinary pipelines and trivially omitted.
  • One detector's score, taken alone. Detectors disagree with each other on the same file routinely, which means at least one is wrong every time it happens.

The case that beats almost everything

A real singer over a generated backing track. The mix averages to something inconclusive, the detector shrugs, and the track passes. This is the single most common way detection fails in practice, and the only fix is to separate the stems and score them independently.

Why hybrids break detectors goes through that case properly.

Common questions

No, and anyone selling you one is wrong. Every individual signal listed here appears in legitimate human records too. Signals only mean something in combination, and even then a trained listener is worse at this than they think. In a 9,000-person survey, 97% could not reliably tell.

They used to. They mostly do not now, and "it sounds cheap" has stopped being a usable test. What survives is structural and acoustic rather than aesthetic.

Rarely. Metadata is stripped routinely by ordinary upload pipelines, and a generator that wants to hide simply does not write it. Watermarks have the same problem: they prove honesty, not dishonesty.

When the tool opens

Three checks a day, free, no account.

We will write once when the benchmark is published and once when the detector opens. Nothing else.