The Augmented Educator

The Augmented Educator

The Noise Floor

Why detecting AI in audio, images, and video fails in the opposite direction from text

Michael G Wagner's avatar
Michael G Wagner
Aug 27, 2026
∙ Paid
Upgrade to paid to play voiceover

This post is going out early as a thank-you to paid subscribers. The full piece opens up to free subscribers next Tuesday. I’m grateful you’re here, and grateful for your commitment to rethinking what education can be. Your support is what keeps this work going.

We need to talk about AI detection again. I know, I have been writing about it a lot lately. But the problem keeps widening, and this time the failure looks different from anything I have covered on The Augmented Educator before.

In my last post, I argued against AI text detection on mathematical grounds. Human writing and machine writing produce overlapping score distributions. Some AI-generated passages will always score as more human than some human-written ones, and vice versa.

Any educator who uses a detector therefore has to set a threshold somewhere inside that overlap. And wherever that dial lands, somebody always pays for it. The two distributions do not permit a clean separation. No threshold can create one. I called the piece There Is Only One Dial and hoped it would close the case.

But, as I came to realize, it closed only a quarter of it.

The other three quarters concern AI-generated images, video, and audio. Education has barely discussed these systems, mostly because they usually operate outside the educator’s view: in stock library upload queues, music distribution pipelines, journal submission portals, and platform moderation systems. Photographers and musicians have been arguing about these tools for years. Most educators have not heard of them.

They should. Because I think these detectors may pose an even greater risk to our students than the text classifiers we have spent so much time debating.

The D’Addario commercial

This piece began, as many of mine do lately, with a YouTube controversy. This time, the trigger was a commercial containing a piece of music that was accused of being AI-generated.

D’Addario is a family-run American manufacturer of music accessories and one of the best-known names in guitar strings. If you play, you have almost certainly had a set of theirs on an instrument. In July 2026, the company launched two extended-range string lines and posted a promotional video built around a high-gain progressive metal track.

Within hours, the comments were filled with accusations. The performance had an odd digital sheen, and the timing was quantized so tightly that it no longer felt played. It lacked the microdynamics that make a guitar sound like a guitar.

The accusation was clear: D’Addario had skipped the musicians and typed a prompt into the AI music generator Suno.

What followed made it worse. D’Addario deleted comments, blocked accounts, and eventually switched comments off completely. The company later explained that the employee who made the track had been doxxed and was being harassed. The moderation was meant to protect him. Whatever the intent, an audience that already suspected a cover-up read the silence as confirmation.

Then a behind-the-scenes video of the project session, posted by D’Addario as proof, drifted out of sync with the commercial, which viewers took as further evidence of fakery.

But that assessment turned out to be wrong. Rhett Shull, a guitarist and producer whose YouTube channel covers gear for a large audience of players, obtained the original Logic Pro session from the company and audited it. The project contained real recorded guitar performances and hand-programmed MIDI. No text prompt ever wrote that song.

What the session did contain, however, was a production chain pushed until it broke: pitch-corrected direct-input guitars, drum compressors stacked in series, every MIDI velocity pinned at maximum, and a dozen synthesizers packed into the midrange. And then, after export, the D’Addario employee added several rounds of automated mastering. Rhett Shull and Steve-san Onotera, another guitarist and YouTuber who posts as samuraiguitarist, both suspected that one of those mastering passes went through Suno.

Mastering is the final stage of music production, when an engineer makes the adjustments that prepare a song for release. A skilled mastering engineer can make a mix sound louder, clearer, and more coherent without ever drawing attention to the work. And some of that work is now handled by automated services such as LANDR and by mastering assistants built into digital audio workstations.

What is usually not known is that Suno’s approach to mastering works differently from these automated systems.

Conventional automated mastering analyzes a finished stereo file, then applies equalization, compression, and limiting to it. Those are the same operations a human engineer would reach for, selected and adjusted by a model.

Suno’s version, on the other hand, runs the audio back through its generative engine and completely rebuilds it. The result is a new waveform rather than a processed copy. And the consequence of that is that the finished waveform is machine-generated, even though the performance underneath it is not.

I find that explanation convincing as an account of why D’Addario’s track sounded artificial, but it remains an educated guess. At the time I am writing this piece, D’Addario has not confirmed which tools were used at that stage, and nothing in the published forensic analysis can prove it either way.

Part of the reason that question is still open is the underlying systemic failure Onotera pointed to.

D’Addario’s communications staff were defending a technical claim they did not understand, with evidence they could not evaluate. And the music professionals watching them did not necessarily understand generative AI any better. The company needed over a week of internal investigation before it could describe its own production chain with reasonable accuracy.

Hold on to that diagnosis. Of all the elements in this story, it is the one with lasting significance.

Thirty-six percent of nothing

In the same video, Onotera described an experiment that should worry any creator deeply. He took a track from his own catalog: human-composed, human-performed, conventionally recorded, with no generative anything anywhere in the chain. He then uploaded it to the AHA Music AI detector, an online tool used across music distribution and content monitoring workflows.

The result was 36 percent AI-generated, with 95 percent confidence. For anyone familiar with the limitations of AI detection, this number should raise a big red flag.

User's avatar

Continue reading this post for free, courtesy of Michael G Wagner.

Or purchase a paid subscription.
© 2026 Michael G Wagner · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture