We need to talk about AI detection again. I know, I have been writing about it a lot lately. But the problem keeps widening, and this time the failure looks different from anything I have covered on The Augmented Educator before.
In my last post, I argued against AI text detection on mathematical grounds. Human writing and machine writing produce overlapping score distributions. Some AI-generated passages will always score as more human than some human-written ones, and vice versa.
Any educator who uses a detector therefore has to set a threshold somewhere inside that overlap. And wherever that dial lands, somebody always pays for it. The two distributions do not permit a clean separation. No threshold can create one. I called the piece There Is Only One Dial and hoped it would close the case.
But, as I came to realize, it closed only a quarter of it.
The other three quarters concern AI-generated images, video, and audio. Education has barely discussed these systems, mostly because they usually operate outside the educator’s view: in stock library upload queues, music distribution pipelines, journal submission portals, and platform moderation systems. Photographers and musicians have been arguing about these tools for years. Most educators have not heard of them.
They should. Because I think these detectors may pose an even greater risk to our students than the text classifiers we have spent so much time debating.
The D’Addario commercial
This piece began, as many of mine do lately, with a YouTube controversy. This time, the trigger was a commercial containing a piece of music that was accused of being AI-generated.
D’Addario is a family-run American manufacturer of music accessories and one of the best-known names in guitar strings. If you play, you have almost certainly had a set of theirs on an instrument. In July 2026, the company launched two extended-range string lines and posted a promotional video built around a high-gain progressive metal track.
Within hours, the comments were filled with accusations. The performance had an odd digital sheen, and the timing was quantized so tightly that it no longer felt played. It lacked the microdynamics that make a guitar sound like a guitar.
The accusation was clear: D’Addario had skipped the musicians and typed a prompt into the AI music generator Suno.
What followed made it worse. D’Addario deleted comments, blocked accounts, and eventually switched comments off completely. The company later explained that the employee who made the track had been doxxed and was being harassed. The moderation was meant to protect him. Whatever the intent, an audience that already suspected a cover-up read the silence as confirmation.
Then a behind-the-scenes video of the project session, posted by D’Addario as proof, drifted out of sync with the commercial, which viewers took as further evidence of fakery.
But that assessment turned out to be wrong. Rhett Shull, a guitarist and producer whose YouTube channel covers gear for a large audience of players, obtained the original Logic Pro session from the company and audited it. The project contained real recorded guitar performances and hand-programmed MIDI. No text prompt ever wrote that song.
What the session did contain, however, was a production chain pushed until it broke: pitch-corrected direct-input guitars, drum compressors stacked in series, every MIDI velocity pinned at maximum, and a dozen synthesizers packed into the midrange. And then, after export, the D’Addario employee added several rounds of automated mastering. Rhett Shull and Steve-san Onotera, another guitarist and YouTuber who posts as samuraiguitarist, both suspected that one of those mastering passes went through Suno.
Mastering is the final stage of music production, when an engineer makes the adjustments that prepare a song for release. A skilled mastering engineer can make a mix sound louder, clearer, and more coherent without ever drawing attention to the work. And some of that work is now handled by automated services such as LANDR and by mastering assistants built into digital audio workstations.
What is usually not known is that Suno’s approach to mastering works differently from these automated systems.
Conventional automated mastering analyzes a finished stereo file, then applies equalization, compression, and limiting to it. Those are the same operations a human engineer would reach for, selected and adjusted by a model.
Suno’s version, on the other hand, runs the audio back through its generative engine and completely rebuilds it. The result is a new waveform rather than a processed copy. And the consequence of that is that the finished waveform is machine-generated, even though the performance underneath it is not.
I find that explanation convincing as an account of why D’Addario’s track sounded artificial, but it remains an educated guess. At the time I am writing this piece, D’Addario has not confirmed which tools were used at that stage, and nothing in the published forensic analysis can prove it either way.
Part of the reason that question is still open is the underlying systemic failure Onotera pointed to.
D’Addario’s communications staff were defending a technical claim they did not understand, with evidence they could not evaluate. And the music professionals watching them did not necessarily understand generative AI any better. The company needed over a week of internal investigation before it could describe its own production chain with reasonable accuracy.
Hold on to that diagnosis. Of all the elements in this story, it is the one with lasting significance.
Thirty-six percent of nothing
In the same video, Onotera described an experiment that should worry any creator deeply. He took a track from his own catalog: human-composed, human-performed, conventionally recorded, with no generative anything anywhere in the chain. He then uploaded it to the AHA Music AI detector, an online tool used across music distribution and content monitoring workflows.
The result was 36 percent AI-generated, with 95 percent confidence. For anyone familiar with the limitations of AI detection, this number should raise a big red flag.
Then he ran the disputed D’Addario track through the same tool. The result was 74.7 percent AI-generated, with 95 percent confidence. Other detectors gave the same file a “Suno match” score of between 73.88 and 91.40 percent.
If the mastering suspicion is right, those tools were picking up a genuine Suno signature sitting on top of a human performance. They may have been correct about the artifact, but they were wrong about everything anybody should ever really care about.
Let’s look at those numbers more closely, starting with that confidence figure, because it is the part people misread the most. It expresses how certain the model is within its own system. It does not independently validate anything. Onotera’s control experiment shows why that distinction is worth making: the tool was highly confident about a result we know was wrong.
A confidently wrong number is more dangerous than an openly uncertain one, because it gives a guess the authority of a measurement.
There is also something hidden in the percentage itself. Audio engineers have a word for it. The noise floor is the bed of hiss underneath a recording, and once a signal falls below it, pulling the two apart gets difficult. In Onotera’s test, 36 percent on a track with no generative involvement at all behaved like the detector’s noise floor. More than a third of the scale had already been used up before any meaningful measurement was taken.
A single control track is not a proper statistical sample. But it does show what that number is not. It is not a literal estimate of how much of a recording was generated with AI. The D’Addario track scored higher, and the tool could not tell anyone what it had found. Generated composition? Generated performance? An automated mastering pass? Or simply the artifacts of very heavy production?
The scale collapses all of those into one figure and hands you a percentage. That is a different problem from the one in my last post, but it is arriving at the same place. There, two overlapping distributions meant no threshold could cleanly separate them. Here, the scale is not measuring what people think it does.
Two ways to be wrong
Research shows that text detectors and media detectors both fail. But they fail in opposite directions, and they take down opposite people on the way. A policy written for one of them will therefore be wrong about the other.
Text is made of discrete tokens. A language model produces words by drawing the next one from a probability distribution, and a detector measures how surprising the resulting sequence is to a reference model. That is perplexity. Alongside it sits burstiness, the variation in sentence length and structure across a passage. Machine text tends to run smooth and evenly paced. Human text tends to be less steady. Those are two of the signals most text detectors lean on.
The failure mode falls straight out of this mechanism. Careful, plain, grammatically conservative writing scores as machine writing. The detectors punish clarity, and they punish it hardest in the people who worked hardest to achieve it. I have written about this problem at length in previous posts on this Substack.
AI images, video, and audio are different mathematical animals. They are continuous, high-dimensional signals, and most current generators build them through diffusion.
Generation starts with static noise and strips it away step by step until a picture or a waveform emerges. Detectors for this kind of content hunt for the traces that denoising leaves behind. Those include reconstruction errors, frequency-domain artifacts, and spectrogram phase relationships that a physical microphone is unlikely to produce.
But those traces are fragile. On the GenImage benchmark, which tests detectors across eight major generators including Midjourney and Stable Diffusion, standard classifiers average 71.7 to 72.5 percent accuracy on architectures they were not trained on. Push a synthetic image through ordinary JPEG compression or scale it down, and reported accuracy falls to between 50.6 and 67 percent.
At the bottom of that range, the detector is barely doing better than a coin flip.
Passive audio detectors show the same collapse, with false negative rates of 20 to 40 percent once a file has been through MP3 or AAC encoding at consumer bitrates. 20 to 40 percent! Between one in five and two in five synthetic files sail straight through.
Active watermarking was supposed to solve the problem by signing content at the moment of generation, and under clean conditions it does indeed work. Tree-Ring watermarking reaches a ROC-AUC of 0.993. ROC-AUC measures how cleanly a classifier separates two classes, so 0.993 is close to perfect.
But researchers then showed that the mark can be stripped using the same publicly available autoencoder the model itself relies on. The score falls to 0.153, and the true positive rate drops from 96.8 percent to a mere 4 percent, with no visible damage to the image.
So text detection primarily fails by accusing the innocent, whereas media detection fails by waving the guilty through.
The tools that sand off the fingerprints
Media detection has a second failure mode: the ordinary enhancement tools creators already use erase the very traces that forensic detectors depend on. This one takes a little technical ground to cover. So, bear with me.
One of the signals a forensic image detector looks for is Photo-Response Non-Uniformity, or PRNU. This is the microscopic pattern of variation between individual pixels on a physical camera sensor. PRNU is unique to that sensor and burned into every frame it captures. It is physical noise produced by one specific piece of hardware.
A generated image has no camera sensor behind it, so some forensic systems treat the absence of that signature as suspicious.
Now think about what a photographer does to a high-ISO file. They use post-production tools such as Adobe AI Denoise, DxO DeepPRIME, Topaz Photo AI, or Luminar Neo. Those tools treat sensor noise as the problem they exist to solve. They suppress it, then reconstruct the fine detail with the help of AI models. The photograph comes out looking better, but it comes out carrying less of the evidence that a camera made it.
The signature that proved it came from a camera has been professionally removed and replaced with the output of a neural network.
Generative upscalers go even further. Tools such as Magnific do more than interpolate the pixels that are already there. They invent the missing detail with a generative model, which can leave behind the same high-frequency artifacts that detectors associate with fully synthetic images.
Audio works the same way, except that the markers being erased are acoustic rather than optical.
What proves a recording happened in a room is the room itself. It is the reverberation profile, the subperceptual noise floor, the breathing, and the continuous way a voice slides from one sound into the next.
Adobe Podcast’s Enhance Speech and Descript’s Studio Sound are designed to strip much of that away. They then rebuild the missing frequency bands with neural vocoders. And a vocoder can leave phase shifts in the upper harmonics and flatten the energy contours. Which is the kind of signature a deepfake voice detector is trained to flag.
Similarly, automated mastering platforms such as LANDR, iZotope Ozone, and eMastered reshape high-frequency harmonics and phase relationships across a stereo master. And if you run human stems through Suno’s generative mastering, the resulting file will pick up the acoustic marks of AI generation, with the instrument dynamics altered along the way.
Human production and machine generation are also converging acoustically. Generative models were trained on commercial music that had already been pitch-corrected, quantized, dynamically limited, and stereo-widened, so they learned to produce exactly those artifacts.
When a human engineer pushes the same techniques far enough, the output starts to arrive at the acoustic profile of generated music.
The repercussions of all of this are already being felt. Researchers have to defend themselves against integrity investigations for running micrographs through routine denoising. And photographers are losing stock library accounts over files their own cameras produced.
There is nothing left in the file to find
Handing an algorithm a finished MP3, JPEG, or MP4 and asking whether a human made it asks the file for a history it no longer carries. A final export is a rendered surface. It holds the statistical residue of whatever touched it last. And in professional work, what touched it last is very often a neural network doing legitimate enhancements.
Most passive commercial detectors reduce the question to a binary classification. They assume each asset is either wholly human or wholly machine-generated. Yet hardly anything created professionally today is exclusively one or the other.
A photograph shot on a physical camera and upscaled by Magnific is a hybrid. So is a song played by a person and mastered by Suno. The composition, the performance, the judgment, and the meaning are human. The surface statistics are not. But the detector reads only the surface, because the surface is all a passive classifier can reach.
Which brings us back to the problem with text classifiers, only with greater force. Post-hoc media detection is not an engineering problem awaiting a better model. It asks a question about the history of a file, and the exported file preserves only a fraction of that history. A better classifier cannot reconstruct a provenance record that nobody kept.
Verification therefore has to move upstream, into the process. Under the C2PA standard, a creator can attach a signed manifest at capture and record every later edit in a tamper-evident chain. A platform could then see that an image came out of a participating camera and was refined with Topaz afterward.
Workings.io aims to apply the same logic to creative work more broadly. It is in closed alpha at the moment. I have been testing it and plan to report on my experience in a future article.
The low-tech version of this approach is straightforward. Save the camera RAW files and the multitrack sessions, in case somebody later needs to inspect them. D’Addario was cleared because the Logic session still existed and somebody could open it. That is the entire mechanism. Not a score. An artifact a person could inspect.
You do not have to use AI tools to be judged by them
A photographer who has never typed a prompt adds AI-like signatures to his work just by denoising a high-ISO file. A podcaster introduces vocoder artifacts into her own voice by cleaning up a recording. An illustrator adds diffusion traces while upscaling a hand-drawn image for print. And a student who writes plainly because English is their third language walks into a 61 percent false-positive rate.
None of them tried to pass off generated work as their own. Some may not even have realized that an enhancement tool was running a generative model. Yet all of them can be scored by AI, and some can be falsely accused because of it.
That is why AI literacy is a baseline professional requirement and not a niche interest for technologists. Learning to use the tools is only half the job. The harder half is being able to explain what an automated score is measuring, and why it may be wrong.
A company with engineers on staff needed more than a week to explain what had happened to its own audio file. A student before an integrity panel gets far less time, far less help, and a panel that usually cannot interpret the score either.
The guitar community was right that something was off about that track. It had been polished into a lifeless sheen, and in a demo whose whole purpose was to let you hear a string respond, that is a real failure. It deserved the criticism. But the crowd reached for the wrong explanation, and when the detectors were brought in to settle it, they appeared to confirm it. But they were wrong.
They were reading the noise floor, not the work.
The images in this article were generated with Nano Banana 2.
If you’d like to go further, the following NotebookLM-generated audio deep dive goes beyond the post into the broader research behind it, drawing on the sources and notes I gathered along the way. This is meant as a companion to the argument, offered as an optional extra rather than a summary of it.
P.S. I believe transparency builds the trust that AI detection systems fail to enforce. That’s why I’ve published an ethics and AI disclosure statement, which outlines how I integrate AI tools into my intellectual work.







