New to ShortSpan? We distil the AI-security research that matters into practitioner takeaways — edited by Ben Williams (NCC Group). Get the weekly email
// Analysis

Partial audio deepfakes evade humans and detectors

Defenses
Partial audio deepfakes evade humans and detectors

New research shows short, high‑quality audio inserts from a 2024 speech synthesiser routinely pass as genuine, even for IT professionals warned about deepfakes. Humans and six pretrained detectors miss different things and neither reliably spots which sentence was altered. False alarms on real audio rise too, eroding trust in authentic speech.

Voice cloning is no longer a parlour trick. This study puts numbers on what many incident responders have felt in their gut: modern, short audio edits are hard to catch and easy to weaponise. The researchers tested three synthesis tools from 2019, 2022 and 2024 and asked 82 IT professionals to judge recordings of nine well‑known English speakers. Even with a warning that deepfakes might be present, the 2024 system’s output slid past people and machines.

The setup is practical and mean. Each recording contained four sentences. In the “full spoof” condition, all four were synthetic. In the “partial spoof”, just one sentence in an otherwise genuine clip was replaced. Replacement text came from ChatGPT‑4 to keep content plausible. Humans were benchmarked against six pretrained anti‑spoofing detectors used in speaker verification research.

Old generators were easy mode. For Real‑Time Voice Cloning (2019) and YourTTS (2022), human F1 landed around 0.90, with near‑unanimous per‑sentence judgements. The 2024 commercial system flipped the script. On full spoofs, human F1 fell to 0.48 and per‑sentence accuracy to about 0.43; only 28% of listeners labelled all four sentences correctly. The partial‑spoof case is the real problem: with one sentence swapped, human F1 sank to 0.29, strict “all correct” fell below 10%, and the altered sentence was accepted as bona fide 77% of the time.

Automated defences did not save the day. The six detectors, trained on ASVspoof‑style data, performed well on older generators but collapsed on the out‑of‑distribution 2024 full spoofs (mean F1 ~0.09 versus humans at 0.48). On partial spoofs from the 2024 tool they posted a higher whole‑recording F1 (~0.40) yet failed the core task: strict localisation at the sentence level was 0.00. Humans and detectors failed in different places, and neither could reliably point to the tampered bit.

Why does this work? One good sentence can hijack meaning while the surrounding bona fide audio carries the prosody and room tone that trick both ears and models. Detectors trained for whole‑clip decisions struggle when the attack hides inside clean context. People, primed to expect fakes, overcorrect: genuine speech was flagged as fake around 17.5% of the time, pushing trust in real audio down.

This is not just a detection gap; it is a governance gap. If a single sentence can swing a message and survive both human scrutiny and common detectors, then voice‑based fraud, fabricated “evidence”, and targeted misinformation scale cheaply. The authors argue for procedural verification, provenance, watermarking and segment‑level detection. The open questions are the ones that matter for policy and procurement: which provenance signals survive common edits, how to standardise segment‑level evaluation, and how institutions should treat audio as evidence when plausible deniability is now the default.

Additional analysis of the original ArXiv paper

📋 Original Paper Title and Abstract

Tracking the Trend in How Speech Synthesizers Deceive People

Authors: Milan Šalko, Anton Firc, Kamil Malinka, Vojtěch Staněk, Martin Perešini, Filip Pleško, and Jakub Reš
Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90% for RTVC and YourTTS to 48% for ElevenLabs, although listeners were explicitly warned that deepfakes were present. For partial spoofing, where only one sentence of an utterance is altered, strict accuracy falls to 9%, and listeners classify the synthetic sentence as bona fide 77% of the time. Humans and detectors fail in complementary ways, and neither reliably localizes short manipulations. Additionally, listeners increasingly mislabel bona fide speech as fake, eroding trust in unmanipulated audio. These findings show that human perception alone is unreliable for the selected modern and partial-spoof conditions and motivate procedural verification, provenance, watermarking, and segment-level detection.

🔍 ShortSpan Analysis of the Paper

Problem

The paper examines how advances in speech synthesis affect human and automated detection of deepfake audio. Earlier studies reported high human detection accuracy but mainly evaluated older synthesizers. The authors test whether modern commercial synthesis and a realistic attack variant called partial spoofing, where a single sentence in an otherwise genuine recording is replaced, undermine human judgement and existing automated detectors. This matters for fraud, disinformation and voice‑based social engineering because short, high‑quality inserts can change meaning while appearing credible.

Approach

The authors ran a controlled questionnaire study with 82 IT professionals listening to recordings for nine well‑known English‑speaking public figures. Three synthesis systems were selected to represent 2019, 2022 and 2024: Real‑Time Voice Cloning (RTVC), YourTTS, and the commercial ElevenLabs. For each speaker participants judged seven recordings composed of four sentence‑level segments: one bona fide, three full‑spoof (all sentences synthetic) and three partial‑spoof (one sentence replaced by a synthetic sentence). Replacement text was generated with ChatGPT‑4 and normalized in post‑processing. Six pretrained automated detectors (two front‑ends, two back‑ends, trained on ASVspoof datasets) were run on the same material. Evaluation used F1, a strict All OK metric (all four segments correct), and Over 50 Percent, with additional analyses of sensitivity (d′), response bias, and inter‑rater agreement.

Key Findings

  • Detection of full‑spoof audio remains high for older tools: RTVC and YourTTS yielded human F1 scores around 0.90 and near‑unanimous per‑sentence accuracy.
  • ElevenLabs full spoofs were far harder: mean human F1 dropped to 0.48, per‑sentence accuracy fell to about 0.43 (near chance), and strict All OK success was 28% despite participants being warned that deepfakes might appear.
  • Partial spoofing is especially deceptive: when a single sentence was inserted from ElevenLabs, mean human F1 fell to 0.29 and All OK rate dropped below 10%; the synthetic sentence was judged bona fide 77% of the time (detected only ~22.65%).
  • Humans and detectors fail in complementary ways: detectors perform well on older generators but collapse on out‑of‑distribution ElevenLabs full spoofs (mean detector F1 ~0.09 versus human 0.48); on partial ElevenLabs detectors show higher whole‑recording F1 (~0.40) but their strict All OK localisation rate is 0.00, so neither humans nor detectors reliably localise short manipulations.
  • Listeners increasingly misclassify genuine audio as fake, with a reported false‑alarm rate of about 17.5%, indicating erosion of trust in bona fide speech.

Limitations

The sample comprises IT professionals and may not generalise to the wider public. Participants were explicitly primed that deepfakes could be present and the experiment used a high fixed prevalence of synthetic content, which can raise suspicion and false positives. Only one model represents each release period and ElevenLabs is the sole commercial system, so results reflect system‑specific differences rather than a general temporal trend. Uncertain responses were excluded and the dataset is not public due to identifiable voices.

Implications

Offensively, modern commercial speech synthesis can produce short inserts that are frequently accepted as genuine, enabling convincing targeted misinformation, fabricated evidence and voice‑phishing attacks that change the meaning of otherwise bona fide messages. Attackers can exploit segment‑level edits to evade both human scrutiny and whole‑recording detectors. The complementary failure modes of humans and detectors mean a single defence is insufficient; adversaries benefit from out‑of‑distribution generators and contextually embedded edits. These results argue that attackers can scale plausible social‑engineering campaigns and that defenders should prioritise provenance, watermarking, procedural verification and segment‑localisation methods over reliance on human perception alone.

// Similar research

Related Research

Get the weekly digest

The few AI-security papers that matter, with the practitioner takeaway. No spam.