Keepnet – AI-powered human risk management platform logo
Menu
HOME > blog > understanding voice generation ai and vishing threats

Understanding Voice Generation AI and Vishing Threats

Discover how voice generation AI is fueling vishing attacks, their implications for communication carriers, and best practices for protection.

Ozan Ucar, Founder and CEO of Keepnet

Understanding Voice Generation AI and Vishing Threats

Vishing, or voice-based phishing, has evolved from crude robocalls to sophisticated scams powered by voice generation AI and deepfake technology. These advancements have made vishing a significant threat to businesses and individuals alike.

The channel is the weak point. In the 2026 Verizon Data Breach Investigations Report, phone-centric simulations show a median click rate of about 2% against about 1.4% for email, so people fail voice-led attempts roughly 40% more often than email ones (p. 50). Generated voice removes the last cue people relied on, which is whether the caller sounds like the person they claim to be.

In this blog, we’ll uncover how AI fuels vishing attacks, the challenges this creates, and strategies to combat these growing threats.

What Is Vishing?

Vishing involves fraudulent phone calls or voice messages designed to trick victims into revealing sensitive information or performing harmful actions. Cybercriminals often impersonate trusted figures, such as executives or bank representatives, to gain credibility and exploit human trust. These scams rely heavily on social engineering tactics to manipulate their victims.

Editor's Note: This article was updated on March 12, 2026.

Deepfakes: The Game-Changer in Vishing

Deepfakes use AI to generate realistic voices or videos that convincingly mimic real individuals. This technology has turned vishing from basic scams into highly sophisticated attacks that are nearly impossible to detect. For example, attackers can convincingly replicate a CEO’s voice to authorize fraudulent transactions or extract confidential data. As deepfake technology becomes more accessible, such scenarios are becoming increasingly common and dangerous.

The Growing Wave of Vishing Attacks

Voice is now a primary channel rather than a fallback. Smishing and voice-led fraud grew 30 to 40 percent quarter over quarter through 2025 (APWG, Phishing Activity Trends Report, Q4 2025, p. 4), and Microsoft reported that AI-assisted phishing reached a 54% click-through rate against 12% for standard attempts (Microsoft Digital Defense Report 2025). A notable example of the risks posed by AI-generated voices occurred when a journalist successfully bypassed a bank’s voice authentication system using an AI-generated voice, exposing critical vulnerabilities in voice ID technologies.

Case Study: The Joe Rogan Deepfake

A vivid example of deepfake misuse occurred in early 2023, when a TikTok ad featured podcast host Joe Rogan seemingly endorsing a male enhancement supplement called Alpha Grind. The ad, created using AI-generated voice and imagery, was entirely fake and published without Rogan’s consent. This misleading video gained millions of views before being removed, illustrating how deepfakes can damage reputations, deceive consumers, and spread misinformation. The incident underscores the urgent need for businesses to implement safeguards against such AI-driven threats.

Vishing-as-a-Service: A Growing Threat

Cybercriminals are increasingly turning to vishing-as-a-service, using AI-powered platforms to automate and scale voice phishing campaigns. While these platforms often have legitimate applications, their misuse has fueled the rise of professionalized cybercrime, making vishing attacks more prevalent and sophisticated.

Challenges in Detecting and Preventing Vishing

AI-generated voices are highly realistic, rendering traditional detection methods ineffective. Many voice authentication systems are now vulnerable, requiring organizations to adopt more advanced tools like biometric analysis and voice pattern recognition to combat these threats.

The 2026 case where no voice cloning was needed

Voice cloning gets the attention, and this page has covered why. It is worth noting a campaign that did the same damage without it. Google Threat Intelligence Group reported in August 2026 on the UNC6671 extortion cluster, whose intrusions begin with an ordinary human caller posing as the internal IT help desk.

There is no synthetic audio in that chain. The caller asks the employee to enrol a FIDO2 passkey or update MFA, reaches them on a personal mobile, and in some campaigns the corporate help desk number is spoofed on the display. The employee is sent to a lookalike enrolment page, and the credential and MFA response are relayed to the real login during the call.

The lesson for defenders is uncomfortable. A programme that teaches people to listen for synthetic speech teaches the wrong test, because a real human voice with a plausible request defeats it completely. The reliable signal is not how the voice sounds but what is being asked for: an identity change, on an inbound call, right now.

Source: Google Threat Intelligence Group, UNC6671 Rebrands: Multi-Brand Vishing Extortion Targets Financial Services and Enterprise Cloud Environments, August 2026.

Best Practices to Defend Against Vishing

  • Strengthen Multi-Factor Authentication (MFA): Combine voice authentication with additional layers, like SMS codes or app-based approvals, to enhance security.
  • Adopt Advanced Biometric Tools: Use technologies that analyze voice cadence and patterns to detect and block AI-generated fraud attempts.
  • Educate Employees: Implement security awareness training to help employees recognize vishing scams, verify unusual requests, and avoid falling victim to social engineering tactics.

Keepnet Extended Human Risk Management Platform and Secure Behavior Management

The rise of voice generation AI has made vishing attacks harder to detect and prevent. Organizations need practical solutions to strengthen their defenses and minimize risks. Keepnet’s tools are tailored to address these challenges:

  • Vishing Simulator: Train employees to identify and respond to voice phishing attacks with realistic, controlled simulations. These exercises replicate advanced scams, improving employee vigilance and decision-making under pressure.
  • Security Awareness Training: Build a cyber-aware workforce with tailored training programs that prepare employees to spot and respond effectively to vishing and other social engineering threats.

Together, these solutions empower organizations to stay ahead of evolving vishing threats and safeguard their most valuable assets.

SHARE ON

twitter
linkedin
facebook

Schedule your 30-minute demo now

You'll learn how to:
tickEnhance your organization's defenses with advanced deepfake detection and biometric tools.
tickEmpower your team through cutting-edge security awareness training programs.
tickImplement multi-factor authentication strategies to strengthen communication security.

Frequently Asked Questions

How does AI voice generation change vishing?

arrow down

It removes the last cue people relied on, which was recognising the voice. A short public recording is enough to produce audio convincing on a phone call.

How much audio does a voice clone need?

arrow down

Modern tools need only seconds of clear speech. Conference talks, podcasts, webinars and voicemail greetings all provide it.

Why are executives the main target?

arrow down

Because their voices are public and their instructions are rarely questioned. 35% of organisations have experienced a deepfake incident, while only 10% of security leaders prioritise deepfake training (Gartner, "6 Ways to Transform Your Cybersecurity Awareness Program", G00840741, March 2026, n=65).

Can people detect a cloned voice?

arrow down

Not reliably, especially over a phone line under time pressure. Detection advice ages quickly, which is why verification through a second channel is the durable control.

What is vishing as a service?

arrow down

Criminal services offering calling infrastructure, scripts and voice generation on subscription, which removes the skill barrier the same way phishing kits did for email.

How do you defend against AI voice attacks?

arrow down

Callback verification on a number the organisation already holds, a second approver on payment and access changes, and voice simulation so employees meet the scenario before the real call.

Does vishing always involve AI voice cloning?

arrow down

No. Google Threat Intelligence Group reported in August 2026 that the UNC6671 extortion cluster runs its calls with human operators posing as internal IT, asking employees to enrol a passkey or update MFA. No synthetic audio is involved. Training that focuses on detecting a cloned voice misses this entirely, which is why the trigger to teach is the request itself rather than the sound of the caller.