
.png)
CrowdStrike President Michael Sentonas said, in an interview clip published on September 1, 2026, that AI has pushed phishing click-through rates past 60 percent, up from roughly 11 to 12 percent a year earlier.
That measures what happens after an email arrives. We have been measuring the half before it. Across more than 20,000 phishing emails collected from production environments, AI-generated phishing bypassed Gmail and Microsoft tier-1 filters 50.3 percent of the time. Human-written phishing bypassed the same filters 28.5 percent of the time.
Roughly half of AI-generated phishing reaches a person, and most of the people it reaches click.
State of the AI Threat in Email: 2025 analyzed more than 20,000 distinct phishing emails from production environments, prepared for the Messaging, Malware and Mobile Anti-Abuse Working Group. Each email was classified on two axes: targeting method (spear phishing or mass phishing) and content origin (likely AI-generated or likely human-written). That produced four cohorts to compare directly.
AI-generated messages made up 22 percent of observed volume, and they were not spread evenly. Of the AI-generated emails, 62.9 percent were targeted spear phishing, which inverts the ratio in the dataset overall. Attackers are spending their AI where precision pays.
One caveat from the paper, repeated here: AI-generation classification is probabilistic. The 22 percent figure covers emails assessed as likely AI-generated from observable features, and it sits between industry estimates that range from under 5 percent to over 80 percent.
The 21.8 point gap between AI-generated and human-written phishing does not come from one variable. Three mechanisms reinforce each other.
Content sanitization. Language models write grammatically correct, professionally toned text with none of the linguistic markers NLP-based spam filters were trained to catch. AI-generated spear phishing in our dataset evaded NLP content filters 93.9 percent of the time, at an average length of 562 words, considerably longer than typical phishing. The filters are looking for tells that stopped being there.
Reputation hijacking. AI spear phishing campaigns overwhelmingly send from compromised legitimate business domains carrying established SPF, DKIM, and DMARC histories. Attacks impersonating telecommunications providers reached the inbox 94.2 percent of the time.
Contextual manipulation. 14.5 percent of evaded emails used a Re: or Fwd: prefix to present as an existing thread. Shipping and delivery lures referencing specific project identifiers had the highest individual success rate in the dataset, 93.2 percent across 2,580 instances.
Among AI spear phishing emails that evaded filters, 72.6 percent passed DMARC. Emails that passed DMARC were 17.8 percent more likely to reach the inbox than those that failed.
Authentication is doing exactly what it was built to do. It verifies infrastructure, not intent. When an attacker operates from a compromised legitimate account, SPF, DKIM, and DMARC authenticate the attack. Egress measured 84.2 percent of phishing attacks passing DMARC, and Darktrace found 62 percent bypassing DMARC verification. In May 2024 the FBI, State Department, and NSA issued a joint advisory on North Korean actors exploiting permissive DMARC policies.
75.7 percent of the domains used for AI spear phishing appeared only in that cohort, never shared with the noisier mass campaigns. AI spear phishing also ran the lowest domain diversity ratio of any cohort at 10.9 percent, meaning a small set of high-value compromised domains used repeatedly. AI generic phishing ran the highest at 14.1 percent, consistent with a burn-and-turn model.
The segmentation is deliberate. Keeping targeted operations off the same domains as mass campaigns preserves their reputation and insulates them from mass-campaign blocklists. It also means domain reputation is least reliable exactly where the risk is highest.
AI spear phishing peaks on Fridays around 17:00 UTC, with 48.1 percent landing during business hours and only 4.5 percent on weekends. AI generic phishing peaks Mondays at 17:00 UTC. Human-written phishing peaks Wednesdays at 15:00 UTC. Semperis found 78 percent of organizations cut SOC staffing by half or more on weekends, which is the window a Friday evening send is aimed at.
On targets, AI spear phishing impersonates Google Workspace at 1.8 times, B2B SaaS brands at 2.1 times, and social media platforms at 2.5 times the rate of human-written spear phishing. Compromised cloud credentials open organizational data, internal mail, and every connected service, which is worth more than a consumer account.
95 emails in the dataset carried specific LLM generation artifacts: 32 with unfilled bracketed template variables such as [Company Name], 25 with GPT-family phrasing signatures, 23 with generic AI patterns, and 15 consistent with Claude output.
That is a small count against 20,000, and it is a floor rather than a rate. It is also the clearest evidence in the set that model output is reaching live campaigns without a human reading it first.
76.4 percent of phishing attacks carried at least one polymorphic feature, and 92 percent of polymorphic attacks used AI. Polymorphism means no two messages share a signature, which removes the assumption blocklists, signature matching, and reputation scoring all depend on: that attacks repeat.
Reputation scoring fails on its own terms once attackers weaponize trusted infrastructure. The EchoSpoofing campaign pushed millions of emails through Proofpoint's own relay carrying valid SPF and DKIM signatures from Disney and IBM domains. Phishing is hosted on Amazon S3, Cloudflare Workers, SharePoint, and Google Drive. IBM benchmarked an AI-built phishing campaign at 5 minutes and 5 prompts, against 16 hours for human experts.
Why does AI-generated phishing beat filters that catch human-written phishing?
Mostly because the text no longer carries the signals those filters were built on. Grammatical errors, awkward phrasing, and short bodies were proxies for malice. AI-generated spear phishing averaged 562 words of clean business prose and evaded NLP content filters 93.9 percent of the time.
Should we stop investing in DMARC?
No. DMARC stops direct domain spoofing and remains worth enforcing. It does not stop mail sent from an account the attacker legitimately controls, which is where the AI spear phishing cohort operates.
Both halves of the funnel moved at once. Messages that would once have been filtered now arrive, and messages that would once have looked wrong now read as normal business correspondence. A 60 percent click rate and a 50.3 percent delivery rate compound.
Download the full report for the four cohort breakdowns, the domain forensics, and sanitized case studies of a BEC invoice attack and a backscatter campaign against an executive. If you want to see how AegisAI's agents reason over live mail, book a demo.
Source: Aegis AI, State of the AI Threat in Email: 2025, March 2026, prepared for M3AAWG. Click-through figure: Michael Sentonas, President, CrowdStrike, interview clip published September 1, 2026.