Can AI-generated phishing emails bypass spam filters?

AI-written phishing emails often don't need to beat your filter — they look like ordinary mail. What the research shows, and how to prepare your team.

Yes — and the uncomfortable part is how: AI-generated phishing emails increasingly don't defeat spam filters at all. They simply look like ordinary email. A peer-reviewed study that put AI-written spear phishing in front of real people measured a 54% click-through rate for fully automated AI emails — on par with human experts and roughly four and a half times higher than generic control emails — because the messages were clean, personal and free of the tells that trigger a filter.

TL;DR

What the research actually shows

The most-cited evidence is a study by researchers from the University of Chicago, Argonne National Laboratory and elsewhere, published as "Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects" (arXiv:2412.00586). It ran a controlled human-subject trial with 101 participants across email groups:

The AI pipeline also scoped targets automatically — searching the web for professional details — at a cost the authors measured in cents and minutes per target. The finding that matters for defenders is not that AI is magical; it is that a fully automated pipeline now matches skilled human attackers at negligible marginal cost. Volume and quality, historically a trade-off, are converging.

Why filters miss the new generation

Legacy filters were built for a specific failure profile: misspelled words, awkward phrasing, suspicious attachments, links to known-bad infrastructure. AI-written email breaks that profile in three ways.

Perfect surface, personalised core. A message that references your actual vendor names, your billing cycle, and a project mentioned on LinkedIn reads like ordinary business correspondence to a machine scanning for anomalies.

No payload to fingerprint. Filter-heavy approaches lean on attachments and links to flagged domains. Modern lures avoid both — a link to a legitimate-looking sign-in page hosted on fresh infrastructure, or simply a reply request that hands the conversation to a human attacker.

Attackers can pre-test. The same tools that write the email can be asked whether a spam classifier would catch it, and the wording adjusted until it would not.

None of this makes filters useless — they still stop the vast majority of bulk campaigns, and older AI-era lures that reuse known infrastructure get caught routinely. The honest statement is narrower: filters were never able to judge whether an email that reads like your supplier's genuine request is actually your supplier, and AI has made that category of email vastly cheaper to produce.

What this changes for small teams

The security model most people grew up with — the filter is the firewall, the user is a filter failure — inverts. When the marginal cost of a convincing, personalised email approaches zero, the number of people who receive one they'd have to detect grows. The measured outcomes (a 54% click rate in the study, achieved by automation) say plainly that untrained people do not detect them.

What actually moves outcomes:

What good practice looks like in 2026

  1. Monthly, varied simulations rather than an annual test. The point is calibration — everyone's internal sense of "this looks wrong" needs refreshing as lures evolve.
  2. Auto-enrol clickers in a short lesson. The teachable moment is the click itself; a three-minute module on the exact lure they fell for beats a quarterly newsletter.
  3. Track the trend, not the event. Click rates on hard templates should fall over months. Human risk reporting turns that into a per-learner score so follow-up reaches the right people.
  4. Keep the technical layer current. MFA, modern email authentication (SPF, DKIM, DMARC) and prompt patching still do the heavy lifting for the majority of spam-tier mail. The training is for what gets through — not a substitute for the filter.

FAQ

Can AI phishing emails bypass Microsoft 365 or Google spam filters? They can avoid the triggers those filters look for — malicious links, attachments, spam-style wording — because AI-written copy reads like normal business email. Whether a specific message lands in spam depends on the sender's infrastructure and reputation as much as the wording. Assume that a determined, well-resourced sender can reach the inbox sometimes, and design your process so that outcome is survivable.

Are AI phishing emails more dangerous than human-written ones? The study above found AI-generated emails performed on par with human expert emails (54% click-through each) and far above generic phishing (12%). The danger is less per-email potency than cost: automation makes expert-quality lures cheap enough to send at scale.

Do spam filters use AI too? Yes, heavily. Modern filters are machine-learning systems and are improving as well. But detection and deception are co-evolving, and the structural gap remains: a filter sees one email; your colleague knows the vendor's actual invoice schedule. The human layer is still where personalised attacks are won or lost.

What's the single best defence against AI-written phishing? Verification processes that don't depend on the reader's judgment for high-stakes actions — call-back rules for payment changes, MFA everywhere, and one-click reporting for everything else. Training raises the odds of the right judgment; process guarantees it.

How often should we run simulations to keep up? Monthly. A year of varied campaigns from one setup (as Cyber Aware's phishing programme runs) keeps cadence without admin overhead, and gives you a measurable trend instead of a single annual datapoint.

One last thing

Your filter has never been the thing standing between a convincing email and an accounts payable click. What has changed is the price of a convincing email, which has fallen to roughly nothing. The response that survives that math is practice plus process: realistic simulations on a monthly cadence, one-click reporting, and a call-back rule that means no single email — however flawless — can move money.

Related guides

Ready to deploy

Same playbook.
Your brand.

Cyber Aware's Human Risk Score works the same way for every MSP partner - under your brand, on your cadence.