Produced by W2D1 Media. Work with us →
Secured Presented by Chainguard

How AI Pen Testing Actually Works (and Where It Breaks)

19 February 2026
Topics AI

AI is starting to change penetration testing, but most people are asking the wrong question. In this episode of Secured, Cole Cornford sits down with Brendan Dolan-Gavitt, AI researcher at XBOW and former NYU professor, to unpack what autonomous pen testing really is, what it can reliably do today, and what still needs humans.

They explore why AI agents are great at scaling the boring parts of testing, like authenticated workflows and broad vulnerability coverage across huge attack surfaces, and why that does not automatically translate to deep, context-aware exploitation. The conversation also gets into the messy parts: AI systems overclaiming “serious” findings, business logic flaws that are hard to verify, audit expectations, and why scope control needs real guardrails, not vibes. From agent traces and validation models to cost curves and creative exfiltration tricks, this episode is a grounded look at where AI helps AppSec and where it can still cause damage if you trust it too much.

Chapters
Transcript Synced · select any line to jump
Proudly presented by
A Day One® show

Secured by Galah Cyber is produced with Day One — the podcast network for founders, operators and technical teams. Want a show like this for your company?

Work with us →
Produced by W2D1 Media

Turn podcasting into pipeline

We're the team behind the Day One Network, and we helped build Blackbird Ventures' Wild Hearts. We help founders, funds and operators build trust, authority and deal flow with a show tailored to their market.

Investors

Win better deals and stay top‑of‑mind with founders.

Book a call →

Founders & Operators

Close more deals and build a category you own.

Book a call →

Sponsors

Reach founders and operators with a show they trust.

Book a call →