Where AI Actually Helps in Vulnerability Research

Matteo Strada (mstreet), 08 June 2026

Where AI Actually Helps in Vulnerability Research

After two engagements in a row where I leaned on an AI coding assistant, one a firmware reverse-engineering pass on a consumer IoT device, one a whitebox source review of a 100 kLOC open-source C++ project, I want to write down what I actually saw it do, what I saw it not do, and how I’d structure the next one differently.

This isn’t a “AI changes everything” post and it isn’t a “AI is overhyped” post. What I have is two data points, opposite enough in shape that the contrast is informative, and concrete enough that I can name the moves that worked and the ones that didn’t. So this is just observation, with opinions where the evidence licenses them.

If you want the source material: the router work is in Teaching the Machine Where to Look and the source-review work is in Four Bugs You Can Reach With nc.


Two engagements, two shapes

The router was a black-box-to-whitebox firmware audit. I dumped the firmware over UART, opened the CGI binaries in Ghidra, and after a handful of manual finds Daniele and I noticed the bugs were all the same shape: a web_get() source flowing into a do_system() sink with no sanitization between them. The interesting move there wasn’t finding one of them. It was scaling the search across every binary on the firmware once we’d named the pattern, and that was achieved with AI. Six or seven additional CVE candidates surfaced from a sweep that would have taken several more evenings by hand.

The source review was the opposite shape. The target was a well-written C++ project with no recurring micro-pattern so every bug had its own context. The PCM L24 overflow lived in a stack-buffer sizing mismatch. The path traversal was a missing canonicalisation step. The SSRF was a fall-through between two input plugins where one’s allow-list didn’t constrain the other. The CR/LF injection was Expat decoding numeric character references before the parser callback. Different code, different surfaces, different invariants in each. There was no abstract pattern to hand AI for it to grep.

What AI did in that engagement was build the reproduction rig around each bug at a speed I could not match: ASan + UBSan build flavours, libFuzzer harnesses for the parser entry points, attacker HTTP servers, payload generators for Xing/Information frames and XSPF and L24 streams, ASan transcript-capture scripts, cross-version verification (git show v0.24.9:<file> against master) so I could confirm each finding reached the released branch. None of that work found a bug. All of it shortened the gap between “I think I see something” and “I have a reproducer”.

The two modes ended up looking very different in practice.

Mode What AI does What humans do
Pattern scaling Sweep many binaries / files for instances of a defined pattern; triage candidate sinks; produce decompilation passes; collate matches Find the bug-class abstraction in the first place; read decompilation candidates and reject false positives; validate against hardware
Workflow plumbing Build harnesses, payload generators, attacker servers, repro scripts, cross-version diffs, build environment automation Read the source and reason about the actual vulnerability chain; design the experiment; verify against the kernel/allocator/CPU behaviour the source alone doesn’t show

These aren’t disjoint as every engagement has some of both, but the proportions invert depending on the target.

Visually, the two pipelines look like this:

PATTERN SCALING                          WORKFLOW PLUMBING
(uniform pattern across many targets)    (heterogeneous bugs in one target)

       HUMAN                                   HUMAN
   ┌──────────┐                            ┌──────────┐
   │  manual  │                            │  manual  │
   │  reads   │  defines the               │  source  │  finds each bug
   │   2-3    │  abstraction               │  reads   │  semantically
   │  files   │                            │          │
   └────┬─────┘                            └────┬─────┘
        │                                       │
        │  "web_get → do_system,                │  "this loop writes
        │   no sanitization"                    │   past its buffer"
        │                                       │
        ▼                                       ▼
         AI                                      AI
   ┌──────────┐                            ┌──────────┐
   │ pattern  │                            │ harness, │
   │  sweep   │  scales the search         │   PoC    │  scales the rig
   │  across  │                            │ servers, │
   │ binaries │                            │ payloads │
   └────┬─────┘                            └────┬─────┘
        │                                       │
        ▼                                       ▼
       HUMAN                                   HUMAN
   ┌──────────┐                            ┌──────────┐
   │ triage,  │                            │  ASan,   │
   │ validate │  filters false positives   │ runtime  │  verifies + reports
   │   on     │                            │   probe  │
   │ hardware │                            │          │
   └──────────┘                            └──────────┘

Same sandwich shape in both modes, human-AI-human, but with very different center filling. On the left, AI is the search engine. On the right, AI is the infrastructure builder. The human bookends are non-negotiable in either case.

What determines the mode

The clearest distinguishing feature I’ve found is whether the bug-class admits a syntactic abstraction.

The router CGIs had a textbook one: web_get(key, body, flag) → … → do_system(fmt, ...), with or without a strchr blacklist between them. That’s a graph you can grep. You can enumerate every call site of the sink, walk back to the source, ask whether sanitization exists, and the answer is mostly a yes/no readable from the decompilation. AI handles that mechanically once you’ve defined the shape.

The C++ project’s bugs didn’t reduce that way. The PCM overflow needed someone to read two adjacent stack-buffer declarations and notice the integer division didn’t round up to the loop’s actual write count. The XSPF injection needed someone to know that Expat decodes numeric character references before invoking the callback. These are semantic observations, not syntactic patterns. An LLM with the source in context can in theory notice some of them, but in practice, at least the way I worked the engagement, it didn’t, and I don’t think it was going to. The humans found every one of these.

The split I’d carry forward: if the bug-class has a syntactic abstraction the human can name, AI scales the search. If the bug-class requires semantic reading of specific code paths, AI scales the reproduction rig instead. Both are real value. They’re just different kinds.

What AI consistently does not do

Define the bug-class abstraction. On the router engagement, the web_get → do_system pattern came from manual reads of three or four CGIs. By the time we asked AI to scale it, the abstraction was already there. Asking AI to discover the pattern from scratch, without a human pointing at it first, didn’t work in any attempt I made. The abstraction step is where the bug-hunter’s intuition lives, and it’s the thing that doesn’t generalize across LLM context windows.

Read kernel/allocator/CPU semantics from source alone. This is the failure mode that cost me a CVE on the second engagement. I had reported a memory allocation issue with a headline number (“192 MiB committed per song”) that turned out to be virtual address space: the new T[N] calls reserved via mmap(MAP_ANONYMOUS) but didn’t touch the pages. The actual resident set was a fraction of that. The right diagnostic was a runtime measurement (/bin/time resident set, strace of the actual mmap behaviour); the wrong diagnostic was static reasoning over the source. An LLM with the source in context would have made, and as far as I can tell does make, the same mistake. POD new T[N] doesn’t emit a constructor loop; the static reading therefore conflates allocation with commitment. The correction came from someone running the binary, not from someone reading the code harder.

Catch dispatch / surface anomalies. On the router, one binary dispatched on a parameter named firewall= instead of the usual page=, and another parsed the POST body as <action>&<arg> instead of key/value pairs. Both were missed on the first AI sweep: the pattern matcher was looking for the canonical shape and silently passed over the variants. Once I noticed and re-explained the dispatch, the sweep caught them. AI is good at applying patterns. It’s notably less good at noticing it’s looking at a different shape and needing to update the pattern.

Write prose the audience won’t dismiss on register. This isn’t AI’s fault, it’s a fact about how technical communities read. AI-generated prose has — em-dash cadence, “It’s important to note”, “demonstrates that”, bullet lists where prose would do that some readers actively look for and discount on. If your disclosure goes to a reviewer who’s hostile to AI prose, every line they spot pushes the whole report toward the rejection pile, regardless of the underlying technical merit. The rational response, if you know your reader has that bias, is to write the customer-facing surface yourself. AI helps you get to the disclosure but it doesn’t write the disclosure.

Validate against hardware or runtime behaviour. Both engagements required physical-device validation of every claim. AI can suggest a PoC, but the PoC isn’t a finding until the daemon actually crashes, the daemon actually returns the file you didn’t authorise, the kernel actually logs the SIGSEGV with the attacker-controlled epc. The hardware-in-the-loop step is the one that turns “candidate” into “claim”, and there’s no AI shortcut for it that I trust.

What AI consistently does well

Headless reverse-engineering automation. Driving Ghidra in headless mode via Jython to walk every binary, enumerate call sites of named sinks, decompile each calling function, extract format strings and source-function references, emit JSON. I wrote a one-page spec; AI produced the script; I diffed the output against a couple of manual reference binaries to validate. Time saved compared to writing the Jython myself: a couple of evenings per engagement, easily. The script then ran across every binary in the firmware tree and produced a triageable JSON I could read alongside the decompiled C.

Fuzz harness scaffolding. libFuzzer harnesses for parser entry points are ~30-60 lines of extern "C" int LLVMFuzzerTestOneInput(...) wiring around the upstream parser, plus a Meson or CMake build target. AI builds these from a one-paragraph description and gets the parser-safety configuration right (e.g. Expat external-entity rejection on XML harnesses) on first cut. Five harnesses in an afternoon is a thing AI can do. Five harnesses in an afternoon is not a thing I can do by hand.

Synthetic payload generators for binary formats. Crafting a valid MP2 with Xing.frames = 0x7FFFFF, a 16 KiB audio/L24 stream, an XSPF with numeric character references planted in <location>, a CONTENT_LENGTH-driven HTTP body for stack overflow sweep tests. These are small but exact byte-layout problems. Any deviation breaks the PoC. AI handles them well, faster than I would, and the output is diff-able against the spec so verification is cheap.

Cross-version verification. “Is this finding reachable at the latest released tag, or only on master?” that’s a git show <tag>:<file> per affected file, structural diff against the master version, decision per finding on whether to claim both versions. Mechanical, easy to delegate, time-consuming to do by hand. AI per-finding cross-versioning was one of the unsung wins of the second engagement.

Cleanup work after the finding is captured. Writing the internal markdown that turns ASan output + decompilation + PoC steps into a coherent draft. Producing the build/run instructions for someone reproducing the bug from scratch. Not the customer-facing surface (see above) but the internal documentation surface where stylistic AI tells don’t matter because nobody outside the team reads it.

Principles I’d carry forward

The two engagements together gave me a small list of rules I’d hold to on the next one:

  1. Decide upfront whether the engagement is pattern-scaling or workflow-plumbing shape. Look at the target. Is the bug-class going to be a recurring syntactic pattern across many files (firmware with many CGIs, kernel drivers with shared subsystem interfaces, web apps with parallel handlers)? Or is each candidate bug going to need semantic reading of its own specific code path (a large C++ application with heterogeneous parsers, a protocol implementation, a compiler)? AI fits in different places in each. Mixing the modes mid-engagement loses time.

  2. Have humans define the abstraction before asking AI to scale. This is the load-bearing rule. The router engagement worked because two evenings of manual reads produced the web_get → do_system pattern, and only then did AI take over the sweep. Asking AI to define the pattern from scratch, before any manual finds, hasn’t worked for me. The intuition step doesn’t transfer.

  3. Validate every AI claim on hardware or with a runtime probe. Static reasoning from source is where the failure modes hide: allocator semantics, kernel-level memory accounting, CPU pipeline behaviour, signal handling. If a finding’s claim is “X bytes committed” or “Y instructions executed” or “Z control-flow reached”, a runtime measurement is the deliverable, not the source argument. I lost a CVE to this. I don’t intend to lose another.

  4. Keep customer-facing prose human. Disclosure threads, CVE requests, blog posts that go to technical communities, write these yourself. The labor cost of hand-writing them is real, but the alternative (AI-shaped prose dismissed on register) is worse. The internal-to-disclosure rewrite is the most time-consuming human-only step of an AI-assisted engagement, and I haven’t found a way to compress it that doesn’t sacrifice the outcome.

  5. Don’t conflate “AI built the harness” with “AI found the bug”. Most of the value I got from AI on the second engagement was infrastructure. None of it was discovery. The split was clean and worth preserving. Conflating them in retrospectives (“AI helped me find this bug”) cheapens both halves, undercounts the human reads that found the bug, overstates the AI contribution that built the rig. They’re different categories.

  6. When the dispatch is non-standard, re-explain the pattern. AI sweeps will silently skip variants that don’t match the canonical shape. The fix is to read the dispatch yourself once you suspect coverage is incomplete, and update the pattern explicitly. The signal that you need to re-explain is when the sweep terminates with fewer findings than the manual sample rate predicted.

  7. Budget for re-reading your own notes. The F-4 mistake I made on the second engagement was visible in my own PoC notes the entire time: VmRSS sat next to VmPeak on the same line, with the resident number a fraction of the virtual. The headline used the wrong one. AI didn’t catch it because AI was not the one writing the headline. Humans didn’t catch it because humans don’t re-read their own measurements when they think they already know what the measurement says. The discipline is to read your own output with fresh eyes before you publish, on the explicit assumption that you may have transcribed the wrong number from your own data.

What I don’t know yet

A few things I’d want more engagements to disambiguate.

How much of the pattern-scaling value on the router was specific to that codebase, with its uniform CGI architecture and shared libwebutil.so? A messier firmware with diverse subsystems might compress the AI value-add closer to “headless decompilation triage” rather than “sink-source sweep across all binaries”.

How much of the source-review value on the second engagement was specific to C++? In a Rust codebase the bug surface is different e.g.most of the memory-safety bugs are gone, the remaining surface is unsafe blocks, FFI, and logic errors. I don’t have intuition for how the labor split sits there.

How AI scales on multi-stage exploit development like taking a memory corruption primitive through info-leak chain into reliable RCE. The second engagement had a finding I never demonstrated end-to-end reliably (libc ASLR + per-process mmap base variation made the chain “demonstrated once, not reproducible”). I don’t know whether AI would have shortened the gap or just hallucinated through it.

How AI handles dispatch surfaces that aren’t HTTP such as IPC, syscall tables, JSON-RPC schemas, binary protocols with their own state machines. The two engagements I have data on were both HTTP-CGI-shaped. Anything more exotic, I’d need to see.

Closing

The honest summary is that AI sits in security research as a companion with very strong technical abilities and very little creative instinct for what to chase next. Excellent at the parts you can specify: harnesses, payload generators, decompilation triage, cross-version diffs, anything you can name precisely. Blind to the part upstream of specification, where you decide what’s interesting and what abstraction is worth defining in the first place. And it will produce confident-sounding answers on anything that requires runtime semantics rather than source-level pattern matching, so the verification responsibility stays on your side of the table.

That’s a useful collaborator if you read the split accurately. AI brings the technical fluency to build whatever you can name precisely. You bring the creative leap that decides what’s worth naming. Hand off the mechanical scale-out and the rig-building. Keep the question of which thread to pull for yourself. And re-read everything before it leaves your hands.

The next engagement will probably have a different shape from both of these, and the split will probably look different again. But the rules above are the ones I’d start from, until something concrete contradicts them.