What Building an AI Vulnerability Hunter Taught Me
I’ve worked in WordPress security for 18 years. We Watch Your Website has been cleaning and protecting sites since 2007, and today we watch close to 4 million of them. I’ve seen just about every way a plugin can go wrong.
So when I started building an AI-assisted analyzer to find vulnerabilities in WordPress plugins, I assumed I’d be the one doing the teaching. I’d tell the machine what bad code looks like, and it would go find it.
That’s not how it went. The project has a static analysis layer that maps every function, input, and dangerous call in a plugin, followed by an AI verification stage that tries to confirm or kill each finding. Every week of building it forced me to rethink something I thought I already knew. First, the question I get asked most. Then the lessons that stuck.
Why not just hand the whole plugin to the smartest model?
It’s the obvious shortcut. Frontier models keep getting better at security work. Even a Mythos-class model, the tier capable enough that access is restricted to vetted organizations, could be pointed at a plugin’s source with one instruction: “find all vulnerabilities.”
We deliberately don’t work that way. The problem isn’t intelligence. It’s the question. “Find all vulnerabilities” asks a model to search, judge, and prove all at once, across tens of thousands of lines, and it leaves no record of what was actually checked.
So we split the work in two.
Deterministic code handles everything that has a right answer. It reads the structure of the plugin and maps every way in: AJAX actions, REST routes, shortcodes, block callbacks, admin pages, core WordPress hooks. It records which user roles can reach each one. It finds every place user input enters, every dangerous function call, and every security check along the path between them. None of that is a judgment call, and none of it is left to a model. The rules behind it came from tens of thousands of real, patched CVEs, turned into code.
The model only gets the part that needs judgment. Instead of “find everything,” it gets a narrow, specific case: here is an entry point an unauthenticated visitor can reach, here is the input they control, here is the path to this unserialize() call, and here are the checks along the way. Is this exploitable? Then a second model is told to prove that it isn’t.
The difference shows up everywhere that matters for a security product:
| “Find all vulnerabilities” | Deterministic first, model second | |
|---|---|---|
| Same answer twice | Not guaranteed; reruns can differ | Static mapping is identical every run; the model judges fixed evidence |
| Knowing what was checked | No way to tell which code the model really examined | Every entry point and dangerous call is enumerated and logged |
| Measuring misses | A miss looks the same as a skip | Every change is scored against 30 known, published CVEs |
| What a finding contains | A narrative | A chain: entry point, access level, input, dangerous call, guards |
| False positives | An open prompt invites confident overclaiming | A skeptic attacks one specific claim against specific evidence |
| Cost per plugin | The whole codebase through a frontier model, every scan | The model sees only the top 10 candidates per plugin |
| Scale | Hard to justify across every plugin on ~4 million sites | Results are keyed by plugin and version and reused everywhere |
This isn’t a knock on the models. A model like Mythos is far better at the judgment step than anything I could write by hand. That’s exactly why we save it for judgment. Every bit of work the deterministic layer takes off the model’s plate makes the model’s job smaller, and a smaller job is one you can check.
It also makes the system debuggable. When the verifier gets something wrong, we can see the exact evidence it was shown and find the gap. Lesson 2 below only happened because of that. With a single “find everything” prompt, a wrong answer is just a wrong answer.
1. Regex is for lists. Logic is for relationships.
I know regex well. I’ve used it for years to hunt malware, and I know exactly where its limits are. Even so, this project showed me a sharper line than I’d drawn before.
Regex and simple pattern matching are excellent when the vocabulary is finite. There’s a known set of dangerous PHP functions. There’s a known set of WordPress hooks. There’s a known set of guard functions like current_user_can() and check_ajax_referer(). Matching those is fast, cheap, and reliable precisely because the list has an end.
The hard question in vulnerability hunting isn’t “is there a dangerous function here?” It’s “does attacker-controlled input actually reach that dangerous function?” That’s a relationship between two things, often separated by wrapper functions, reassigned variables, and multiple candidate inputs.
We hit this wall twice. The analyzer paired the wrong input with the wrong output, we fixed it with a better pattern, and a few weeks later the same bug came back in a different shape. That was the tell. When you’re writing a pattern for every possible way code can connect two things, you’re never done.
The fix was to stop asking patterns to understand code. We now walk the actual structure of the code (its syntax tree) and follow assignments and calls with real logic. Patterns still handle the bounded lists. Logic handles the connections.
The takeaway: if the set of things you’re matching has an end, use a pattern. If you’re trying to describe how things relate, a pattern will eventually lie to you.
2. When the AI misses a real bug, check what you fed it
The AI side of the analyzer doesn’t just look for bugs. It argues with itself. One role builds the case that a vulnerability is real. A second role, the skeptic, tries to tear that case apart. A third role reads both and makes the call.
I built it that way on purpose. An AI that’s only asked to find bugs will find bugs, whether they’re there or not. Doubt has to be part of the process, not an afterthought.
That design produced a frustrating result early on. We’d feed the pipeline a plugin with a known, confirmed vulnerability, and the verifier would reject it. My first instinct was the same one most people have: the model isn’t smart enough. Try a bigger one.
So we tested two different vendors’ models side by side. Both rejected the same real bugs. When we traced exactly why, the answer wasn’t the model at all. The evidence we handed it was wrong or incomplete. One chain claimed the wrong entry point into the plugin. Another left out the one piece of context that proved an attacker could get there. The skeptic did its job. It refused to sign off on a case that, as presented, didn’t hold up.
Once we fixed what went into the model, the verdicts came out right. Over 17 test cases, the model we chose correctly handled 11, with fewer false alarms than the alternative, at a little over half the cost. But the bigger win was learning where the real problem lived.
The takeaway: a model can only reason about the evidence you give it. Before you upgrade the model, audit the input.
3. A security tool that changes its answer is a broken tool
This one was a quiet bug with loud consequences.
Deep in the analysis layer, the code walked through a collection of items whose order wasn’t guaranteed. In Python, that order can shift from one run to the next. Most of the time it didn’t matter. Sometimes it did. The same plugin, analyzed twice, could come back with different results.
For a lot of software, that’s an annoyance. For this project, it’s a deal breaker. The plan is that every vulnerability the analyzer finds becomes a protection rule, stored by plugin and exact version. Every other site we watch that runs that same plugin version gets the same rule. A finding on one site protects thousands.
That only works if the same input produces the same output, every single time. If the analyzer finds a bug on Monday and misses it on Tuesday, you don’t have a protection system. You have a coin flip with extra steps.
We fixed it by forcing a stable order everywhere the analysis walks through data. It added nothing visible to the results. It made every result trustworthy.
The takeaway: in security tooling, determinism isn’t a nice-to-have. If you can’t reproduce a finding, you can’t defend it, and you can’t build on it.
A few more, in brief
A nonce is not authorization. We mined thousands of real access-control vulnerabilities from public disclosures. In roughly 55% of them, the developer had added a nonce or a login check but never checked whether the user was allowed to do the action. A nonce proves the request came from your site. It doesn’t prove the person sending it should be able to delete posts or change settings. Only a capability check does that.
Reachable is not exploitable. An attacker being able to trigger a function is one finding. An attacker controlling the data that function acts on is a different, much more serious one. When we scored on reachability alone, one plugin produced 18 candidates tied for first place, with the real bug sitting right next to routine database housekeeping. Ranking by where the data actually comes from separated them.
“Safe” depends on where the value lands. Input that was sanitized isn’t safe if something decodes it later. An escaping function that’s right for the body of a page can be wrong inside an HTML attribute. The analyzer had to learn that the correct defense depends on context, the same way a careful developer does.
Measure, don’t guess. We built a scorecard against 30 real, published vulnerabilities and ran every change against it. It replaced “I think this helps” with a number. It also produced a humbling moment: a correct fix made the score drop, because it removed a hit that had only been right by luck. Without measurement, we’d have rolled back the better code.
Don’t use a discovery engine to find known bugs. If a vulnerability already has a CVE, the fastest path is a direct rule, not a fresh analysis. The analyzer’s real job is finding what nobody has reported yet. That includes custom plugins and AI-generated, vibe-coded plugins that will never get a CVE because no researcher will ever look at them.
Fix for the ecosystem, not the test case. Every benchmark plugin is a test case. A fix that only helps one plugin pass isn’t a fix. Several of our best improvements came from recognizing a pattern shared across thousands of plugins, like the boilerplate scaffolding many developers start from.
Pick your tools with your own test, and set the rule before you run it. Public leaderboards measure code generation and bug fixing. None of them measure careful, adversarial verification. So we wrote down the decision rule first, then ran both options on our own cases. Deciding the rule upfront keeps you honest about the result.
Where this is headed
The goal hasn’t changed since day one. Find vulnerabilities before attackers do, and protect sites before a patch exists. Not just in the popular plugins with millions of installs, but in the obscure, custom, and AI-generated ones that nobody else is looking at.
The time between a vulnerability being disclosed and being exploited keeps shrinking. AI is making attackers faster. The only answer is for defenders to get faster too, and to do it with tools they can actually trust.
If there’s one thread running through every lesson above, it’s this: AI is a powerful partner, but it doesn’t replace rigor. It amplifies it. Give it bad evidence and it will confidently reach the wrong conclusion. Give it good evidence, make it argue with itself, and measure everything, and it becomes something genuinely useful.
I expected to teach the machine. It turns out we taught each other.
