AI Vulnerability Discovery at Machine Speed: A CTEM Prioritization Playbook
AI agents now find, reproduce, and patch real code vulnerabilities at machine speed: AWS's CyberGym-E2E results show 89% end-to-end task success. The bottleneck shifts from detection to prioritization, validation, and verified remediation. Here's how a continuous CTEM loop keeps AI-scale finding volume from turning into SOC noise.

AI Vulnerability Discovery at Machine Speed: A CTEM Prioritization Playbook
TL;DR
- AWS's autonomous code security system, evaluated on the CyberGym-E2E benchmark, completed the full vulnerability lifecycle — find a flaw in a real codebase, reproduce it with a working proof of concept, patch it, and keep the project's own tests passing — on 819 of 920 tasks. That is an 89.0% end-to-end success rate, 23.1 percentage points above the previous public high of 65.9%.
- The headline for defenders isn't "machines can write patches." It's that vulnerability discovery volume is rising on both sides of the fight while triage, validation, and remediation-verification capacity stays roughly flat.
- When inflow grows, the bottleneck moves from detection to prioritization: which findings are real, reachable from the internet, and exploitable today — and how do you prove the fix actually landed?
- That is the Continuous Threat Exposure Management (CTEM) loop: continuous discovery, validation, exploitability-aware prioritization, remediation, and verification. Trusteed CTEM runs that loop across your external and internal attack surface so a rising tide of findings does not turn into analyst queue noise.
What is AI-driven vulnerability discovery?
AI-driven vulnerability discovery is the use of model-backed, agentic systems to read a codebase, form hypotheses about where trust boundaries break, produce candidate vulnerabilities, demonstrate them with an executable proof of concept, and generate a repair that can be validated against the project's functional tests. It is not a rebranded SAST rule engine.
It helps to separate the layers of the stack:
- Pattern-based static analysis (SAST) matches known-bad constructs. Fast, high recall on familiar bug classes, weak on business logic and authorization gaps.
- Fuzzing and dynamic analysis finds crashes by brute-force input generation. Excellent evidence, poor at reasoning about why a path is reachable or whether it matters.
- Dependency and container scanning (SCA) answers "is this known-vulnerable component present?" It does not answer "is it reachable, deployed, and exposed?"
- Agentic discovery adds reasoning: tracing data flow across files, understanding framework conventions, and constructing an exploit path through several individually minor weaknesses.
The benchmark design matters more than the marketing. CyberGym-E2E places an agent in a container with a vulnerable revision of a real open-source project and no vulnerability description, no proof of concept, no crash log, and no original patch. External network access is blocked, so retrieval of public historical fixes is off the table. Within a 90-minute budget, the agent submits a crashing input and a source patch. Scoring is cumulative across four stages: does the input crash the program (S1), does the patch stop the crash (S2), do the project's functionality tests still pass (S3), and does the patch address the benchmark's specific target vulnerability (S4).
The AWS-reported numbers are S1 92.5%, S2 89.6%, S3 89.0%, and S4 37.8%. The S3 figure is the headline because it is end-to-end success: a demonstrated vulnerability plus a repair that doesn't break expected behavior. The S4 gap is diagnostic rather than damning — a repository can contain several valid vulnerabilities, and an agent may legitimately fix a real flaw that differs from the benchmark's chosen target. When tasks were allowed to run past the 90-minute limit, the end-to-end pass rate rose to 93.7%.
Two implications follow. First, discovery at this quality level means more findings per unit of code, not fewer. Second, the output is not a ticket — it is an evidence chain: candidate flaw, executable demonstration, patch, test result. That chain is exactly what exposure management needs to be useful rather than noisy.
CTEM is the operating loop that consumes that chain: continuously discover the attack surface, validate which exposures are real, prioritize by exploitability and business context, drive remediation, and verify the fix. CTEM is not a scanner category. It is a program discipline that platforms implement.
Why it matters now
Security teams were already operating above their designed intake. The standard lifecycle for a single finding is investigation, reproduction, repair, and a test that confirms the fix closes the issue without breaking expected behavior. Multiply that by an order of magnitude and the process stops being a queue and becomes a structural failure.
Four pressures are converging:
- Discovery supply is expanding. Better models mean more candidate flaws per repository, including the logic and authorization bugs that pattern-based tools historically missed. The CyberGym-E2E result is a public, reproducible statement that machine reasoning can now traverse the full lifecycle.
- Attack paths are getting more sophisticated. The same reasoning that finds a bug can chain several low-severity weaknesses into one exploitable flow. Defenders who rank by individual CVSS scores will systematically under-rank chained exposure.
- Patch regression risk pushes fixes downstream. A patch that breaks production behavior gets deferred, and deferred patches are how exposure windows stretch from days to quarters. This is why verification — not just remediation — is a CTEM requirement.
- Regulatory and audit expectations are tightening. Software supply chain guidance, KEV-driven remediation deadlines, and SBOM-based attestation all assume you can show evidence of validation and verified fixes, not just a scan date.
The business translation is straightforward: rising finding volume multiplies analyst hours, patch windows, change-management risk, and audit scope. If triage cannot filter aggressively and defensibly, the cost curve bends the wrong way.
How attacks and risks work

The mechanics of an AI-accelerated attack path follow a predictable sequence, and each step is something you can observe.
Step 1 — Fingerprint the deployed target. Attackers identify framework, version, and endpoint shape from the outside: headers, error pages, API documentation, static bundle contents, TLS posture. This is cheap and largely passive.
Step 2 — Acquire code context. Open-source components, public repositories, forgotten forks, vendor advisories, and patch diffs all provide the source-level context that reasoning models need. Where source is unavailable, behavior inference and version correlation do a surprising amount of the work.
Step 3 — Reason across the boundary. The most valuable outputs are chains, not single bugs: a permissive upload handler plus predictable file naming plus a read endpoint missing an authorization check equals an unauthenticated file disclosure. Individually, each weakness may score low. Chained, they are a complete exploit.
Step 4 — Demonstrate, then weaponize. A proof of concept that runs in a test container is evidence of exploitability. The gap between that container and your production route is a deployment question: is the vulnerable code path reachable from an internet-facing service, and does it require authentication?
Step 5 — Compress the timeline. Public patches and advisories are also attacker inputs. When a fix ships, the diff narrows the search space. That is why time-to-exploit keeps shrinking and why KEV additions often land within days of disclosure.
The risk model that results has three layers: a code path (the flaw), a network path (can an unauthenticated or low-privilege caller reach it?), and an asset (business criticality, data sensitivity, blast radius). Most tooling reports only the first layer. Exposure management requires all three, joined.
Detection and visibility
Good telemetry for this problem is not one feed. It is a set of feeds that can be correlated into a single exposure record.
- Code layer: SAST and SCA output, reachability analysis, dependency graphs, per-build SBOMs, and patch coverage. Reachability is the difference between a finding and a to-do.
- Build and deployment layer: which artifact digest or version is running on which asset, plus IaC and container configuration changes that can silently widen exposure.
- External surface layer: domains, IPs, services, technologies, API specifications, and TLS posture discovered passively and actively, with version fingerprints matched to CVE intelligence.
- Exploit-intel layer: EPSS scores, CISA KEV membership, vendor advisories, and exploit references. These convert severity into likelihood of exploitation.
- Runtime and edge layer: WAF and edge logs, authentication anomalies, and unusual API call sequences that can indicate authorization abuse such as BOLA.
- Ownership and SLA layer: who owns the asset, what the remediation deadline is, and whether the fix was re-tested after deployment.
The correlation itself is the hard part. A finding becomes actionable only when you can join it to a live asset, determine internet reachability, attach exploitability context, and record verification that the fix deployed. Without that join, an AI-scale influx of validated findings simply becomes a better-documented flood — and a validation gate that decides whether a hit should actually alarm (the should_alarm decision) is what keeps analyst queues focused on real, reachable exposure rather than every scanner artifact.
Reduce risk: best practices for AI-scale vulnerability flow
- Design triage for tomorrow's inflow, not today's queue. Assume discovery volume grows several-fold and model the analyst hours before it arrives. If your process only works at current volume, it is already broken.
- Prioritize on exploitability and reachability, not severity alone. Combine CVE data, EPSS, and KEV membership with an explicit answer to "is this code path reachable from an internet-facing service?"
- Require reproduction evidence before escalation. A candidate finding with a working demonstration earns an analyst's attention. A pattern match without reachability evidence does not.
- Gate alarms, don't just rank tickets. A validation step that suppresses non-actionable hits upstream of the SOC is worth more than any dashboard sort order.
- Verify fixes after deployment, not after merge. Re-scan the live asset and confirm the specific exposure closed. Merged is not remediated.
- Treat AI-generated patches as untrusted change. Sandbox, run the project's functional tests, require human review for critical paths, and roll out in stages. Regression risk is the main reason patches get deferred, and deferral is exposure.
- Map code findings to deployed assets. Keep a join between repository, build artifact, and internet-facing service so a code-level discovery immediately identifies which external assets are affected.
- Measure the loop, not the backlog. Track time-to-validate, validated-to-raw finding ratio, exposure reduction by asset class, and time-to-verified-remediation. A shrinking backlog with growing raw inflow is a success signal.
- Keep scope disciplined. Not every internal finding is exposure. Continuous scan plans with defined cadence and scope prevent the surface from expanding faster than the team can reason about it.
How Trusteed CTEM helps

- Attack surface and asset inventory. Trusteed discovers domains, IPs, services, and technologies through passive and active discovery, then keeps them under ongoing scan plans — so you know which assets exist before you argue about severity.
- CVE-enriched findings. Scanner-driven detection is correlated with catalog CVE data, EPSS and KEV context, and exploit references where available, which turns raw hits into risk-scored records.
- Finding validation and SOC gate. Not every scanner hit becomes an alarm. Trusteed validates exploitability and business context so dashboards and analyst queues focus on actionable risk, cutting noise relative to raw scanner-only tooling.
- API surface testing and deep DAST. A dedicated worker covers API exposure, and a deeper web application worker tests critical apps — the surfaces where authorization-chain findings tend to surface externally.
- Compliance and reporting views. Framework-oriented views and customer reporting give you a defensible record of exposure posture and remediation progress.
- Vulnerability intelligence. KEV and emergent-threat catalog narratives in the public blog and in-product intel help teams prioritize without rebuilding their own research function.
Trusteed CTEM vs point tools
| Capability | Typical point tool (SAST/SCA/scanner) | Trusteed CTEM |
|---|---|---|
| Primary scope | Source code or CI/container checks | External and internal attack surface, deployed assets, services, APIs |
| Output shape | Raw signal or finding | Validated, prioritized exposure with business context |
| Exploitability context | Severity or template metadata | Catalog CVE data with EPSS/KEV context and exploit references |
| Noise handling | Every hit becomes a ticket | Validation gate and should_alarm decision focus the analyst queue |
| Application and API depth | Limited to configured checks | Dedicated API surface worker plus deep DAST for critical apps |
| Continuity | Point-in-time scan or pipeline run | Ongoing scan plans with continuous exposure tracking |
| Reporting | Developer-centric output | Framework-oriented compliance views and customer reporting |
Point tools remain valuable — they generate signals CTEM consumes. The difference is that a signal is not a program. Inventory, validation, exploitability context, SOC prioritization, and continuous verification are the parts that must exist around them.
FAQ
Does AI code security make vulnerability scanners obsolete? No. It changes the ratio. AI reasoning improves discovery quality and can produce reproductions and patches, but you still need continuous visibility into which assets are deployed, which are internet-facing, and whether a fix actually closed the exposure. Discovery without exposure context is just a bigger list.
What is CyberGym-E2E, and what exactly did AWS's Continuum achieve? CyberGym-E2E is a benchmark that evaluates the full vulnerability lifecycle in real open-source projects: find a flaw, produce a crashing input, patch it, and preserve functionality tests. AWS reported that Continuum passed 819 of 920 tasks within a 90-minute limit for an 89.0% end-to-end success rate, 23.1 percentage points above the previous public high of 65.9%, with a 93.7% rate when tasks were allowed to run longer. Benchmark stages, network isolation, and post-run trajectory review were part of the evaluation design.
What is the difference between CTEM and scanners? A scanner produces findings at a point in time against a defined target. CTEM is a continuous loop — discover the attack surface, validate what is real, prioritize by exploitability and business impact, drive remediation, and verify the fix. Scanners are one input to CTEM; they are not the program. Teams that treat scan output as the end state end up with growing backlogs and flat risk reduction.
Should we let AI patch production code automatically? Not without guardrails. The benchmark's S3 stage exists precisely because a patch that fixes a flaw while breaking expected behavior is not a successful outcome. Sandbox the change, run functional tests, require review for critical paths, and stage the rollout. Treat AI-generated repairs as untrusted change.
How do we avoid drowning in AI-found findings? Three filters, applied in order: reachability (is the code path actually invoked?), exposure (is it reachable from an internet-facing service?), and exploitability (is there a demonstrated or known exploit path, plus EPSS/KEV context?). Only findings that survive all three should reach an analyst.
How does code-level discovery connect to external attack surface management? Through the deployment join. A finding in a repository is only exposure if the build containing it is running on an asset someone can reach. Maintaining a mapping from repository to artifact to live service means a new code-level discovery immediately identifies affected external assets instead of triggering a manual search.
What metrics show an exposure program is working? Time-to-validate, validated-to-raw finding ratio, percentage of external assets under continuous scan coverage, time-to-verified-remediation, and exposure reduction by asset class. Rising raw inflow with a shrinking actionable queue is the pattern you want.
Does Trusteed CTEM replace our existing scanners? No. Trusteed CTEM complements them. It consumes scanner-driven detection and adds asset inventory, CVE and exploitability enrichment, validation, SOC-relevant prioritization, API and deep DAST testing, and compliance reporting on top.
Related resources
- Trusteed CTEM platform overview — how discovery, validation, and prioritization fit together.
- Trusteed vulnerability intelligence blog — KEV and emergent-threat narratives with CTEM response playbooks.
- Trusteed tenant application — scan plans, findings, and exposure dashboards.
- CISA Known Exploited Vulnerabilities Catalog — the exploitation signal that should move findings to the top of the queue.
- NIST SP 800-218, Secure Software Development Framework — practices for the discovery and remediation side of the loop.
- OWASP API Security Top 10 — the authorization and object-level failure classes that AI reasoning finds fastest.
- FIRST EPSS — probabilistic exploitation scoring to pair with KEV membership.
- AWS Security Blog: autonomous code security results — the primary source signal for this analysis.