What is under the hood of A42's AI Pentest, and why it outperforms manual penetration testing.
This summer, we’ve launched AI Pentest. Built on a multi-agent architecture and the playbooks of elite white-hat hackers from the Gov & Defence Tech space, it delivers broader coverage and results up to 10× faster than a traditional manual pentesting.
In this article, we'll walk through how AI Pentest works and why it's more effective than the conventional approach.
The Multi-Agent Architecture of AI Pentest
At the core of our approach are two principles: distributed work and deep specialization. A dedicated orchestrator agent is responsible for the pentest workflow: it launches scripts and directs a "team" of more than 130 specialized agents. The entire process is supervised by a human: an experienced penetration tester who remains accountable for the outcome.
1. The orchestrator agent: the brain of the system
Its job is strategic planning, understanding the system's architectural context, and coordinating the specialized agents. The orchestrator analyzes the application like an architect: it builds a mental map of the system, tracks the relationships between pages, databases, and microservices, and then assigns precise, targeted tasks to the narrowly focused "infantry."
2. The specialized agents
Each of the 130+ agents is a narrow-domain expert, and they all operate in parallel, exchanging information through the orchestrator in real time.
For example, the injection agent knows more than 1,000 techniques for bypassing filtering and executing SQL, NoSQL, LDAP, XML, and OS injections, while the business-logic agent verifies that the application's logic is implemented correctly, checks request routing and data integrity, and uncovers logic flaws — such as price manipulation or ID tampering.
Other agents take on more complex vectors. The JWT agent, for instance, dissects authentication mechanisms (JWT, OAuth, Basic Auth) and attempts privilege escalation, while the infrastructure-attack agent focuses on the hardest challenges — from CRLF Injection and HTTP Request Smuggling desync attacks to identifying points for RCE, XXE, RFI, and LFI.
This approach delivers two advantages at once. Narrow specialization ensures a depth of testing that a single general-purpose agent can't match, because models are designed to conserve tokens, while distributing tasks guarantees that each check is executed more thoroughly. Parallel execution delivers speed, since dozens of attack vectors are tested simultaneously rather than one after another. This is exactly how AI Pentest combines broader coverage with dramatically faster results, making it a superior alternative to manual testing.
Human in the loop
Given the current state of model development, we believe it's prudent to keep a human in the loop. Our experienced penetration tester provides an added layer of oversight, ensuring the agents don't stray from the methodology, and validates the results.
The AI Pentest Methodology, Step by Step
A42's AI Pentest fully reproduces the best practices of manual testing and Red Teaming, but scales them through the power of AI systems.
Stage 1. Passive reconnaissance
The first step is finding all of the client's subdomains. We use standard tools to gather data about the target only from open sources, without touching the infrastructure directly. We start this as early as the proposal stage. First, it lets us see the real scale of the infrastructure — clients often don't know how many subdomains they actually have. Second, it helps us spot the subdomains that are most business-critical and therefore need deep penetration testing.
Once the subdomains are mapped, we look at domain parameters — WHOIS, DNS records, IP ranges, and SPF/DKIM/DMARC email policies. We check metadata and HTML comments, work out the technology stack (Wappalyzer, BuiltWith, WhatWeb), and capture the "fingerprints" of web applications and frameworks.
Why. This is needed to build a full map of the external attack surface — it means every asset reachable from outside — because in the real world around 80% of cyberattacks start by exploiting perimeter vulnerabilities. This happens because organizations regularly expose more than they think: outdated subdomains, test environments, framework versions in response headers, and config files left open to the public. Passive reconnaissance shows the full picture of the attack surface exactly as an attacker sees it, and helps decide what needs a closer look. Without this map, everything that follows would be patchy.
Stage 2. Active reconnaissance
This is where direct interaction with the infrastructure begins. We run port scanning (Nmap, Masscan, Nessus) to find open ports and externally reachable services. Banner grabbing collects the technical details of those services: software names and versions, protocols, and configuration details.
Why. To find hidden entry points and Shadow IT: services and hosts that run inside the infrastructure but aren't under the IT team's control. These assets are often left out of patching and monitoring which makes them weak links. The software versions collected during banner grabbing are then checked against known vulnerabilities, which leads into the next stage.
Stage 3. Automated scanning
Using the collected data, we launch automated vulnerability scanning. The AI agents work like industry tools such as OWASP ZAP, Burp Suite, Nessus, Nikto, Acunetix, Nuclei, and Naabu, finding and rating vulnerabilities across web applications, APIs and network components.
Why. To quickly find known vulnerabilities. In the real world these are what cause most incidents, not rare one-off attacks. Just as important, this stage prioritizes findings and filters out false positives, so that further testing focuses on what really matters and the client gets a report with zero false positives.
Stage 4. Deep testing by AI agents
As noted above, deep testing is applied to the client's most important subdomains, for example those tied to payments or core business processes. At the proposal stage, we use technical tools to identify exactly which assets need deep scanning, and then confirm that scope with the client. The client can also choose different subdomains or add and remove specific parts of the deep scan.
At this stage, the AI agents look for vulnerabilities and check whether they can actually be exploited, repeating the actions of an experienced attacker against a detailed checklist that covers the OWASP Top 10 and goes beyond it.
Specifically, this stage tests:
- Identity management and authentication: whether roles and registration work correctly, password policies, lockout mechanisms, and authentication bypass attempts, plus testing of JWT, OAuth, and Basic Auth.
- Authorization: attempts at privilege escalation and getting around access-control logic.
- Session management: cookie protection, CSRF resilience, and login bypass potential.
- Input validation: looking for injections (SQL, LDAP, XML, OS, NoSQL) and XSS.
- Business logic: correct request routing, integrity checks, and handling of file uploads and data flows.
- More complex vectors: CRLF injections, HTTP Request Smuggling, file-upload attacks, RCE, RFI, LFI, XXE, password attacks, testing internal and external APIs for authentication resilience, injection protection, and handling of non-standard HTTP methods.
Why. To test the less obvious threats and to separate real, confirmed risks from ones that are only theoretical. Every hypothesis is tested in practice under controlled conditions. A vulnerability isn't just logged. It's exploited on terms agreed with the client, to prove its real impact while doing no harm to their infrastructure.
The Report
The pentest results are delivered as a structured report, with a detailed description of every vulnerability found, supporting evidence, and specific recommended fixes for each one. The report is prepared in line with current international standards and regulations: ISO 27001, SOC 2, DORA, and PCI DSS.
Why. The technical team gets exact remediation steps, each tied to a specific finding. Management gets the findings ranked by severity, which shows which vulnerabilities to fix right away and which can wait for a planned release. The same report can also serve as evidence during security audits.
Full Re-test (optional)
The re-test is optional. It covers the full testing scope and costs 50% of the pentest price.
Why. We think this approach serves our clients best. Fixing one vulnerability can sometimes create a new one by accident, and a partial re-test misses these cases, so it can leave risks in place. Some audit standards also require companies to fix the most critical vulnerabilities within very tight deadlines. Our re-test gives a company a simple way to prove to an auditor that it has met this requirement at half the price of a full pentest.
Final thoughts
Modern cybersecurity isn't about a once-a-year, check-the-box review. It's about continuous visibility into your external perimeter and ongoing control of your infrastructure. A42's multi-agent AI Pentest architecture turns security from a bureaucratic bottleneck into a transparent, efficient process, giving developers clear, validated instructions for closing real risks.



