Prancer Blog / SwarmHack Deep Dive
Deterministic + LLM Pentesting: Why the Evidence Gate Matters
How deterministic execution, optional frontier-model exploration, model routing, and a mandatory evidence gate work together.
SwarmHack Team · 2026-04-29 · 10 min
The previous parts (1, 2) showed an autonomous engagement: 11 findings, 35 crown jewels, 6 minutes, zero humans, and bit-for-bit reproducible execution across 11 runs. The next question is not whether to choose deterministic systems or frontier models. It is how to combine them without turning generated output into security “proof.”
- Where deterministic execution is non-negotiable
- Where frontier models add meaningful exploratory reach
- Why every model-generated hypothesis must pass an evidence gate
- How private, open-weight, and sovereign deployments preserve control
Two Natures of Attack, One Evidence Gate
Prancer combines two complementary operating layers:
1. Deterministic core — versioned playbooks, GOAP A* planning, real packet execution, reproducible findings, and air-gap-ready operation without a model. 2. Model-assisted exploration — optional frontier-model reasoning for reconnaissance, attack hypotheses, prioritization, and paths beyond fixed playbooks.
The model changes what gets *considered* or *tried*. It does not change what counts as *proven*. A finding reaches the report only after deterministic execution captures target evidence under the same authorization and safety controls.
This distinction matters. LLM-only pentesting inherits unpredictable cost, non-deterministic coverage, hallucination risk, and data-governance concerns. A hybrid architecture can use model creativity without asking the model to be the system of record.
1. Use Models Where Exploration Benefits
Frontier models are valuable when the search space is ambiguous: understanding unfamiliar application behavior, proposing attack-chain hypotheses, correlating a disclosure flood, or choosing among several plausible next moves.
Prancer can route eligible workloads to the most appropriate approved model. Depending on the engagement, that may be an Anthropic or OpenAI frontier model, an open-weight model hosted privately, or a model inside a sovereign-AI environment. Customers can also disable the model layer entirely.
The deterministic core remains responsible for execution. A proposed SQL injection path becomes a finding only when a specialist agent sends the payload, reads the target response, and captures replayable evidence.
2. Token Economics Still Matter
A model-led pentest against a moderately complex target can consume substantial tokens, and exploratory cost varies with target complexity, response size, and the number of attempted branches. That makes an LLM-only loop difficult to budget for continuous testing.
SwarmHack avoids putting every payload and response through a model. The compiled execution core handles repeatable work locally, while model routing reserves inference for workloads where adaptive reasoning adds value. Deterministic-only operation carries no model metering at all.
What Anthropic Mythos Demonstrates
Anthropic's Mythos research demonstrates the value of deep model-led exploration for difficult vulnerability discovery. It also illustrates why exploratory search and continuous validation have different cost and repeatability profiles.
Prancer's architecture connects those strengths: use frontier reasoning to expand the search when appropriate, then move candidate paths through deterministic execution and evidence validation. The customer chooses the model and deployment boundary; the evidence standard does not change.
3. Reproducibility Belongs in the Execution Layer
Generative inference is probabilistic. Even tightly controlled model settings can produce different hypotheses or paths across runs. That variation can be useful during exploration, but it cannot be the basis for confirming that a vulnerability exists or that a patch worked.
SwarmHack therefore reproduces findings in the deterministic core. Its CMDI agent sends a known arithmetic marker and checks for the expected result. The same principle applies across exploit classes: replay the actual request, inspect the actual response, and grade the finding from captured evidence.
A model may discover a candidate path on Tuesday. SwarmHack must reproduce that path deterministically before it is reported, and it re-runs the same exploit after remediation to prove closure.
4. Hallucinations Become Hypotheses, Not Findings
Language models can generate plausible but incorrect statements. In an ungated system, that can become fabricated vulnerabilities, drifting severity, skipped coverage, or prose that sounds like evidence.
Prancer treats model output as an input to testing—not as the verdict. Confidence comes from observable execution:
| Method | Confidence | Evidence basis |
| -------- | -----------: | ---------------- |
| Marker-based command injection | 0.99 | Arithmetic value reflected by the target |
| Reflected-payload XSS | 0.90 | Payload appears unencoded in the response |
| Tautology SQL injection | 0.75 | Measured response-change observation |
| In-band XXE | 0.60 | Target file content detected |
A reviewer can replay the payload and inspect the response. If captured target output is missing, an Exploited label is automatically downgraded. Critical severity remains reserved for proven evidence.
5. Deployment Choice Solves Different Governance Needs
There is no single acceptable model boundary for every customer. Prancer supports multiple operating patterns:
- Deterministic-only for air-gapped, classified, OT, or tightly regulated environments.
- Private open-weight models for organizations that want adaptive exploration without sending workloads to a third-party inference service.
- Sovereign AI where residency and jurisdiction requirements govern model hosting.
- Approved frontier models where the customer permits them and their capabilities best fit the workload.
Model, version, and operating parameters can be recorded with the engagement. Exploitation still runs inside the customer's environment, and the evidence gate applies regardless of model choice.
6. Stateful Execution Is Still the Backbone
Real exploitation spans hundreds of requests and multiple agents: test payload variants, fingerprint the target, extract credentials, reuse them, establish a bounded pivot, and validate downstream impact.
Models can help formulate or prioritize that path. The swarm carries it out through shared state, specialist tooling, scope enforcement, and deterministic exploit logic. A credential discovered by one agent becomes authorized input for the next; generated prose never substitutes for the network interaction.
Side by Side
| Dimension | Deterministic-only SwarmHack | Ungated LLM-only pentesting | Prancer hybrid mode |
| ----------- | ------------------------------ | ----------------------------- | --------------------- |
| Exploration | Versioned playbooks and GOAP A* | Broad, adaptive, probabilistic | Models broaden search; agents execute |
| Reproducibility | Same execution path is replayable | Outputs and coverage can vary | Findings reproduce in the deterministic core |
| Evidence | Captured target output | May rely on generated interpretation | Mandatory captured evidence gate |
| Cost | Fixed infrastructure profile | Variable token usage | Model routing only where useful |
| Data control | Air-gap ready | Often tied to an external API | Frontier, private, open-weight, sovereign, or disabled |
| Patch validation | Original exploit re-run | A new model run may differ | Original deterministic exploit re-run |
| Reporting authority | Execution evidence | Model judgment | Execution evidence only |
So What Is AI-Native Pentesting?
AI-Native does not mean delegating truth to a chatbot. It means using the right intelligence for each layer:
| Capability | Mechanism |
| ----------- | ----------- |
| Goal planning and execution | GOAP A* plus versioned specialist playbooks |
| Adaptive exploration | Optional frontier, private, open-weight, or sovereign models |
| Workload selection | Model routing across eligible models |
| Cross-agent intelligence | Shared engagement state and credential propagation |
| Finding authenticity | Captured target output and fail-closed evidence grading |
The model can be disabled. The evidence gate cannot. Prancer combines adaptive exploration with deterministic execution so creativity expands coverage without weakening proof.
Wrapping Up the Series
Across three posts we followed one validated engagement end to end:
- Part 1 — the architecture and external scan
- Part 2 — credential correlation, lateral movement, and tunnel pivoting
- Part 3 (this post) — how deterministic execution and optional LLM exploration work together under one evidence gate
One command. Full kill chain. Reproducible evidence. Customer-controlled deployment. That is autonomous penetration testing built for both exploration and proof.