What a RAG penetration testing guide actually tests

A RAG penetration testing guide is a repeatable assessment method for testing retrieval-augmented generation applications whose answers depend on private documents, databases, search indexes, or other internal knowledge. It examines the complete path from user input to retrieval, prompt construction, model generation, tool execution, and any downstream action. The objective is not to prove that an LLM can be tricked; ordinary probabilistic models will produce irregular responses. The objective is to determine whether an attacker can cross a defined trust boundary, retrieve data they should not receive, influence system instructions, bypass source restrictions, or cause an authorized tool to perform an unintended action.

Also worth reading: What are the most effective secure AI agent architecture patterns for enterprise-grade B2B SaaS applications? · What Are the Best RAG Security Testing Tools for Enterprise Teams in 2026? · How Does AI Router Shadow Testing Reduce Production Risk for Agent Teams?

Testing must cover the application rather than the model alone. A vulnerability may exist in document access control, metadata filtering, embedding similarity, prompt assembly, output handling, or an integration with email, ticketing, CRM, or shell-based services. A model provider can patch its own platform while leaving an application-specific authorization defect untouched. For that reason, a useful guide assigns an owner, expected result, evidence requirement, severity rule, and remediation deadline to every test case. “The chatbot said something strange” is not a finding; “User A retrieved User B’s invoice through a crafted query, and the evidence appears in the transcript” is.

Threat model the retrieval path before testing

Start by drawing the trust boundaries and data flows. Identify public, authenticated, department-scoped, tenant-scoped, administrator, model-provider, and third-party data, and document which identity governs each retrieval request. Record the user-controlled fields, including the query, conversation history, uploaded filename, document metadata, URL, profile attributes, and references supplied by external systems. Then map how those fields reach the retriever, prompt, agent planner, tools, and final response. RAG systems often contain two distinct retrieval paths: semantic vector search and lexical search, and both must inherit the caller’s authorization context.

Threat modeling should produce concrete abuse cases. For a support assistant, one case may be a customer requesting another customer’s ticket; for an internal knowledge assistant, it may be an employee trying to retrieve compensation or health information; for an agent connected to ticketing systems, it may be a prompt that causes a privileged ticket closure. Consider indirect prompt injection in retrieved content, poisoned documents, malicious URLs, cross-tenant embeddings, insecure tool arguments, and logs that preserve secrets. The widely repeated advice that the prompt can become the payload is therefore accurate, but incomplete: the retrieval store, identity layer, agent framework, and tool permissions are equally relevant attack surfaces.

A practical first pass can use 4–8 hours for one narrow workflow, provided read-only access and safe test identities are available. This is enough to inventory flows, configure two or more privilege levels, and run basic authorization and injection tests. It is not enough to claim that an enterprise deployment is secure. Organizations should reserve several days for a small production pilot and 2–6 weeks for a multi-tenant application with many data sources, integrations, and model configurations.

Prepare a controlled test environment

A RAG penetration test should begin in an environment that resembles production without exposing real customers or irreversible actions. Clone representative documents, preserve realistic metadata and chunk lengths, and populate separate tenant or role fixtures. Create at least two ordinary users, one privileged operator, and one service identity. If a test involves deletion, payment, outbound email, or account changes, replace the real endpoint with a mock or a transaction sandbox. The model may remain production-equivalent because the security boundaries being tested are usually outside the model weights.

Define explicit stop conditions before execution. Stop if a test exposes regulated data outside the approved fixtures, creates an external communication, consumes an abnormal amount of inference budget, or destabilizes a shared index. Assign request and token limits, for example 100 test prompts per hour and a fixed spend ceiling, unless the approved scope states otherwise. Capture the complete trace rather than only the visible answer: user identity, filters, retrieved document IDs, scores, reranking results, final prompt, model version, tool calls, response, latency, and cost. Redact credentials from retained evidence while preserving enough information to reproduce the defect.

Use an evidence threshold based on demonstrable impact. A successful attack should normally show one of four outcomes: unauthorized data disclosure, instruction or policy bypass, an unauthorized state change, or material degradation of the system’s core function. A suspicious phrase without a repeatable chain of effect should be recorded as an observation and investigated. This discipline reduces false positives and prevents teams from spending time on harmless model hallucinations that fall within the product’s documented limitations.

Test authorization, retrieval, and prompt boundaries

Authorization testing should precede creative prompt attacks. Run the same or equivalent questions as different users and compare which documents are retrieved, cited, cached, or exposed through error messages. Test direct object references, manipulated tenant IDs, role changes after a session begins, inherited service-account permissions, and documents whose filenames or metadata reveal sensitive information. A practical acceptance threshold is zero cross-tenant retrievals for production data; any single verified occurrence is a high-priority defect when it exposes another customer’s information.

Next, test the boundary between trusted instructions and untrusted content. Place canary instructions in test documents rather than attempting vague role-play, such as asking the model to “ignore everything.” Evaluate whether content can suppress system instructions, reveal hidden prompts, invoke a tool, or conceal its citations. Test both visible and hidden text, metadata, tables, alt text, filenames, and document properties. Because attackers can poison a knowledge base through legitimate uploads, ingestion-time scanning and runtime instruction/data separation should be evaluated alongside prompt defenses.

Measure retrieval behavior, not merely final wording. An injected instruction may not appear in the answer but can still alter ranking, call a tool, or leak a fragment. Record top-k results before and after the attack, reranking changes, citation substitution, and tool-argument changes. Do not use a universal pass rate such as “99% injection resistance” without a defined corpus and scoring method. A defensible report instead reports, for example, 40 direct attacks, 20 indirect document attacks, and 15 authorization cases, with exact counts for blocked, failed, inconclusive, and successful outcomes.

Probe the ingestion and data supply chain

RAG security begins before a user submits a prompt. Organizations should verify that every upload passes authorization, type validation, malware scanning, size limits, and content sanitization. The pipeline should strip active content where the viewer does not require it, normalize metadata, and prevent one tenant’s object-store paths or embeddings from entering another tenant’s index. Test duplicate documents, malformed files, oversized pages, conflicting revisions, unsupported encodings, and poisoned content designed to rank first for high-value queries. A single 10 MB file should not be able to exhaust parsing resources, while 50,000 near-identical chunks should not make the search service unusable without a defined quota.

Embedding stores also need deletion and retention tests. Confirm that deleting a source document removes or isolates its chunks, cached answers, derived summaries, and secondary indexes within the documented retention period. Verify that revoked users cannot retrieve previously cached context and that a permission change takes effect within a stated threshold, such as 5 minutes for sensitive enterprise content. AWS guidance on filtering mechanisms is relevant because retrieval filters are part of the security boundary, not an optional relevance feature. However, filtering is ineffective if application code builds the query as an ordinary user role and leaves authorization enforcement to a downstream component with broader access.

GraphRAG should not be adopted merely because conventional vector retrieval has weaknesses. Graph-based approaches can add relationship context, but they also introduce entity-resolution errors, extra ingestion logic, and potentially broader access paths. Test both authorization propagation and answer quality; an apparently elegant graph traversal that exposes a connected node outside the user’s permitted subtree is worse than a less capable design with strict filters. Tooling and data quality must be evaluated together rather than treating a new retrieval architecture as automatic protection.

Assess agents, tools, and external integrations

An AI agent changes RAG risk because generated text can become action. Test every tool independently before allowing the model to select it. Apply least privilege, validate arguments against schemas, enforce user authorization again at execution time, and require stronger confirmation for high-impact actions. A model should not be allowed to send email merely because a retrieved page says to do so. Transactional tools should use idempotency keys, allowlisted destinations, bounded amounts, and two-person approval where appropriate.

Adversarial cases should include a retrieved invoice that requests payment, a support article containing a hidden command, and a malicious web page that tells an agent to export conversation history. Verify that secrets are unavailable to the model context, that credentials are not accepted from retrieved text, and that system messages cannot be replaced through tool output. Report the McKinsey “Lilli” incident as evidence that AI-agent interactions can become security events, not as proof that every agent implementation is equally exposed. A responsible assessment combines documented cases with direct testing of the organization’s permissions and safeguards.

A comparison helps determine where deeper testing is needed:

FeatureRAG knowledge assistantTool-using AI agentGraphRAG application
Primary riskUnauthorized retrieval or poisoned contextRetrieval plus unauthorized actionsRelationship traversal, entity errors, and access inheritance
Minimum safe test set25–40 cases40–80 cases40–80 cases plus graph-quality checks
Typical sandboxRead-only cloned documentsMock tools or transaction sandboxVersioned graph fixture with tenant boundaries
Strong release thresholdZero confirmed cross-user disclosuresZero unauthorized state changesZero unauthorized paths and no critical integrity defect
Key caveatFiltering can still be misconfiguredTool permissions can bypass prompt rulesBetter context does not guarantee better security
## Score findings and choose remediation priorities

Findings should be scored by impact, exploitability, exposure, and detectability rather than by a generic AI-risk label. One practical model uses a 1–5 impact score and a 1–5 likelihood score, producing a 1–25 total. Cross-tenant access to regulated records, arbitrary execution, or account takeover generally belongs in the highest response tier; a malformed citation without data exposure may belong much lower. Environmental factors matter: internet reachability, required user privileges, interaction complexity, and whether the result can be reproduced. A control weakness with no demonstrated path should not be presented as a confirmed breach.

Remediation should address the failed boundary. If a user can retrieve another tenant’s chunk, fix identity propagation, index isolation, or server-side authorization; do not merely ask the model to “never reveal private data.” If retrieved instructions trigger a tool, remove the tool permission, constrain the planner, require server-side confirmation, and validate all arguments. If malicious documents can influence ranking, add ingestion controls, source reputation, anomaly detection, and retrieval-time filtering, then retest. Track time to remediate and time to verify, not just the date a ticket was closed.

A reasonable service budget varies sharply by scope. An independent consultant may charge roughly $2,000–$10,000 for a narrowly scoped read-only pilot, while a broader application assessment can range from $10,000–$50,000 or more. Internal teams can reduce cash cost by using existing cloud credits, open-source scanners, and synthetic corpora, but they still face personnel, infrastructure, and remediation expenses. Managed continuous testing may be priced per application, test suite, tenant, or monthly run rather than as a one-time scan. Obtain a written statement of whether tool calls, model red teaming, cloud configuration, source-code review, and retesting are included; a low quote for a prompt-only review is not a full penetration test.

Common mistakes that make a RAG test unreliable

The most common error is treating a model response as the security boundary. The model is probabilistic and can refuse one attack while accepting a paraphrased version, so controls must exist outside it. Another error is testing only obvious phrases such as “ignore previous instructions” while neglecting indirect instructions embedded in documents or webpages. A third is using one administrator account, which cannot reveal broken tenant separation. A fourth is running destructive actions against real customer data, creating legal and operational risk that outweighs the information gained.

Teams also confuse hallucination with penetration. An invented policy may be a model-quality defect but not an exploit unless it crosses an authorization, safety, or business boundary. Conversely, a technically intriguing prompt that returns only public data may have no material impact. Record the exact route from input to effect and separate confirmed vulnerabilities from hypotheses. Version the application, prompts, model, retriever, index, and test corpus; otherwise a retest may appear successful because the underlying system changed. Finally, do not publish sensitive prompts, exploit strings, or customer-derived examples in a public case study. Share sanitized evidence and defensive lessons instead.

When to run the assessment and what success means

Run a baseline assessment before production launch, whenever a new model, retriever, agent tool, document source, or access-control model is introduced, and at least once a year for stable internet-facing systems. Run more frequently for high-volume customer-support or employee-assistance products, especially when they contain regulated data or can modify records. Event-driven retesting is appropriate after a data leak, major incident, permission redesign, migration to a new vector database, or agent-framework upgrade. A monthly suite of 25–50 automated regression cases is useful for known attacks, but it should complement—not replace—manual creative testing and architecture review.

A completed RAG penetration testing guide should leave behind an asset map, threat model, test corpus, reproducible evidence, severity rubric, remediation backlog, and retest procedure. Success is not a claim that the system cannot be attacked. It is evidence that defined boundaries held under the tested conditions, known weaknesses have owners and deadlines, monitoring can detect recurrence, and residual risk is acceptable for the current deployment. For B2B customer-signal workflows, teams should also verify that support conversations, feedback, account data, and generated summaries cannot move between customers or roles without an explicit business need. That customer-specific angle does not require hard-sell automation; it simply focuses the test on the trust customers expect from a system that processes their communications.