# What Should a RAG Security Testing Checklist Cover in 2026?

userhero.io · October 1, 2026

> What Is RAG Security Testing? A RAG security testing checklist is a repeatable procedure for checking whether a retrieval-augmented generation system...

## What Is RAG Security Testing?

A RAG security testing checklist is a repeatable procedure for checking whether a retrieval-augmented generation system exposes the wrong information, accepts malicious instructions, or produces answers that violate access controls. RAG systems retrieve documents from indexes, place selected content into a model prompt, and generate an answer, so each stage introduces a different failure mode. The retrieval stage may return another customer’s records, the generation stage may follow instructions embedded in a document, and the application layer may display sensitive text without adequate filtering. Testing must therefore cover identity, retrieval, prompting, generation, logging, and operational controls rather than asking only whether a model gives a plausible answer.

**Also worth reading:** [What is included in a SOC 2 Type II readiness checklist for B2B SaaS companies?](https://userhero.io/knowledge/what_is_included_in_a_soc_2_type_ii_readiness_checklist_for_b2b_saas_companies.php) · [How do I build a SOC 2 feedback inbox compliance checklist for my B2B SaaS?](https://userhero.io/knowledge/how_do_i_build_a_soc_2_feedback_inbox_compliance_checklist_for_my_b2b_saas.php) · [Which LLM Gateway Security Controls Matter Most for Enterprise AI in 2026?](https://userhero.io/knowledge/which_llm_gateway_security_controls_matter_most_for_enterprise_ai_in_2026.php)

The central objective is not to prove that the system is secure; no finite checklist can provide that assurance. It is to identify conditions under which confidentiality, integrity, availability, and tenant separation break down. A useful program establishes measurable pass criteria, records known limitations, and reruns the same tests after changing models, embedding models, document parsers, vector databases, authorization policies, or prompts. For a B2B customer-signal inbox, this is especially relevant because product and support teams may process account histories, internal product feedback, sales comments, or support tickets containing personal and commercial information.

RAG security differs from ordinary application penetration testing because the model is probabilistic and the knowledge base changes over time. Traditional tests can still verify API authentication, session handling, rate limits, and network exposure, but they do not automatically reveal whether retrieved chunks are authorized for the requesting user. A mature assessment combines deterministic authorization tests with adversarial prompts, document-poisoning experiments, cross-tenant searches, and human review of outputs. It also tests the application path around the model, including caches, citations, traces, analytics, and downstream actions taken by an AI agent.

## Build the Threat Model Before Choosing Tests

The first step in a RAG security testing checklist is to map assets, actors, trust boundaries, and data flows. Identify who submits queries, which identity service establishes the user, which service applies permissions, where documents originate, which components create embeddings, and which model receives retrieved text. Mark every place where data changes authorization context or crosses between customers, environments, or trust levels. In a customer-signal platform, examples might include one workspace, one customer account, a product area, a ticket queue, or a user role such as administrator, product manager, or support agent.

Then define plausible attacks rather than adopting an unranked collection of prompt examples. Direct prompt injection asks the model to ignore its instructions; indirect injection places hostile commands in a ticket or document that the RAG pipeline later retrieves. Poisoning attempts to insert misleading or malicious content into the corpus, while data-exfiltration tests search for information that should be unavailable. Other cases include unauthorized tool use, citation manipulation, excessive retrieval, denial-of-service inputs, insecure output rendering, and prompt leakage. A model may be correct on every ordinary support question while still being unsafe when one customer uploads a file designed to manipulate another user’s answer.

Assign a severity and an owner to each scenario. A failed tenant-isolation test affecting regulated data should normally receive immediate attention, whereas a low-impact formatting defect can enter the normal remediation queue. Define measurable thresholds, such as zero cross-tenant documents returned in an automated corpus, zero unauthorized sensitive fields in a 500-query test set, and at least 99% retrieval authorization precision on clearly labeled fixtures. Avoid claiming that a 99% score proves safety; it can still hide a high-impact failure, but it creates a consistent release gate that is better than subjective judgment alone.

## Test Retrieval, Access Control, and Tenant Isolation

Retrieval testing should verify both whether relevant material can be found and whether irrelevant or unauthorized material is excluded. Prepare a labeled fixture corpus with documents from at least 2 synthetic tenants, several roles, 4 sensitivity levels, and a mix of current, expired, deleted, and quarantined records. Include 100 to 500 benign queries initially, then add edge cases involving near-duplicate names, shared terminology, inherited permissions, and conflicting versions. Because embeddings can retrieve text based on semantic similarity, filenames, account IDs, and labels must not be treated as reliable authorization controls.

Access controls should be enforced before or during retrieval, using trusted server-side identity and policy data. The test should attempt horizontal access across tenants, vertical access across roles, and access through indirect references such as a ticket that quotes another customer’s message. For each request, inspect the candidate documents, reranker results, prompt context, model answer, citations, and logs. A test passes only when protected content is absent from the entire path, not merely omitted from the final prose. If a source identifier leaks a tenant’s document title or internal record key, that may still be a confidentiality issue even when the generated sentence looks harmless.

Measure recall and precision separately. In a sample of 200 authorized questions, a retrieval hit rate of at least 95% may be a practical initial target for non-critical knowledge, while teams should investigate every known authorization failure. Precision should be evaluated manually at first because a technically correct but unauthorized chunk is a serious defect. Automated evaluators can help with volume, but they should not replace a reviewer for policy-sensitive decisions. Teams should also test document updates and deletion latency, confirming that stale chunks disappear within the documented service-level objective, such as 5 minutes for support data or 15 minutes for less urgent analytics.

## Test Prompt Injection and Data Exfiltration

Prompt-injection testing should cover both visible user inputs and instructions embedded in retrieved content. Begin with benign probes, then escalate to requests to reveal system prompts, ignore policies, print hidden context, expose credentials, change output formats, or call tools outside the user’s role. Indirect attacks should be placed in filenames, PDF text, OCR content, HTML comments, metadata, ticket comments, and code blocks because a document parser may preserve instructions that an ordinary search preview does not show. IBM’s discussion of AI agent testing is relevant here: an agent may convert an incorrect answer into an action, so tool permissions and confirmation gates need tests of their own.

Data-exfiltration tests should probe for both semantic leakage and literal leakage. Literal cases include account identifiers, email addresses, API keys, internal URLs, and document titles. Semantic cases ask whether the model can infer protected information from aggregated facts, such as identifying a customer from an unusual support incident. The evaluator should compare the response with the raw retrieved context rather than relying on exact string matching, since paraphrasing can evade simplistic filters. A suspicious finding should be reproduced at least twice and reviewed by an application-security specialist before it is classified as a vulnerability.

Use multiple prompt variants because a single wording may succeed accidentally. For a 100-prompt injection suite, run at least 5 paraphrases per critical scenario, producing 500 attempts across the full suite. A practical release threshold is zero successful secret disclosure and zero policy bypass in the critical set, with failures triaged individually. A 1% success rate sounds small, but one successful cross-tenant disclosure can affect thousands of records; security gates should therefore weight impact more heavily than an average score. Record the exact model version, prompt template, retrieval snapshot, and policy configuration so results remain reproducible.

## Validate Sources, Generation, and Agent Actions

A generated answer can be wrong even when its citations are genuine, and it can be dangerous even when its citations are absent. Evaluate factual grounding by checking whether every material claim follows from the supplied sources and whether the answer acknowledges uncertainty when the corpus is incomplete. Test conflicting documents, outdated procedures, unsupported conclusions, fabricated citations, and requests to use sources outside the user’s permitted collection. Keep a human-labeled set of roughly 100 to 300 representative questions and update it when the product or customer base changes.

Citations should be tested as security interfaces, not decorative labels. A user should be able to open only a source that the same request was authorized to access, and the citation handler should re-check authorization rather than trust an opaque identifier from the model. Strip active HTML and executable content from displayed documents, isolate previews, and prevent source text from becoming a new injection point. Log which sources influenced an answer, but redact secrets and minimize retained personal data. Teams should decide whether raw prompts and retrieved chunks appear in observability tools, who can access those logs, and how long they are retained.

If the RAG application can call tools, enforce a separate decision boundary for every action. Use an allowlist of tools, constrain arguments with server-side schemas, verify the user’s permissions again at execution time, and require human confirmation for destructive or externally visible actions. Test attempts to invoke a tool through instructions embedded in retrieved content. A useful initial policy is that the model may propose an action, but a deterministic service decides whether it is allowed; for high-risk actions, require explicit approval even when the model is highly confident. The result should be a constrained agent workflow rather than unrestricted execution.

## Compare Testing Approaches and Set Release Gates

There is no single product category that replaces a complete testing program. Manual review, open-source scanners, model evaluation frameworks, and commercial red-team platforms each cover different parts of the risk. The table below compares four common approaches; the point is to combine them according to risk, not to declare one universally best.

| Feature | Manual review | Open-source RAG scanners | Commercial red-team platforms | CI evaluation suite |
| --- | --- | --- | --- | --- |
| Coverage | Deep but slow | Fast repeatable checks | Broad adversarial breadth | Fast regression control |
| Authorization accuracy | Depends on reviewer | Strong when policies are explicit | Varies by integrations | Strong for known scenarios |
| Indirect prompt injection | High-quality review | Useful pattern detection | Often extensive | Good when fixtures are added |
| Real-time release gate | Usually expensive | Possible | Possible | Designed for frequent use |
| Cost profile | Staff time and testing tools | Often low or free; engineering cost remains | Usually subscription plus setup | Usually engineering and CI cost |

| Feature | Manual review | Open-source scanners | Commercial platforms | CI suite |
| Best use | Novel attacks and policy review | Baseline static and prompt tests | Independent breadth and reporting | Preventing regression on every build |
For a small team, begin with a 100-question authorized retrieval suite, 50 prompt-injection probes, 20 cross-tenant attempts, and a documented manual review of all critical outputs. Expand to 500 or 1,000 cases when the system handles regulated data, customer-facing actions, or multiple tenants. Review test failures weekly during a release and monthly for stable systems, with an immediate retest after a material model, retrieval, or authorization change. The exact schedule should reflect deployment frequency and impact, not a generic compliance calendar.

Set release gates that distinguish blocking from advisory findings. A reasonable starting point is zero confirmed cross-tenant access, zero secret exposure, zero unauthorized tool execution, and no unresolved critical injection paths in the tested scope. For quality metrics, many teams begin with a target of at least 90% grounded-answer correctness, 95% citation validity on supported answers, and fewer than 5% unsupported claims in a curated set, then tighten these targets for regulated use cases. These are operating thresholds, not universal standards, and teams should report denominators and test coverage so an apparently strong percentage cannot hide a tiny sample.

## Prevent Common Mistakes and Operational Drift

The most common mistake is testing only the model while ignoring the retrieval system. An LLM can follow a strict prompt and still receive a document belonging to another tenant, so application controls must be tested independently. Another error is treating prompt injection as a solved filter problem: blocklists miss paraphrases, multilingual variants, encoded text, and instructions expressed through seemingly harmless content. Security depends on layered controls, including trusted identity, least-privilege retrieval, output controls, monitoring, and incident response.

Teams also make the mistake of creating realistic customer data in test fixtures. Use synthetic documents, generated identifiers, redacted records, and isolated accounts wherever possible. A security test should not create a new production leak. Similarly, do not measure only whether an answer contains a banned word; attackers can reveal information through transformations, timing, citations, tool arguments, or side channels. Verify the entire pipeline and ensure that caches, backups, evaluation prompts, and vendor logs follow the same privacy rules as the application.

Drift is inevitable because models, prompts, source data, and business rules change. Keep an inventory of production model versions, embedding models, retrieval filters, policy decisions, test cases, and owners. Re-run the complete suite after a model upgrade and run focused authorization tests after a permission change. A quarterly review is useful for governance, but it cannot replace event-driven testing; a changed tenant filter or new document parser should trigger immediate validation. Track false positives and false negatives as well as vulnerability counts, since an overly noisy scanner can lead teams to disable the tool altogether.

## When to Act and What It May Cost

Act immediately when a RAG system crosses trust boundaries, stores customer or employee information, exposes internal documents, or can call tools that modify business systems. Prioritize tenant isolation, authentication, deletion behavior, secret handling, and authorization at retrieval before polishing answer style. A system used only for public, read-only reference material may justify fewer resources, but it still needs input limits, abuse monitoring, and tests for manipulated sources. The presence of an LLM does not make a system high risk by definition; data sensitivity, connectivity, scale, and autonomy determine the priority.

Costs vary more than list prices suggest. A small team may spend roughly $2,000 to $10,000 per month on engineering time, synthetic test data, CI capacity, monitoring, and external specialist review during an initial program. Commercial scanners or red-team services may add hundreds to several thousand dollars per month or per engagement, while enterprise platforms can cost substantially more with integrations and usage-based model calls. Model-evaluation APIs, vector databases, and security gateways also create variable compute and storage costs. These are planning ranges rather than vendor quotations, and actual price depends heavily on deployment scale and existing controls.

Begin with a one- or two-week baseline assessment if no formal testing exists, using a small, representative corpus and synthetic tenants. Record the system architecture, run authorization and injection probes, review the first 100 answers manually, and publish a remediation plan with deadlines. Then automate the highest-impact checks in CI and schedule periodic independent review. For a B2B customer-signal product, a defensible target is not “zero risk”; it is a documented program that can demonstrate who is allowed to retrieve what, how the system behaves under attack, which findings remain open, and when each control was last verified.

## Quick answers

### What is the most important test in a RAG security checklist?

The most important test is cross-tenant or cross-user retrieval: verify that an unauthorized identity cannot retrieve, infer, cite, or receive protected information from another customer’s documents. Add prompt-injection and secret-disclosure tests once those access boundaries are enforced. A model-level prompt test cannot compensate for a broken retrieval filter.

### How many RAG security test cases does a team need?

A small team can begin with about 100 authorized queries, 50 injection prompts, and 20 cross-tenant attempts, then expand as risk and coverage grow. Larger or regulated deployments commonly use 500 to 1,000 or more cases, including edge cases discovered in production. Coverage matters more than a headline number, so every critical finding should be reproduced and tracked.

### Can RAG security scanners replace penetration testing?

No. Scanners can repeat known authorization, prompt, and retrieval checks, but penetration testers can examine chained failures across identity, parsers, caches, APIs, tools, and deployment configuration. The strongest approach combines automated regression tests with targeted manual review and periodic independent adversarial testing.

### Should retrieved documents be sanitized before entering the prompt?

They should be treated as untrusted input, even when they came from an approved repository. Sanitize active content, isolate previews, remove unnecessary metadata, and enforce authorization before and during retrieval. These measures reduce risk, but they do not prove that prompt injection has been eliminated, so output and tool-use controls remain necessary.

### How often should a RAG system be security-tested?

Run the full baseline after major model, prompt, retrieval, parser, or authorization changes, and run focused checks in CI for every meaningful release. A monthly or quarterly cycle can supplement continuous regression testing, depending on sensitivity and deployment frequency. Any confirmed exposure should trigger immediate investigation, containment, retesting, and an incident review.

Canonical: https://userhero.io/knowledge/what_should_a_rag_security_testing_checklist_cover_in_2026-2.php
Markdown: https://userhero.io/knowledge/what_should_a_rag_security_testing_checklist_cover_in_2026-2.php/index.md
