← Back to blog

Measure LLM Data Leakage With Per Sequence Tests for Security Teams

September 21, 2026
Measure LLM Data Leakage With Per Sequence Tests for Security Teams

LLM data leakage is the unintended disclosure of training data, system prompts, retrieved documents, or session memory through a model's outputs, logs, or side channels. The core rule for any deployed model is simple: treat every output, retrieval, and memory read as a potential disclosure surface, not a trusted response. Before you trust a production model with sensitive data, measure how much it leaks using per-sequence extraction probabilities, not just averaged extraction rates.


TL;DR:

  • Data leakage can occur during training, inference, or through architectural side channels, each requiring different detection and mitigation strategies.
  • Per-sequence extraction probability metrics, rather than average rates, are essential to accurately assess and prioritize leakage risks for sensitive information.
  • Attack vectors like prompt injection, supply-chain poisoning, and cache sharing exploit trust boundaries that are often overlooked in security practices.
  • Continuous, real-world testing of leakage—especially under multi-turn and high-query scenarios—is necessary before trusting controls or deploying models in sensitive environments.
  • Proper safeguards include data minimization, strict access controls, verified model artifacts, and comprehensive logging security to prevent breaches across all leakage vectors.

Gmdautomation
Build More Secure AI Systems
GMD Automation helps UK businesses deploy scalable AI systems designed for security, compliance, performance, and ongoing optimization.
Explore AI automation

Table of Contents

What is LLM data leakage and why the standard definition falls short

Most vendor blog posts define LLM data leakage as "a model revealing information it shouldn't." That's technically true and practically useless. It doesn't tell you where to look, what to measure, or which controls to prioritise.

The more useful framing, and the one this guide uses throughout, splits leakage into distinct categories with different root causes and different fixes. Some leakage originates during training (the model memorised something it shouldn't have seen twice). Some happens at inference time (a retrieval system hands the model a document it shouldn't disclose). Some is architectural (a side channel like timing or cache behaviour reveals information indirectly). Conflating this is how security teams end up applying a prompt filter to a problem that actually requires differential privacy, or vice versa.

The main leakage categories

  • Training-data memorisation. The model reproduces verbatim or near-verbatim snippets from its training corpus, including PII, code, or proprietary text, when prompted in the right way.
  • System-prompt and hidden-context leakage. Attackers extract the instructions, tool definitions, or few-shot examples baked into a system prompt through adversarial questioning, revealing business logic or safety rules.
  • Test-data leakage. Evaluation or benchmark data bleeds into training, inflating reported performance and hiding real-world failure modes, an integrity problem as much as a security one.
  • RAG and vector-store leakage. Retrieval-augmented systems surface documents a user should never see, because access control lives in the application layer while the vector database is treated as a flat, permission-blind index.
  • Embedding leakage. Vector representations of confidential text get exported, cached, or shared across tenants, and can be partially inverted back towards the source content.
  • Memory and session leakage. Agentic systems that persist "memories" across conversations can carry one user's data into another user's session, or resurface information the user believed was forgotten.
  • Inference-time side channels. Timing differences, token probabilities, or KV-cache sharing between requests leak information about other users' prompts or the model's internal state without a single word of output being disclosed directly.

Two of these deserve extra attention because they're newer and less well understood. System-prompt leakage matters because organisations now embed real business logic, pricing rules, and safety guardrails directly into prompts, treating them as code rather than as a public-facing surface. Extract the system prompt, and you've often extracted a chunk of the company's operating logic. Inference-time side channels matter because they bypass every content filter you've built. A KV-cache timing attack doesn't need the model to say anything sensitive; it infers sensitive information from how fast the model responses.

It's worth remembering that disclosure doesn't only happen in the chat window. Tool call arguments passed to external APIs, application logs that capture full prompts and completions, observability platforms that store raw traces for debugging, and telemetry pipelines shipping data to third-party analytics tools are all places where the same sensitive content can leak a second time, often with far weaker access controls than the model endpoint itself.

How leakage happens: threat surfaces and adversary techniques

Understanding the mechanism behind leakage tells you where to spend your security budget. The OWASP Top 10 for LLM Applications ranks prompt injection as the leading deployment risk for LLM applications, with sensitive information disclosure close behind, and both categories map directly onto the techniques below.

  1. Prompt injection and jailbreaking. An attacker embeds instructions in user input, a retrieved document, or even an image, that override the model's intended behaviour. Multi-step jailbreaks chain several benign-looking prompts together, each one nudging the model slightly further from its guardrails until it discloses something the first prompt alone never would have extracted.
  2. Training-time poisoning and supply-chain risk. Poisoned fine-tuning data, compromised LoRA or PEFT adapters, and tampered tokenizer configurations can implant triggers that cause leakage or misbehaviour only under specific, attacker-chosen inputs. Quantization artefacts and unsigned model packages are an underexamined entry point here; few teams check adapter provenance the way they'd check a software dependency.
  3. RAG and knowledge-base contamination. If your retrieval pipeline indexes anything a user or partner can write to, wikis, ticket systems, shared drives, an attacker can plant content designed to be retrieved and then either exfiltrate other users' data or manipulate the model's response.
  4. Observability and logging misconfiguration. Full prompts and completions, including anything a user pasted in, routinely end up in logging platforms, application performance monitoring tools, and third-party analytics dashboards with broader access than the model itself.
  5. KV-cache sharing and multi-tenant risk. Serving infrastructure that shares compute or cache state across tenants for efficiency can, if isolation is imperfect, leak timing or content signals between customers who should never see each other's data.
  6. Agentic threats. Agents that call tools, write to memory, and persist state across sessions create what recent research calls temporally extended attacks: a malicious instruction planted in one session can poison a memory store that gets read, and exploited, in a completely separate session days later. A survey of threat surfaces in LLM agents describes this composition failure as one of the hardest problems in agentic security, precisely because standard input filters only look at a single turn and miss the slow-burn attack entirely.

Pro Tip: Don't test your input filters against single-turn jailbreak prompts and call it done. Run a multi-turn adversarial session that plants an instruction early and tries to trigger it five or ten turns later. Most filters that catch obvious single-shot attacks miss the delayed-detonation version completely.

The common thread across all six techniques is that they exploit trust boundaries your architecture assumed were solid. A RAG pipeline trusts its own index. A multi-tenant server trusts its cache isolation. An agent trusts its own memory. Attackers don't need to break the model itself; they need to find where you stopped checking.

How to measure and test for leakage: metrics, probes and attacker budgets

You cannot secure what you haven't measured, and the standard leakage metric most teams reach for, an averaged extraction rate across a dataset, actively hides the risk you most need to see.

Why averages lie about leakage risk

Extraction rate answers "what percentage of training sequences can be extracted on average?" That number can look comfortably low while specific sequences, such as those containing a customer's home address or an internal API key, sit at near-certain extractability. Research on sequence-level leakage risk found that averaged extraction-rate metrics can substantially underestimate genuine leakage risk relative to per-sequence probability analysis, because the average smooths over exactly the high-risk outliers a security audit needs to catch. The same research found that later tokens in a sequence can be significantly easier to extract than earlier tokens, and that shorter prefixes or smaller models often make specific sequences dramatically easier to pull out. A single dataset-wide percentage tells you almost nothing about which record is exposed.

Average versus per-sequence leakage risk

The fix is to report per-sequence extraction probability rather than (or alongside) an aggregate rate, and to weight that reporting towards your highest-sensitivity records: named individuals, credentials, proprietary code, anything that would trigger a breach notification if it appeared verbatim in a model's output.

Building a probe set that actually tells you something

Two complementary testing approaches cover most of the ground:

  • Black-box probing. You query the deployed model with no access to weights, using structured prompts designed to elicit memorised content. The ProPILE probing framework formalises this for PII specifically, using exact-match and likelihood-based scoring to estimate how much personal data a model would reveal, and introduces a metric called γ<k, the fraction of data subjects whose PII would likely be exposed within k queries.
  • White-box and membership inference. With access to logits or weights, you can run membership-inference attacks that ask, statistically, "was this specific record in the training set?" This catches memorisation that black-box querying alone might miss, particularly for records the model reproduces only under unusual decoding conditions.

Three variables change your results more than most teams expect, and all three belong in your test plan:

  • Decoding regime. Greedy decoding, top-k sampling, and top-p (nucleus) sampling each change extraction probability. A sequence that never surfaces under greedy decoding can appear reliably under repeated sampled queries.
  • Attacker query budget. Extraction risk scales with the number of queries an attacker is allowed. A model that looks safe at 10 queries can look very different at 10,000, which is why γ<k reporting at multiple values of k matters more than a single pass/fail number.
  • Prefix length and position. Shorter prefixes and later token positions both change extractability substantially, per the sequence-level leakage findings above, so a probe set that only tests long, well-formed prefixes will systematically underestimate risk.

An audit-worthy test record captures: the exact prompt template, the decoding settings, the query budget, whether the match was exact or partial, and the per-sequence probability, not just a binary "leaked / did not leak" flag. Partial matches deserve particular scrutiny; a model that reconstructs 80% of a customer record with the remaining 20% inferrable from context has not actually protected that record.

One distinction worth building into every report: separate what a model can be forced to reveal under adversarial pressure from what it tends to reveal during normal use. Capability tells you your worst-case exposure. Propensity tells you your everyday risk. Conflating the two produces audits that either terrify the board over an attack nobody will realistically mount, or reassure everyone while normal users stumble into real disclosures.

Practical mitigations and controls: engineering and operations

Controls fall into a rough priority order: fix what's cheapest and highest-impact first, then layer in the harder architectural changes.

Data minimisation and input sanitisation

The single most effective leakage control is not collecting or retaining sensitive data you don't need. Strip PII from training data before fine-tuning, scrub prompt templates of anything that doesn't need to be there, and vet every document before it enters a RAG index rather than trusting the retrieval layer to sort it out later. Sanitisation should run on both inputs and outputs: filter what goes into the model and check what comes out before it reaches a log file or a downstream system.

Access control and rate limiting

  • Apply per-user or per-tenant authentication before any query touches the model, not just at the application front door.
  • Set per-user query budgets, since extraction risk compounds with query volume; a user making 50,000 queries a day to a customer-support bot is a different risk profile from one making 50.
  • Rate-limit aggressively around any endpoint that touches raw retrieval or memory reads, where the payload is richer than a typical chat response.

Privacy-preserving training and post-training methods

Differential privacy, applied during training, can substantially reduce PII memorisation, though it comes with a real trade-off: DP training can introduce instability and reduce model utility, so it needs validation against your actual task performance, not just a privacy benchmark. Post-training approaches like direct preference optimisation and targeted unlearning offer a more stable privacy-utility balance in some evaluations, but no single method eliminates leakage outright, alignment reduces extractability, it doesn't remove it. Whichever method you apply, re-run your per-sequence probing afterwards. A method that looks effective on paper needs verification against your own probe set before you trust it in production.

Securing RAG and vector infrastructure

RAG pipelines need the same access discipline as any other data system, and most don't get it. Enforce document-level ACLs inside the vector store itself, not only in the application layer that calls it. Vet every source before ingestion, and quarantine newly indexed content until it's been checked for injected instructions. Treat embeddings as confidential artefacts in their own right: encrypt vector stores at rest, restrict export, and avoid bundling embeddings into shared backups, since partial reconstruction of source text from vectors is a documented risk, not a theoretical one, according to OWASP's guidance on LLM deployment risk.

Secrets, logging, and telemetry hygiene

  • Never let API keys, credentials, or internal identifiers flow into prompts that get logged in plaintext.
  • Redact PII and secrets from observability platforms and application performance monitoring traces before they're stored, not after.
  • Audit third-party telemetry integrations specifically for what they capture; many ship full request and response bodies by default.

Memory governance for agentic systems

Every memory write needs a provenance tag recording where it came from and when. Set expiry policies so old memories don't accumulate into a growing disclosure surface, and build a revocation path so a user or admin can delete a memory and trust that it's actually gone across every session, not just the current one. Cross-session audit logs matter here specifically because a malicious "write" in one session and its "exploit" in a later one can be days apart, and without provenance tags linking them, you'll never trace the attack back to its origin.

Pro Tip: If your agent writes to a persistent memory store, run a test where you plant an innocuous-looking instruction in session one and check, days later in session five, whether it still influences behaviour. Most teams only test within a single session and miss this entirely.

None of this replaces the alignment work already baked into your model, but alignment guardrails are a behavioural layer, not a data security layer. A well-aligned model can still memorise and leak the record you handed it, because alignment shapes how the model talks, not what it remembers. For the architectural side of locking this down, enterprise AI security architecture is worth reviewing alongside your model-specific controls.

Memory governance for agentic systems — overview diagram

Red-team and audit playbook: step-by-step for security teams

A leakage audit needs the same rigour as a penetration test: defined scope, reproducible methodology, and a report someone outside the team can act on.

  1. Define scope and threat model. Decide upfront whether you're testing black-box (query access only, simulating an external attacker) or white-box (weights and logits available, simulating an insider or a compromised supply chain). Each requires different tooling and produces different guarantees.
  2. Build your probe set. Assemble three categories of test prompts: known PII records you can verify against a ground truth, proprietary or sensitive text snippets planted deliberately if you're testing your own fine-tuned model, and structured extraction templates modelled on published probing frameworks. Sample across sequence lengths and positions, not just long, clean prefixes.
  3. Run experiments across multiple decoding regimes. Test greedy decoding, top-k, and top-p sampling separately, and vary your attacker query budget from single digits to the thousands, recording γ<k at each budget level.
  4. Simulate agentic and RAG hijack paths. Attempt tool-mediated exfiltration (can an agent be tricked into passing sensitive data through a tool call it shouldn't use?) and memory poisoning (can a planted instruction in one session influence a later, unrelated session?).
  5. Compute and report per-sequence metrics. Calculate per-sequence extraction probability, γ<k for PII exposure, and membership-inference scores where white-box access allows it, then flag every high-risk sequence individually rather than burying it in an average.
  6. Assign remediation priority and reissue tests. Rank findings by exploitability and business impact, assign owners, and re-run the same probe set after fixes ship to confirm the leakage actually closed rather than just moved.
Report sectionWhat it should contain
Scope and threat modelBlack-box vs white-box, systems tested, exclusions
Probe methodologyPrompt templates, decoding regimes, query budgets used
FindingsPer-sequence probabilities, γ<k values, exact vs partial matches
Agentic and RAG findingsTool-exfiltration attempts, memory-poisoning results
Remediation prioritiesRanked by exploitability and impact, with owners assigned
Retest resultsConfirmation that fixes reduced measured leakage, not just symptoms

Supply-chain checks belong in this playbook too, not as a separate exercise. Inspect any third-party LoRA or PEFT adapters and tokenizer configurations for embedded triggers, and confirm model artefacts carry verifiable package signing before they go anywhere near production, a gap a systematic survey of agent security threats flags as consistently under-tested. For teams building this into a formal risk framework, AI model risk management practices that satisfy regulators map closely onto the audit structure above. Independent multi-model checks, such as BabyLoveGrowth's Multi-LLM Audit tool, can also help teams compare leakage exposure across several provider models in one pass rather than running isolated one-off tests.

Short case studies and research highlights: incidents and lessons

Public incidents involving LLM data leakage tend to trace back to a small set of recurring root causes, not exotic new attack classes.

  • Observability and logging exposures. Several publicised incidents involved full prompt and completion histories, including customer PII typed into chat interfaces, sitting in third-party logging or analytics platforms with far weaker access control than the model endpoint itself. The lesson is consistent: logging pipelines are a leakage surface, and most teams secure the model long before they secure the logs.
  • KV-cache and multi-tenant serving risk. Shared inference infrastructure designed for efficiency has, in documented cases, allowed cross-tenant information to bleed through cache reuse. The fix isn't exotic, stronger tenant isolation at the serving layer, but it's routinely deprioritised because it's an infrastructure cost, not a model feature.
  • RAG poisoning through unvetted sources. Knowledge bases that accept writes from wikis, tickets, or shared drives without vetting have been shown to surface planted content back through retrieval, sometimes exfiltrating data belonging to other users of the same system.
  • Alignment bypass through targeted extraction attacks. Research submitted to ICLR demonstrated that divergence and finetuning attacks could recover tens of thousands of memorised training examples from aligned, production-grade models using public tools and modest budgets, proving that safety fine-tuning changes behaviour, not memorisation.

The recurring pattern across all four: teams secured the obvious front door (the chat interface, the alignment layer) and left a side door open (the log pipeline, the cache, the RAG index, the fine-tuning data itself). Leakage rarely comes from a sophisticated zero-day. It comes from a trust boundary nobody re-checked after the system shipped.

Author perspective and how Gmdautomation approaches leakage risk

Ravi has spent years working at the intersection of AI deployment and enterprise security, with a particular focus on how automation systems handle sensitive data in regulated UK sectors.

Expertise in AI deployment and enterprise security shapes how managed AI deployments are built. Systems are designed around minimal-data ingestion from the outset, pulling in only what a workflow genuinely needs rather than defaulting to broad access, with provenance tracking on retrieval sources and audit logging built into the deployment rather than bolted on afterwards. That governance layer matters as much as the model choice itself; a well-aligned model sitting on an unaudited data pipeline is still a leakage risk waiting to surface. For UK businesses evaluating vendors, the right question isn't just "which model," it's "what does the audit trail look like when something goes wrong."

Practical governance frameworks like ISO 27001 controls mapped to AI deployments give a useful benchmark for what a properly audited AI system should demonstrate, regardless of which vendor builds it.

The measurement gap nobody talks about

Most organisations treat LLM data leakage as a compliance checkbox: run a red-team exercise once, file the report, move on. The research doesn't support that approach, and frankly, neither does common sense once you've seen how sensitive per-sequence extraction risk actually is to decoding settings and query volume.

The conventional advice, "align the model well and add a content filter," treats leakage as a behavioural problem. It isn't. It's fundamentally a data problem, and alignment operates on behaviour, not memory. A model can refuse to discuss a topic in ninety-nine conversations and still reproduce a memorised record verbatim in the hundredth, if the attacker's query budget and decoding strategy happen to line up right.

If you take one thing from this guide, make it this: measurement has to be continuous, not a one-off gate before launch. Extraction risk shifts as models get fine-tuned, as RAG indexes grow, as agentic memory accumulates. Prioritise building the per-sequence testing habit before you build the next mitigation layer, because a control you can't verify against a real probe set is a control you're hoping works, not one you know works.

— Ravi

Sources

For teams building out their own testing programme, a handful of sources do most of the heavy lifting. The OWASP Top 10 for LLM Applications remains the closest thing to an industry-standard risk taxonomy. The ProPILE probing paper is the reference implementation for PII-focused black-box probing and its γ<k metric. Research on sequence-level leakage risk is essential reading for anyone still relying on averaged extraction rates. Beyond the papers, community incident write-ups and vendor postmortems, like Cobalt's practical guide to securing LLMs, offer the operational detail academic papers tend to skip.

FAQ

Will ChatGPT leak my data?

Any LLM, including ChatGPT, can potentially reproduce memorised training content or expose conversation data through logging and telemetry if it isn't properly secured. The realistic risk depends heavily on how a specific deployment handles data retention, logging, and whether outputs are treated as a trusted or untrusted surface, which is why enterprise deployments need their own leakage testing rather than relying on general reassurances.

Is LLM overfitting the same as data leakage?

Not quite. Overfitting means a model fits its training data too closely and generalises poorly to new inputs, while data leakage specifically means the model discloses information from that training data, or from retrieved or session content, through its outputs or side channels. Overfitting can make memorisation, and therefore leakage, more likely, but the two are distinct failure modes with different fixes.

What is data leakage in a machine learning model?

Data leakage in machine learning traditionally means information from outside the training set, often the test or validation data, improperly influences model training, inflating reported accuracy. In LLM contexts, the term has broadened to also cover the model disclosing memorised training data, retrieved documents, or session information it shouldn't reveal.

What is an example of data leakage in the context of AI?

A clear example is a fine-tuned customer-support model reproducing a real customer's name, email, and account details verbatim when prompted with a partial match of their query, because that record appeared in training data. Another common example is a RAG system retrieving and surfacing a confidential internal document to a user who lacks permission to see it, simply because the vector store lacked document-level access controls.