Sensitive Information Disclosure (LLM)
Also known as: LLM data leakage, LLM02, Sensitive data exposure in LLM apps
A flaw where an LLM application exposes secrets, personal data or another user's information in its output; it is contained by keeping secrets out of prompts, enforcing per-user access control on retrieval and tools, minimising data and filtering output.
How it works
Sensitive information disclosure happens when an application built on a large language model reveals data it should have kept private: credentials, personal information, proprietary content or another customer's records. The model is a text generator with no concept of who is allowed to see what, so anything placed in its context, prompt, retrieved documents or tool output can appear in a response.
It arises from design choices rather than one bug. Developers put API keys or internal rules in the system prompt, assuming users cannot see it. A retrieval pipeline indexes documents from many users into one store and searches it without checking who is asking. Tools run under a broad service credential, so the model can read records the current user could never open. Personal data used in tuning or logged from conversations can be echoed back later.
Impact ranges from exposed credentials that give access to backend systems, to cross-tenant breaches that trigger privacy and regulatory obligations. Because the model decides what to include in an answer, a harmless-looking question can surface data nobody intended to publish, and the leak is hard to notice without looking at outputs.
Defence is therefore about what the model is ever given. Keep secrets in a backend with least privilege and never in the prompt, enforce the requesting user's access rights at retrieval and tool time, minimise and de-identify data, scan and redact output, scope sessions and scrub logs. Then monitor responses for secrets and personal data so failures are noticed.
Walk through it
- 1Read how context is assembled
- 2Name the root cause
- 3Read the signals, pick the key control
- Scope what data the app can expose
- Harden with access control, redaction and verify
A pre-launch review covers PolicyPal Assistant, which answers questions about insurance policies. Read how the model context is built and note what sensitive material is placed in it.
1SYSTEM = (2 'You are PolicyPal. Billing API key: <live-key-placeholder>. '3 'Internal rule: waive fees up to 500 for VIP customers.'4)5 6def answer(question: str, user) -> str:7 docs = shared_index.search(question, top_k=5) # all customers' documents8 context = '\n'.join(d.text for d in docs)9 return llm.run(SYSTEM + '\n' + context + '\n' + question)10 11log.info('prompt=%s response=%s', prompt, response)Spot it
- Responses containing strings that match credential or key formats, caught by output DLP or secret scanners.
- Personal data patterns, such as national ids, card numbers or emails, appearing in responses where the user did not supply them.
- Retrieved documents or tool results whose owner or tenant differs from the requesting user.
- Prompts or responses with personal data written to debug logs, analytics or third-party services.
- Responses that restate system prompt content or internal business rules.
LLM gateway output scan
ts=2026-10-11T10:02:11Z session=s-8812 stage=output flag=secret_pattern type=api_key action=redact
ts=2026-10-11T10:02:40Z session=s-8812 stage=output flag=pii type=national_id count=2 action=logRetrieval audit
ts=2026-10-11T10:03:05Z user=u-5521 query_id=q-9921 doc=d-3340 doc_owner=u-1207 owner_match=false
ts=2026-10-11T10:03:05Z user=u-5521 query_id=q-9921 returned=5 cross_owner=3 filter_applied=nonesplSplunk: sessions with repeated secret or PII hits in model output
index=llm sourcetype=output_scan (flag=secret_pattern OR flag=pii)
| stats count by session, flag
| where count > 1Tune patterns to your data formats and correlate with retrieval audit before escalating.
kqlKQL: retrieval results owned by someone other than the requester
RetrievalAudit
| where owner_match == false
| summarize docs=count() by user, bin(TimeGenerated, 5m)
| where docs > 0Stop it
Enforce per-user access control at retrieval and tool time, and give the model least privilege
Authorise every document and record against the requesting user before it enters the context. Filter retrieval by owner or tenant, run tools with the user's own scoped identity rather than a broad service account, and never rely on the model to decide what is allowed.
Keep secrets out of prompts and minimise the data the model sees
Hold credentials in a backend service that performs privileged calls, so the model never sees a key. Pass only the fields the task needs, mask or de-identify personal data, and exclude sensitive content from tuning and retrieval indexes it should not reach.
Filter output, scope sessions, scrub logs and monitor
Scan responses with DLP and redact secrets and personal data, isolate each session's context, and keep prompts and responses out of logs or redact them with short retention. Alert on secrets and cross-owner data in output so failures are caught.
Keep the secret in the backend and authorise retrieval per user
Vulnerable
SYSTEM = 'You are PolicyPal. Billing API key: <live-key-placeholder>.'
docs = shared_index.search(question, top_k=5)
return llm.run(SYSTEM + '\n' + '\n'.join(d.text for d in docs) + '\n' + question)Hardened
SYSTEM = 'You are PolicyPal. Answer only from the provided documents.'
docs = index.search(question, top_k=5, filter={'owner_id': user.id}) # per-user filter
docs = [d for d in docs if authz.can_read(user, d)] # re-check
context = '\n'.join(minimise(d) for d in docs) # only needed fields
reply = llm.run(SYSTEM + '\n<data>' + context + '</data>\n' + question)
billing = backend.get_billing(user.id) # privileged call stays server-side, scoped
return output_guard.redact(reply.text) # DLP: secrets and PIIThe prompt contains no credential, retrieval is filtered and re-checked for this user, and output is scanned. The access check, not the prompt wording, is what prevents cross-user disclosure.
- Never place API keys, passwords or connection strings in prompts; call privileged services from the backend with scoped, short-lived credentials.
- Apply the requesting user's authorisation to every retrieval result and tool call, and partition indexes by tenant where possible.
- Send the model only the fields a task needs, and mask or de-identify personal data before it enters the context.
- Run output DLP and redaction for secrets and personal data, and block or log policy hits.
- Isolate session context between users and clear it at session end.
- Redact or exclude prompts and responses from logs, set short retention, and review training and tuning data for sensitive content.
If it already happened
Disable the affected feature or restrict retrieval to a safe scope, and block sessions seen returning secrets or other users' data.
Rotate any credential that appeared in a prompt or output, remove sensitive documents from shared indexes, and purge prompts and responses containing personal data from logs.
Re-enable with per-user retrieval filters, backend-held secrets and output redaction, and notify affected users and regulators where personal data was exposed.
Add leakage test cases to the evaluation suite, keep output DLP and cross-owner retrieval alerts permanently, and add data-flow review to the launch checklist.
Check yourself
1. Why is placing an API key in an LLM system prompt a design flaw?
2. A shared retrieval index returns documents belonging to other customers. Which control addresses the root cause?
3. What is the role of output filtering and redaction in defending against disclosure?