# Encrypted but Still Sensitive: Why AI Teams Should Redact at Intake

> Encryption controls who can read sensitive data. Redaction decides whether you keep it at all. For AI evaluation, testing, and analytics, that difference turns raw calls and documents from a liability into a durable asset.

_Topic: Security · 8 min read_

Encryption protects data in transit and at rest. It does not change what is inside the file.

That distinction matters when teams begin building AI products from call recordings, email threads, scanned PDFs, spreadsheets, images, and support documents. Those files can contain card numbers, Social Security numbers, account numbers, dates of birth, addresses, and details that identify a person in combination. If the raw values remain, the files remain highly sensitive, even when they are encrypted.

The better question is not only "Can we secure this data?" It is also "Do we need to keep these specific values at all?"

**In short:** encryption controls access to sensitive data, while redaction reduces the sensitive data you retain in the first place. For AI evaluation, testing, and analytics, that difference determines whether useful source material becomes a long-term liability or a durable asset.

## What is the difference between structured and unstructured data?

Structured data lives in predictable fields. A customer database might have columns named `first_name`, `date_of_birth`, `card_number`, and `account_id`. The schema tells you where to look, so access controls, field-level encryption, tokenization, and masking rules can be applied with reasonable precision.

Unstructured data is different. It is the information that arrives as a raw file or a conversation:

- A customer on a billing call reads a card number aloud in several groups.
- An email thread has a phone number in the signature and an account number in an attachment.
- A scanned PDF includes handwritten notes.
- A spreadsheet holds sensitive values in comments, hidden sheets, or free-text cells.
- An image shows an address or a face, or it carries GPS metadata.

There may be structure inside these files, but it is not exposed as reliable database columns. A person can recognize that "my number is 4111..." is sensitive in the middle of a conversation. A storage system only sees an audio file, and a document repository only sees a PDF.

This is why raw files need a different intake process. You cannot rely on a label such as `card_number` when the value appears in a transcript, a screen capture, a scanned page, or an attachment inside an email chain.

The U.S. Department of Health and Human Services makes a similar point in its [de-identification guidance](https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html). Identifiers must be removed whether they appear in standardized fields or in free text, as long as they are recognizable as identifiers.

## Why does encryption alone leave a retention problem?

Encryption is essential. It helps prevent unauthorized parties from reading data when systems are configured and operated correctly. But encryption does not reduce the sensitivity of the underlying content.

Consider a call recording that contains a full card number, an expiration date, a CVV, and an address. Encrypting the recording protects it, but the complete raw recording must still be treated as sensitive data throughout its lifecycle.



> [Figure: Encryption locks the file but leaves the sensitive values inside it. Redaction removes the values and keeps the context that makes the file useful.]



That has real consequences:

1. **Retention decisions become harder.** Every copy, backup, archive, test environment, and downstream system needs the same level of care, because the sensitive values are still present.

2. **Future use becomes constrained.** A team may want to use historical calls to evaluate an agent, train an extraction model, investigate a recurring issue, or create a benchmark. If the raw data is too risky to make available, the useful context stays trapped with the sensitive values.

3. **Deletion becomes the default answer.** When a raw data store becomes too risky or costly to govern, organizations often delete it. That may reduce risk, but it also removes the evidence and patterns needed to improve systems later.

4. **A breach puts more at stake.** Encryption is a strong control, but it is not a reason to retain data indefinitely. A compromised credential, an overly broad permission, or a misrouted export does far more damage when the underlying files contain information that never needed to be stored.

NIST describes de-identification as a way to limit disclosure risk while preserving the ability to perform meaningful analysis. Its guidance in [SP 800-188](https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf) also stresses defining privacy, usability, and access objectives before choosing a transformation method.

Encryption answers the question "Who can read this?" Redaction helps answer the question "Why are we keeping this at all?"

## What should remain after sensitive values are removed?

Removing sensitive values does not mean removing all value from a file.

For many AI use cases, the model does not need a real card number. It needs to understand that a card number was provided, where it appeared in the interaction, and what happened next. It does not need a person's real Social Security number to learn that an agent should not repeat it back, and it does not need a home address to identify the steps in an intake workflow.

A useful redacted record can preserve all of the following:

- It keeps the type of each sensitive value, such as `[CARD_NUMBER]` or `[DATE_OF_BIRTH]`.
- It keeps the position of that value in the document or conversation.
- It keeps the timestamps of sensitive spans in audio or video.
- It keeps the surrounding language and workflow context.
- It keeps the sequence of the real interaction, including interruptions, corrections, and follow-up questions.

For example, this transcript preserves the pattern without retaining the digits:

> Caller: "It starts [CARD_NUMBER]."
>
> Agent: "Mm-hm."
>
> Caller: "[CARD_NUMBER]. Sorry, that last group was wrong."

That interaction is useful for testing whether an agent recognizes a payment event, follows the right workflow, and avoids exposing sensitive information. The real digits add risk without adding any signal for that evaluation.

This is especially important for realistic test data. Generic sample data rarely captures how people actually speak, write, hesitate, correct themselves, or drop identifiers into the middle of an otherwise routine request. Redacted source material retains the shape of real work. Teams can use it to create evaluation sets, build benchmarks around their own document types, and generate realistic synthetic data without carrying the original values forward.

## How does redaction support data retention and future AI work?

Retention is often treated as a legal or security exercise. It is also a product decision.

If a team stores raw files first and decides how to sanitize them later, it creates a backlog that is difficult to work through. The files may be spread across buckets, file shares, ticketing systems, and vendor tools. Some may lack clear ownership, some may be unreadable by the tools chosen for the cleanup, and some may be deleted before anyone can preserve their useful structure.

Redacting before storage changes that path.



> [Figure: Redacting at intake means sensitive values never reach storage. Anything the system cannot process confidently goes to a person instead of being assumed safe.]



Instead of retaining an encrypted archive full of personal data, teams retain files that are designed for the uses they expect, such as operations review, quality analysis, AI evaluation, model benchmarking, and controlled synthetic data generation. The information that makes a file sensitive is removed at intake, while the information that makes it useful stays available.

This does not eliminate the need for governance. Redaction is one control among many. Teams still need clear retention policies, access boundaries, audit trails, quality checks, and a review path for files that cannot be processed automatically. They also need to consider indirect identifiers and re-identification risk, not only obvious values such as names and card numbers.

What redaction does is reduce the amount of sensitive data that governance must carry forward.

A practical policy starts with a simple principle: retain the minimum information necessary for the intended use. If a model evaluation needs conversation flow, retain the conversation flow. If an analytics workflow needs document categories and extracted fields, retain those. If the original card number, account number, or Social Security number is not necessary, do not store it.

## What should teams do before storing data for AI?

Start at the point where data enters your environment, and ask these questions for each source:

1. **What sensitive values can appear here?** Include direct identifiers, financial data, health information, addresses, credentials, and metadata.

2. **Is the source structured, unstructured, or both?** A CRM export and a scanned attachment in the same workflow need different controls.

3. **What must remain useful after processing?** Define the evaluation, analytics, support, or product use case before choosing what to remove.

4. **Can sensitive values be detected and redacted before storage?** Build this step into intake rather than leaving it for a cleanup project after the data has spread.

5. **What happens when automation cannot confidently process a file?** Flag it for review, and never silently treat an unreadable file as safe.

6. **How will you verify the result?** Keep an audit trail of processing, sample outputs for review, and maintain test cases that reflect the files your teams actually handle.

The goal is not to make every file anonymous in every context. The goal is to keep data useful for a defined purpose while reducing the information that creates unnecessary retention, privacy, and security risk. For the wider picture of protecting AI data once it is spread across systems, see [How Do You Secure AI Application Data When It Is Spread Across Multiple Systems?](/blog/securing-ai-data-across-systems)

## What should you do next?

Before your next AI project stores a new bucket of calls, documents, email, or images, map what sensitive values can enter and decide what the model truly needs to retain. Redact at intake, preserve the context that supports useful evaluation, and keep your future data foundation in your control.

[See how SecurePass redacts data before you store it](/securepass).

## Sources

- [U.S. Department of Health and Human Services, Guidance Regarding Methods for De-identification of Protected Health Information](https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html)
- [NIST SP 800-188, De-Identifying Government Datasets](https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf)
- [NIST SP 800-122, Guide to Protecting the Confidentiality of Personally Identifiable Information](https://csrc.nist.gov/pubs/sp/800/122/final)
