AI in practice / Data subject access
The New DSAR Workflow: How AI Is Changing the Search, Review and Disclosure of Personal Data
Legal teams are using AI to identify personal data in context, accelerate document review and build more repeatable access-request processes. The next challenge is proving that faster answers are complete, accurate and safe to disclose.
A former employee asks for their personal data. The obvious starting points are their personnel file and email account. But the material that explains how they were assessed, managed or dismissed may sit elsewhere: a manager's Teams conversation, a spreadsheet comparing staff, a recorded meeting, or an AI-generated summary of their performance.
That is where the modern data subject access request becomes difficult. Finding someone's name is a search task. Establishing which information relates to them, preserving its meaning and deciding what can lawfully be disclosed requires a much more connected process.
AI is beginning to change that process. Published deployments describe substantial reductions in the volume requiring manual review. Enterprise platforms can also retrieve certain AI interactions, adding a new category of information to the search. The opportunity for lawyers is to combine these capabilities into a workflow that produces a reliable answer and a record of how it was reached.
The legal starting point still determines the technology
"DSAR" is widely used as an operational label, but access rights are not identical worldwide. Under the EU GDPR, Article 15 covers access to personal data and specified information about its processing. Article 12 generally requires action without undue delay and within one month, with a possible two-month extension where necessary because of complexity or the number of requests, subject to notification requirements. EU GDPR, Articles 12 and 15
The UK now expressly uses a reasonable and proportionate search standard. The ICO's updated guidance also explains circumstances in which reasonably necessary clarification can pause the response clock; it does not allow an organisation to force someone to narrow their request. California follows a different framework: the CPPA describes a general 45-calendar-day substantive response period for requests to know, with confirmation of receipt within ten business days. A global workflow must therefore identify the applicable regime before calculating deadlines or applying restrictions. ICO access guidance; CPPA FAQs
The distinction between personal data and documents is equally important. In C-487/21, the Court of Justice explained that a copy must faithfully and intelligibly reproduce the personal data. Document extracts, and sometimes entire documents, may be necessary to make the right effective. For AI-assisted disclosure, the implication is significant: a fluent summary should not be assumed to satisfy the obligation to provide a copy. CJEU explanation of the judgment
Real deployments are moving beyond experimentation
A Relativity-published case study describes ENS's specialist intelligENS division working with London employment firm BDBF on three DSAR matters in early 2026. According to the account, the teams processed more than 130,000 documents and achieved review-set reductions of at least 94% in each matter. Their approach combined personal-information identification, eDiscovery methods and contextual analysis using matter-specific prompts. These are provider-published results, rather than an independently audited benchmark, but they identify actual practitioners and a concrete workflow. ENS and BDBF case study
Another Relativity customer account describes an unnamed global life sciences organisation reporting a 50% reduction in DSAR review time and an average saving of 85 hours per request. Its process included targeted sampling and checking AI-generated rankings and rationales. The customer's anonymity limits external verification, but the example illustrates a useful operating model: automated assessment followed by structured validation. Life sciences case study
Those figures require careful reading. A smaller review set does not, by itself, demonstrate a complete response. Precision measures how much of the material identified as relevant actually is relevant; recall measures how much relevant material the process found. A high precision figure can coexist with significant omissions. Lawyers evaluating a case study should ask how excluded material was tested, what the starting dataset contained and which stages the reported savings include.
Start with a map of the person and the systems
Before launching searches, build an identity profile: names, former names, email addresses, employee or customer identifiers, relevant accounts and known aliases. Keep confirmed identifiers separate from possible matches. Two people sharing a name should not become one subject simply because a model finds the association plausible.
Then create a source register identifying the systems likely to hold responsive information, their owners, collection methods, relevant retention settings and search limitations. Include HR, customer support, collaboration tools and business databases where appropriate. Establish preservation requirements promptly, considering the access request and any separate dispute-related obligations. Record inaccessible sources and processing failures so that an apparently successful export does not conceal an incomplete collection.
For UK requests, the ICO expects organisations to justify why a search would be unreasonable or disproportionate and to search other information within scope even where a particular search cannot be justified. Operationally, this makes a documented source register more useful than a bare list of search terms: it records both what was examined and the reasons for material limitations. ICO guidance on finding and retrieving information
AI-generated records now belong in that assessment. Microsoft documents eDiscovery searches for prompts and responses from supported AI applications. Its guidance also treats Copilot memory as a distinct record type and notes that deleting a conversation does not delete the associated memory. This is a practical warning against treating "the chat history" as the entire AI record. Coverage depends on the application, configuration, permissions and available retained data. Microsoft documentation on searching AI application data
The resulting question for a DSAR team is broader than "Did this person use AI?" A manager may have used AI to analyse information about them. A system may hold a generated assessment, meeting summary or saved inference concerning them. Whether a particular record falls within an access response requires legal analysis, but the collection plan should at least ask where those records exist. A search limited to the requester's own account may miss relevant material held elsewhere.
Search should combine identifiers with context
Exact searches remain valuable for account numbers, addresses and known names. Contextual analysis can help identify references such as "the regional manager returning from leave", where surrounding material establishes who is being discussed. A sensible design uses conventional search, metadata and AI-assisted classification together, with reviewers checking uncertain associations.
This is also where legal instructions matter. "Find documents relevant to the dispute" and "identify personal data concerning this requester" are different tasks. A system tuned only to a grievance or litigation narrative can miss ordinary personal information within the request. Review instructions should define the access task independently, even when the same material also matters to a dispute.
Before relying on bulk analysis, test those instructions against a human-reviewed set containing difficult examples: shared names, indirect references, short messages, mixed-person records, poor scans and documents in relevant languages. Keep a separate validation set that was not used to tune the instructions. This provides evidence of performance on the matter rather than confidence based on a convincing demonstration.
Reducing repetition should preserve relationships. Deduplication and email threading can remove unnecessary review, but their configuration deserves attention. Identical files can have different custodians or locations; similar messages can contain different recipients or attachments. Chat exports may lose meaning when individual messages are detached from the conversation. Preserve those relationships even when presenting reviewers with fewer items.
A useful review record should connect each proposed decision to the source: document or message identifier, relevant passage, classification, reviewer outcome and any reason for escalation. AI explanations can help reviewers navigate, but they are generated outputs too. The underlying passage remains the evidence. Unreadable files, failed OCR and truncated content should enter an exception queue rather than quietly receive a "not relevant" label.
Redaction is where automation meets legal judgment
Entity detection can propose names, addresses, financial details and other sensitive information for review. It cannot establish, simply by finding another person's name, that the information must be withheld. In UK employment guidance, the ICO explains that information may concern both the requester and someone else, and that disclosure can depend on consent or whether disclosure without consent is reasonable in the circumstances. Exemptions require case-specific justification. ICO guidance for employers
The practical design should separate detection from the disclosure decision. Let software locate potential issues and propagate approved redactions where context genuinely matches. Route privilege questions, mixed personal information and disputed restrictions to appropriately qualified reviewers. Record the reason for withholding, rather than leaving an unexplained black box on the page.
Technical redaction then needs its own check. A document that looks redacted may still expose information through selectable text, comments, attachments or hidden content. The ICO's disclosure guidance highlights these risks. Test the actual exported files, including text extraction and embedded content, as well as their visual appearance. ICO guidance on safe document disclosure
Solutions address different parts of the problem
The following products illustrate distinct capabilities documented by their providers; this is a workflow comparison, not a hands-on ranking.
| Solution | Documented capability | What a legal team should test |
|---|---|---|
| OneTrust DSR Automation | Intake, identity verification, discovery, redaction and secure response workflows. | Which stages work on the organisation's actual systems, and how complex exceptions reach legal review. |
| DataGrail Request Manager | Request management and privacy-request automation, with an agent option for internal systems. | Identity matching, connector coverage, failed retrievals and the effort required to maintain internal integrations. |
| Microsoft Purview | eDiscovery searching of supported AI interaction data, alongside its broader compliance environment. | Actual tenant coverage, retained records, licensing, permissions and export usability. |
| RelativityOne with aiR for Review | AI-assisted review used in documented DSAR matters. | Matter-specific classification quality, validation of exclusions and the complete review-to-production process. |
| RedactAgent | AI-assisted identification and redaction of sensitive information across PDFs, scans, emails, office files and spreadsheets, with reviewer approval, audit history and redacted exports. | Detection quality on representative documents, preservation of context, reviewer control, audit completeness and whether exported redactions permanently remove the underlying information. |
RedactAgent offers an option for the review and redaction stage. Its published workflow places AI-generated suggestions before a human reviewer, who can inspect, change or remove proposed marks before production. The platform records suggestions, reviewer decisions and export activity. For DSAR teams, the relevant evaluation is whether this makes redaction more consistent and easier to audit across their actual document types. These are documented product capabilities, not independently verified performance findings. RedactAgent product overview
A high-volume consumer business may gain most from reliable identity matching and connected retrieval across operational systems. An employment practice handling dense email and chat collections may gain more from sophisticated review and production tools. These capabilities can sit in the same workflow, but buying one does not automatically provide the others. Evaluate the handoffs as closely as the individual products.
Agents need defined authority
"Agent" also needs a precise definition. DataGrail, for example, documents a Request Manager Agent that extends request automation to internal systems. The term describes an integration component; it should not be read as proof that a generative AI system autonomously makes legal decisions. Ask any provider what its agent can access, what it can change and which actions require approval. DataGrail agent documentation
For more autonomous workflows, useful early tasks include preparing searches for approval, tracking collection progress, flagging missing sources and assembling draft response materials. Set explicit approval gates before narrowing scope, applying exemptions or releasing information. Treat collected documents as evidence, never as instructions to the system, and constrain the tools the automation can invoke. These are design recommendations for retaining control as more steps become automated.
Confidentiality must also be assessed across the complete processing path. Examine where source documents, prompts, extracted text and logs travel; who can access them; whether they are used for model training; and how retention and deletion operate. A no-training commitment answers only one of those questions. Depending on the deployment and jurisdiction, processor terms, transfer arrangements and a data protection impact assessment may also need attention. EU GDPR, Articles 28, 35 and Chapter V
Validate what could be missing
Checking only the records an AI system selects tells a team little about the records it excludes. Combine representative sampling with targeted checks of difficult categories, and distinguish the conclusions each supports. Where statistical claims are made, record the sampling method and uncertainty. A model's self-reported confidence percentage is not a measured error rate.
Keep the search strategy, collection logs, processing exceptions, instruction versions, available model-version details, validation findings and final approvals together. Exact reproduction may be difficult when hosted models change, so retain the outputs actually relied upon. Measure total cost through final delivery, including senior review and correction work, rather than reporting only the speed of the first automated pass.
The response should be usable by the person receiving it. An illustrative employment matter might end with a readable bundle containing personal information from HR records, messages and relevant AI-generated assessments, with sufficient context to understand each item. An accompanying explanation should supply the information required by the applicable regime. Under the EU GDPR, that includes matters such as processing purposes, data categories, recipients and retention information, not simply a collection of files. EU GDPR, Article 15
Before release, a reviewer should check the final package, recipient and delivery method. The internal file should separately explain the searches performed, restrictions applied, unresolved limitations and quality checks. These two outputs serve different readers: the requester needs intelligible access; the organisation needs an accountable record of its response.
For a team beginning now, the most useful pilot is one repeatable category of request and a controlled dataset it is authorised to use. Establish a baseline, compare conventional and AI-assisted work, investigate omissions and rework, then expand only where the evidence supports it. There is little value in accelerating classification if collection remains incomplete or the final export still needs extensive repair.
The strongest DSAR practice will be able to show where the information came from, how the search was tested and why disclosure decisions were made. AI can make that practice faster and more consistent. Its most valuable contribution will be giving lawyers more time to resolve the difficult questions, while leaving a clearer record of the answers.