Data Privacy in the AI Era: The Four Risks Indian Corporations Have Already Inherited
By Privy
Sep 29, 2026

Data Privacy in the AI Era: The Four Risks Indian Corporations Have Already Inherited

Most conversations about AI governance in India treat the risk as future tense. Something to address as regulation matures. Something to revisit when enforcement begins.

The four risks documented in this whitepaper are not future concerns. They are present in the operating environments of Indian corporations right now, accumulating inside deployed models, vendor contracts, and AI pipelines that were built for speed, not for accountability.

According to NASSCOM's 2026 findings, only 30% of Indian businesses have responsible AI practices in place. Only 46% track data provenance. Only 36% take active steps to reduce model bias. These are not figures about companies unaware of AI. They describe companies deploying AI at scale without the legal and governance architecture to do so safely.

"AI is no longer experimental," writes Malcolm Gomes, Chief Operating Officer of IDfy. "It is making decisions that affect customers, employees, businesses, and risk, right now, across sectors. Governance has not always kept pace."

This blog documents the four breakpoints where that gap shows up, each with its legal exposure under the DPDP Act, 2023, and the DPDP Rules, 2025.

How Personal Data Enters the AI Risk Cycle

Personal data does not move in a straight line. From the moment it is collected to the moment it informs a decision, it passes through 4 stages, each with its own actors, systems, and governance requirements.

At Stage 1, data collection, the terms of engagement are set: what data is taken, for what purpose, and on what consent. At Stage 2, data sharing and procurement, that data begins to travel into vendor platforms, third-party AI tools, and external infrastructure. At Stage 3, model training, it is consumed, transformed, and embedded into algorithmic architecture in ways that cannot always be reversed or traced. At Stage 4, deployment and decisioning, it produces outputs that affect real people through credit decisions, hiring filters, fraud flags, and insurance outcomes.

Each stage carries a distinct risk. And each risk maps to specific provisions of the DPDP Act.

ai governance framework

Risk I: The Purpose Limitation Failure

Consider three everyday scenarios. A customer shares a mobile number to receive an OTP. A loan applicant uploads income documents for a creditworthiness assessment. An employee submits a bank account number for payroll. In each case, the act of collection is bounded by the stated purpose and the individual's reasonable expectation that their data will be used for precisely what they were told.

When that same data is later fed into a customer sentiment classifier, an HR attrition tool, or a fraud detection pipeline, it is being processed for a purpose that was never consented to. This is a purpose limitation failure: the invisible migration of personal data from its original, lawful context into AI systems that serve entirely different ends.

This rarely reflects deliberate deception. Data science teams building models draw on whatever internal datasets are accessible and statistically useful, while the legal teams maintaining consent records operate in a separate silo. The two do not systematically speak to each other. The result is that every time a model is retrained on accumulated operational data, there is a real probability that it is being retrained on data whose original consent does not extend to that use.

The scale of this problem is documented. A 2024 audit recognised by MIT, covering 14,000 web domains, found that 45% of the tokens comprising C4, one of the most widely used AI training datasets globally, are now restricted by the Terms of Service of the sites from which the data was scraped. The DPDP Rules, 2025 require specific, granular consent for the use of personal data in AI training pipelines, not the general "service improvement" language that most privacy notices still carry.

The financial exposure across Sections 6 and 8 alone reaches up to ₹250 crore for the security failure and ₹200 crore for the notification lapse. The Star Health breach of August 2024, where 7.24 terabytes of data covering 31.2 million policyholders, including Aadhaar numbers and medical diagnoses, were exfiltrated through AI-powered Telegram chatbots, illustrates what happens when health data collected for insurance purposes migrates into uncontrolled AI-adjacent pipelines.

Risk II: The Vendor Blindspot

Most Indian corporations do not build the AI systems they rely on. A human resources function deploys a third-party screening tool that scores candidates. A finance team routes accounts payable through a vendor-hosted AI platform. A customer service operation runs on a large language model licensed from a global provider. In each case, personal data enters an infrastructure that the deploying corporation does not control, governed only by a contract almost certainly not written with AI-era data flows in mind.

The standard data processing agreement was designed for a world in which a processor received data, performed a defined task, and returned or deleted it. AI vendors do not uniformly operate this way. Many use client interaction data to retrain or fine-tune their own models. Some store prompts and responses in vendor-controlled infrastructure that the deploying organisation has never mapped.

The Cisco 2026 Data and Privacy Benchmark Study, drawing on 5,200 professionals across 12 global markets, found that only 55% of organisations require contractual terms defining data ownership, responsibility, and liability with their generative AI providers, and 70% acknowledge risk exposure from how their vendors handle proprietary and customer data.

When a vendor is involved, the corporate instinct is to assume the vendor bears the risk. Indian law takes the opposite position. Under Section 8(7) of the DPDP Act, the data fiduciary remains responsible for ensuring that data processors, including AI vendors, process data only on instruction and only with adequate safeguards in place. The contract does not transfer the liability. It either documents the safeguards or it exposes their absence.

International enforcement has already made this unambiguous. The Irish Data Protection Commission's €310 million fine against LinkedIn in October 2024 rested on the use of behavioural signals (scroll speed, dwell time, click patterns) to profile users for advertising without valid consent. The mechanism is structurally identical to how Indian enterprises currently deploy behavioural analytics AI. Organisations building unlawful training pipelines today are building assets they may be required to destroy tomorrow.

Risk III: The Inferential Privacy Risk

A widely held assumption in Indian corporate governance is that privacy risk is directly proportionate to the sensitivity of the data an organisation holds. AI attacks the foundation of that assumption through attribute inference.

A trained AI model, when queried with publicly available, non-sensitive inputs, produces sensitive personal attributes as outputs: attributes that were never collected, never consented to, and which the deploying organisation may not realise it is generating. A model trained on a data principal's purchase patterns, utility payment timings, and mobile recharge frequency may infer their cash flow status or financial distress, affecting their access to credit.

The 2025 research report "Disparate Privacy Vulnerability: Targeted Attribute Inference Attack Defenses" demonstrated formally that machine learning models deployed in healthcare and finance contexts can be queried using innocuous, publicly known attributes to infer whether an individual carries a specific health condition, financial vulnerability, or behavioural tendency.

India's digital lending sector provides the most immediate illustration. In 2025, it processed over ₹1.5 lakh crore in disbursals. NBFCs and fintechs have deployed AI credit scoring models trained on transactional, behavioural, and demographic data. These models routinely generate outputs functionally equivalent to health risk scores, financial distress indicators, and social vulnerability profiles, inferences that no borrower consented to and for which the lender holds no lawful basis under the DPDP Act. The RBI's October 2024 bulletin on AI in banking explicitly flagged the risk of AI algorithms producing discriminatory outcomes.

The EDPB's Opinion 28/2024 established the international benchmark: AI models trained on personal data are not automatically anonymous. Under the DPDP Act, inferred sensitive data carries the same legal obligations as collected sensitive data, and the fact that the organisation did not intend to produce it is legally irrelevant.

Retrieval Augmented Generation systems compound this risk. RAG tools ingest sensitive enterprise documents into vector databases where they lose their original access control settings. Documents from HR systems, legal repositories, and financial records become retrievable by any query that matches their semantic content. In 2025, documented incidents included the "EchoLeak" vulnerability in Microsoft 365 Copilot's RAG pipeline, where a crafted email caused the AI to retrieve and exfiltrate sensitive enterprise data without any employee interaction.

Risk IV: The AI Audit Trail Problem

With every passing day, an AI system is quietly making a decision that can adversely affect a person: a loan application rejected, a job candidate filtered out, an insurance claim denied. That person has both a moral and a legal interest in understanding why. Under Section 12 of the DPDP Act, read with Rule 14 of the DPDP Rules, 2025, that interest has the force of law. Data principals have the right to obtain information about processing that affects them and to seek grievance redressal.

In the overwhelming majority of Indian corporate AI deployments today, the answer to that "why" is unavailable because no record exists from which it could be constructed.

The root cause is architectural. Deep learning and ensemble models, the dominant architectures in commercial AI, are not inherently interpretable. They produce outputs without preserving the intermediate reasoning that generated them. Corporations deploying these systems have optimised for output accuracy while neglecting the governance infrastructure needed to reconstruct decision logic when a specific decision, on a specific date, is challenged.

This absence has consequences across three registers. Regulatory: the EDPB's 2025 guidance on LLMs identifies the absence of logging and audit mechanisms as a high-severity privacy risk, concluding that the inability to reconstruct processing activities makes regulatory accountability structurally impossible. Legal: a formal analysis of 168 algorithmic decision-making cases identified a "two-gate" problem in AI contestation: a procedural gate requiring evidentiary access to the decision record, and a doctrinal gate requiring substantive liability rules. Most AI deployments fail at gate one. Institutional: a US court in October 2024 preserved constitutional claims against a state agency for deploying an opaque AI housing classification algorithm, holding that transparency and accountability in AI implementation are necessary to protect individual rights. Judicial tolerance for algorithmic opacity in high-stakes decisions is contracting globally.

In the Indian BFSI context, the consequences are immediate. SEBI's June 2025 consultation paper explicitly requires regulated entities to maintain audit trails and explainability records for AI systems affecting investors. The RBI's FREE-AI framework requires regulated entities to implement model lifecycle documentation sufficient for independent audit.

AI coding tools introduce a specific dimension of this problem. Research from the Chinese University of Hong Kong demonstrated that GitHub Copilot and Amazon CodeWhisperer can be induced, through carefully crafted prompts, to surface hard-coded credentials and personal data memorised from training datasets. Repositories using AI coding assistants exhibit secret leakage rates 40% higher than those developed without them. When a breach later occurs, the absence of logs recording which AI tool generated which code means the organisation cannot trace the origin of the vulnerability.

Why This Sits on the Board's Desk, Not the Compliance Team's

These four risks share a property that makes them particularly dangerous: they are invisible in the ordinary course of business. A purpose limitation failure does not surface on a compliance dashboard unless specifically documented. Vendor accountability gaps appear only when something goes wrong. Inferential privacy exposure accumulates silently inside deployed models. Audit trail deficits become apparent only when a decision is contested, and the record does not exist.

Under the DPDP Act, accountability is explicitly institutional. Liability rests with the data fiduciary as an organisation, not with individual engineers or external vendors. Section 8 assigns the obligations of purpose limitation, security safeguards, and breach notification to the fiduciary itself. Section 9 elevates this further for Significant Data Fiduciaries, through board-accountable DPOs, mandatory audits, and Data Protection Impact Assessments for high-risk processing. Delegation does not dilute liability. Regulatory penalties can extend to ₹250 crore.

"Privacy architecture is no longer a compliance layer," writes Tridib Mukherjee, Chief Data Scientist and AI Officer at IDfy. "It is the foundation that allows AI decisions to be trusted, questioned, and defended."

The four risks, Purpose Limitation Failure, the Vendor Blindspot, Inferential Privacy Risk, and the AI Audit Trail Problem, are inherently cross-functional. Each represents a systemic governance exposure that must be owned at board level. Each requires solutions at the architectural, contractual, and operational layers, not one confined to privacy policies and consent notices.

Building the AI Trust Stack

The AI Trust Stack is the governance model that addresses these four risks across the full AI lifecycle. It structures accountability across 5 stages: data collection, data protection, model creation and validation, safe AI runtime, and agentic AI.

At the data collection stage, the controls are discovery and classification, ownership and lineage establishment, and consent and purpose validation. At the data protection stage: Privacy Enhancing Technologies, data minimisation, and access and retention enforcement. At model creation and validation: fairness and explainability validation, security posture assessment, and deployment readiness. At safe AI runtime: inference-time controls and prompt and execution boundary governance. At the agentic AI stage: verifiable agent identity, delegated authority, decision rights, and execution provenance with continuous oversight.

Every stage produces verifiable evidence for governance, compliance, and responsible AI operations. This is the distinction between assurance narratives and governance artefacts. Directors need the latter.

As Sreenidhi Srinivasan, Partner at Ikigai Law, writes in the whitepaper's closing note: "Standard procurement diligence is no longer sufficient for AI deployments. Organisations must ensure that contracts address training use, data segregation, deletion capabilities, audit rights, and incident reporting obligations in technically enforceable terms. A purpose limitation clause must be supplemented with access controls, governance processes, and infrastructure that support it."

How Privy by IDfy Helps Address These Risks

Privy by IDfy is built to address exactly the four risk categories this whitepaper documents.

Consent governance tracks purpose at the point of collection and flags when data is used in a context the original consent does not cover, addressing the purpose limitation risk before model retraining propagates it. Third-party risk management extends vendor due diligence to AI-specific contractual requirements: restrictions on model training, inference jurisdiction disclosure, and audit rights, closing the vendor blind spot gap.

Privacy impact assessments with a specific trigger for high-risk AI processing, including alternative data inference, address the inferential privacy risk before models go to production. And incident management with tamper-proof, immutable audit trails tied to model version, input feature set, and decision timestamp directly addresses the audit trail problem that makes AI decisions legally indefensible when challenged.

Privy's AI compliance copilot scans live digital journeys and AI pipelines continuously, flagging where personal data enters AI systems without documented consent coverage, where vendor tools are processing data outside their contracted scope, and where AI outputs could constitute inferred sensitive attributes under the Act.

Dpdp compliance

Conclusion

The Cisco 2026 Privacy Benchmark records that 99% of surveyed organisations with mature privacy programmes report measurable business benefits: greater agility, stronger customer trust, and operational efficiency. The organisations that grasp that AI privacy governance is a foundational design question, not a compliance checkpoint, will be the ones that prevail.

"Privacy in the age of AI is not a compliance checkbox bolted onto the end of a product development cycle," the whitepaper concludes. "It is a foundational design question that belongs at the very start of one."

The four risks documented here are already operating. The window to address them architecturally, before a regulator or a contested decision forces the issue, is now. If your organisation wants to assess its exposure across these four risk categories, or see how Privy by IDfy structures the AI Trust Stack for your sector, write to shivani@idfy.com.

FAQ's

What are the four AI privacy risks facing Indian corporations under the DPDP Act? 

Purpose Limitation Failure (data used in AI systems beyond its original consent scope), the Vendor Blindspot (accountability for third-party AI tools that remains with the data fiduciary, not the vendor), Inferential Privacy Risk (AI models generating sensitive attributes from non-sensitive data without a lawful basis), and the AI Audit Trail Problem (the structural absence of records explaining how AI decisions were made).

Does the DPDP Act apply to AI-generated inferences?

 Yes. Under the DPDP Act, inferred sensitive data carries the same legal obligations as collected sensitive data. The fact that an organisation did not intend to produce sensitive inferences is legally irrelevant.

Is a data fiduciary responsible for AI systems operated by a vendor? 

Yes. Under Section 8(7) of the DPDP Act, the data fiduciary remains responsible for ensuring that data processors, including third-party AI vendors, process data only on instruction and only with adequate safeguards in place. The vendor contract does not transfer the liability.

What does an AI audit trail need to contain under DPDPA? 

Under Section 12 of the DPDP Act and Rule 14 of the DPDP Rules, 2025, data principals have the right to obtain information about processing that affects them. A defensible AI audit trail needs to record model version, input feature set, decision timestamp, and decision output in an immutable format that can reconstruct the reasoning behind a specific decision when challenged.

What is the AI Trust Stack? 

The AI Trust Stack is a governance model that structures accountability across 5 stages of the AI lifecycle: data collection, data protection, model creation and validation, safe AI runtime, and agentic AI. At each stage, controls produce verifiable evidence for governance, compliance, and responsible AI operations.

What does the NASSCOM 2026 responsible AI data show? 

Only 30% of Indian businesses have responsible AI practices in place, only 46% track data provenance, and only 36% take active steps to reduce model bias. These are companies deploying AI at scale, not companies unaware of it.

Search Here

Reach out to us

Explore More

What IDfy and MIT’s Research Reveals About Enterprise Privacy & DPDP Rules
DPDP Rules

Feb 23, 2026

What IDfy and MIT’s Research Reveals About Enterprise Privacy & DPDP Rules

How AI Is Transforming Data Privacy, Governance, and Compliance in India
DPDP Rules

Mar 05, 2026

How AI Is Transforming Data Privacy, Governance, and Compliance in India

DPDP for Board: A Leadership Framework
DPDP Rules

Aug 11, 2026

DPDP for Board: A Leadership Framework

Share