Subscribe
Learn Library

When AI Meets Data-Protection Law: Are They Natural-Born Enemies

Content Factory imported article: When AI Meets Data-Protection Law: Are They Natural-Born Enemies.

ads
2026-07-29Go Next Marketer14 min read

A little while ago, a friend asked me: look, AI keeps getting smarter, and it has an ever-bigger appetite — for data. But there happens to be a law whose whole job is to govern data. It says: if you want to use someone's data, you have to tell them, you have to get their consent, and you can't use more than necessary.

So is this law a guardrail for AI, or a straitjacket(a tightening magic headband in the Chinese classic Journey to the West — an inescapable constraint that tightens whenever the wearer steps out of line)around its head?

I said, that's a good question. But the answer may surprise you.

The two of them have to learn to dance together. Guardrail and engine — you can't do without either one.

1. First, a story that sent a chill down my spine

  1. The US presidential election.

A company found roughly 320,000 voters and sent them a survey. A hundred and twenty questions about your personality, your politics. Fill one in and they'd pay you two to five dollars.

Since when does anyone give away money for nothing?

To take the survey you had to log in with your social-media account. The moment you did, the company didn't just grab your account info — likes, posts, friends — it scooped up all of your friends' account info too.

320,000 people unlocked the data of 50 million.

And then? The company used those 320,000 people's "personality plus account behaviour" as teaching material and trained a model. What could the model do? Look at someone's likes and infer their personality — extrovert or introvert, anxious or confident, impulsive or thoughtful.

Once the model was built, it was applied to those 50 million people. Every one of them was given a "psychological profile."

The final step: in a few swing states, find the voters who "might change their vote because of a particular ad" and serve them precisely targeted political ads. The ads didn't appeal to reason; they went after emotions, after biases, after weaknesses the voters didn't even know they had.

Nobody knew they'd been played.

When the story broke, the world was stunned. It thrust a concept that had been quietly academic — profiling — right into the spotlight.

What is profiling? Plainly put: using the data you've already made public to infer the secrets you haven't.

2. Why is AI so hungry

To grasp how frightening this is, you first have to understand one thing: AI is raised on data.

What is machine learning? Let me give you an analogy.

Imagine you want to teach a child to recognise cats. How? You show her ten thousand photos of cats: "This is a cat." Then ten thousand photos of dogs: "This is not a cat." After enough looking, the child learns — next time she sees an animal she's never seen before, she can tell whether it's a cat.

That, in essence, is machine learning. You feed it a pile of examples (this is called the training set), each one labelled with the right answer. A learning algorithm (the "teacher") finds the patterns in the examples and builds a model. That model can then be used to judge new, never-before-seen things.

Interestingly, learning comes in three flavours:

Supervised learning — the teacher guides every step, and every example is labelled with the answer. This is the most popular kind today.

Reinforcement learning — no right answers are given, just a treat for getting it right and a smack for getting it wrong. The system figures it out on its own.

Unsupervised learning — tell it nothing and let it find the patterns and clusters in the data all by itself.

Whatever the flavour, one thing holds: the more examples, the more accurate the learning.

That's why AI is so hungry. It has to eat data. Mountains of it. One person's data isn't enough — it needs data from millions. A single behaviour isn't enough — it needs the accumulation of days and years.

And so a self-reinforcing loop emerges: the smarter AI gets, the more data it demands; the more data it has, the smarter it gets.

This hunger has powered the so-called "big data" industry. Every click you make, every change in your location, every record of your heartbeat — all of it has become feed for AI.

3. The nature of data: you think it's privacy, but it's actually a commodity

This is the unsettling part.

Before AI came along, your likes, your browsing trails, your shopping records — this "data exhaust" — nobody cared much about any of it. Once AI arrived, it suddenly became valuable.

Why? Because through AI, these seemingly trivial pieces of data can be assembled into a complete picture of you — your personality, your income, your health, your politics, your weak spots.

One study stuck with me: knowing only a person's birthday, zip code, and gender was enough to pick them back out of a pile of "anonymised" medical data.

This is called re-identification. You think you've been anonymised, but you haven't. Against an adversary holding oceans of data, a few seemingly innocent bits of information, combined, can pin you down.

A governor of Massachusetts was picked out exactly this way. Researchers took his publicly known birthday, zip code, and gender, matched them against a "de-identified" hospital records database, and bam — a match.

Whether a piece of data counts as "personal data" no longer depends on what the data itself looks like, but on the environment it sits in. This point is crucial, and we'll come back to it again and again.

4. Algorithms inherit bias

AI has another flaw, more hidden: it can pick up bad habits.

How so?

Supervised learning teaches the machine using past examples. But what if those past examples are themselves biased?

Consider this. A company has been hiring for years, and its HR has a heavy bias — it doesn't like hiring people of a certain ethnicity. You take those historical hiring records as a training set and use them to train a résumé-screening system. What does the system learn? It learns "how HR chose in the past," not "who's actually good at the job." It will inherit HR's bias, intact.

More insidiously, bias can spread indirectly. Suppose people of that ethnicity mostly live in certain neighbourhoods. As long as the system uses "home address" as a feature, it can still discriminate — it doesn't need to know your ethnicity; knowing where you live is enough.

A scholar who studies algorithmic discrimination once put it beautifully. The gist: an algorithm assigns a person a score, and that score can ruin a life. Yet when that person wants to push back, a vague "this doesn't seem right" doesn't count for anything. To overturn it, you need hard evidence.

The algorithm itself never has to prove that it's right.

And that's not fair.

Of course, it's not hopeless. Some have pointed out that algorithms are actually easier to regulate than people. Human bias is hidden in the brain, invisible; algorithmic bias is hidden in data and code, which can at least be audited, tested, and corrected. A biased algorithm may still be fairer than an even more biased human.

The key question is: who audits it? How? And once audited, can it actually be fixed?

5. Surveillance capitalism: when your experience becomes someone else's asset

The data-plus-AI combo has spawned a new economic model. It has a name: surveillance capitalism.

What is surveillance capitalism? Let me put it in plain terms.

Over the past two centuries, capitalism has, several times, "turned things that shouldn't be sold into commodities." Human labour became "wages" that could be bought and sold. Land became "real estate" that could be bought and sold. The medium of exchange became "finance" that could be bought and sold.

The fourth time, it came for your experience.

What you clicked on, how long you lingered, what you searched for late at night, who you gave a like to — these were once your private experiences. Now they are recorded, analysed, predicted, and finally turned into "a tradable asset that can predict and influence your behaviour."

You became the product.

What's scarier is that this logic doesn't only run in commerce. Governments can use it too. There's something called a "social credit" system that rolls your financial records, political activity, social connections, and any legal violations into a single score. Score high, and you can get into good schools, live in nice housing, take the high-speed train; score low, and every door slams in your face.

On paper it's there to promote social trust. But think about it: when a system uses carrots and sticks to precisely tune your every action — will you still be yourself? Or will you become a calculating actor, performing for the system?

So here's the question: some AI applications, even if they're accurate, even if they're unbiased, don't deserve to be built.

Some systems, even when technically perfect, should not exist in the first place. Judging sexual orientation from a face, for example. Judging criminal tendency from physiognomy, for example. Asking "is it accurate?" isn't enough. You also have to ask "should it exist at all?"

6. That law: what does it actually say

All right, enough scene-setting. Now to the main event — how exactly does that data-protection law govern AI?

The law is, of course, the GDPR — Europe's General Data Protection Regulation — and its core logic comes down to four principles:

First, you have to be clear about what you want the data for. This is called purpose limitation. You can't say "I'm taking your health data to treat you" and then turn around and sell it to an insurance company.

Second, take as little as you can. This is called data minimisation. If you need an email address to send a verification code, then taking the email is enough — don't quietly make off with my address book too.

Third, you have to let people know you're using their data, and what for. This is called transparency. No operating in the shadows.

Fourth, you can't let a machine make important decisions about a person entirely on its own. This is called the restriction on solely automated decision-making. For example, if a bank turns down your loan, it can't just be because a machine said "no" and that's the end of it. There has to be a way for a human to step in, an explanation to be given, an appeal to be possible.

Sounds reasonable, right?

But here's the rub: each of these four principles sits in tension with how AI actually works.

7. Where the conflict lies

Take purpose limitation first.

What is AI best at? Finding patterns you never thought of in a pile of data collected for some other purpose. Health data collected to treat illness — once AI analyses it, it can predict whether you'll come down with a certain disease next year. And that prediction can be used to set your insurance premium.

When the data was first collected, who could have foreseen it being used that way?

Then there's minimisation.

AI needs to eat huge amounts of data to be accurate. Tell it to "take as little as possible" and it can't learn well. But leave it unconstrained and it wants to grab everything. What's the answer?

And then automated decision-making.

Today, plenty of decisions — credit approvals, résumé screening, insurance pricing — are run by AI in the background. The law says that decisions which are wholly automated and have a significant effect on an individual are, in principle, prohibited. But what does "wholly automated" actually mean? Does a human pressing "confirm" count? And how do you define "significant effect"?

These grey areas are exactly where the conflicts live.

GDPR's four principles each sit in tension with how AI actually works

8. The good news: the conflict isn't a dead end

But hold on. After digging into it, I reached a conclusion that came as a relief: these conflicts are not unbreakable dead ends.

How do you break them?

Look at purpose limitation first. The law actually leaves a door open: data can be used for a new purpose, as long as the new purpose is "compatible" with the original one. What does "compatible" mean? That calls for flexible judgement. Using health data for medical research is generally compatible; using it for individually targeted advertising is where you'd better be careful.

Then look at minimisation. Minimisation doesn't necessarily mean "take less data." It can mean "lower the degree to which the data is tied to your identity." For instance, through pseudonymisationthe data is still there, but your name has been swapped for a token, so it can't be linked straight back to you). That way AI can still learn the patterns, and your privacy is still protected.

As for automated decision-making, the law doesn't ban it outright either. It leaves a number of exceptions: where a contract requires it, where the law allows it, where you've consented — all fine. But the precondition is that there are "suitable safeguards" — at the very least, the person concerned has to be able to find a human to review the decision, to express their view, and to appeal.

Plainly put: the law wants to put rules around AI, not to choke it to death. The rules have give in them, but give isn't the same as having no rules at all.

A guardrail guides AI's path without trapping it

9. The bad news: the rules are too vague

And here's exactly where the trouble lies: the rules exist, but they're far too vague.

What does "compatible" mean? What does "suitable" mean? What counts as a "significant effect"? What's a "reasonable inference"? These words are everywhere in the law, yet they have no precise definition.

For big companies, that's no problem — keep a team of lawyers on retainer and grind through it slowly. But for small companies, for founders, it's a nightmare.

You want to build a product that uses AI to help doctors read scans. You're not sure whether you count as "automated decision-making." You're not sure whether you need to carry out a "Data Protection Impact Assessment." You're not sure whether your "I consent" button is designed freely enough. And then you hear that the fines for breaking the rules can run as high as €20 million, or 4% of global revenue.

Which small company can stomach that?

The result: vague law + massive fines = no one dares to move. The rules were meant to protect people, but vague rules end up protecting the big companies and scaring off the small players.

That's not what the people who wrote the law wanted to see.

10. The way forward

So what do we do? Tear the law up and start again?

No need.

The conclusion from the research is that this law doesn't need a major rewrite. Its principles are right, and its direction is right. What's really missing is concrete guidance on how to make it land.

Who provides the guidance?

The regulators. They understand best how the law applies. They should issue detailed rules, cases, checklists — which scenarios require an assessment, what counts as compliant, what counts as a violation. Ideally in the form of "soft law"(guidelines and best practices that carry authority but aren't binding statutes)— authoritative yet flexible enough to keep up with technological change.

Industry associations and standards bodies. Codes of conduct and certification schemes let honest businesses prove themselves, and let consumers recognise them.

All of society. Which AI applications are acceptable? Which shouldn't exist? This isn't a technical question; it's a values question. The judgement can't be left to engineers alone, and it can't be left to bureaucrats alone. There has to be open debate, and civil society has to take part.

And one more thing matters enormously: collective redress.

Right now, if your data has been abused and you want to sue, you have to spend your own time, your own money, hire your own lawyer. But the harm from data violations is often "small loss per person, but millions of victims."

One person suing has no leverage; millions suing together have power.

Some places have already begun pushing for class-action mechanisms. That's a good direction. Without effective enforcement, even the best law is a paper tiger.

Finally

Back to the question at the start: is this law a guardrail for AI, or a straitjacket(the inescapable constraint from the opening)?

My answer is: it's a guardrail. But right now that guardrail is still a draft. The lines are too thick, the gaps too wide — it neither stops what it should stop, nor frees what it should free.

Drawing it finer, drawing it truer, so that AI can run without hurting people — that is an unavoidable reckoning for our generation.

Because it comes down to a fundamental choice: what kind of society do we want to build? One in which a person's experience is respected, or one that treats a person's experience as raw material?

That choice shouldn't be made by technology. It shouldn't be made by capital. And it shouldn't be made by government alone.

It should be made by all of us, together.

When AI Meets Data-Protection Law: Are They Natural-Born Enemies | Go Next Marketer