How will the GDPR impact machine learning?(oreilly.com)
oreilly.com
How will the GDPR impact machine learning?
https://www.oreilly.com/ideas/how-will-the-gdpr-impact-machine-learning
8 comments
i want to share your optimism but i wish there were at least some provisions that security/law enforcement/the Nsa etc will respect some of that privacy.
Also, consider making your site compliant.
Also, consider making your site compliant.
Pity the biggest abusers of our privacy - credit reference agencies - will slip through the net. How this racket can be classified as legitimate interest is beyond me.
Legitimate interest or not, if a UK based credit reference agency loses data in the same way that Equifax, for instance, did I think they will find GDPR has the teeth to deal with them.
It's not breaches that worry me, it's their right to exist in the first place. Credit reference agencies put Facebook in the shade when it comes to abuse of privacy and acting as a law unto themselves.
But wouldn't "deal with them" basically mean work with them and give them a lot of help and support while they try to reform their practices and prevent a recurrence, because of course they will demonstrate that none of their screwup was willful?
Agree.
Not very long ago there were a bunch of articles posted on HN about how ML and deep neural networks produce knowledge, without humans being able to explain all the steps or any of the steps, and asked the question if this knowledge is still good.
The GDPR answers this question for business decisions regarding humans: No. It's not good enough. Your business is absolutely allowed to use ML for business decisions, but if a person asks how you reached that decision, you can no longer hide behind the machine, you can't treat it as some kind of oracle in a black box that spits out correct answers without any human knowing how the mind of the oracle works.
And I think that's an important foundation to have going forward. It's like how math tests work, you can't just print the correct answer, you have to show your work as well. Because the alternative is a dystopian "the computer is never wrong" future, and that's gonna suck for everyone.
Not very long ago there were a bunch of articles posted on HN about how ML and deep neural networks produce knowledge, without humans being able to explain all the steps or any of the steps, and asked the question if this knowledge is still good.
The GDPR answers this question for business decisions regarding humans: No. It's not good enough. Your business is absolutely allowed to use ML for business decisions, but if a person asks how you reached that decision, you can no longer hide behind the machine, you can't treat it as some kind of oracle in a black box that spits out correct answers without any human knowing how the mind of the oracle works.
And I think that's an important foundation to have going forward. It's like how math tests work, you can't just print the correct answer, you have to show your work as well. Because the alternative is a dystopian "the computer is never wrong" future, and that's gonna suck for everyone.
You do not need to say how you reached a decision you at most may be required to identify which information was used to reach the decision and how it was obtained.
Businesses may be required to state the reason why a decision was made under other regulation but in those cases those are essentially one liners that do not provide any real explanation e.g. credit score does not meet our treshold.
There is however a real problem with GDPR and ML and by real I mean it’s not you’re getting fined out of business but that no one has a good answer for it and that is if I use your information to train a model and you then tell me to stop and delete all your information do I need to retrain my network? (While most would say no it’s not that simple)
Because for example there are other questions like does a trained model constitutes anonymization or not? And again before you say “yes” there are ways to reconstruct the training data from models e.g. http://www.deeplearningbook.org/contents/autoencoders.html
There’s also the issue of using the data for decision making once a request to stop processing has been issued primarily does processing happens only during training or also during inferencing?
Businesses may be required to state the reason why a decision was made under other regulation but in those cases those are essentially one liners that do not provide any real explanation e.g. credit score does not meet our treshold.
There is however a real problem with GDPR and ML and by real I mean it’s not you’re getting fined out of business but that no one has a good answer for it and that is if I use your information to train a model and you then tell me to stop and delete all your information do I need to retrain my network? (While most would say no it’s not that simple)
Because for example there are other questions like does a trained model constitutes anonymization or not? And again before you say “yes” there are ways to reconstruct the training data from models e.g. http://www.deeplearningbook.org/contents/autoencoders.html
There’s also the issue of using the data for decision making once a request to stop processing has been issued primarily does processing happens only during training or also during inferencing?
> There is however a real problem with GDPR and ML [...]
You are pretty much just repeating the questions and issues from the original article now. It addresses them better than I can.
You are pretty much just repeating the questions and issues from the original article now. It addresses them better than I can.
I’ll go read the article then :)
But again the fact that a blog post says X isn’t enoguh and while I have no doubt that eventually sense will prevail it still doesn’t fix the current problem and that is that the lack of prescriptive guidance from the 28 DPAs is problematic because it’s not a small risk.
What people don’t seem to understand is that most people who have problems with the GDPR don’t have problems with its principles but rather with the fact that it has been rolled out without almost any clear guidelines and rulings. If you can’t be 100% sure that your business model is compliant with the GDPR it can be a pretty big risk to take on and this ambiguity is much more disruptive than the requirements themselves.
BTW the machine learning problem isn’t unique to GDPR, HIPAA also has this problem most companies elected to classify their trained models as patient data to be on the safe side unless they were only trained with public information.
But again the fact that a blog post says X isn’t enoguh and while I have no doubt that eventually sense will prevail it still doesn’t fix the current problem and that is that the lack of prescriptive guidance from the 28 DPAs is problematic because it’s not a small risk.
What people don’t seem to understand is that most people who have problems with the GDPR don’t have problems with its principles but rather with the fact that it has been rolled out without almost any clear guidelines and rulings. If you can’t be 100% sure that your business model is compliant with the GDPR it can be a pretty big risk to take on and this ambiguity is much more disruptive than the requirements themselves.
BTW the machine learning problem isn’t unique to GDPR, HIPAA also has this problem most companies elected to classify their trained models as patient data to be on the safe side unless they were only trained with public information.
[deleted]
China is not going to care. Their complete obliviousness to privacy has made a lot of their machine learning projects move ahead at a breathtaking pace. I was at a presentation by a Chinese surveillance camera company last year and the amount of automated behaviors they can recognize that allow automatic crime detection, etc. is astounding! All the data processing will probably end up running through China as Chinese courts will ignore any GDPR violation fines for local companies, etc.
Why on earth would a Chinese court consider GDPR? GDPR is a EU thing.
If you want to do business with the EU then GDPR will probably apply in some way but not for say US to China relations.
However the GDPR is a damn fine set of standards to live up to - bear in mind that you personally are a person. Would you not want your rights as an individual protected in the same way. Go on ... live a little (and be decent).
GDPR may be what everyone needs.
If you want to do business with the EU then GDPR will probably apply in some way but not for say US to China relations.
However the GDPR is a damn fine set of standards to live up to - bear in mind that you personally are a person. Would you not want your rights as an individual protected in the same way. Go on ... live a little (and be decent).
GDPR may be what everyone needs.
I don't think it matters whether a Chinese court would consider it, the point is that if a Chinese company is doing business in the EU, or to any significant extent with EU nationals, then GDPR is an issue for them. Under threat of fines and worse for their EU subsidiaries.
No idea how much that's really an issue, but I'm pretty sure lots of people order from AliBaba in the EU (I have).
No idea how much that's really an issue, but I'm pretty sure lots of people order from AliBaba in the EU (I have).
If EU citizens go to the Chinese to give up their rights they are probably giving their consent, but good luck providing any sort of service to EU citizens without actually having a legal presence inside the union.
The API call that analyzes what's going on in an image will get sent to China and analyzed there by no-name state owned mega surveillance corp and the result sent back like "35/m/$50k income/55% criminality likelihood/60% likely homosexual/Woman standing next to him is 35% likely to be mistress/subject is currently engaged in 30% flirting behavior 20% shopping for clothes/Target with ad for greek vacation/" and they'll all say it's just a model. Pinky swear there's no personal data that was used to look that up in China.
It doesn't matter if that lookup is done in China or in the EU or elsewhere. To be compliant a company has to be able to list all third parties that get access to personal data it collects, and what they in turn do with it and further third parties that get access to it.
If a company uses a Chinese mega surveillance corp API, they still have to disclose it, and they can't just hide behind a "the computer says no" response if they use the results of that API call to make a business decision. The GDPR gives the data subject the right to know why and how the business decision was made, and gives the subject the right to appeal.
If a company uses a Chinese mega surveillance corp API, they still have to disclose it, and they can't just hide behind a "the computer says no" response if they use the results of that API call to make a business decision. The GDPR gives the data subject the right to know why and how the business decision was made, and gives the subject the right to appeal.
So there was a 5 petabyte learning set we trained a 20 layer deep neural net and a value came out the other end. Here's all the math we did to figure that out (several gigs of float computations). We don't even know what it's doing exactly. What does GDPR say about that? Is feeding data into a deep neural nets illegal in Europe now because they lack explainability?
If you took someone's picture and ran it through neural style, would that be illegal because you couldn't tell them exactly why it painted their nose blue while imitating Leonardo Da Vinci's artistic style? Is Google auto identification of objects in personal images illegal now because they can't explain how a deep neural net works and classified their friend as something non-human by accident? This has actually happened.
If you took someone's picture and ran it through neural style, would that be illegal because you couldn't tell them exactly why it painted their nose blue while imitating Leonardo Da Vinci's artistic style? Is Google auto identification of objects in personal images illegal now because they can't explain how a deep neural net works and classified their friend as something non-human by accident? This has actually happened.
> So there was a 5 petabyte learning set we trained a 20 layer deep neural net and a value came out the other end. Here's all the math we did to figure that out (several gigs of float computations). We don't even know what it's doing exactly. What does GDPR say about that?
That depends completely on what you are using the value for. Recommending five funny articles - GDPR does not apply. Denying an insurance claim - GDPR most certainly applies, and you have to be able to explain what factors went into the decision. You can't have unaccountable oracle boxes.
Note that a perfectly valid answer could be something like "We've analyzed your posts on social media and we've categorized you as having severe anger issues, which is why we're denying this auto collision insurance claim, because you are most likely at fault given situations like this". The GDPR says you have to be able to explain your decision like that. It doesn't forbid you from making these decisions.
> If you took someone's picture and ran it through neural style, would that be illegal because you couldn't tell them exactly why it painted their nose blue
This is not a business decision, don't be ridiculous.
That depends completely on what you are using the value for. Recommending five funny articles - GDPR does not apply. Denying an insurance claim - GDPR most certainly applies, and you have to be able to explain what factors went into the decision. You can't have unaccountable oracle boxes.
Note that a perfectly valid answer could be something like "We've analyzed your posts on social media and we've categorized you as having severe anger issues, which is why we're denying this auto collision insurance claim, because you are most likely at fault given situations like this". The GDPR says you have to be able to explain your decision like that. It doesn't forbid you from making these decisions.
> If you took someone's picture and ran it through neural style, would that be illegal because you couldn't tell them exactly why it painted their nose blue
This is not a business decision, don't be ridiculous.
The "posts" -> "you have severe anger issues" decision sounds like a black box to me with the way NLP models have been going.
The Chinese Corp API seems like the same thing. "We looked at you and decided you look like someone who wants to go on a Greek vacation."
The Chinese Corp API seems like the same thing. "We looked at you and decided you look like someone who wants to go on a Greek vacation."
If I ignore GDPR for the next year or so, waiting until this all simmers down to a level of abstraction below the surface level I care about, will I be any worse off?
I think that Jacques Mattheij has some thoughts about this that you can think about it in the abstractions that he gives. Note that IANAL and YMMV.
EDIT: Also Jacques is not a Lawyer either.
1: https://jacquesmattheij.com/gdpr-hysteria
EDIT: Also Jacques is not a Lawyer either.
1: https://jacquesmattheij.com/gdpr-hysteria
ryanwaggoner(3)
The most rational piece on GDPR I've yet seen.
So of course you are being downvoted.
So of course you are being downvoted.
I have been downvoted for both reasons on HN. I don't feel bad about it at all I think its funny.
I think in the larger cultural context, the public at large is getting more aware of and concerned about data collection. If you don’t have large amounts of PII on EU residents then simmer. If you might, don’t be the one they make an example of.
In the threads like 'Removing Monal from the EU' [1] most folks say you'll get a warning before any enforcement is done. You could try rolling the dice on that.
1 - https://news.ycombinator.com/item?id=17095217
1 - https://news.ycombinator.com/item?id=17095217
you can wait a reasonable amount of time before adequate measures must be taken for appropriate procedures.
Depends on the size of the company you're doing work for, really.
Also on the type of data. If you are using super sensitive data (eg health profiles/histories of people) you better not risk because you are gonna get hit hard.
If your company is okay with the ethical/legal reality of breaking the law for a year, why not just keep at it?
[deleted]
Companies outside of GDPR-land will have a competitive advantage over those inside of it, especially in the area of machine learning.
From the article: "...one of the first major distinctions the GDPR makes about ML models is whether they are being deployed autonomously, without a human directly in the decision-making loop. If the answer is yes—as, in practice, will be the case in a huge number of ML models—then that use is likely prohibited by default."
If users explicitly consent to it, then it's permitted under GDPR, but what percentage of users are going to do that? Most will be hitting the "decline" button by default on all websites, even if a given application of ML will be beneficial to them.
This is a showstopper for most ML in the EU. If you have an ML startup in the EU....either close up shop or move.
From the article: "...one of the first major distinctions the GDPR makes about ML models is whether they are being deployed autonomously, without a human directly in the decision-making loop. If the answer is yes—as, in practice, will be the case in a huge number of ML models—then that use is likely prohibited by default."
If users explicitly consent to it, then it's permitted under GDPR, but what percentage of users are going to do that? Most will be hitting the "decline" button by default on all websites, even if a given application of ML will be beneficial to them.
This is a showstopper for most ML in the EU. If you have an ML startup in the EU....either close up shop or move.
If you are using ML to make automated decisions for say insurance risk then you will have to say so explicitly. There is a pay-off for a customer to agree, and that is convenience. If all insurance companies are doing this then a customer will likely have to consent, or not get a fast insurance decision out of hours that they seem to want. I think the big difference under GDPR is that the user will be able to request that the decision is made by a human, but if this is the case they will not get to set up their insurance at 10PM for tomorrow morning like they can in the automated system. It is completely reasonable that a customer would have to wait a couple of days for this.
In practice I wonder if insurance companies (for instance) are actually doing a live ML evaluation on customers, or do they just use ML to develop some 'bandings' that they fit you in to? They could still do this of course, because they can use anonymous data to build a model.
I don't see where the competitive advantage you mention lies? Yes a company in the US could use ML on US citizens, but not EU ones. Presumably a company in France could set up a server in AWS US-West region and do ML on US citizens, but not EU citizens? Surely it is a matter of where the data-subject lives, not where the company is based?
In practice I wonder if insurance companies (for instance) are actually doing a live ML evaluation on customers, or do they just use ML to develop some 'bandings' that they fit you in to? They could still do this of course, because they can use anonymous data to build a model.
I don't see where the competitive advantage you mention lies? Yes a company in the US could use ML on US citizens, but not EU ones. Presumably a company in France could set up a server in AWS US-West region and do ML on US citizens, but not EU citizens? Surely it is a matter of where the data-subject lives, not where the company is based?
I don't see where the competitive advantage you mention lies?
Well, if you can't perform ML on data about the population that is most accessible to you, you are at a competitive disadvantage to those who can. I can't think of many EU companies that control large amounts of data on US citizens. Though I'm sure they exist, there are many more US companies that would have larger volumes of data on US citizens.
Well, if you can't perform ML on data about the population that is most accessible to you, you are at a competitive disadvantage to those who can. I can't think of many EU companies that control large amounts of data on US citizens. Though I'm sure they exist, there are many more US companies that would have larger volumes of data on US citizens.
So you are arguing that US firms operating in the US are in competition with UK firms operating in the UK? I don't see how they are in competition at all?
Humans are humans, regardless of what country they live in. If I have a larger dataset than you, and I can do more things with it, then I can predict human behavior in a given situation better than you can. That gives me a competitive advantage over you. I may be able to then go compete with you in the EU without ever having to violate the ML provisions of GDPR, because I know what humans - whether in the EU or US - will do.
Granted, there are some country-specific behaviors that this will not apply to, but then nobody will be finding out what those are in the EU because it's now illegal to do there.
Granted, there are some country-specific behaviors that this will not apply to, but then nobody will be finding out what those are in the EU because it's now illegal to do there.
Is it true that users can consent their way into it?
Yes, there's an exception in the GDPR for explicit consent to ML:
The regulation identifies three areas where the use of autonomous decisions is legal: where the processing is necessary for contractual reasons, where it’s separately authorized by another law, or when the data subject has explicitly consented.
But, again, nobody's going to consent - after encountering countless permission popups on websites each and every day, most will just be hitting "decline" without reading what is being asked.
The regulation identifies three areas where the use of autonomous decisions is legal: where the processing is necessary for contractual reasons, where it’s separately authorized by another law, or when the data subject has explicitly consented.
But, again, nobody's going to consent - after encountering countless permission popups on websites each and every day, most will just be hitting "decline" without reading what is being asked.
Users will still have to consent for ML treatment of their data if they want to use ML based services. Base whatever service you offer around it, and they'll have to either accept or get no service. The only ones losing on this are the intermediaries who up to now could slurp whatever data they wanted and sell it to third parties, while now they'll have to offer something of value to the user or get no data.
That's a novel claim, usually security experts are always complaining that people just click "accept" and ignore all the warnings of broken crypto and all that. How foolish they will feel when they realize that all that was missing was putting "GDPR" in the popup and people will automatically stop ignoring warnings!
The difference is that before, they had to click accept in order to receive whatever service they were trying to obtain. Now they don't. Also, the sheer number of consent dialogs they will encounter will rise exponentially.
> The difference is that before, they had to click accept in order to receive whatever service they were trying to obtain. Now they don't.
This doesn't receive nearly enough attention.
This doesn't receive nearly enough attention.
all it takes is a celebrity tweet to tell them to click "decline" and it becomes a meme.
I want to live in a world where the individual has control over their data, but I don't think GDPR will help create that world.
Imagine you are a company that makes $500k+/year off of user data. You are not just going to stop doing it because of GDPR, you are going to spend up to $500k on lawyers figuring out how to get around it.
Imagine you are a company that makes $500k+/year off of user data. You are not just going to stop doing it because of GDPR, you are going to spend up to $500k on lawyers figuring out how to get around it.
That world will come around when a majority of people will understand that running untrusted code on your computer is just madness (i.e. JavaScript). I wish some mainstream browser like Firefox created an easy UI for JS whitelists, and/or a dialog box asking consent for JS execution per FQDN, like it does for location etc. Among the big players, only Google has major reasons to not do such a thing, but FF, IE and Safari can try.
Is it possible to cover GDPR requirements with another line in an EULA that specifies consent?
No, you have to ask and receive consent explicitly for the processing you want to do, not as part of some larger package.
a modal pop-up dialog asking for consent will be required?
yes, that's the intent of the law
I can't wait to issue denial of service requests to every EU company when it comes into effect. I am going to send a nightmare letter to all of them.
Does GDPR apply to animals too? I mean if its going to be hard to manage human data, perhaps replacing some with monkeys would be OK.
GDPR applies to "natural persons", which I think mostly excludes animals (I base this on us killing animals all the time and it doesn't count as murder).
I am extremely proud to be associated with the GDPR (I'm still a European for now). It is an absolute belter of a set of regulations. If you read it, it is actually pretty concise for a legal thingie. It is also very prescriptive which is pretty odd for a legal thingie. It is up there with the "quietly enjoy your own property" basic right that is sort of enshrined in English Law nowadays (IANAL but a search would tell you what I'm on about).
If complying with GDPR is considered a problem then good luck with monetising that weakness.