Employees are feeding sensitive data to ChatGPT, raising security fears(darkreading.com)
darkreading.com
Employees are feeding sensitive data to ChatGPT, raising security fears
https://www.darkreading.com/risk/employees-feeding-sensitive-business-data-chatgpt-raising-security-fears
348 comments
We saw these same fears with the release of Gmail. Why would you trust your email to Google?!! Aren't they going to train their spam filters on all your data? Aren't they going to sell it, or use it to sell you ads?
Corporations constantly put their most sensitive data in 3rd party tools. The executive in the article was probably copying his company strategy from Google docs.
Yes, there are good reasons for concern, but the power of the tool is simply too great to ignore.
Banning these tools will go the same way as prohibition did in the US, people will simply ignore it until it becomes too absurd to maintain and too profitable to not participate in.
Companies which are able to operate without these fears will move faster, grow more quickly, and ultimately challenge companies restricted to operate without.
Now I think the article should be a wake-up call for OpenAI. Messaging around what is and what is not used for training could be improved. Corporate accounts for Chat with clearer privacy policies would be great and warnings that, yes, LLMs do memorize data and you should treat anything you put into a free product on the web as fair game for someone's training algorithm.
Corporations constantly put their most sensitive data in 3rd party tools. The executive in the article was probably copying his company strategy from Google docs.
Yes, there are good reasons for concern, but the power of the tool is simply too great to ignore.
Banning these tools will go the same way as prohibition did in the US, people will simply ignore it until it becomes too absurd to maintain and too profitable to not participate in.
Companies which are able to operate without these fears will move faster, grow more quickly, and ultimately challenge companies restricted to operate without.
Now I think the article should be a wake-up call for OpenAI. Messaging around what is and what is not used for training could be improved. Corporate accounts for Chat with clearer privacy policies would be great and warnings that, yes, LLMs do memorize data and you should treat anything you put into a free product on the web as fair game for someone's training algorithm.
This is the issue with a tool so powerful, you can't just tell people not to use it, or to use it responsibly. Because there's too much incentive for them to use it.
If it saves hours of a persons' workday, and they're not seeing any of the harm caused from data leakage, there's no incentive for them to not use it.
Which is why a private option is so critical. To not fight against human nature, means providing an ability to use the tool in a safe way.
Which is why a private option is so critical. To not fight against human nature, means providing an ability to use the tool in a safe way.
We published an internal policy for AI tools last week. The basic theme is: "We see the value too, but please don't copypasta our intellectual property until we get a chance to stand up something internal."
We've granted some exceptions to the team responsible for determining how to stand up something internal. Lots of shooting in the dark going on here, so I figured we would need some divulgence of our IP against public tools to gain traction.
We've granted some exceptions to the team responsible for determining how to stand up something internal. Lots of shooting in the dark going on here, so I figured we would need some divulgence of our IP against public tools to gain traction.
I think there's more fear of OpenAI leaking data than say, Airtable or Notion or Github or AWS/S3 or Cloudflare or Vercel or some other company that has gobs of a company's data. Microsoft also has gobs of data: anything on Office and Outlook is your company data — but the fear that they'll leak (intentional or accidental) is somehow more contained.
If we want to be intellectually honest with ourselves, we can either be fearful and have a plan to contain data from ALL of these companies, OR, we address the risk of data leaks through bugs as an equal threat. OpenAI uses Azure behind the scenes, so it'll be as solid (or not solid) as most other cloud-based tools IMO.
As for your data training their data: OpenAI is mostly a Microsoft company now. Most companies use Microsoft for documentation, code, communications, etc. If Microsoft wanted to train on your data, they have all the corporate data in the world. They would (or already could!) train on it.
If there's a fear that OpenAI will train their model on your data submitted through their silly textbox toy, but NOT through training on the troves of private corporate data, then that fear is unwarranted too.
This is where OpenAI should just get a "corporate" tier, charge more for it, and is basically make it HIPAA/SOC2/whatever compliant, and basically do that to assuage the fears of corporate customers.
If we want to be intellectually honest with ourselves, we can either be fearful and have a plan to contain data from ALL of these companies, OR, we address the risk of data leaks through bugs as an equal threat. OpenAI uses Azure behind the scenes, so it'll be as solid (or not solid) as most other cloud-based tools IMO.
As for your data training their data: OpenAI is mostly a Microsoft company now. Most companies use Microsoft for documentation, code, communications, etc. If Microsoft wanted to train on your data, they have all the corporate data in the world. They would (or already could!) train on it.
If there's a fear that OpenAI will train their model on your data submitted through their silly textbox toy, but NOT through training on the troves of private corporate data, then that fear is unwarranted too.
This is where OpenAI should just get a "corporate" tier, charge more for it, and is basically make it HIPAA/SOC2/whatever compliant, and basically do that to assuage the fears of corporate customers.
Not only do you have to worry about employees directly sharing data, but many companies are also just wrappers around GPT. Or they may use your data in the future to roll out new AI services.
While this is not a new problem -- employees share sensitive data with Google all the time -- the data leakage will be more clear than ever. With ads-based tracking and Google search, the leakage was very indirect. With generative AI, it can literally regurgitate memorized documents.
The security risk goes beyond data exfiltration. Folks are already trying to teach the AI incorrect information by spamming it with something like 2+2 = 5.
Data exfiltration + incorrect data injection are super underrated risks to mass adoption of generative AI tech in the B2B world...
While this is not a new problem -- employees share sensitive data with Google all the time -- the data leakage will be more clear than ever. With ads-based tracking and Google search, the leakage was very indirect. With generative AI, it can literally regurgitate memorized documents.
The security risk goes beyond data exfiltration. Folks are already trying to teach the AI incorrect information by spamming it with something like 2+2 = 5.
Data exfiltration + incorrect data injection are super underrated risks to mass adoption of generative AI tech in the B2B world...
No one cares about security because there is no consequence for getting it wrong. Look at all the major breaches ever. And look specifically at the stock price of those companies. They took small short term hits at best.
Worst case the CISO gets fired and then they all play musical chairs and end up in new roles.
Heck, even Lastpass, ostensibly a security company, doesn't seem particularly affected by their breach.
My point is, especially with ChatGPT, where it can reasonably 10x your productivity, most people will be willing to take the risk.
Worst case the CISO gets fired and then they all play musical chairs and end up in new roles.
Heck, even Lastpass, ostensibly a security company, doesn't seem particularly affected by their breach.
My point is, especially with ChatGPT, where it can reasonably 10x your productivity, most people will be willing to take the risk.
We went pretty quickly from:
No way I’m giving Google any of my data! I will use 5 different browsers in incognito mode and never log in.
To ->
Sure I will login with my name and email and feed you as much of my most personal thoughts and data as I can dear ChatGPT!
No way I’m giving Google any of my data! I will use 5 different browsers in incognito mode and never log in.
To ->
Sure I will login with my name and email and feed you as much of my most personal thoughts and data as I can dear ChatGPT!
This is nothing new at all. How many people have Grammarly plugins installed? They are advertising aggressively, so I'd think it is the new hotness. Don't tell me Grammarly is not hoovering up all of the Slack, Word, Docs, and Gmail data that everyone sends it, and holding on for some future purpose. We'll see.
This is one of the reasons Databricks created Dolly, a slim LLM that unlocks the magic of ChatGPT. A homegrown LLM that can tap into/query the datasets of all the data in an organizations Data Lakehouse will be hugely powerful.
I am working with customers that are looking to train a homegrown LLM that they host and have blocked access to ChatGPT.
https://www.datanami.com/2023/03/24/databricks-bucks-the-her...
https://news.ycombinator.com/item?id=35288063
I am working with customers that are looking to train a homegrown LLM that they host and have blocked access to ChatGPT.
https://www.datanami.com/2023/03/24/databricks-bucks-the-her...
https://news.ycombinator.com/item?id=35288063
This is a user led data leak that ranks up there with Facebook and LinkedIn asking for email passwords to “look for your contacts to add”.
In my experience most corporate employees just take the path of least resistance. It is not uncommon for people to paste non public data into websites just to do json formatting, and paste base64 strings to random websites just to decode them. So just telling people not to do something won't accomplish much. Most corporate employees also somehow think they know better than the policy.
Any company that doesn't want to feed data into ChatGPT should need to proactively block both ChatGPT and any website serving as a wrapper over it.
Any company that doesn't want to feed data into ChatGPT should need to proactively block both ChatGPT and any website serving as a wrapper over it.
I believe there were FUD pieces like this when internet search engines were rolled out, and again when social media became popular. I suppose its universal for new technologies.
I had an interview awhile ago at a place where during the phone screen "they can't talk about their tech stack in detail" so I looked on linkedin and figured out their entire tech stack before on the onsite interview. Come on guys, according to linkedin, you have an entire department of people doing AWS with Terraform and Ansible, you don't have to pretend you can't say it in public.
I had an interview awhile ago at a place where during the phone screen "they can't talk about their tech stack in detail" so I looked on linkedin and figured out their entire tech stack before on the onsite interview. Come on guys, according to linkedin, you have an entire department of people doing AWS with Terraform and Ansible, you don't have to pretend you can't say it in public.
ChatGPT Business Edition seems pretty obvious and I'd surprised if OpenAI isn't already working on it. Separate models for each customer, data silos and protection. The infra is already there on Azure.
For fun I once just made a blank from with a submission button and a giant text field.
It was quite amazing what people would submit unprompted, so I'm not at all surprised that people would feed sensitive data into ChatGPT. The next cycle will be that ChatGPT gets - surprise - trained on that data, and may start using fragments of it - which may well still be sensitive enough to cause trouble - as its output.
Don't paste confidential information into a textbox, in fact don't trust anybody or any company with your confidential information unless there is a strong contractual relationship backed up by penalties if it gets broken. And even then: the ultimate responsibility is yours, you may be able to recover some $ for damages but your reputation may well be toast.
It was quite amazing what people would submit unprompted, so I'm not at all surprised that people would feed sensitive data into ChatGPT. The next cycle will be that ChatGPT gets - surprise - trained on that data, and may start using fragments of it - which may well still be sensitive enough to cause trouble - as its output.
Don't paste confidential information into a textbox, in fact don't trust anybody or any company with your confidential information unless there is a strong contractual relationship backed up by penalties if it gets broken. And even then: the ultimate responsibility is yours, you may be able to recover some $ for damages but your reputation may well be toast.
when it first came out and my boss was behind himself about how cool it was, he was feeding it all of his emails with other businesses to have it clean them up. boggled my mind.
Meanwhile over at Github Copilot...
Hahahahahahaha
Hahahahahahaha
"In one case, an executive cut and pasted the firm's 2023 strategy document into ChatGPT and asked it to create a PowerPoint deck."
There's really not much you can do here. This is complete lack of very basic common sense. Having someone like this in your business, particularly at the executive level, is a liability regardless of ChatGPT.
There's really not much you can do here. This is complete lack of very basic common sense. Having someone like this in your business, particularly at the executive level, is a liability regardless of ChatGPT.
Given that they use all the labor of the Internet without attribution, we should assume that they will use every additional drop of data we give to them for their own ends.
This is scary, but it doesn't surprise me even in the slightest. ChatGPT is useful for so many things that it's extremely tempting to convince yourself that you should trust it.
For example, I was having some issues with my LTO-6 drive recently, and I had to finagle through a bunch of arcane server logs to diagnose it. I had the idea of simply copypasting the logs into ChatGPT and having it look at them, and it quickly summarized the logs and told me what things to look for. It didn't directly solve the problem, but it made the logs 100x more digestible and I was able to figure out my problem. It made a problem that probably would have taken 2-3 hours of Googling take about 20 minutes of finagling.
I'm not doing anything terribly interesting or proprietary on my home server, so I didn't really have any reservations sharing dmesg logs with it, but obviously that might not be the case in a company. Server logs can often have a ton of data that could be useful for a competitor (whether it should be there or not), and someone not paying attention to what they're pasting into ChatGPT could easily expose that data.
For example, I was having some issues with my LTO-6 drive recently, and I had to finagle through a bunch of arcane server logs to diagnose it. I had the idea of simply copypasting the logs into ChatGPT and having it look at them, and it quickly summarized the logs and told me what things to look for. It didn't directly solve the problem, but it made the logs 100x more digestible and I was able to figure out my problem. It made a problem that probably would have taken 2-3 hours of Googling take about 20 minutes of finagling.
I'm not doing anything terribly interesting or proprietary on my home server, so I didn't really have any reservations sharing dmesg logs with it, but obviously that might not be the case in a company. Server logs can often have a ton of data that could be useful for a competitor (whether it should be there or not), and someone not paying attention to what they're pasting into ChatGPT could easily expose that data.
This was my first concern when it came to IDE plugins.
It's alright, I just told it that I don't consent to my data being used. Checkmate openAI!
OpenAI. The heist of the century. I am waiting for A.I. generated blockbuster in the near future.
This cycle happens regularly and it seems often times the service provider wises up and charges for extra controls.
Yammer pre-Microsoft and nowadays Blind — lots of “insider” information seemingly posted.
As usage goes up the target size, and opportunity cost, both go up.
Yammer pre-Microsoft and nowadays Blind — lots of “insider” information seemingly posted.
As usage goes up the target size, and opportunity cost, both go up.
I’m curious if anyone’s employer has set up their own LLM. My employer has a couple of A100 sitting around which could easily host a couple instance of 65B LLaMA or Alpaca. Convincing upper management to allow me is the hard part.
Funnily enough I wrote a cautionary comment on this just 2 days ago :
https://news.ycombinator.com/item?id=35299695
https://news.ycombinator.com/item?id=35299695
Let’s not forget that we’re also feeding in all our code into OpenAI Codex.
my understanding of GPT is that the only vector for your data to get "into the model" is if it's used in fine tuning/RLHF. My guess is if you do the thumbs up or thumbs down, the session probably will be, but otherwise probably not. Still wouldn't put in private employer data primarily because of the other exposure risks - it's obviously not stored securely on the OpenAI side. But besides typical IT risk, the big unknown is whether or not the model will spit out what you put into it in somebody else's session. and my understanding is, that's only possible if your conversation is used for RLHF.
I guess another way to say that is, OpenAI (or another service provider with a better security track record) could broker this service in the cloud, with guarantees around not using the session data for RLHF, not storing session data, stronger auth (OpenAI has had a couple of incidents that show that they have pretty lax security in their backend), etc. and could make a killing selling or re-selling ChatGPT to businesses.
I guess another way to say that is, OpenAI (or another service provider with a better security track record) could broker this service in the cloud, with guarantees around not using the session data for RLHF, not storing session data, stronger auth (OpenAI has had a couple of incidents that show that they have pretty lax security in their backend), etc. and could make a killing selling or re-selling ChatGPT to businesses.
Seems like a temporary problem. Surely OpenAI will have a version which runs in a customers public cloud VPC, orchestrated by OpenAI.
First thing that went through my mind as I read the headline was Zuckerberg comments on people posting their info on Facebook.
- there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model
- OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away if it hasn’t already happened from an exploit of their systems (that likely weren’t built with security as the top priority as opposed to scale and performance) that could leak a monumental amount of data from users.
In the most innocent case they could leak the personal info of naive users. But largely if Linkedin is any indication, the business world is filled with dopes who genuinely believe the AI is free thinking and better than their employees. For every org that restricts ChatGPT use, there are fifty others that don’t, most of which have at least one of said dopes who are ready to upload confidential data at a moments notice.
Wouldn’t even put it past military personnel putting S/TS information into it at this point. OpenAI should include more brazen warnings against providing this type of data if they want to keep up this facade of “we can’t release it because ethics” because cybersecurity is a much more real liability than a supervised LM turning into terminator.