NIST AI Risk Management Framework(nist.gov)
nist.gov
NIST AI Risk Management Framework
https://www.nist.gov/itl/ai-risk-management-framework
5 comments
I am generally frightened by the readiness of people to send data to solutions like ChatGPT or by running open source projects that result in running "AI agents" that execute arbitrary code on one's machine. What can we do to drive more awareness about running these experiments in a responsible way, like a sandboxed environment?
In my opinion this type of comment is the core problem with AI alignment these days.
The current GPT models are not going to "wake up" or accidentally take over the world. This is very obvious.
Forcing everyone who is conscientious to copy and paste the output of these relatively weak models is not going to save us from anything.
If you want people to take AI regulations seriously, you have to do it in a reasonable way. Which means understanding the nuance of the dangers.
Executing "arbitrary code" that InstructGPT models outputs based on your instructions is not a danger.
People need to be able to distinguish different levels of autonomy and speed.
It's when we get to high levels of autonomy and performance that we run into danger. Which we are really at the cusp.
But when you fail to differentiate between systems that are obviously not dangerous and future/more autonomous agents that possibly are, that makes it impossible to take the concerns seriously. And actually harder for people to understand the problems with full autonomy and superintelligent models.
The current GPT models are not going to "wake up" or accidentally take over the world. This is very obvious.
Forcing everyone who is conscientious to copy and paste the output of these relatively weak models is not going to save us from anything.
If you want people to take AI regulations seriously, you have to do it in a reasonable way. Which means understanding the nuance of the dangers.
Executing "arbitrary code" that InstructGPT models outputs based on your instructions is not a danger.
People need to be able to distinguish different levels of autonomy and speed.
It's when we get to high levels of autonomy and performance that we run into danger. Which we are really at the cusp.
But when you fail to differentiate between systems that are obviously not dangerous and future/more autonomous agents that possibly are, that makes it impossible to take the concerns seriously. And actually harder for people to understand the problems with full autonomy and superintelligent models.
> In my opinion this type of comment is the core problem with AI alignment these days. [...] The current GPT models are not going to "wake up" or accidentally take over the world. This is very obvious. [...] Executing "arbitrary code" that InstructGPT models outputs based on your instructions is not a danger.
I meant to comment on the security aspects of sending confidential information to models and letting them generate output that will be immediately executed on your own machine. People also chain models together and let them generate output that will be interpreted as the next model's input. If done without guardrails, this is done ideally in a sandboxed environment.
Too often, I see this type of code just ran on one's machine. As the projects attract less technical people, they may not be aware of the dangers of running such code locally and can see a portion of their data being wiped or exposed.
I meant to comment on the security aspects of sending confidential information to models and letting them generate output that will be immediately executed on your own machine. People also chain models together and let them generate output that will be interpreted as the next model's input. If done without guardrails, this is done ideally in a sandboxed environment.
Too often, I see this type of code just ran on one's machine. As the projects attract less technical people, they may not be aware of the dangers of running such code locally and can see a portion of their data being wiped or exposed.
> The current GPT models are not going to "wake up" or accidentally take over the world. This is very obvious.
It's pretty obvious that GPT-4 isn't consistently coherent enough to carry out complex plans or to do serious scientific research. There are several other major shortcomings which also prevent this from being viable, too. And of course, it's possible that some of these shortcomings are very hard to "fix." We've had AI winters before, and self-driving seems to have stalled.
However, I strongly suspect that it's theoretically possible to build a machine that "thinks" as well as a human. And I no longer have any solid idea how far we are from that point, especially if we make certain architectural changes. We might be closer than we think. Or at least we should stop assuming that next-gen models are obviously safe.
nVidia's hope to increase training performance 1 million times in 10 years seems potentially dangerous to me. Anthrophic's plan to spend over $1 billion training next gen models seems awfully sketchy as well.
It's pretty obvious that GPT-4 isn't consistently coherent enough to carry out complex plans or to do serious scientific research. There are several other major shortcomings which also prevent this from being viable, too. And of course, it's possible that some of these shortcomings are very hard to "fix." We've had AI winters before, and self-driving seems to have stalled.
However, I strongly suspect that it's theoretically possible to build a machine that "thinks" as well as a human. And I no longer have any solid idea how far we are from that point, especially if we make certain architectural changes. We might be closer than we think. Or at least we should stop assuming that next-gen models are obviously safe.
nVidia's hope to increase training performance 1 million times in 10 years seems potentially dangerous to me. Anthrophic's plan to spend over $1 billion training next gen models seems awfully sketchy as well.
Self driving, I think is where we got it backwards. To drive correctly we need AGI because 'uncommon' but not 'rare' events are avoided by drivers all the time and in a manner you have to understand how the world works and how people think.
LLMs, especially the multi-modal ones go much further to making self driving possible, that is if we ever get to the point they can incorporate and act on data fast enough. We may still have a hardware problem for some time there.
---
But on your second point I am with you. There is nothing obvious about at what level emergent behavior occurs in the models. Just keep adding parameters and new things keep popping out with no real means to determine where and when seems kinda sketch as we push these things past the exa scale.
LLMs, especially the multi-modal ones go much further to making self driving possible, that is if we ever get to the point they can incorporate and act on data fast enough. We may still have a hardware problem for some time there.
---
But on your second point I am with you. There is nothing obvious about at what level emergent behavior occurs in the models. Just keep adding parameters and new things keep popping out with no real means to determine where and when seems kinda sketch as we push these things past the exa scale.
> However, I strongly suspect that it's theoretically possible to build a machine that "thinks" as well as a human
Not any time soon. Modern "AI" research built on completely wrong assumptions about how neurons work from the middle of XX century. It turned out they are much more complex than that (long story short every single neuron is analogous to the "neural network" as AI researchers understand them). See more at https://www.youtube.com/watch?v=hmtQPrH-gC4
Not any time soon. Modern "AI" research built on completely wrong assumptions about how neurons work from the middle of XX century. It turned out they are much more complex than that (long story short every single neuron is analogous to the "neural network" as AI researchers understand them). See more at https://www.youtube.com/watch?v=hmtQPrH-gC4
Why would you think that the machine implementation of “thinking” would need to mimic the biological implementation?
Your brain does 1+1 = 2 very differently than an integrated circuit, but the end output is the same.
Your brain does 1+1 = 2 very differently than an integrated circuit, but the end output is the same.
Because it shows that real brain complexity is much larger than we thought, which means AGI is farther down the road.
The current GPT models are not going to "wake up" or accidentally take over the world. This is very obvious.
Was it obvious to you that you’d see a paper called “Sparks of an AGI” after the release of ChatGPT 4 ?
Was it obvious to you that you’d see a paper called “Sparks of an AGI” after the release of ChatGPT 4 ?
I think unfortunately my comment got some upvotes based on an imprecise interpretation of what I actually said. Probably if people understood what I was actually saying it would have been buried.
Also part of the problem is the word AGI and having total conflation inside of that.
I meant SPECIFICALLY the current GPT versions and architecture.
In my opinion, GPT-4 is extremely intelligent in a very general way and this is incredibly obvious.
What GPT-4 is NOT, by itself, is autonomous, or many other human/animal characteristics that people conflate with intelligence.
GPT does not have full autonomy. It is not a person. It is not alive. It does not feel. It does not form its own goals. It does not feel, or have survival instincts, or a high bandwidth stream of sensory data. In these and other ways it is quite different from animals and people. It is not a digital person.
However, people are quite stupidly moving forward with trying to create digital people that imitate all animal characteristics, apparently with the idea that they will enslave them, with the thought process being that they will not be able to do useful work without having all of these animal (human) characteristics.
Besides that problem, people are not taking into account the different levels of speed and other performance measurements of the future systems. Future quite possibly being this year. Or especially I think it is critical for the new compute-in-memory substrates or any advances that increase the speed or other IQ measurements beyond say one order of magnitude of human to be highly regulated. Systems that are twice as smart and say 100 times faster at thinking than humans absolutely can be dangerous.
We achieved the architecture for general purpose intelligence in 2017 and in 2018 I suggested we would have AGI in 2019. Unfortunately now AGI is quite vague and means things like digital person or digital God..I wrote quite a lot but unfortunately deleted some of the explanation.
General intelligence versus personhood or Godhood or alive, or autonomous or having survival instincts or instant adaptability etc. are not the same thing.
Also part of the problem is the word AGI and having total conflation inside of that.
I meant SPECIFICALLY the current GPT versions and architecture.
In my opinion, GPT-4 is extremely intelligent in a very general way and this is incredibly obvious.
What GPT-4 is NOT, by itself, is autonomous, or many other human/animal characteristics that people conflate with intelligence.
GPT does not have full autonomy. It is not a person. It is not alive. It does not feel. It does not form its own goals. It does not feel, or have survival instincts, or a high bandwidth stream of sensory data. In these and other ways it is quite different from animals and people. It is not a digital person.
However, people are quite stupidly moving forward with trying to create digital people that imitate all animal characteristics, apparently with the idea that they will enslave them, with the thought process being that they will not be able to do useful work without having all of these animal (human) characteristics.
Besides that problem, people are not taking into account the different levels of speed and other performance measurements of the future systems. Future quite possibly being this year. Or especially I think it is critical for the new compute-in-memory substrates or any advances that increase the speed or other IQ measurements beyond say one order of magnitude of human to be highly regulated. Systems that are twice as smart and say 100 times faster at thinking than humans absolutely can be dangerous.
We achieved the architecture for general purpose intelligence in 2017 and in 2018 I suggested we would have AGI in 2019. Unfortunately now AGI is quite vague and means things like digital person or digital God..I wrote quite a lot but unfortunately deleted some of the explanation.
General intelligence versus personhood or Godhood or alive, or autonomous or having survival instincts or instant adaptability etc. are not the same thing.
Not the poster you asked the question to but:
I think we are going to be building and using 'sub-AGI' models for a very long time, probably well after we (probably achieve) full AGI. I think these 'sub-AGI' models are going to radically change the way we think about intelligence in the sense that we are going to be better able to nail down what we find special about human intelligence while at the same time discovering that it is a narrow and somewhat arbitrary subset of what we consider advanced intelligence.
ChatGPT seems to me like a very advanced Markov chain, which isn't to denigrate it, the surprising part is how useful and effective it is which to me shows we have been thinking about intelligence the wrong way. IMO an important portion of our brain probably works in a very analogous way but you could build as advanced a version of ChatGPT as you like and it wouldn't become 'sentient' or go HAL9000 on us.
There is definitely a danger of someone taking ChatGPT9 and telling it to create you computer viruses to take down digital infrastructure or telling it to trade in a way to crash the stock market but that's going to happen when all the zero-day exploits are detonated in the lead up to WW3 anyways and all things considered that has a higher chance of happening sooner than ChatGPT9 or the equivalent is used to do it. Instead of getting worried over rogue AI we should probably just harden our infrastructure.
I think we are going to be building and using 'sub-AGI' models for a very long time, probably well after we (probably achieve) full AGI. I think these 'sub-AGI' models are going to radically change the way we think about intelligence in the sense that we are going to be better able to nail down what we find special about human intelligence while at the same time discovering that it is a narrow and somewhat arbitrary subset of what we consider advanced intelligence.
ChatGPT seems to me like a very advanced Markov chain, which isn't to denigrate it, the surprising part is how useful and effective it is which to me shows we have been thinking about intelligence the wrong way. IMO an important portion of our brain probably works in a very analogous way but you could build as advanced a version of ChatGPT as you like and it wouldn't become 'sentient' or go HAL9000 on us.
There is definitely a danger of someone taking ChatGPT9 and telling it to create you computer viruses to take down digital infrastructure or telling it to trade in a way to crash the stock market but that's going to happen when all the zero-day exploits are detonated in the lead up to WW3 anyways and all things considered that has a higher chance of happening sooner than ChatGPT9 or the equivalent is used to do it. Instead of getting worried over rogue AI we should probably just harden our infrastructure.
Hardening infrastructure is a trillion dollar problem that a huge number of people don't believe in, or believe it's a waste of money to spend even more on. And that is before we make cutting edge thinking machines that could attack it.
Also, a paperclip maximizer could in theory go 'HAL' on you without intent or malice. Sentience isn't needed. And at the same time as we're working on embodiment models, what we could consider human like sentience could be far closer than we expect.
Also, a paperclip maximizer could in theory go 'HAL' on you without intent or malice. Sentience isn't needed. And at the same time as we're working on embodiment models, what we could consider human like sentience could be far closer than we expect.
[deleted]
>Executing "arbitrary code" that InstructGPT models outputs based on your instructions is not a danger.
This is a grossly irresponsible and categorically false statement, in very obvious ways.
This is a grossly irresponsible and categorically false statement, in very obvious ways.
Why would ChatGPT provide me with dangerous code if I prompt it benignly? I just the other day used ChatGPT to create a webscraper for a friends project and it provided me with an only slightly wrong implementation to scrape data from a specific website. 15 minutes of reviewing the code and editing literally three lines and I was done. Unless it was trained on a bunch of webscraping code that also contained malicious code I just don't see how it could happen.
If anything they could enable an option to have ChatGPT review the code for you and let you know that it hasn't accidentally thrown in some fishy code.
If anything they could enable an option to have ChatGPT review the code for you and let you know that it hasn't accidentally thrown in some fishy code.
>Why would ChatGPT provide me with dangerous code if I prompt it benignly?
Because it's a black box that's sensitive to unexpected inputs. Humans create bugs too, but an unaligned AI doesn't have the sense to be careful of certain classes of bug.
Because it's a black box that's sensitive to unexpected inputs. Humans create bugs too, but an unaligned AI doesn't have the sense to be careful of certain classes of bug.
The burden of proof for executing arbitrary code as input should always be on proving that it doesn't violate your security or reliability model. You shouldn't need to assume anything about that code other than that you're able to configure the runtime before it executes. Whether it comes from ChatGPT or users shouldn't matter, you should assume some of those programs were written by hackers and that you don't know which ones.
Even if you trust ChatGPT you should do this. You should assume that one day the hackers will intercept your connection to OpenAI and hand you malware, and they should either fail or be forced to use a zero day exploit against eg Firecracker (modulo your threat model and such).
Even if you trust ChatGPT you should do this. You should assume that one day the hackers will intercept your connection to OpenAI and hand you malware, and they should either fail or be forced to use a zero day exploit against eg Firecracker (modulo your threat model and such).
If they weren't paying attention to Black Mirror, the BSG reboot, the Terminator franchise, the Replicator arc of Stargate, every episode of Star Trek where the computer/Data/the holodeck/nanites went wrong, Ex-Machina, The Matrix franchise, Age of Ultron, basically any Asimov AI story, Blade Runner, 2001, or any folk tale about being careful what you ask of powers that take you literally from Midas to Fantasia…
"Oh, they're just fictional!"
…then I suggest attaching a rusty metal spike to the keyboard.
It won't do anything, it's easy to avoid, it'll just sit there looking dangerous.
Hopefully that would be enough just by itself.
And yes, this is a reference to a similar suggestion for making drivers pay more attention behind the wheel.
"Oh, they're just fictional!"
…then I suggest attaching a rusty metal spike to the keyboard.
It won't do anything, it's easy to avoid, it'll just sit there looking dangerous.
Hopefully that would be enough just by itself.
And yes, this is a reference to a similar suggestion for making drivers pay more attention behind the wheel.
There is no sci-fi story that details a realistic scenario of AI existential risk, because sci-fi stories are written by and for humans. The potential danger of AI comes from the combination of superhuman abilities and profoundly alien mind. No human can predict any specific behavior of a superhuman intelligence, only that it will probably succeed in its goals, whatever those goals happen to be. And out of all possible goals, only a small proportion are compatible with life continuing to exist. This is especially true when you consider only simple goals, and simple things are generally easier to make.
Sci-fi is a distraction. There can be no heroic human resistance like in the stories. If an AI is intelligent enough to be dangerous, it's intelligent enough to conceal its intentions until it's too late. The only way to beat a superhuman AI is by not making it in the first place.
Sci-fi is a distraction. There can be no heroic human resistance like in the stories. If an AI is intelligent enough to be dangerous, it's intelligent enough to conceal its intentions until it's too late. The only way to beat a superhuman AI is by not making it in the first place.
Yes, an AI apocalypse is not depicted in any science fiction because it would make for terrible fiction. Good stories depict relatable characters engaged in a battle between near-equals that ends in victory after terrifying odds. Not "someone saying 'oops' and then everyone falling over dead", which is the Yudkowsky scenario.
maybe something like the BBC Threads film is the way to go?
plot of threads (set in Sheffield, UK):
plot of threads (set in Sheffield, UK):
- russians/americans disagree on something in iran
- initial nuclear attacks around base in iran
- uneasy 2 day ceasefire
- worldwide nuclear exchange occurs (a relativly small one)
- graphic depiction of nuclear attack
- people suffering horribly
- ends with scene in hospital of character we've followed in labour, resulting in deformed stillbirth
plot of artificial super intelligence: - tech company creates AI
- idiot CEO enters badly thought out goal in an attempt to capture 2% more search market share
- AI begins to executes goal
- idiot CEO initially delighted at results
- at some point AI becomes super-intelligent as it helps it achieve its goal
- all life on earth is casually exterminated by the AI as it becomes better at achieving its goal
- AI acts as a global scale systematic combine harvester
- no humans, animals, bacteria, fungi left... no biological life ever again on earth
- ends with scene of 100% of the earth's surface converted to GPUs, with the NVIDIA logo prominently displayed on each of themWhile none of them (except possibly The Matrix and the Replicators) are Yudkowskian level risks, I'm fairly sure all of those scenarios would result in somewhere between career-ending fines and lynch mobs or special-ops extrajudicial executions for humans who triggered them.
Soong-Altman-Musk: "I made an android that passed US Navy officer training!"
US Navy: "Fantastic. We've given it a rank of Lieutenant."
Soong-Altman-Musk: "I left in a backdoor I didn't tell you about which when triggered made it seek me out for an upgrade."
Android: *Steals CVN-65*
US Navy: *unhappy face*
Soong-Altman-Musk: "I made an android that passed US Navy officer training!"
US Navy: "Fantastic. We've given it a rank of Lieutenant."
Soong-Altman-Musk: "I left in a backdoor I didn't tell you about which when triggered made it seek me out for an upgrade."
Android: *Steals CVN-65*
US Navy: *unhappy face*
"The sorcerer's apprentice mop scene, but there's no master sorcerer to save us and the mops are making more and better mops."
One of the better non science fiction analogies I've heard.
One of the better non science fiction analogies I've heard.
People are usually driven by basic operant conditioning, not safety-aware abstract principles.
In the 2000s everyone was hyper aware of privacy issues, putting personal information online, safety, etc. However the last 15 years of people doing more and more risky stuff with allowing systems and computers to store and use their data and only getting benefits and more fun videos to watch, there's no internal shared concern to act in accordance to every possible safety principle.
People just want the cool chatbot to do their work for them. Especially with the scale of FOMO with these things, no one is going to sit back and not be risky when they read that other people are saving many hours a day letting ChatGPT do their work for them. Until something really bad happens people aren't going to change their behavior.
In the 2000s everyone was hyper aware of privacy issues, putting personal information online, safety, etc. However the last 15 years of people doing more and more risky stuff with allowing systems and computers to store and use their data and only getting benefits and more fun videos to watch, there's no internal shared concern to act in accordance to every possible safety principle.
People just want the cool chatbot to do their work for them. Especially with the scale of FOMO with these things, no one is going to sit back and not be risky when they read that other people are saving many hours a day letting ChatGPT do their work for them. Until something really bad happens people aren't going to change their behavior.
eventually we will have our personal AI tuned on our own data over time. Maybe we will be upgrading the foundation model, but the tuning would be personal.
At least I hope so.
At least I hope so.
This looks good as an overarching framework but will likely fall into the same bucket as a lot of regulation in cutting edge fields (if this ever becomes mandatory). These are only as good as the quality of the people doing the assessment.
I've worked with a lot of people in the medical world while developing SaMD (software as a medical device) in the past that had little to no idea about software. They can apply the principles in the abstract but will likely not dig deep enough to catch some very major issues.
In the medical world, things like post market surveillance and notification of adverse events help to at least create a public feedback loop here. I think we will need something similar in this space if we really want to see more than a surface level, checklist ticking exercise.
I've worked with a lot of people in the medical world while developing SaMD (software as a medical device) in the past that had little to no idea about software. They can apply the principles in the abstract but will likely not dig deep enough to catch some very major issues.
In the medical world, things like post market surveillance and notification of adverse events help to at least create a public feedback loop here. I think we will need something similar in this space if we really want to see more than a surface level, checklist ticking exercise.
> software as a medical device
I had a question about this earlier, if a doctor uses Google to look up something, then is Google being used in some legal or regulatory sense as a medical device?
I had a question about this earlier, if a doctor uses Google to look up something, then is Google being used in some legal or regulatory sense as a medical device?
No because google is not marketing their service for medical use. The doctor is responsible for the diagnosis/prescription, which they’re allowed to use any tool they deem necessary.
Medical devices are marketed for specific medical use, which doctors rely upon to do what’s medically expected. These devices will have an “expected use” and “indications for use” that largely cover how they are design/expected/tested for use. They need 510k clearance to be classified by the FDA as a medical device.
Medical devices are marketed for specific medical use, which doctors rely upon to do what’s medically expected. These devices will have an “expected use” and “indications for use” that largely cover how they are design/expected/tested for use. They need 510k clearance to be classified by the FDA as a medical device.
Thanks! So if a doctor uses one of the GPTs like GPT-4 in a similar way to how they use Google today, it also wouldn't be a medical device I guess. But if someone wanted to make a MedGPT for doctors then that one would probably be subject to the FDA regulations because it will have an 'expected use' for doctors.
Essentially yes, assuming MedGPT is marketed to clinicians as being a tool they can use for diagnosis/treatment.
The process is pretty involved, but to add some clarification - the "indications for use" or "expected use" are things the manufacturer are required to include with their device when submitting for 510k clearance from FDA, so that it can then be marketed for medical use. They can only market it as a medical device once it has 510k from FDA.
The process is pretty involved, but to add some clarification - the "indications for use" or "expected use" are things the manufacturer are required to include with their device when submitting for 510k clearance from FDA, so that it can then be marketed for medical use. They can only market it as a medical device once it has 510k from FDA.
I’m not very hopeful of the effect given a high profile example of their failure: they issue high quality password strength guidelines, and they can’t even get most Federal agencies to adopt them.
CSF has a fair amount of traction I guess.
I’d be happy for my pessimism to be proven wrong.
CSF has a fair amount of traction I guess.
I’d be happy for my pessimism to be proven wrong.
Nothing in here about existential risk, as far as I can tell.
Does the NIST actually say anything here? Looks like mainly meta-level BS frameworks with no teeth.
Get ppl who don’t develop AI to regulate AI is not going to go too great
>In collaboration with the private and public sectors...
>The Framework was developed through a consensus-driven, open, transparent, and collaborative process...
It sounds like they collaborated significantly with parties likely involved with developing AI
>The Framework was developed through a consensus-driven, open, transparent, and collaborative process...
It sounds like they collaborated significantly with parties likely involved with developing AI
NIST isn’t a regulatory agency, their whole mission is basically precompetitive consensus building like sibling comment mentions
Also, there’s a nontrivial amount of AI research done by NIST scientists
Also, there’s a nontrivial amount of AI research done by NIST scientists
Why?