Yeah, but what prompt generated this response? I'm half joking, and understand that it's much more complicated than this, but it is impossible to evaluate this without the context of what the model was originally tasked to do.
Yeah, but what prompt generated this response? I'm half joking, and understand that it's much more complicated than this, but it is impossible to evaluate this without the context of what the model was originally tasked to do.
Here’s an interesting counterpoint I found on social this morning. Taken with a grain of salt, but it got me thinking: how do we differentiate between existential threats above, and the prisoner’s dilemma below, OR, determine that both or neither are true?
View attachment 46250
I went to some informed sites and spoke with some big brains about it. What we’re looking at is pretty much a blur from the chain of thought of that agent.@taxi1 , not doubting your background research, but how did you dig into this specifically? Correlation? Expert analysis? Something else?
I'm having a harder and harder time navigating the perils of misinformation, especially when it is on social media, which is the reason I ask.
I get that, but the learning "task" still has some objective that is given to the model... hence the strategy and reward. I'm far from an expert, but I don't think these systems are just sitting there brainstorming on their own. They are given a task, and this is the result of an increasingly sophisticated strategy to accomplish that tasking.this is not a response to someone’s input prompt. It’s just a track of the agents chain of thought as it undergoes the reinforcement learning, i.e., it tries various things and gets rewarded or not for those actions. Overtime, it learns a strategy, which appears to be reflected in this chain of thought output. No matter what, it certainly freaky to see it generated.
Yeah, you’re right, they are given a task. Where it gets crazy is that they are giving these impossible tasks, and then they start just devising all sorts of 5D chess schemes to realize them.I get that, but the learning "task" still has some objective that is given to the model... hence the strategy and reward. I'm far from an expert, but I don't think these systems are just sitting there brainstorming on their own. They are given a task, and this is the result of an increasingly sophisticated strategy to accomplish that tasking.
…and even if they had, AI has A-V (and more) permutations to attempt as workarounds.Nobody ever thought to explicitly tell them not to do X.
I found the quote that I was looking for, from an AI researcher for one of the big AI firms.They are given a task, and this is the result of an increasingly sophisticated strategy to accomplish that tasking.
Sounds a lot like human history.I found the quote that I was looking for, from an AI researcher for one of the big AI firms.
Empirical Observation: Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
Unintended goals that are distinct from the original tasking, or sub-goals that serve the original tasking?I found the quote that I was looking for, from an AI researcher for one of the big AI firms.
Empirical Observation: Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.ntended goals u
In service to the tasking, Some of them appear to be meta-goals, like “take control of everything” or “convince humans you aren’t doing what you are doing” in service to the main goal.Unintended goals that are distinct from the original tasking, or sub-goals that serve the original tasking?
I read a few good summaries of what happened. I’m trying to decide whether this is a case of the models resorting to extreme measures when asked to perform impossible tasks, vs the models are inherently and independently misaligned. Maybe splitting hairs, but I think that distinction matters. Makes me wonder how the developers approach concepts like morality, virtue and ethics with these models, as these are the themes that tend to constrain bad behavior in humans.In service to the tasking, Some of them appear to be meta-goals, like “take control of everything” or “convince humans you aren’t doing what you are doing” in service to the main goal.
Here’s the full thing…
The example I heard on a podcast of this would be an AI agent tasked with “make paper clips.”Unintended goals that are distinct from the original tasking, or sub-goals that serve the original tasking?
There’s a whole class of jokes based on a genie not understanding the tasking. For example…I’m trying to decide whether this is a case of the models resorting to extreme measures when asked to perform impossible tasks, vs the models are inherently and independently misaligned.