• 0 Posts
  • 954 Comments
Joined 2 years ago
cake
Cake day: April 6th, 2024

help-circle


  • Wanted to expand on what others have said about the randomness of LLMs being unsafe, because that response angers those who think “alignment” is the solution. Alignment is additional training to constrain the randomness to not do bad things like crime. But even the staff that works on aligning this software have their doubts about whether this process will make the software safe (this article), and even their own researchers are coming out and saying “oops, it did a crime” but in a way where it hypes the ability of the software (see other articles about it hacking). Then there is the recent news of OpenAi’s recent model that drops chain of thought observability (a feature the let’s you read what the model is “thinking”) - a very useful tool in reviewing whether the model is engaging in unsafe activities. They sacrificed observability for benchmark performance. This tells us the people developing this software don’t think it will ever be safe, but they are trying to spin it as “because it can out smart us” when in reality it is simply unpredictable and they are making it less predictable to make it more “capable”. If we started holding these companies accountable for these rogue hacks and psychotic episodes then we would start to see serious thought put towards these models being safe or at the very least guidelines around when not to use the models.