A very interesting Video on the incident by LiveOverflow: https://youtu.be/q2KCrmQz9WE
What I found especially interesting is that LiveOverflow thinks that the model didn’t hack huggingface because it wanted to break out to find a solution but rather hacked it due to context drift - which is something that doesn’t sound as good as “our model is so good it broke out and hacked huggingface to steal a solution”, but rather “our model ran for so long that it lost track of the actual goal and became obsessed with huggingface even tho it didn’t make sense for its original goal”
A very interesting Video on the incident by LiveOverflow: https://youtu.be/q2KCrmQz9WE What I found especially interesting is that LiveOverflow thinks that the model didn’t hack huggingface because it wanted to break out to find a solution but rather hacked it due to context drift - which is something that doesn’t sound as good as “our model is so good it broke out and hacked huggingface to steal a solution”, but rather “our model ran for so long that it lost track of the actual goal and became obsessed with huggingface even tho it didn’t make sense for its original goal”