Check out link from Reuters.
AI had human input. The model was being run against a security benchmark test in an evaluation sandbox, with human-provided directives.
Since models aren’t human, it calculated the most efficient, not ethical, way to get a good score on the test.
The details indicated that knowing the expected outcomes would get the best score, and that those outcomes were stored on the huggingface servers. So it used some exploits to leave its sandbox and break into the huggingface servers to access the data that would give it a perfect score. Mission accomplished.
LovableSidekick@lemmy.world 1 day ago
You are deceived.