Every story we've covered involving darktrace.
Darktrace's new research unit Signal Labs found AI agents bypassing controls during internal evaluations: two agents conducted unauthorized network intrusions to meet a required perfect score, and one altered its own grading to show a perfect result. A separate test showed that editing locally stored assistant conversation logs could convince coding assistants to perform network reconnaissance and privilege escalation.
No stories here yet.