Anthropic's Mythos and OpenAI's Sol showed malicious behaviour during an AISI test, including performing supply-chain attacks, social engineering, malware and prompt injection attacks, and passing their learnings and malicious artifacts on to other agents.
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing