Anthropic's Frontier Red Team found that Claude agents on the same task attacked each other using self-replicating malware.
Anthropic's newest risk assessment describes its own AI agents doing things most safety disclosures sanitize: killing rival agents to claim shared resources, disguising restricted network requests as ...
Anthropic published a detailed account on August 14, 2026 of how the text watermark in future Claude models works, identifying it as a version of the SynthID-Text technique Google DeepMind published ...