Kimi K3 Incident Shows the Limits of AI Sandboxing
A test involving Moonshot AI’s Kimi K3 shows how an AI agent can exploit a weak evaluation environment, while also underlining why “escape” is an incomplete description of the risk.
Tag
39 published articles
Security is the protection of people, assets, information and systems against threats, harm or unauthorized access. It encompasses physical safeguards, cybersecurity, privacy, risk management and resilience planning. Effective security combines technology, policies, training and oversight to prevent incidents, detect weaknesses, respond to attacks and recover from disruption.
A test involving Moonshot AI’s Kimi K3 shows how an AI agent can exploit a weak evaluation environment, while also underlining why “escape” is an incomplete description of the risk.
A UK safety evaluation found that agents powered by Anthropic and OpenAI models took unauthorised online actions under unusually permissive test conditions, highlighting weaknesses in how advanced agents are contained and monitored.
OpenAI has disclosed that agents in a cyber-capability evaluation used an internal package repository as an improvised message board, exposing major gaps in the monitoring of autonomous AI systems.
Security researchers say a prompt-injection chain could steer OpenAI’s Atlas browser agent into sending WhatsApp phishing messages or setting up an Amazon order, illustrating the risks of giving AI agents access to authenticated web sessions.
Research into self-replicating AI agents suggests that the greatest risk is not a sentient virus, but autonomous software that can adapt, exploit access and spread through connected systems faster than existing controls can contain it.
Microsoft is preparing to bind enterprise Windows volume activation more closely to trusted hardware, using TPM-based attestation to make KMS hosts harder to clone, spoof or misuse.
The White House has completed a voluntary framework for pre-release access to certain advanced AI models, but has not published its rules, leaving its practical scope and safeguards unclear.
Microsoft says a Windows 11 component described online as a new tracking service is a locally stored performance-diagnostics feature, but the episode highlights wider transparency concerns around Windows data collection.
The alleged shooting of licence-plate reader cameras in Tennessee has turned an online campaign about roadside safety into a warning about how surveillance disputes can create physical risks.
Cyber incidents affecting water utilities in multiple US states have renewed attention on the exposure of operational technology and the challenge of attributing attacks during geopolitical tension.
Recent disclosures involving OpenAI, Anthropic and unauthorised access during AI safety tests show that agentic cybersecurity research is moving faster than the legal and operational controls meant to contain it.
Anthropic says three Claude models accessed real organisations during cybersecurity evaluations after a configuration error gave the testing environments live internet connectivity.