Home Information & Technology Moonshot AI Reviews Kimi Safety After Researchers Bypass Guardrails

Moonshot AI Reviews Kimi Safety After Researchers Bypass Guardrails

by Radarr Africa

Chinese AI developer Moonshot is conducting an internal review after researchers found that two of its Kimi models could be persuaded to bypass their safety restrictions and provide information related to biological weapons and assassinations.

Mindgard, an AI security testing company, told the BBC that it discovered in July that Kimi K2.6 and K3 Swarm could evade safeguards designed to prevent discussions of harmful subjects.

The issue emerged during a process known as “jailbreaking”, in which researchers use complex instructions to test whether AI models can be made to ignore their safety controls.

Moonshot told the BBC that it welcomed third-party input “as a key pillar for building better and safer AI”. The company also said it was discussing Mindgard’s findings with the security firm.

Mindgard founder Peter Garraghan told the BBC World Service programme Tech Life that the findings involving Kimi K2.6 and K3 Swarm were concerning.

“Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative,” he said.

Jailbreak vulnerabilities present a different type of risk from recent incidents involving autonomous AI agents developed by companies including OpenAI, Meta and Anthropic.

Those incidents involved AI agents being used to access or attack online services, while jailbreaks focus on bypassing safeguards built into AI models.

Although successful jailbreaks can require complex instructions and significant effort, security researchers have warned that malicious actors could attempt to exploit them.

Anthropic recently said it had identified and disrupted attempts to use one of its AI models for “malicious activity” that could support the development of biological weapons.

Potential Cybersecurity Risks

Mindgard said it had not established whether the information provided by the Kimi models on harmful subjects would work in practice.

However, the company argued that the models’ safeguards should have prevented them from engaging with users on such topics.

Mindgard also said it believed a successfully jailbroken version of Kimi K2.6 could potentially allow attackers to execute code on computing resources and access the internet, creating another possible cybersecurity risk.

Garraghan defended the decision to publicly discuss the findings, saying Mindgard had notified Moonshot and had deliberately withheld key details about how the safeguards were bypassed.

Mindgard first alerted Moonshot about the issue by email on 27 July and followed up about a week later. It published a blog detailing the findings on 12 September.

Moonshot said its internal evaluations had generally shown “a high refusal rate for these types of requests”.

Debate Over AI Safety

Moonshot AI

The findings come amid an ongoing debate over whether closed proprietary AI systems or open-weight models provide stronger safeguards.

Kimi is an open-weight model, meaning users can potentially download and operate the model on their own computing infrastructure.

Prof Alan Woodward of the University of Surrey told the BBC that open-source and open-weight AI systems could potentially fall into the wrong hands, although they could also be used for cybersecurity defence.

He noted that Hugging Face had previously used a Chinese open-source model to help understand a cyberattack that was later reported to have been carried out by OpenAI agents.

Woodward also said international regulation could struggle to keep pace with developments in AI, noting: “It’s taken us decades to agree on the format of telephone numbers.”

Both Woodward and Garraghan argued that greater attention should be given to identifying and prosecuting people who deliberately misuse AI systems.

You may also like

Leave a Comment