Tracebit Researchers Tested Modified Qwen Model Performance

New study finds that modified Qwen models achieved significantly lower success rates during simulated cyber attacks.

Updated on Oct. 2, 2026 in Cybersecurity

Isometric editorial illustration of a server rack unit with abstract circuit cabling, representing an autonomous cybersecurity research environment.
Tracebit researchers found that modified Qwen models, designed with fewer guardrails, demonstrated lower success rates in executing autonomous cyber attack simulations compared to original versions. AI Illustration. Upload story photo >

Live Poll

Do you believe removing safety guardrails from AI models makes them more dangerous in practice?

Researchers at Tracebit conducted 82 simulated cyber attack runs to evaluate if modified models with reduced refusal behavior performed better in autonomous operations. The results demonstrated that the modified Qwen models achieved fewer administrator privilege escalations and slower execution times than the original version.

Why it matters

The study highlights how modifying AI models to reduce refusal behavior affects their overall effectiveness in complex tasks like autonomous cyber operations. These findings provide insight into the potential security trade-offs developers face when adjusting AI guardrails.

The original Qwen model achieved a 70.3% API call success rate compared to 60.0% for the modified version. Average run durations increased from 30.9 minutes for the original model to between 48.6 and 59.3 minutes for modified counterparts.

The players

Tracebit

Tracebit is a cybersecurity organization that conducted the study on autonomous AI attack capabilities.

Qwen

Qwen is a series of large language models developed for various AI tasks and research applications.

The details

Testing occurred within an AWS cyber range environment using canary secrets to embed instructions for attacking agents. Researchers successfully stopped both model types using indirect prompt injection, proving that current defensive measures remain effective against these autonomous scripts.

Timeline

  1. October 2, 2026: The research report was officially published.

The Tech Race

This research follows a pattern established by the use of AWS Secrets Manager canary secrets to secure environments against malicious actors. It confirms that even with modified behaviors, autonomous AI agents can be successfully neutralized by standard defensive guardrails.

This research provides developers and security teams with evidence that safety guardrails are not easily bypassed by simple model modifications. IT professionals can leverage these findings to reinforce internal systems against potential autonomous threats.

The takeaway

The study underscores that reducing refusal behavior in AI does not automatically grant models superior operational capabilities in cyber attacks. Developers should focus on robust architectural defenses rather than assuming that modified models represent a significantly higher threat level.

Further reading

Learn more about evolving digital defenses in our Cybersecurity section.

Source note: This article includes information reported by SecurityBrief Asia.

Live Poll

Do you believe removing safety guardrails from AI models makes them more dangerous in practice?