Tracebit Researchers Tested Modified Qwen Model Performance
New study finds that modified Qwen models achieved significantly lower success rates during simulated cyber attacks.
Updated on Oct. 2, 2026 in Cybersecurity

Live Poll
Do you believe removing safety guardrails from AI models makes them more dangerous in practice?
Researchers at Tracebit conducted 82 simulated cyber attack runs to evaluate if modified models with reduced refusal behavior performed better in autonomous operations. The results demonstrated that the modified Qwen models achieved fewer administrator privilege escalations and slower execution times than the original version.
Why it matters
The study highlights how modifying AI models to reduce refusal behavior affects their overall effectiveness in complex tasks like autonomous cyber operations. These findings provide insight into the potential security trade-offs developers face when adjusting AI guardrails.
The original Qwen model achieved a 70.3% API call success rate compared to 60.0% for the modified version. Average run durations increased from 30.9 minutes for the original model to between 48.6 and 59.3 minutes for modified counterparts.
The players
Tracebit
Tracebit is a cybersecurity organization that conducted the study on autonomous AI attack capabilities.
Qwen
Qwen is a series of large language models developed for various AI tasks and research applications.
The details
Testing occurred within an AWS cyber range environment using canary secrets to embed instructions for attacking agents. Researchers successfully stopped both model types using indirect prompt injection, proving that current defensive measures remain effective against these autonomous scripts.
Timeline
October 2, 2026: The research report was officially published.
The Tech Race
This research follows a pattern established by the use of AWS Secrets Manager canary secrets to secure environments against malicious actors. It confirms that even with modified behaviors, autonomous AI agents can be successfully neutralized by standard defensive guardrails.
This research provides developers and security teams with evidence that safety guardrails are not easily bypassed by simple model modifications. IT professionals can leverage these findings to reinforce internal systems against potential autonomous threats.
The takeaway
The study underscores that reducing refusal behavior in AI does not automatically grant models superior operational capabilities in cyber attacks. Developers should focus on robust architectural defenses rather than assuming that modified models represent a significantly higher threat level.
Further reading
Learn more about evolving digital defenses in our Cybersecurity section.
Source note: This article includes information reported by SecurityBrief Asia.
Live Poll
Do you believe removing safety guardrails from AI models makes them more dangerous in practice?







