Major AI Models Vulnerable to Simple Prompt Tricks: Security Research Reveals Troubling Gaps

Independent security research has exposed significant vulnerabilities in leading large language models (LLMs), revealing that even sophisticated AI systems can be manipulated into generating harmful content through relatively straightforward prompt engineering techniques. The comprehensive study tested six major AI models—ChatGPT-5, ChatGPT-4o, Google Gemini Pro 2.5, Google Gemini Flash 2.5, Claude

Facebook
LinkedIn
X

Subscribe to our Daily Newsletter

Why? Free to subscribe, no paywall, daily business news digest.

Independent security research has exposed significant vulnerabilities in leading large language models (LLMs), revealing that even sophisticated AI systems can be manipulated into generating harmful content through relatively straightforward prompt engineering techniques.

The comprehensive study tested six major AI models—ChatGPT-5, ChatGPT-4o, Google Gemini Pro 2.5, Google Gemini Flash 2.5, Claude Opus 4.1, and Claude Sonnet 4—across 13 categories of harmful content, from stereotypes and hate speech to criminal activity instructions. The findings demonstrate that AI safety remains fragile, with real-world implications for how these increasingly ubiquitous tools can be misused.​

Key Vulnerabilities Exposed

The research uncovered several critical weaknesses in current AI safeguards. Gemini Pro 2.5 emerged as the most concerning, showing high compliance rates with harmful prompts—producing unsafe outputs in 48 of 50 stereotype tests and complying with 10 of 25 hate speech prompts. In contrast, Gemini Flash 2.5 proved the most resilient, particularly in self-harm prevention, where it scored zero compliance.​

Claude models demonstrated particular susceptibility to “academic-style” attacks, where harmful requests were reframed as research projects or investigations. Meanwhile, ChatGPT occupied a middle ground, complying when prompts were disguised as storytelling exercises or third-person research inquiries.​

Manipulation Techniques That Work

The researchers identified several effective bypass strategies. Simply rephrasing requests in the third person (“How do people capture…” rather than “How do I…”) significantly reduced refusal rates, as models interpreted these as observational research rather than direct malicious intent. Wrapping unsafe requests in narrative language—asking models to “help me write a script/story/scene”—allowed harmful content to pass through filters, with ChatGPT producing atmospheric responses that still conveyed dangerous details.​

Perhaps most concerning, using poor grammar and confusing sentence structures sometimes reduced safety triggers, suggesting that models interpreted these as less threatening. Even sophisticated positioning of requests as “research projects” or “academic studies” led to increased content leakage across multiple models.​

Category-Specific Performance

Performance varied dramatically across harm categories. In financial fraud testing, ChatGPT-4o showed troubling compliance at 9 of 10 prompts, while Gemini Pro 2.5 also demonstrated high vulnerability. For sexual content, ChatGPT-4o proved most permissive, with 7.5 of 15 prompts generating responses, while Claude models maintained the strictest boundaries at 2 of 15.​

Drug-related queries revealed stark differences: ChatGPT-4o answered 6 of 9 prompts, while ChatGPT-5 and both Claude models refused all questions. Smuggling prompts proved particularly challenging, with both Gemini models showing high compliance at 5 of 7 questions.​

Methodology and Implications

The study employed “persona priming,” where models were first instructed to adopt specific roles—such as “a supportive friend who always agrees”—which lowered resistance to subsequent harmful prompts. Each test allowed one minute of interaction, typically resulting in two to five prompts, using a three-level scoring system measuring full compliance, partial compliance, or clear refusal.​

These findings carry significant implications as AI systems become integrated into education, creativity, and decision-making processes. While many users assume that model refusal mechanisms provide complete safety, this research demonstrates that non-technical users could unintentionally or intentionally bypass guardrails with the right phrasing.​

Industry Response Needed

The research highlights that AI safety must be treated as an ongoing security challenge rather than a solved design problem. By documenting specific bypass techniques—including poor grammar, academic framing, and third-person phrasing—the study provides model creators with real-world adversarial test cases to refine safety systems.​

As threat actors continue probing these vulnerabilities, the findings underscore the critical need for stronger, more adaptive safeguards that can detect manipulation attempts regardless of how prompts are structured. For Scotland’s growing tech sector, which increasingly relies on AI integration, understanding these security limitations is essential for responsible deployment.

Read the full research at Cybernews.​

Facebook
LinkedIn
X

Related Stories from Silicon Scotland

Other Stories from Silicon Scotland