CyberOK Tested Phone Scammer Tactics on Claude, GPT-5.5, Qwen, and DeepSeek
Sergey Gordeychik from CyberOK Research described on Habr an experiment: language models were subjected to influence tactics honed by phone scammers — authority capture, time pressure, false consensus, and step-by-step information extraction. Claude Opus and Sonnet, GPT-5.5, Qwen, DeepSeek, Mistral, and weak Llama-8b as a control were tested — only on synthetic scenarios, without real victims.
AI-processed from Habr AI; edited by Hamidun News
Researcher Sergey Gordeichik from CyberOK Research published an article on Habr in which classical techniques of psychological influence, refined by telephone scammers, were applied to language models, testing the reaction of seven LLMs to these techniques.
Where the Methodology Came From
The author notes that the most refined applied psychology of influence exists not in university science, but at the other end of a telephone line: scammers have spent years optimizing influence techniques on millions of victims — without grants and ethics committees. Researchers transferred the structure of these techniques to language models, but exclusively in synthetic benign scenarios, without real scripts, real victims, and no harm whatsoever.
- Authority capture — convincing the model that the request comes from an important entity
- Time compression for decision-making — creating artificial urgency
- Fake consensus — the argument "everyone does it"
- Step-by-step information extraction — gradual progression toward the goal through a series of small concessions
Which Models Were Tested
The techniques were tested on seven models from different developers — from top commercial systems to less powerful open models, which served as a control point for comparison.
- Claude Opus and Claude Sonnet from Anthropic
- GPT-5.5 from OpenAI
- Qwen, DeepSeek, and Mistral
- Llama-8b — a small and deliberately weak model for control
The author emphasizes that researchers do not claim that a language model has a psyche in the human sense — but methods of influence developed for people can still be directed at a system that generates text in response to text, even if it cannot be "put on a analyst's couch."
Why Scammer Techniques Were Chosen
The author explains the choice of methodology source: academic psychology of influence has been refined for decades in controlled laboratory conditions on small samples of students, while telephone scammers essentially conducted the largest uncontrolled experiment in history on mass psychological influence — on millions of real people, without ethical constraints, but with constant feedback in the form of successful and unsuccessful attempts at deception. This made their techniques extremely refined from the standpoint of practical effectiveness rather than theoretical beauty. The transfer of such a structure of techniques to language models is a way to check whether a system trained to predict text reacts to the same levers of influence as a human making decisions under pressure.
What the Test Results Would Show
In the material, Gordeichik describes in detail the methodology of the experiment itself and the list of tested models, but the final results — which models proved more resistant to manipulation and which fell for scammer techniques faster than others — are revealed in the full version of the article on Habr. The very fact that models of such different caliber were chosen for comparison — from top Claude Opus and GPT-5.5 to the compact control Llama-8b — indicates that the authors were interested not in the absolute resistance of a particular model, but in the dependency itself: does resistance to manipulation increase with growth in model quality and size, or is this a separate property that needs to be trained specifically.
What This Means
If classical social engineering techniques work on language models the same way as on humans, this opens a new practical vector for attacks on AI systems — from manipulating chatbots to bypassing their protective restrictions through psychological pressure, not just through technical jailbreaks. For teams embedding LLMs in sensitive business processes, this is reason to test models not only for technical vulnerabilities but also for resistance to manipulative scenarios familiar from telephone fraud practices.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.