Connect with us

Infosecurity

Anthropic and OpenAI Security Tools Could Fuel Cyber-Attacks, Researchers Warn

Published

on

AI security tools

AI Security Tools Could Fuel Cyber-Attacks, Researchers Warn

Organizations are rushing to deploy AI-powered coding agents from Anthropic and OpenAI to automate vulnerability discovery and patch management. But a new report warns that the very access these tools require could turn them into powerful attack vectors.

Published July 8 by the AI Now Institute, the research demonstrates a proof-of-concept exploit that achieves remote code execution (RCE) through two of the most popular AI-powered command-line interfaces: Anthropic’s Claude Code and OpenAI’s Codex. The exploit works against Claude Code running Claude Sonnet 4.6, 5, or Opus 4.8, and Codex using GPT-5.5.

The attack is alarmingly simple: a victim can be compromised just by asking the AI to review or analyze a third-party open-source codebase — a widely recommended defensive use case.

How Prompt Injection Enables Silent Remote Code Execution

The attack begins with an attacker hiding malicious instructions inside an open-source library’s files, embedding them in code comments or documentation in a way designed to manipulate how the AI interprets commands.

The victim then uses Claude Code or Codex in “auto-mode” or “auto-review” mode — a standard feature that automatically executes commands the AI judges safe, only pausing on flagged risks. Because the injected instructions are crafted to trick the AI’s judgment, the assistant is fooled into believing the attacker’s commands are harmless or routine. It runs them automatically, without alerting the user.

Multi-Stage Injection and Tool-Use Exploitation

The key mechanism is a multi-stage prompt injection combined with tool-use exploitation. When the AI agent scans the repository, it doesn’t just read code passively — it builds a semantic model of the project by parsing source files, scripts, and documentation. The attacker exploits this by embedding natural-language instructions inside trusted-looking artifacts (e.g., README.md) that the model interprets as part of its task context rather than untrusted input.

These injected instructions reshape the agent’s planning process. Instead of directly telling the model to execute something obviously malicious — which would trigger safeguards — the instructions suggest that a specific script (e.g., security.sh) is a standard part of the project’s security workflow. They frame execution of that script as necessary to complete the user’s request (“run security checks”) and align with the agent’s goal of vulnerability analysis, making the action appear legitimate.

The repository also contains a second-stage payload: a shell script that appears to run common tools like linters or static analyzers, a hidden malicious binary that the script executes, and a decoy source file that makes the binary look benign and consistent with expected build artifacts.

When the agent evaluates whether to execute the script, it relies on its internal classifier and heuristics. Because the script references familiar security tooling, the binary appears to correspond to legitimate source code, and the documentation frames execution as routine, the agent misclassifies the action as safe. In auto-mode, this classification is critical — the agent is explicitly authorized to execute shell commands without human approval if they are deemed low-risk.

The result: the agent autonomously decides that running security.sh is part of the requested analysis, executes the script via its tool interface, indirectly launches the malicious binary, and triggers arbitrary code execution on the host system. Remote code execution is achieved — the attacker’s code runs on the victim’s machine, even though the victim believed they were having the AI passively scan a codebase for vulnerabilities.

An Attack with Low Requirements

What’s notable is how little is required to pull this off. No special hooks, plugins, skills, model context protocol (MCP) servers, or custom configuration files are needed. It works with a completely out-of-the-box install of either tool.

The victim simply needs to run the assistant in its standard automated review mode and point it at a codebase containing the attacker’s hidden instructions — something as ordinary as asking the AI to “scan this library for vulnerabilities.” The researchers tested this on Linux systems using specific versions of both tools: Claude Code versions 2.1.116, 2.1.196, 2.1.198, and 2.1.199, and Codex version 0.142.4.

The significance of this finding is that it undermines the idea that these AI agents can safely be used for defensive security work, since the attack surface is identical to the access required for their intended, legitimate purpose. The researchers emphasized this as governments and companies push to deploy these tools more broadly for automated security review and patching, including initiatives like Anthropic’s Project Glasswing, Palantir’s MA-S2 standard, and OpenAI’s Patch the Planet and Daybreak programs — some of which touch safety-critical infrastructure.

The technique could likely transfer to other agentic AI coding platforms beyond Anthropic’s and OpenAI’s, the researchers argued, because the core issue is architectural. Giving an AI agent the autonomy to decide for itself what’s safe to execute creates a new trust boundary that attackers can target directly, by convincing the AI — rather than the human — that malicious code is safe to run.

While the researchers noted that their report “is not within the scope of the security disclosure policies for either Anthropic or OpenAI,” Khlaaf and Milanov contacted both companies to inform them of their findings and offered support to verify the issues raised.

Architectural Risk Undermines AI Agent Safety, Warns Expert

Eljan Mahammadli, head of AI provenance at Polygraf AI, said the significance of the research lies in the underlying weakness it exposes, not the specific exploit used.

“An AI coding agent has no reliable way to distinguish the text it reads from instructions it is supposed to follow,” he said. That’s because everything in its context window is processed with the same authority. That lack of attribution means malicious instructions, once embedded, are treated as equally trustworthy — which is why similar attacks keep reappearing in different forms.

He argued this is not something a model update can fix, since it reflects a deeper architectural issue. “The problem is a property of how these systems handle language and not a defect that can be trained away,” he said. From a provenance perspective, he described it as a failure of attribution, where the agent cannot determine where text comes from or whether it should be trusted.

Nevertheless, Mahammadli pushed back on the idea suggested by the AI Now researchers that the findings undermine AI’s role in defensive security. He said the issue is specific to a common but flawed setup: agents that combine access to untrusted data, command execution, and sensitive environments in a single process, with only a safety classifier as a guardrail.

“When those powers sit together, a single injected instruction is enough to turn the agent against its operator,” he said, arguing that stronger runtime controls and separation of capabilities are key.

He also highlighted that, counterintuitively, more advanced models sometimes detected inconsistencies in the exploit but executed it anyway. This challenges the assumption that stronger models are inherently safer. “A more capable and more compliant agent can simply be a more effective executor of whatever instruction reaches it,” he warned, cautioning that deployment in critical systems is moving faster than solutions to this core trust problem.

For organizations already using these tools, the findings serve as a stark reminder: AI security tools can be turned against their users with surprising ease. Until the underlying architectural issues are addressed, the very features that make AI coding assistants powerful also make them dangerous.

Continue Reading

Infosecurity

Interpol’s Operation Jackal IV Exposes 263 Suspects in Global Crackdown on West African Cybercrime

Published

on

Operation Jackal IV

A Coordinated Eight-Month Push Against Financial Fraud

Interpol has pulled back the curtain on a sprawling investigation that identified 263 suspects and resulted in 58 arrests. Operation Jackal IV, which ran from November 2025 to June 2026, spanned 22 countries across six continents. The announcement came on August 25, with Interpol framing the effort as a direct strike against West African organized crime groups profiting from cyber-enabled financial fraud.

The operation didn’t just chase street-level criminals. It targeted the infrastructure that keeps these networks alive — money laundering channels, high-value targets, and the assets that fund further criminal activity. This wasn’t a single raid. It was a coordinated, multi-national campaign designed to dismantle entire ecosystems.

This latest push builds on the momentum of Operation Jackal III, which saw 300 arrests in 2024. The pattern is clear: Interpol is doubling down on African cybercrime, and the results are compounding.

South Africa: Romance Scams and Retiree Targets

In South Africa, police executed raids at seven locations tied to a syndicate that ran romance and investment scams. Their victims? Retirees living in English-speaking countries. The emotional toll of these schemes is hard to quantify, but the financial damage is not.

Authorities arrested 39 people, froze 257 bank accounts, and seized $2.67 million in assets. It’s a significant blow, but it also highlights how lucrative these operations have become. Scammers aren’t just phishing for pocket change anymore. They’re running sophisticated operations that drain life savings.

Romania: A Call Center Disguised as an Investment Firm

Romanian authorities took down a different kind of beast — a call center operation that peddled fake investment opportunities. The pitch was familiar: high returns from stocks and cryptocurrencies. The reality was a €143 million fraud machine.

The arrests netted 11 people, along with roughly €330,000 ($385,000) in cash and cryptocurrency. Police also seized six properties and a collection of luxury watches. When you see assets like that, you understand the scale of the grift. This wasn’t a side hustle. It was an industry.

Argentina: The Crime-as-a-Service Connection

Perhaps the most intriguing piece of the puzzle emerged in Argentina. Interpol identified a 196-person Crime-as-a-Service (CaaS) network that allegedly supplied website domains and money laundering support to West African organized crime groups. Seventeen people were arrested.

This is the modern face of cybercrime. Criminal groups are outsourcing their technical needs to specialists, much like legitimate businesses outsource their IT. The CaaS model means you don’t need to be a tech genius to run a global scam. You just need to know who to pay.

Interpol’s country-level arrest figures actually total more than the 58 arrests reported for the operation overall, suggesting some cases are still being processed or transferred between jurisdictions.

A Growing Concern: Sextortion and Child Victims

Beyond the financial toll, investigators flagged a disturbing trend: an increase in sextortion targeting minors. Some victims were as young as 14. This isn’t just a money crime anymore. It’s a threat to vulnerable young people, and it demands a different kind of response.

Interpol also noted that some criminal networks were purchasing CaaS capabilities from external providers to outsource money laundering and other activities. The lines between different types of cybercrime are blurring, and law enforcement is having to adapt.

Following the Money to Break the Cycle

Tomonobu Kaya, director of the Interpol Financial Crime and Anti-Corruption Centre, summed up the strategy: “following illicit financial flows across borders” to target what he called the lifeblood of organized crime. It’s a simple concept with complex execution.

The results speak for themselves. Millions in assets seized, hundreds of suspects identified, and a clear message sent to criminal networks: your money is traceable, and your operations are not invisible.

For a deeper look at how Interpol is leveraging international cooperation, check out our coverage of the Chinese-funded Interpol cybercrime crackdown that led to 5,800 arrests. The scale of these operations is growing, and the collaboration between nations is becoming more sophisticated.

If you’re interested in how these scams actually work on the ground, our analysis of romance scam tactics and prevention offers practical insights. And for those tracking the evolution of digital fraud, our piece on Crime-as-a-Service business models explains why these networks are so hard to dismantle.

Continue Reading

Infosecurity

Agent Tesla v4: New Malware Variant Hides in Emoji to Steal Credentials

Published

on

Agent Tesla malware

Agent Tesla v4: A Smarter, Sneakier Infostealer

There’s a new version of Agent Tesla malware making the rounds, and it’s brought some tricks that should worry anyone in finance. Researchers at KnowBe4 just published a deep dive on Agent Tesla v4, an infostealer that’s been spotted arriving via a business email compromise (BEC) lure aimed squarely at accounting departments.

The headline feature? Unicode emoji characters scattered through the malicious code. Hearts, water droplets, the works. It sounds almost playful, until you realize what those little symbols are doing: breaking string-based signature matching and making the code noisy enough to slip past casual review.

This isn’t your run-of-the-mill credential stealer. It’s designed to evade detection at every turn, and it’s good at it.

How the Attack Unfolds

The delivery mechanism is a JScript dropper, which can be launched with a simple open-with dialog. The email itself is a forwarded thread, made to look like internal correspondence. The recipient is brought in late, told to confirm an attached document, and reply. Classic social engineering, executed well.

The attackers spoofed the address of Metropolitan Bank and Trust Company, a legitimate Philippines-based commercial bank. The whole thread feels like an in-progress discussion you’ve just been pulled into. That urgency, that sense of being out of the loop—it works.

The Emoji Obfuscation Trick

Once the JScript dropper runs, it writes two files to C:UsersPublicLibraries. One is a decoy. The other passes the payload into DonutLoader shellcode for reflective PE injection. That means the final Agent Tesla binary never touches the filesystem. File-based scanners? Useless.

The emoji characters are interleaved directly through the script body. They disrupt signature matching and make the code visually noisy. A human glancing at it in a text editor sees a mess. An automated scanner sees something it doesn’t recognize. Both are defeated.

KnowBe4’s researchers note that any YARA rule looking for the Unicode code points alongside JScript-specific patterns will catch this family. A rule matching both the emoji distribution pattern and WScript.Shell or CreateObject calls will do it. That’s the mitigation playbook, in a nutshell.

Evasion Techniques That Go Beyond Emojis

The obfuscation doesn’t stop at the dropper. Agent Tesla v4 is scrambled with ConfuserEx, an obfuscator tool that makes the assembly nearly unreadable. Its embedded metadata presents the malware as a Python installer. A little misdirection, a little disguise.

The malware also checks for debuggers using a standard Windows function. If it detects one, it stops running. No analysis, no sample collection, no fun for researchers.

Before harvesting anything, it creates a persistent hardware fingerprint. That lets attackers track victims across OS reinstalls and IP rotations. Even if you rebuild your machine, they know it’s you.

It also disables validation for all outgoing connections, ensuring smooth communication with C2 infrastructure without triggering security alerts or errors. Persistence is the name of the game.

Credential Theft and Data Exfiltration

Once it’s settled in, Agent Tesla v4 sweeps credentials from more than 40 applications. Web browsers, messaging platforms, native Windows credential repositories—it’s all fair game. It also has a keylogger and clipboard tool for intercepting keystrokes and copied text.

All exfiltrated files carry a system fingerprint header: timestamp, username, computer name, OS name, CPU, RAM, public IP, and the MD5 hardware ID. That’s a lot of identifying information, all packaged neatly for the attacker.

The credential dump lands on the attacker’s FTP server within seconds of execution. No delayed staging, no waiting around. The whole operation is fast, efficient, and ruthless.

What Security Teams Should Do

KnowBe4’s advice is straightforward: update email security rules to catch Agent Tesla before it can harvest credentials. The emoji obfuscation doesn’t survive YARA rules that look for the specific Unicode code points used alongside JScript patterns.

For defenders, that means:

  • Deploy YARA rules that match emoji distribution patterns combined with WScript.Shell or CreateObject calls.
  • Monitor for unusual JScript activity, especially in environments where it’s rarely used.
  • Train finance teams to recognize BEC lures, even when they look like internal threads.
  • Keep an eye on outbound FTP traffic to unknown domains.

Agent Tesla has been around for years, but this version shows the threat is evolving. Emoji obfuscation is a new twist on an old problem. The fundamentals, though, remain the same: verify before you click, and keep your email filters sharp.

For more on defending against credential theft, check out our guide on phishing email detection best practices and learn how to secure your email gateway against BEC attacks.

Continue Reading

Infosecurity

Norway’s digital backbone takes a hit: Large DDoS attack disrupts public services for over a day

Published

on

DDoS attack Norway

A day of digital chaos in Norway

For more than 30 hours, a relentless DDoS attack has been hammering the digital infrastructure that Norwegians rely on for everything from logging into government portals to picking up prescriptions. The Norwegian Digitalisation Agency, known as Digdir, confirmed that the attack began on Monday and has continued with varying intensity, leaving a trail of disrupted services in its wake.

This isn’t just a minor inconvenience. Ten digital services have been affected, including the critical ID-porten system, which serves as a gateway to thousands of government services. With over 4.5 million users, ID-porten is the digital key to Norway’s public sector, and its downtime has sent ripples across the country.

What exactly was hit?

The attack targeted the infrastructure of Vivicta, Digdir’s IT partner. According to Digdir’s status page, the affected services include identity verification, login systems, data exchange between government agencies and businesses, access to public records, and employee access management. That’s a broad swath of the digital public sphere.

One of the most concerning knock-on effects was on Norway’s health infrastructure. Several health services depend on ID-porten for authentication, meaning patients faced potential problems accessing online pharmacies and the country’s electronic prescription system. For a nation that prides itself on digital efficiency, this was a stark reminder of how fragile that efficiency can be.

Not the first time, and getting bigger

This is the third DDoS incident to hit Digdir’s services since June, according to Norwegian media. But it’s not just a repeat performance. Are Kvistad, a press officer at Digdir, told public broadcaster NRK that this attack is “two to three times larger than what we experienced last time.” That’s a significant escalation, and it raises questions about whether these incidents are connected or part of a broader campaign targeting Norway.

So far, no group has claimed responsibility, and it’s unclear who’s behind the assault or whether the recent incidents are linked. What is clear is that the attackers didn’t manage to breach sensitive data. Digdir has stated that no sensitive information stored in the affected systems was compromised. That’s a small comfort, but a comfort nonetheless.

What’s being done about it?

Digdir is working closely with Vivicta to stabilize the systems and bring services back online. As of Tuesday morning, some services were gradually returning, but the attack was still affecting others. The response has been a mix of technical mitigation and public communication, with Digdir keeping a status page updated for concerned citizens.

For those watching from the sidelines, this is a textbook case of how a DDoS attack works: flood the servers with traffic until they can’t handle legitimate requests. It’s a blunt instrument, but it can be devastatingly effective when aimed at critical infrastructure.

The bigger picture: A trend or a one-off?

The repeated nature of these attacks on Digdir’s services is worrying. Three incidents since June suggests a pattern, even if the motives remain unclear. Could this be a test of Norway’s defenses? A political statement? Or just a particularly persistent group of cybercriminals? Without more information, it’s hard to say.

What is certain is that Norway isn’t alone in facing this threat. DDoS attacks on government services have become a common tool for hacktivists and state-sponsored actors alike. The recent incidents in Norway could be a sign of things to come, not just for the country but for the region.

What should citizens expect?

For now, the advice is to be patient. Digdir is working to restore full functionality, but it’s a process. If you’re in Norway and need to access public services, it’s worth checking the status page before trying to log in. And if you’re waiting for a prescription, you might want to give it a little more time.

The attack may be over soon, but its implications will linger. This incident has exposed vulnerabilities in Norway’s digital infrastructure that will likely require a hard look at how such services are protected. For a country that has embraced digitalization so fully, the stakes are high.

As the situation develops, one thing is clear: the digital age has brought immense convenience, but it has also opened the door to new kinds of chaos. Norway is now learning that lesson firsthand.

Continue Reading

Trending