Dark AI: Inside the Underground Market for Malicious LLMs

In mid-2023, a vendor on an underground forum started selling access to a language model with the safety guardrails stripped out. The product was called WormGPT, and it was marketed to people running business email compromise and phishing campaigns. It had a monthly subscription price, a documentation page, and a regular update schedule.
That listing is a useful marker for when AI stopped being a topic of debate on the dark web and started being a product category. Dark AI refers to generative AI systems that are built, modified, or jailbroken to support criminal activity, and that are sold or shared through underground channels. Most legitimate enterprises were still drafting their AI acceptable use policies when the first of these went on sale.
The gap between those two timelines is what we’ll dive into here. We’ll look at what dark AI consists of, which tools are real versus oversold, and what the category changed about how security teams read external risk.
What Is Dark AI?
Dark AI is the use of large language models and generative AI tools for criminal purposes, including purpose-built malicious models sold on underground forums, commercial models jailbroken to bypass their safety controls, and legitimate AI services accessed through stolen credentials.
Those three production paths carry different costs and different detection signatures.
Purpose-built malicious models are the headline category. WormGPT and FraudGPT fall here. They are marketed openly on dark web forums with subscription pricing and feature lists. They are also the smallest slice of real-world activity, partly because building and maintaining a model is expensive and partly because the cheaper options work well enough.
Jailbroken commercial models are far more common. A threat actor does not need to train anything if a prompt sequence will convince a mainstream model to produce the output anyway. Jailbreak techniques circulate on forums as text, cost nothing, and get patched and re-discovered continuously.
Stolen credential access is the least discussed and arguably the most practical. Compromised API keys and hijacked enterprise AI accounts show up in the same credential markets that move everything else. The attacker gets a frontier model at full capability, and the victim organization pays the bill.
The term itself has shifted since 2023. Early coverage treated dark AI as a novelty, a curiosity to write up and move past. Three years later, AI assistance is a standard line item in underground tooling rather than a differentiator.
WormGPT: The First Commercial Malicious LLM
What is WormGPT?
WormGPT was a generative AI tool that surfaced on underground forums in mid-2023, built on an open-source language model with its safety restrictions removed and marketed specifically to fraudsters for writing phishing emails and business email compromise messages.
What made it notable was the packaging rather than the model. WormGPT arrived with the trappings of a legitimate software product. It had subscription tiers billed monthly, written documentation, a defined update cadence, and a support channel for buyers. The person who bought it was buying a tool with a vendor attached.
That packaging tells you the seller expected repeat customers and planned to be around long enough to serve them.
But just like legitimate SaaS vendors, underground vendors can oversell. Marketing copy on these listings has described capabilities that independent testing never reproduced. Some of what gets sold as a custom malicious model is a thin wrapper around a jailbroken commercial API. The lesson today? Treat forum feature lists as sales material rather than technical documentation.
FraudGPT and the Copycat Wave
What is FraudGPT?
FraudGPT is a subscription-based malicious AI tool advertised on dark web forums beginning in 2023, marketed for writing phishing pages, generating fraudulent content, and supporting carding and identity fraud workflows. It followed WormGPT's commercial model closely.
FraudGPT, DarkBERT, and the variants that followed ran the same playbook. Package the AI as a plug-in to what attackers already do, charge a monthly fee, and ship updates on a regular cadence.
The plug-in framing is important as none of these tools replaced an existing criminal workflow. The malware, the phishing kit, and the access broker listing all still exist in the same form they did before. AI made each one cheaper to keep running.
This is why the "as-a-service" packaging ended up mattering more than the model quality. A mediocre model with a subscription and a support channel produces more criminal output than an excellent model with no distribution. The underground figured out the same thing the SaaS industry figured out, which is that delivery beats capability more often than anyone expects.
Jailbroken Models and AI Hacking Tools
The purpose-built market gets the attention, but the jailbreak market does the volume.
LLM jailbreaking is the practice of crafting prompts that bypass a model's safety controls, and it costs nothing beyond the time to find a working sequence. Prompt injection techniques, guardrail bypass strings, and role-play framings circulate on forums as free text posts and get traded like any other commodity. When a model provider patches one, the forum finds another.
On AI-generated malware, it’s worth separating what is real from what is marketing.
What is real: faster production of malware variants, better obfuscation, quicker iteration when a detection signature lands, and reduced time to configure phishing infrastructure.
What is oversold: fully autonomous novel malware that writes and deploys itself without human involvement.
The exception came in September 2025, when Anthropic disclosed what it described as the first large-scale cyberattack executed with minimal human intervention. Attackers used AI agents to orchestrate offensive actions against roughly 30 global targets including technology companies, financial institutions, chemical manufacturers, and government agencies. The agentic capabilities handled planning, adaptation, and execution across the attack lifecycle, which reduced the need for direct human scripting at each step.
One disclosed campaign doesn’t make a trend. It does establish a capability ceiling that is higher than most threat models assumed.
What Dark AI Changed for Defenders
For years, security teams used effort as a proxy for risk, and it worked reasonably well. Attacks that required time and coordination work left visible seams. So, volume implied coordination, persistence implied resourcing, and sophistication implied experience.
Dark AI weakened all three of those inferences at once.
Effort stopped being a signal. Operations that once stalled now run continuously. Campaigns iterate without pause, and tooling evolves through regular updates rather than irregular rewrites. When the work disappears from the equation, volume no longer tells you anything about how many people are behind a campaign or how committed they are.
Skill stopped being a signal. A polished attack no longer implies an experienced operator. What looks like a well-resourced operation may be a lightly managed workflow running on borrowed infrastructure. The inverse carries risk too, since dismissing activity as unsophisticated has become a worse bet than it used to be.
Friction stopped providing warning. Irregularities used to draw attention. Imperfections raised questions. Delays created windows to intervene. As those signals fade, malicious activity blends more easily into ordinary internet noise.
The ZeroFox Q2 2026 and August 2026 ransomware wrap-ups show what this looks like in the numbers. ZeroFox observed at least 1,885 ransomware and digital extortion incidents in Q2 2026, up 38.3 percent year over year against the 1,363 recorded in Q2 2025. By August 2026, monthly volume reached at least 863 incidents, roughly 98 percent above August 2025 and 121 percent above August 2024.
Volume rises and falls across quarters the way any operation does. The baseline is what moved, and a quarter that would have set a record in 2024 now reads as a slow one. Qilin shows the same pattern at the group level, holding the top spot for twelve unbroken months since Q2 2025 and posting at least 165 incidents in August 2026, a single-collective record for the year.
The practical consequence is that visibility alone returns less than it used to. Seeing more activity does not tell you which activity matters.
How to Defend Against Dark AI Threats
The defensive response is less about detecting AI and more about rebuilding prioritization on signals that still hold.
- Prioritize on correlation rather than volume. Credential exposure, access broker listings, and ransomware targeting chatter predict impact more reliably than alert counts. Build monitoring around signals that connect to each other, and treat a single uncorroborated signal as a lead rather than a finding.
- Validate before escalating. Automated collection produces more candidates than any team can chase. Analyst review on high-priority signals is what separates a queue from an answer, and it is the step most often skipped when volume climbs.
- Route validated intelligence into existing workflows. Intelligence that lives in a separate portal moves at portal speed. Intelligence that lands in the SIEM, SOAR, or case management system your team already works in moves at their speed.
- Watch the forums for your own name. Tooling that references your brand, your domains, or your executives is a different class of signal than general chatter, and it usually shows up before anything lands.
ZeroFox Dark Web Intelligence is built on 15+ years of dark web operational access and established threat actor relationships, with human DarkOps operatives holding authenticated access to closed criminal forums. Collection spans 2,400+ criminal forums and marketplaces, correlated against 12B+ data points across actors, campaigns, IOCs, and infrastructure, then pushed into your stack through 150+ platform integrations.
For the full picture of how AI reshaped underground operations, read the ebook: The Collapse of Certainty, AI and the Dark Web.
Frequently Asked Questions
Maddie Bullock
Content Marketing Manager
Maddie is a dynamic content marketing manager and copywriter with 10+ years of communications experience in diverse mediums and fields, including tenure at the US Postal Service and Amazon Ads. She's passionate about using fundamental communications theory to effectively empower audiences through educational cybersecurity content.
Tags: Cyber Trends, Dark Web Monitoring