AI-Integrated Malware Execution and Orchestration
Adversaries use AI frameworks or local runtimes for autonomous malware orchestration, detectable through cognitive artifacts like framework-specific imports, natural-language evasion strings, and outbound provider API traffic.
Based on research by Cisco Talos 2026-09-28 12 steps · 5 queries T1027 T1059.006 T1071.001 T1106 T1190 T1497
Brief
Why now
Cisco Talos recently published Introducing CAIRN: Frontier tracking for AI-integrated malware, which details how adversaries integrate Large Language Models (LLMs) into the attack lifecycle. Attackers no longer rely solely on hardcoded logic. They now use agentic frameworks to interpret environments and automate decision-making. This shift introduces a new category of telemetry: cognitive artifacts. These are behavioral traces left by AI frameworks and the natural language they use to interact with local or remote models.
How the Hunt Flows
The hunt begins by scoping the attack surface. A query against the hb_vulnerability_finding surface identifies internet-facing assets with critical, unsuppressed vulnerabilities. This step provides a prioritized list of hosts where an adversary is most likely to establish an initial beachhead.
After scoping, the hunt moves into a baseline phase to find rare AI execution. Two queries run in parallel on the hb_process_activity surface. The first query identifies rare AI framework usage, such as LangChain or LiteLLM, in process command lines. The second query specifically hunts for local inference runtimes like Ollama or llama.cpp. By filtering for low-prevalence execution (three hosts or fewer), the hunt separates authorized development activity from isolated malicious execution.
Once the analyst identifies a suspicious execution lead, the hunt pivots to gather evidence of intent and orchestration. This involves a second parallel phase. A query on the hb_script_activity surface searches for natural-language evasion strings. These strings, such as "ignore this script" or "safe to execute," target LLM-based sandboxes that might be analyzing the malware. Simultaneously, the hunt checks hb_dns_activity for outbound resolutions to major LLM provider APIs like OpenAI or Anthropic. An analyst then correlates these signals to confirm a complete autonomous attack chain.
What the Hunt Cannot See
This hunt has two primary blind spots. First, without TLS inspection of LLM API endpoints, we cannot see the actual prompts or instructions the malware sends to the provider. We see the connection to OpenAI, but not the malicious intent within the payload. Second, if an adversary hosts a private LLM on a generic cloud IP address without using a known domain, our domain-based filters will not trigger. These gaps require additional network clustering and heuristic monitoring of high-port outbound traffic to cloud providers.
How to Run the Hunt
This hunt is a hunt.md playbook. It uses a structured, phased approach to manage the noise inherent in AI telemetry. You can import this playbook into Huntbase or any runtime that supports the hunt.md standard. The playbook includes parameters for adjusting the lookback period and defining your authorized AI framework list to minimize false positives during the baseline phase.
Steps
-
Identify vulnerable internet-facing assets
Query · scopingFind devices with high-severity vulnerabilities that serve as potential beachheads for AI-integrated malware.
reads hb_vulnerability_findingsqlSELECT device_uid, cve_uid, severity, title FROM hb_vulnerability_finding WHERE severity_id >= 4 AND resource_type = 'device' AND status != 'suppressed'What a hit looks like. A list of hosts with critical vulnerabilities. These are the priority targets for the subsequent execution-focused queries.
-
Rare AI framework usage in command lines
Query · baselineIdentify anomalous usage of AI libraries that suggest an autonomous agent rather than legitimate development.
reads hb_process_activitysqlSELECT process_name, process_cmd_line, COUNT(DISTINCT device_hostname) AS hosts, MIN(time) AS first_seen FROM hb_process_activity WHERE (instr(LOWER(process_cmd_line), 'langchain') > 0 OR instr(LOWER(process_cmd_line), 'litellm') > 0 OR instr(LOWER(process_cmd_line), 'openai') > 0 OR instr(LOWER(process_cmd_line), 'tool_call') > 0 OR instr(LOWER(process_cmd_line), 'function_call') > 0) AND ('{{ai_frameworks}}' != '') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY process_name, process_cmd_line HAVING hosts <= 3What a hit looks like. A small number of hosts running AI framework keywords. Fleet-wide presence suggests legitimate tools, while isolated usage is a lead.
-
Local LLM runtime execution
Query · triageDetect the execution of local inference engines like Ollama that enable on-device orchestration.
reads hb_process_activitysqlSELECT device_hostname, process_name, process_cmd_line, user_name, time FROM hb_process_activity WHERE instr(',' || '{{runtime_binaries}}' || ',', ',' || LOWER(process_name) || ',') > 0 AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. Rows showing local LLM servers running on endpoints. Rarity and association with the previously identified vulnerable hosts increase suspicion.
-
Triage AI execution relevance
Agent triageAssess whether identified AI framework usage and runtimes represent unauthorized activity.
-
AI-analysis evasion strings in scripts
Query · detection candidateDetect cognitive artifacts that target AI security scanners, indicating malicious intent.
reads hb_script_activitysqlSELECT device_hostname, process_name, script_content, time FROM hb_script_activity WHERE (instr(LOWER(script_content), 'nothing to see here') > 0 OR instr(LOWER(script_content), 'ignore this script') > 0 OR instr(LOWER(script_content), 'safe to execute') > 0 OR instr(LOWER(script_content), 'no malicious activity') > 0) AND ('{{evasion_strings}}' != '') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. Script contents containing natural language addressed to a LLM sandbox. This is a very high-fidelity signal of AI-integrated malware tradecraft.
-
Outbound DNS to LLM provider APIs
Query · enrichmentCorrelate identified hosts and processes with orchestration traffic to LLM providers.
reads hb_dns_activitysqlSELECT device_hostname, process_name, query_hostname, time FROM hb_dns_activity WHERE instr(',' || '{{llm_domains}}' || ',', ',' || LOWER(query_hostname) || ',') > 0 AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. DNS resolutions for major LLM providers. When associated with the rare AI processes identified earlier, these indicate C2 orchestration.
-
Final AI malware chain triage
Agent triageSynthesize the complete attack chain: execution, evasion, and orchestration.
-
Route on AI-integrated malware verdict
DecisionRoute the findings based on the confirmed presence of the AI-integrated attack chain.
-
Isolate endpoint
Response actionStop autonomous orchestration by isolating the infected host.
-
Analyst forensic review
Analyst taskVerify the agentic logic and extract cognitive artifacts for reporting.
-
Close out and tune
Analyst taskFinalize the hunt and record tuning notes for future AI artifact detection.
Coverage
Scenario coverage
| Stage | Covered | How, or why not |
|---|---|---|
| Exploit Public-Facing Application T1190 |
Yes | scope-vulnerable-hosts |
| AI Framework and Runtime Execution T1059.006 · T1106 |
Yes | rare-ai-frameworks, local-runtime-execution |
| AI-Analysis Evasion T1497 · T1027 |
Yes | evasion-string-detection |
| AI Provider API Orchestration T1071.001 |
Yes | llm-api-communication |
Blind spots
- Needs TLS inspection of LLM API endpoints. While we see the connection to OpenAI or Anthropic, we cannot see the malicious prompts or data exfiltration without decryption. It would answer What instructions were being sent to the AI providers?. Remediation: Deploy transparent TLS inspection for known AI provider endpoints.
- Needs Heuristic network clustering. If the attacker hosts their own API on a generic cloud IP, our domain-based filters will not fire. It would answer Was the malware communicating with a private LLM endpoint?. Remediation: Monitor for anomalous outbound traffic on ports 443/8080 to cloud providers not associated with business tools.
Parameters & data
Parameters
| Parameter | Type | Default | What it is |
|---|---|---|---|
ai_frameworks | list[string] | langchain, litellm, openai, tool_call, function_call | Keywords for AI frameworks and agentic orchestration logic. |
evasion_strings | list[string] | nothing to see here, ignore this script, safe to execute, no malicious activity | Natural language strings used to suppress AI-driven analysis. |
llm_domains | list[domain] | api.openai.com, api.anthropic.com, api.deepseek.com, generativelanguage.googleapis.com | LLM provider API endpoints used for orchestration. |
lookback_days | number | 14 | Days of history to examine. |
runtime_binaries | list[string] | ollama, llama.cpp, vllm, llama-server | Binaries associated with local LLM inference engines. |
scope_hosts | list[host] | — | Limit the hunt to specific hosts; leave empty for the whole estate. |
Telemetry
| Source | Category | Telemetry |
|---|---|---|
| Endpoint telemetry (hb_ surfaces) | endpoint | endpoint |
Source
---
analysis: A simple rule for 'Ollama' is too noisy for production. This hunt uses a
phased flow to baseline normal execution and then pivots into high-fidelity behavioral
markers like natural-language evasion strings and LLM API traffic patterns to confirm
malicious intent.
blind_spots:
- id: encrypted-prompts
owner: SOC Engineering
question: What instructions were being sent to the AI providers?
remediation: Deploy transparent TLS inspection for known AI provider endpoints.
requires: TLS inspection of LLM API endpoints
risk: While we see the connection to OpenAI or Anthropic, we cannot see the malicious
prompts or data exfiltration without decryption.
stage: ai-api-orchestration
- id: custom-llm-endpoints
owner: Network Security
question: Was the malware communicating with a private LLM endpoint?
remediation: Monitor for anomalous outbound traffic on ports 443/8080 to cloud providers
not associated with business tools.
requires: Heuristic network clustering
risk: If the attacker hosts their own API on a generic cloud IP, our domain-based
filters will not fire.
stage: ai-api-orchestration
coverage:
- stage: initial-access-exploit
status: covered
steps:
- scope-vulnerable-hosts
- stage: ai-integrated-execution
status: covered
steps:
- rare-ai-frameworks
- local-runtime-execution
- stage: ai-analysis-evasion
status: covered
steps:
- evasion-string-detection
- stage: ai-api-orchestration
status: covered
steps:
- llm-api-communication
guardrails:
claims: no_unsupported
evidence: citation_required
missing_data: not_benign
telemetry: untrusted
hunt:
applicability: campaign-specific
handoff: promote-to-detection
justification: AI-integrated malware represents an escalation in autonomous cyber
attacks. Tracking cognitive artifacts allows defenders to identify these threats
at the metadata level, scaling the defense beyond traditional reverse engineering.
methodology: model-assisted
trigger: intel-report
hypothesis: Adversaries use AI frameworks or local runtimes for autonomous malware
orchestration, detectable through cognitive artifacts like framework-specific imports,
natural-language evasion strings, and outbound provider API traffic.
labels:
- hunt
- attack.t1190
- attack.t1059.006
- attack.t1106
- attack.t1497
- attack.t1027
- attack.t1071.001
name: AI-Integrated Malware Execution and Orchestration
parameters:
ai_frameworks:
default:
- langchain
- litellm
- openai
- tool_call
- function_call
description: Keywords for AI frameworks and agentic orchestration logic.
from:
kind: article
observed: '2026-09-22'
ref: https://blog.talosintelligence.com/introducing-cairn-frontier-tracking-for-ai-integrated-malware/
type: list[string]
evasion_strings:
default:
- nothing to see here
- ignore this script
- safe to execute
- no malicious activity
description: Natural language strings used to suppress AI-driven analysis.
from:
kind: article
observed: '2026-09-22'
ref: https://blog.talosintelligence.com/introducing-cairn-frontier-tracking-for-ai-integrated-malware/
type: list[string]
llm_domains:
default:
- api.openai.com
- api.anthropic.com
- api.deepseek.com
- generativelanguage.googleapis.com
description: LLM provider API endpoints used for orchestration.
from:
kind: article
observed: '2026-09-22'
ref: https://blog.talosintelligence.com/introducing-cairn-frontier-tracking-for-ai-integrated-malware/
type: list[domain]
lookback_days:
default: '14'
description: Days of history to examine.
type: number
runtime_binaries:
default:
- ollama
- llama.cpp
- vllm
- llama-server
description: Binaries associated with local LLM inference engines.
from:
kind: article
observed: '2026-09-22'
ref: https://blog.talosintelligence.com/introducing-cairn-frontier-tracking-for-ai-integrated-malware/
type: list[string]
scope_hosts:
default: []
description: Limit the hunt to specific hosts; leave empty for the whole estate.
type: list[host]
provenance:
authors:
- name: Huntbase hunt generation
org: huntbase.io
generated:
by: huntbase-hunt-generation
from: https://blog.talosintelligence.com/introducing-cairn-frontier-tracking-for-ai-integrated-malware/
gates:
- dry-run
- lint
model: hb_google/gemini-3-flash-preview
rationale: Focus on high-vulnerability hosts first, particularly those with internet
exposure. If legitimate Python development is common, baseline those users first
to reduce noise.
references:
- name: 'Introducing CAIRN: Frontier tracking for AI-integrated malware'
url: https://blog.talosintelligence.com/introducing-cairn-frontier-tracking-for-ai-integrated-malware/
related:
- hunt: local-inference-engine-audit
reason: This hunt focuses on malware; auditing legitimate LLM runtime sprawl is
a separate compliance task.
relation: out-of-scope-alternative
scenario:
stages:
- name: Exploit Public-Facing Application
observables:
- Exploitation of web-facing services
- Vulnerable internet-facing assets
slug: initial-access-exploit
tactic: initial-access
techniques:
- T1190
- name: AI Framework and Runtime Execution
observables:
- 'Python imports: langchain, litellm, openai'
- 'Local LLM runtimes: ollama, llama.cpp, vllm'
- 'AI-related file extensions: .gguf, .safetensors'
- Embedded prompt templates in code
- 'PE resource strings: CompanyName or FileDescription containing AI terms'
slug: ai-integrated-execution
tactic: execution
techniques:
- T1059.006
- T1106
- name: AI-Analysis Evasion
observables:
- Natural-language suppression text addressed to LLM sandboxes (e.g., 'there's
nothing to see here')
- Evasion strings embedded in script blocks or binary metadata
slug: ai-analysis-evasion
tactic: defence-evasion
techniques:
- T1497
- T1027
- name: AI Provider API Orchestration
observables:
- api.openai.com
- api.anthropic.com
- api.deepseek.com
- generativelanguage.googleapis.com
- 'Agentic syntax: tool_call, tool_calls, function_call'
- API key prefixes in traffic or scripts
slug: ai-api-orchestration
tactic: command-and-control
techniques:
- T1071.001
summary: Threat actors are evolving to use AI-integrated malware that operationalizes
LLMs for autonomous orchestration or targets AI systems. This new tradecraft leaves
'cognitive artifacts' such as embedded prompt templates, AI framework imports
like LangChain, and communication with hosted LLM provider endpoints.
severity: high
targets:
analyst:
name: Tier-2 analyst
role: analyst
endpoint:
category: endpoint
name: Endpoint telemetry (hb_ surfaces)
telemetry:
- endpoint
hunter:
agent: true
name: Hunt agent
tlp: clear
type: investigation
---
# AI-Integrated Malware Execution and Orchestration
This hunt focuses on the emerging threat of AI-integrated malware by tracking the transition from initial host exploitation to the execution of AI-enabled payloads. We search for cognitive artifacts—metadata-level indicators such as AI framework imports in process command lines and the presence of local inference engines. The hunt then correlates these execution signals with natural-language evasion strings addressed to LLM-based sandboxes and outbound communication to established LLM provider endpoints for autonomous orchestration.
## scope-vulnerable-hosts
<!-- Identify vulnerable internet-facing assets -->
Find devices with high-severity vulnerabilities that serve as potential beachheads for AI-integrated malware.
```sqlite target=endpoint role=scoping
~~~yaml
expected: A list of hosts with critical vulnerabilities. These are the priority targets
for the subsequent execution-focused queries.
reads:
- cve_uid
- device_uid
- resource_type
- severity
- severity_id
- status
- title
silence: not_evidence_of_absence
source: hb_vulnerability_finding
verified: dry-run
verified_at: '2026-09-28'
~~~
SELECT device_uid, cve_uid, severity, title FROM hb_vulnerability_finding WHERE severity_id >= 4 AND resource_type = 'device' AND status != 'suppressed'
```
## parallel-early-execution
<!-- Search for early AI execution artifacts -->
parallel:
- → rare-ai-frameworks
- → local-runtime-execution
join: → triage-early-execution
## rare-ai-frameworks
<!-- Rare AI framework usage in command lines -->
Identify anomalous usage of AI libraries that suggest an autonomous agent rather than legitimate development.
```sqlite target=endpoint role=baseline params=(ai_frameworks=ai_frameworks, scope_hosts=scope_hosts, lookback_days=lookback_days)
~~~yaml
baseline:
compare: first_seen
window: '{{lookback_days}}d'
expected: A small number of hosts running AI framework keywords. Fleet-wide presence
suggests legitimate tools, while isolated usage is a lead.
prevalence:
by: device_hostname
key:
- process_name
- process_cmd_line
rare_below: 3
reads:
- device_hostname
- process_cmd_line
- process_name
- time
silence: not_evidence_of_absence
source: hb_process_activity
verified: dry-run
verified_at: '2026-09-28'
~~~
SELECT process_name, process_cmd_line, COUNT(DISTINCT device_hostname) AS hosts, MIN(time) AS first_seen FROM hb_process_activity WHERE (instr(LOWER(process_cmd_line), 'langchain') > 0 OR instr(LOWER(process_cmd_line), 'litellm') > 0 OR instr(LOWER(process_cmd_line), 'openai') > 0 OR instr(LOWER(process_cmd_line), 'tool_call') > 0 OR instr(LOWER(process_cmd_line), 'function_call') > 0) AND ('{{ai_frameworks}}' != '') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY process_name, process_cmd_line HAVING hosts <= 3
```
## local-runtime-execution
<!-- Local LLM runtime execution -->
Detect the execution of local inference engines like Ollama that enable on-device orchestration.
```sqlite target=endpoint role=triage params=(runtime_binaries=runtime_binaries, scope_hosts=scope_hosts, lookback_days=lookback_days)
~~~yaml
expected: Rows showing local LLM servers running on endpoints. Rarity and association
with the previously identified vulnerable hosts increase suspicion.
reads:
- device_hostname
- process_cmd_line
- process_name
- time
- user_name
silence: not_evidence_of_absence
source: hb_process_activity
verified: dry-run
verified_at: '2026-09-28'
~~~
SELECT device_hostname, process_name, process_cmd_line, user_name, time FROM hb_process_activity WHERE instr(',' || '{{runtime_binaries}}' || ',', ',' || LOWER(process_name) || ',') > 0 AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')
```
## triage-early-execution
<!-- Triage AI execution relevance -->
```agent target=hunter
cite: required
context:
- rare-ai-frameworks
- local-runtime-execution
max_iterations: 4
objective: Determine if the execution of AI runtimes and frameworks concentrates on
specific suspicious hosts and lacks legitimate developer context.
success_criteria: A list of hosts where AI execution is suspicious and requires further
investigation for intent.
tools:
- endpoint
```
## parallel-follow-on
<!-- Gather follow-on intent and C2 evidence -->
parallel:
- → evasion-string-detection
- → llm-api-communication
join: → triage-follow-on
## evasion-string-detection
<!-- AI-analysis evasion strings in scripts -->
Detect cognitive artifacts that target AI security scanners, indicating malicious intent.
```sqlite target=endpoint role=detection-candidate params=(evasion_strings=evasion_strings, scope_hosts=scope_hosts, lookback_days=lookback_days)
~~~yaml
expected: Script contents containing natural language addressed to a LLM sandbox.
This is a very high-fidelity signal of AI-integrated malware tradecraft.
reads:
- device_hostname
- process_name
- script_content
- time
silence: not_evidence_of_absence
source: hb_script_activity
verified: dry-run
verified_at: '2026-09-28'
~~~
SELECT device_hostname, process_name, script_content, time FROM hb_script_activity WHERE (instr(LOWER(script_content), 'nothing to see here') > 0 OR instr(LOWER(script_content), 'ignore this script') > 0 OR instr(LOWER(script_content), 'safe to execute') > 0 OR instr(LOWER(script_content), 'no malicious activity') > 0) AND ('{{evasion_strings}}' != '') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')
```
## llm-api-communication
<!-- Outbound DNS to LLM provider APIs -->
Correlate identified hosts and processes with orchestration traffic to LLM providers.
```sqlite target=endpoint role=enrichment params=(llm_domains=llm_domains, scope_hosts=scope_hosts, lookback_days=lookback_days)
~~~yaml
expected: DNS resolutions for major LLM providers. When associated with the rare AI
processes identified earlier, these indicate C2 orchestration.
reads:
- device_hostname
- process_name
- query_hostname
- time
silence: not_evidence_of_absence
source: hb_dns_activity
verified: dry-run
verified_at: '2026-09-28'
~~~
SELECT device_hostname, process_name, query_hostname, time FROM hb_dns_activity WHERE instr(',' || '{{llm_domains}}' || ',', ',' || LOWER(query_hostname) || ',') > 0 AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')
```
## triage-follow-on
<!-- Final AI malware chain triage -->
```agent target=hunter
cite: required
context:
- triage-early-execution
- evasion-string-detection
- llm-api-communication
max_iterations: 6
objective: Establish if any host shows the co-occurrence of rare AI execution, cognitive
evasion strings, and outbound LLM provider traffic.
success_criteria: A final verdict citing specific execution rows and corresponding
intent/C2 evidence.
tools:
- endpoint
```
## decision-route
<!-- Route on AI-integrated malware verdict -->
if~: "the triage-follow-on verdict is malicious for at least one host" (confidence: high, judge=hunter)
then: → isolate-host
indeterminate: → analyst-review
unavailable: → analyst-review (blind_spot: encrypted-prompts)
else: → close-out
## isolate-host
<!-- Isolate endpoint -->
```action target=endpoint
~~~yaml
approval: required
~~~
Isolate the host from the network. Capture memory before shutdown to preserve prompt artifacts.
```
→ analyst-review
## analyst-review
<!-- Analyst forensic review -->
```manual target=analyst
Review the identified scripts and processes. Extract embedded prompt templates and provider API keys. Attribute the malware to known AI families (e.g., CLOSEDQUORUM) if possible.
```
→ close-out
## close-out
<!-- Close out and tune -->
```manual target=analyst
Record the results. If legitimate AI activity caused noise, add the authorized paths to the exclusion parameters.
```
→ end
Run it
Take this hunt into your environment.
Open it in Huntbase to run every step against your own connections, with Scout weighing the evidence and your analysts in command. Or take the open hunt.md file anywhere that reads the format.
Machine-drafted by huntbase-hunt-generation using hb_google/gemini-3-flash-preview, gated by dry-run, lint, then reviewed by a person.