AI-Driven Persistence and Automated Data Theft
An adversary is using compromised AI runtimes to maintain persistence and automate the exfiltration of sensitive data to AI skill repositories.
Based on research by ESET Research 2026-09-29 11 steps · 3 queries T1041 T1190 T1543 T1566
Brief
Why this hunt?
Recent analysis in Cyberthreats are moving faster than SMBs: Readiness must accelerate (https://www.welivesecurity.com/en/business-security/cyberthreats-moving-faster-smbs-readiness-must-accelerate/) highlights that attackers move faster than defensive governance. As businesses adopt AI tools, they often leave runtimes and API keys exposed. This hunt addresses the emergence of persistence techniques where malicious code operates within AI interpreters.
How the hunt flows
The first phase uses HTTP telemetry in hb_http_activity to find leads. A query identifies hosts making over 100 requests to AI service domains within the lookback window. This step distinguishes automated agent behavior from manual user chats.
Once an analyst confirms the traffic looks automated rather than human, the hunt gates into a parallel forensic phase. This phase correlates endpoint activity with network volume to build a high-confidence verdict.
On the endpoint, the hunt examines hb_process_activity for anomalies. It specifically looks for processes where the binary is no longer on disk but the parent is a known AI runtime or browser. This indicates fileless execution characteristic of a hijacked runtime.
Simultaneously, the hunt queries hb_network_connection to aggregate outbound traffic to AI providers. It flags any host sending more than 50MB of data. Large uploads to a skill repository often signal exfiltration rather than simple prompt requests.
The final triage step combines these signals. A host showing high-frequency HTTP requests, a fileless process anomaly, and a network volume spike receives a malicious verdict for containment.
Blind Spots
This hunt relies heavily on HTTP-level visibility. If an adversary uses an encrypted tunnel or a non-HTTP protocol to reach an AI repository, the initial scoping query misses the host. Furthermore, without endpoint-level on_disk state, an analyst cannot easily differentiate between a legitimate transient process and a malicious fileless injector.
Steps
-
High-frequency AI service interactions
Query · scopingIdentify hosts with unusual volumes of HTTP traffic to AI providers as a lead for automated agent abuse.
reads hb_http_activitysqlSELECT device_hostname, url_hostname, COUNT(*) AS request_count, user_agent, MIN(time) AS first_request, MAX(time) AS last_request FROM hb_http_activity WHERE instr(',' || '{{ai_service_domains}}' || ',', ',' || LOWER(url_hostname) || ',') > 0 AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, url_hostname, user_agent HAVING request_count > 100 ORDER BY request_count DESCWhat a hit looks like. Hosts with high request counts suggest automated activity rather than manual chat. Silence indicates no large-scale AI service usage observed in HTTP logs.
-
Assess traffic automation
Agent triageDetermine if the lead traffic represents an automated AI agent or legitimate user interaction.
-
Gate: Pursue forensics?
DecisionAvoid expensive endpoint and network-wide queries if the HTTP leads are benign.
-
AI runtime process anomalies
Query · detection candidateIdentify processes with no binary on disk running under AI-related parents, characteristic of PromptSpy.
reads hb_process_activitysqlSELECT device_hostname, process_name, process_cmd_line, parent_process_name, on_disk, user_name, time FROM hb_process_activity WHERE ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND (instr(',' || '{{ai_parent_processes}}' || ',', ',' || LOWER(process_name) || ',') > 0 OR instr(',' || '{{ai_parent_processes}}' || ',', ',' || LOWER(parent_process_name) || ',') > 0) AND (on_disk = 0 OR on_disk = 'false') AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. A process launched from a browser or AI runtime that has since been deleted from disk. This is a high-confidence signal for fileless execution.
-
Massive data exfiltration to AI
Query · baselineEstablish a baseline and identify hosts sending unusually large volumes of data to AI domains.
reads hb_network_connectionsqlSELECT device_hostname, dst_endpoint_hostname, SUM(traffic_bytes) AS total_out_bytes, COUNT(*) AS session_count FROM hb_network_connection WHERE ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND instr(',' || '{{ai_service_domains}}' || ',', ',' || LOWER(dst_endpoint_hostname) || ',') > 0 AND direction = 'outbound' AND state_kind = 'log' AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, dst_endpoint_hostname HAVING total_out_bytes > 50000000What a hit looks like. Hosts sending more than 50MB to AI domains. Silence suggests no bulk exfiltration occurred via these endpoints during the window.
-
Final intrusion triage
Agent triageCorrelate the HTTP automation leads with process anomalies and exfiltration volume to reach a verdict.
-
Route for containment
DecisionDetermine whether to trigger immediate containment or manual investigation.
-
Isolate compromised host
Response actionStop further exfiltration from the AI-hijacked host.
-
Analyst confirmation
Analyst taskProvide human oversight for suspicious or indeterminate cases.
-
Close out report
Analyst taskDocument the findings and update governance notes for AI usage.
Coverage
Scenario coverage
| Stage | Covered | How, or why not |
|---|---|---|
| AI-Enhanced Phishing and Exploitation T1190 · T1566 |
Yes | ai-service-leads, assess-traffic-volume |
| AI Engine Runtime Persistence T1543 |
Yes | ai-runtime-anomalies |
| Automated Data Exfiltration T1041 |
Yes | exfiltration-volume |
| Execution via Malicious AI Skills T1204.002 |
Out of scope | Not examined by this hunt; belongs to a separate hunt. |
Blind spots
- Needs hb_http_activity or proxy logs. A host connecting to a malicious AI skill repository via an SSH tunnel would be missed in the initial gate. It would answer Whether the host initiated the connection via a non-HTTP protocol or encrypted tunnel..
- Needs hb_process_activity with on_disk state. Without endpoint-level 'on_disk' data, fileless persistence within an AI interpreter is indistinguishable from legitimate process execution. It would answer Whether the malicious runtime hijack was fileless or executed from memory..
Parameters & data
Parameters
| Parameter | Type | Default | What it is |
|---|---|---|---|
ai_parent_processes | list[string] | chrome.exe, firefox.exe, msedge.exe, gemini.exe, python.exe, python3.exe | Common parent processes for AI runtimes and browsers. |
ai_service_domains | list[domain] | openai.com, anthropic.com, gemini.google.com, huggingface.co, mistral.ai | Known AI provider and skill repository domains. |
lookback_days | number | 14 | Days of history to examine. |
scope_hosts | list[host] | — | Hosts identified in the scoping lead; leave empty to scan the full estate. |
Telemetry
| Source | Category | Telemetry |
|---|---|---|
| Endpoint telemetry (hb_ surfaces) | endpoint | endpoint |
| Network telemetry | network | network |
| Web server / proxy logs | siem | network |
Source
---
analysis: A standard rule might detect large uploads, but this hunt pivots between
high-frequency HTTP traffic, fileless process state (on_disk = 0), and network baseline
deviations to distinguish an intrusion from legitimate AI usage.
blind_spots:
- id: no-http-telemetry
question: Whether the host initiated the connection via a non-HTTP protocol or encrypted
tunnel.
requires: hb_http_activity or proxy logs
risk: A host connecting to a malicious AI skill repository via an SSH tunnel would
be missed in the initial gate.
stage: initial-access-ai-enhanced-phishing-and-exploitation
- id: no-endpoint-telemetry
question: Whether the malicious runtime hijack was fileless or executed from memory.
requires: hb_process_activity with on_disk state
risk: Without endpoint-level 'on_disk' data, fileless persistence within an AI interpreter
is indistinguishable from legitimate process execution.
stage: persistence-ai-runtime-abuse
coverage:
- stage: initial-access-ai-enhanced-phishing-and-exploitation
status: covered
steps:
- ai-service-leads
- assess-traffic-volume
- stage: persistence-ai-runtime-abuse
status: covered
steps:
- ai-runtime-anomalies
- stage: exfiltration-automated-data-theft
status: covered
steps:
- exfiltration-volume
- reason: Not examined by this hunt; belongs to a separate hunt.
stage: execution-malicious-ai-skills
status: out_of_scope
guardrails:
claims: no_unsupported
evidence: citation_required
missing_data: not_benign
telemetry: untrusted
hunt:
applicability: campaign-specific
handoff: promote-to-detection
justification: AI adoption is outpacing governance; this hunt identifies the misuse
of AI runtimes and exfiltration channels that bypass traditional signature-based
detection.
methodology: model-assisted
trigger: intel-report
hypothesis: An adversary is using compromised AI runtimes to maintain persistence
and automate the exfiltration of sensitive data to AI skill repositories.
labels:
- hunt
- attack.t1041
- attack.t1190
- attack.t1566
- attack.t1543
- execution
- exfiltration
- initial access
- persistence
name: AI-Driven Persistence and Automated Data Theft
parameters:
ai_parent_processes:
default:
- chrome.exe
- firefox.exe
- msedge.exe
- gemini.exe
- python.exe
- python3.exe
description: Common parent processes for AI runtimes and browsers.
from:
kind: manual
observed: '2026-06-15'
ref: promptspy-persistence-research
type: list[string]
ai_service_domains:
default:
- openai.com
- anthropic.com
- gemini.google.com
- huggingface.co
- mistral.ai
description: Known AI provider and skill repository domains.
from:
kind: article
observed: '2026-05-01'
ref: eset-research-2026
type: list[domain]
lookback_days:
default: '14'
description: Days of history to examine.
type: number
scope_hosts:
default: []
description: Hosts identified in the scoping lead; leave empty to scan the full
estate.
type: list[host]
provenance:
authors:
- name: Huntbase hunt generation
org: huntbase.io
generated:
by: huntbase-hunt-generation
from: https://www.welivesecurity.com/en/business-security/cyberthreats-moving-faster-smbs-readiness-must-accelerate/
gates:
- dry-run
- lint
- critic
model: hb_google/gemini-3-flash-preview
rationale: Begin by scanning all endpoints for high-frequency AI domain traffic. Focus
on departments likely to use AI, such as Engineering (coding assistants) and Marketing
(content generation), as they are primary targets for AI-based persistence and data
theft.
references:
- name: 'Cyberthreats are moving faster than SMBs: Readiness must accelerate'
url: https://www.welivesecurity.com/en/business-security/cyberthreats-moving-faster-smbs-readiness-must-accelerate/
related:
- hunt: malicious-ai-skill-execution
reason: This hunt focuses on runtime hijacks (PromptSpy), while execution of malicious
plugins belongs in a hunt targeting container and dependency scanning.
relation: out-of-scope-alternative
scenario:
stages:
- name: AI-Enhanced Phishing and Exploitation
observables:
- automated social media reconnaissance
- personalized phishing messages in local languages
- rapid exploitation of N-day vulnerabilities in public applications
- prompt injection via chatbots
slug: initial-access-ai-enhanced-phishing-and-exploitation
tactic: initial-access
techniques:
- T1190
- T1566
- name: Execution via Malicious AI Skills
observables:
- malicious AI skills/plugins
- download of secondary malware payloads
- suspicious agent instructions causing unintended actions
slug: execution-malicious-ai-skills
tactic: execution
techniques:
- T1204.002
- name: AI Engine Runtime Persistence
observables:
- abuse of Google Gemini runtime for persistent execution
- AI-powered spyware (PromptSpy) persistence mechanisms
slug: persistence-ai-runtime-abuse
tactic: persistence
techniques:
- T1543
- name: Automated Data Exfiltration
observables:
- data exfiltration over existing C2 channels
- AI-driven classification and extraction of stolen data
- high-volume traffic to AI skills repositories
slug: exfiltration-automated-data-theft
tactic: exfiltration
techniques:
- T1041
summary: Threat actors are leveraging AI to automate victim reconnaissance, accelerate
exploit development for public-facing applications, and conduct high-volume personalized
phishing. The intrusion lifecycle involves the deployment of malicious AI skills
or plugins that execute malware and exfiltrate data via automated classification
agents.
severity: medium
targets:
analyst:
name: Tier-2 analyst
role: analyst
endpoint:
category: endpoint
name: Endpoint telemetry (hb_ surfaces)
telemetry:
- endpoint
hunter:
agent: true
name: Hunt agent
network:
category: network
name: Network telemetry
telemetry:
- network
web:
category: siem
name: Web server / proxy logs
telemetry:
- network
tlp: clear
type: investigation
---
# AI-Driven Persistence and Automated Data Theft
This hunt identifies the misuse of AI ecosystem components, such as hijacked Google Gemini runtimes and malicious AI skill repositories. It uses a gated flow to first identify hosts with high-frequency AI service interactions before performing deeper process-level and network-volume forensics. The hunt focuses on detecting 'PromptSpy' style persistence where malicious code runs within an AI interpreter and exfiltrates data at scale.
## ai-service-leads
<!-- High-frequency AI service interactions -->
Identify hosts with unusual volumes of HTTP traffic to AI providers as a lead for automated agent abuse.
```sqlite target=web role=scoping params=(lookback_days=lookback_days, ai_service_domains=ai_service_domains)
~~~yaml
expected: Hosts with high request counts suggest automated activity rather than manual
chat. Silence indicates no large-scale AI service usage observed in HTTP logs.
reads:
- device_hostname
- url_hostname
- user_agent
- time
silence: evidence_of_absence
source: hb_http_activity
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT device_hostname, url_hostname, COUNT(*) AS request_count, user_agent, MIN(time) AS first_request, MAX(time) AS last_request FROM hb_http_activity WHERE instr(',' || '{{ai_service_domains}}' || ',', ',' || LOWER(url_hostname) || ',') > 0 AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, url_hostname, user_agent HAVING request_count > 100 ORDER BY request_count DESC
```
## assess-traffic-volume
<!-- Assess traffic automation -->
```agent target=hunter
cite: required
context:
- ai-service-leads
max_iterations: 3
objective: Review the request counts and user agents from ai-service-leads. Mark hosts
as suspicious if they show persistent, high-frequency requests or use non-standard
browser user agents.
success_criteria: A per-host verdict of suspicious or benign.
tools:
- endpoint
- network
- web
```
## gate-decision
<!-- Gate: Pursue forensics? -->
if~: "The assessment identifies at least one host where AI interaction is likely automated or malicious." (confidence: medium, judge=hunter)
then: → forensic-fan-out
indeterminate: → analyst-confirmation
unavailable: → analyst-confirmation (blind_spot: no-http-telemetry)
else: → close-out-report
## forensic-fan-out
<!-- Endpoint and network fan-out -->
parallel:
- → ai-runtime-anomalies
- → exfiltration-volume
join: → final-triage
## ai-runtime-anomalies
<!-- AI runtime process anomalies -->
Identify processes with no binary on disk running under AI-related parents, characteristic of PromptSpy.
```sqlite target=endpoint role=detection-candidate params=(lookback_days=lookback_days, ai_parent_processes=ai_parent_processes, scope_hosts=scope_hosts)
~~~yaml
expected: A process launched from a browser or AI runtime that has since been deleted
from disk. This is a high-confidence signal for fileless execution.
reads:
- device_hostname
- process_name
- parent_process_name
- on_disk
- time
silence: not_evidence_of_absence
source: hb_process_activity
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT device_hostname, process_name, process_cmd_line, parent_process_name, on_disk, user_name, time FROM hb_process_activity WHERE ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND (instr(',' || '{{ai_parent_processes}}' || ',', ',' || LOWER(process_name) || ',') > 0 OR instr(',' || '{{ai_parent_processes}}' || ',', ',' || LOWER(parent_process_name) || ',') > 0) AND (on_disk = 0 OR on_disk = 'false') AND time >= datetime('now', '-{{lookback_days}} days')
```
## exfiltration-volume
<!-- Massive data exfiltration to AI -->
Establish a baseline and identify hosts sending unusually large volumes of data to AI domains.
```sqlite target=network role=baseline params=(lookback_days=lookback_days, ai_service_domains=ai_service_domains, scope_hosts=scope_hosts)
~~~yaml
baseline:
compare: prior_equal_window
window: '{{lookback_days}}d'
expected: Hosts sending more than 50MB to AI domains. Silence suggests no bulk exfiltration
occurred via these endpoints during the window.
prevalence:
by: device_hostname
key:
- dst_endpoint_hostname
rare_below: 3
reads:
- device_hostname
- dst_endpoint_hostname
- traffic_bytes
- direction
- state_kind
- time
silence: evidence_of_absence
source: hb_network_connection
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT device_hostname, dst_endpoint_hostname, SUM(traffic_bytes) AS total_out_bytes, COUNT(*) AS session_count FROM hb_network_connection WHERE ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND instr(',' || '{{ai_service_domains}}' || ',', ',' || LOWER(dst_endpoint_hostname) || ',') > 0 AND direction = 'outbound' AND state_kind = 'log' AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, dst_endpoint_hostname HAVING total_out_bytes > 50000000
```
## final-triage
<!-- Final intrusion triage -->
```agent target=hunter
cite: required
context:
- assess-traffic-volume
- ai-runtime-anomalies
- exfiltration-volume
max_iterations: 6
objective: Review the correlated evidence. Confirm a malicious verdict if a host shows
suspicious traffic automation (from assess-traffic-volume), a fileless process anomaly
(from ai-runtime-anomalies), and a matching exfiltration spike to an AI domain (from
exfiltration-volume).
success_criteria: A verdict of malicious, suspicious, or benign citing specific event
rows.
tools:
- endpoint
- network
- web
```
## final-route
<!-- Route for containment -->
if~: "The final triage verdict is malicious for at least one host." (confidence: high, judge=hunter)
then: → isolate-compromised-host
indeterminate: → analyst-confirmation
unavailable: → analyst-confirmation (blind_spot: no-endpoint-telemetry)
else: → close-out-report
## isolate-compromised-host
<!-- Isolate compromised host -->
```action target=endpoint
~~~yaml
approval: required
~~~
Isolate the host immediately. Revoke all active session tokens for AI services (OpenAI, Gemini) used on this device.
```
→ analyst-confirmation
## analyst-confirmation
<!-- Analyst confirmation -->
```manual target=analyst
Review the cited process and network volume rows. Verify if the 'on_disk = 0' process correlates with the network spike to the AI provider. Determine if this represents a rogue AI agent or a legitimate developer activity.
```
→ close-out-report
## close-out-report
<!-- Close out report -->
```manual target=analyst
Summarize the findings. Note any legitimate automated AI tasks that should be excluded from future runs. Update the corporate AI policy if shadow AI usage was discovered.
```
→ end
Run it
Take this hunt into your environment.
Open it in Huntbase to run every step against your own connections, with Scout weighing the evidence and your analysts in command. Or take the open hunt.md file anywhere that reads the format.
Machine-drafted by huntbase-hunt-generation using hb_google/gemini-3-flash-preview, gated by dry-run, lint, critic, then reviewed by a person.