Shadow AI Usage and Prompt Injection Exposure
Employees are bypassing corporate AI controls by using personal accounts to process sensitive documents, or external attackers are exploiting public-facing AI applications to extract internal data.
Based on research by Huntress 2026-10-04 12 steps · 4 queries T1190 T1566
Brief
Why Now?
104.18.37.228 172.64.150.31 Recent research titled Companies Push AI Use But Skip Training and Official Policy (https://www.huntress.com/blog/llm-security-report) highlights a significant security gap: 43% of workers lack AI training, and 25% admit they would use personal accounts to process sensitive contracts. This creates a dual risk of data exfiltration to unmanaged platforms and the exposure of internal AI assets to external prompt injection. 104.18.37.228 172.64.150.31
How the Hunt Flows
104.18.37.228 172.64.150.31 The hunt begins by establishing the external attack surface. It queries exposed asset data to identify any public-facing applications using AI frameworks, Llama, or GPT-based notebooks. This identifies what an attacker could see and potentially target for prompt injection. 104.18.37.228 172.64.150.31 Next, the hunt shifts focus to internal behavior. The lead query scans HTTP activity for traffic directed at known personal AI domains like Claude, ChatGPT, and Gemini. This provides a list of hosts and users engaging in shadow AI usage outside of enterprise-managed instances. 104.18.37.228 172.64.150.31 Once we identify hosts with personal AI traffic, the hunt gates deeper investigation on an analyst verdict. For confirmed shadow AI users, the hunt triggers two parallel queries. One looks for sensitive file access—specifically documents containing keywords like 'contract', 'salary', or 'confidential'—occurring within an hour of the AI access. The second query calculates network traffic volume to AI-related IP addresses to find high-byte transfers that suggest file uploads. 104.18.37.228 172.64.150.31 Finally, an analyst triages the results. By combining the record of sensitive file access, the timing of the web traffic, and the volume of data sent, the analyst confirms whether a policy violation or data exfiltration event occurred. 104.18.37.228 172.64.150.31
What the Hunt Cannot See
104.18.37.228 172.64.150.31 This hunt has two primary blind spots. First, without TLS inspection or endpoint proxy logs, we cannot see the specific text inside the AI prompts; we only see that a connection occurred. Second, without EDR telemetry for clipboard events, we cannot definitively prove a user copy-pasted text from a document into a browser window. We rely on temporal proximity to infer this action. 104.18.37.228 172.64.150.31
Steps
-
Establish external AI attack surface
Query · scopingIdentify public-facing AI applications that attackers might target for prompt injection.
reads hb_exposed_assetssqlSELECT domain_or_ip, asset_type, product, port, discovered_at FROM hb_exposed_assets WHERE (LOWER(product) LIKE '%ai%' OR LOWER(product) LIKE '%llama%' OR LOWER(product) LIKE '%gpt%' OR LOWER(product) LIKE '%chat%' OR LOWER(product) LIKE '%notebook%' OR LOWER(product) LIKE '%langchain%')What a hit looks like. A list of IP addresses or domains hosting AI-related software exposed to the internet.
-
Lead: Detect personal AI platform usage
Query · baselineFind internal hosts accessing personal AI domains that lack enterprise data protections.
reads hb_http_activitysqlSELECT device_hostname, actor_user_name, url_hostname, COUNT(*) as request_count, MIN(time) as first_access, MAX(time) as last_access FROM hb_http_activity WHERE (instr(',' || '{{personal_ai_domains}}' || ',', ',' || LOWER(url_hostname) || ',') > 0) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, actor_user_name, url_hostnameWhat a hit looks like. Users and hosts accessing personal AI tools. Silence indicates absence of traffic to these domains.
-
Evaluate AI lead
Agent triageDetermine if the HTTP traffic indicates personal account usage rather than corporate instances.
-
Gate: Confirm personal usage
DecisionOpen the expensive file and network volume queries only when personal AI usage is confirmed.
-
Correlate sensitive document access
Query · enrichmentIdentify if hosts accessing personal AI were touching sensitive files within a one-hour window of the access.
reads hb_file_activitysqlSELECT device_hostname, actor_user_name, file_name, file_path, time FROM hb_file_activity WHERE (LOWER(file_name) LIKE '%contract%' OR LOWER(file_name) LIKE '%confidential%' OR LOWER(file_name) LIKE '%salary%') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time BETWEEN datetime('{{ai_access_timestamp}}', '-1 hour') AND datetime('{{ai_access_timestamp}}', '+1 hour')What a hit looks like. File touch events on sensitive documents occurring near the time of AI platform access.
-
Verify data exfiltration volume
Query · enrichmentExamine network traffic to AI IPs to find large data uploads that confirm document exfiltration.
reads hb_network_connectionsqlSELECT device_hostname, dst_endpoint_ip, SUM(traffic_bytes) as total_bytes, MIN(time) as start_time FROM hb_network_connection WHERE (instr(',' || '{{personal_ai_ips}}' || ',', ',' || dst_endpoint_ip || ',') > 0) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, dst_endpoint_ipWhat a hit looks like. High total_bytes counts to known AI infrastructure indicating potential file uploads.
-
Triage AI risks
Agent triageAssess the combination of personal AI usage, sensitive file access, and traffic volume.
-
Route on risk verdict
DecisionDirect exfiltration findings to isolation and others to remediation.
-
Isolate suspect host
Response actionHalt suspected exfiltration and preserve browser evidence.
-
Analyst remediation
Analyst taskHarden exposed assets and verify policy compliance.
-
Hunt close-out
Analyst taskFinalize documentation and findings.
Coverage
Scenario coverage
| Stage | Covered | How, or why not |
|---|---|---|
| Prompt Injection against Public-Facing Chatbots T1190 |
Yes | external-ai-exposure |
| Unauthorized Processing of Sensitive Documents | Yes | sensitive-file-activity |
| Data Exfiltration via Personal AI Accounts | Yes | personal-ai-access, network-data-volume |
Blind spots
- Needs TLS inspection or endpoint proxy logs. Without TLS decryption, we see the connection but cannot confirm the presence of sensitive text in the prompts. It would answer What specific text was included in the AI prompts?. Remediation: Enable TLS decryption for AI endpoints.
- Needs EDR with clipboard telemetry. We rely on temporal proximity; without clipboard logs, we cannot definitively prove data movement. It would answer Did the user copy-paste content from the file into the browser?. Remediation: Deploy endpoint monitoring that captures clipboard events.
Parameters & data
Parameters
| Parameter | Type | Default | What it is |
|---|---|---|---|
ai_access_timestamp | string | 2026-10-02T12:00:00Z | The UTC timestamp from the lead query result to center the file-activity window. |
lookback_days | number | 14 | Days of history to examine for AI usage and file access. |
personal_ai_domains | list[domain] | chatgpt.com, claude.ai, gemini.google.com, openai.com, anthropic.com | Domains associated with personal-tier AI services. |
personal_ai_ips | list[ip] | 104.18.37.228, 172.64.150.31 | Observed IP addresses for personal AI services to verify data volume. |
scope_hosts | list[host] | — | Optional list of hosts to narrow the hunt, populated from the lead results. |
Telemetry
| Source | Category | Telemetry |
|---|---|---|
| Endpoint telemetry (hb_ surfaces) | endpoint | endpoint |
| Network telemetry | network | network |
| Web server / proxy logs | siem | network |
Source
---
analysis: A rule flags AI domains; the hunt pivots to file activity within a narrow
temporal window to prove intent, and evaluates the external exposure that standard
internal rules miss.
blind_spots:
- id: no-tls-visibility
owner: Network Engineering
question: What specific text was included in the AI prompts?
remediation: Enable TLS decryption for AI endpoints.
requires: TLS inspection or endpoint proxy logs
risk: Without TLS decryption, we see the connection but cannot confirm the presence
of sensitive text in the prompts.
stage: exfiltration-shadow-ai-usage
- id: no-clipboard-logs
owner: Security Operations
question: Did the user copy-paste content from the file into the browser?
remediation: Deploy endpoint monitoring that captures clipboard events.
requires: EDR with clipboard telemetry
risk: We rely on temporal proximity; without clipboard logs, we cannot definitively
prove data movement.
stage: collection-sensitive-file-access
coverage:
- stage: initial-access-prompt-injection
status: covered
steps:
- external-ai-exposure
- stage: collection-sensitive-file-access
status: covered
steps:
- sensitive-file-activity
- stage: exfiltration-shadow-ai-usage
status: covered
steps:
- personal-ai-access
- network-data-volume
guardrails:
claims: no_unsupported
evidence: citation_required
missing_data: not_benign
telemetry: untrusted
hunt:
applicability: campaign-specific
handoff: keep-as-periodic-hunt
justification: 43% of workers lack AI training and 25% admit they would use personal
accounts for sensitive contracts; this hunt secures IP and mitigates the external
attack surface.
methodology: model-assisted
trigger: intel-report
hypothesis: Employees are bypassing corporate AI controls by using personal accounts
to process sensitive documents, or external attackers are exploiting public-facing
AI applications to extract internal data.
labels:
- hunt
- attack.t1190
- attack.t1566
- collection
- exfiltration
- initial access
name: Shadow AI Usage and Prompt Injection Exposure
parameters:
ai_access_timestamp:
default: '2026-10-02T12:00:00Z'
description: The UTC timestamp from the lead query result to center the file-activity
window.
type: string
lookback_days:
default: '14'
description: Days of history to examine for AI usage and file access.
type: number
personal_ai_domains:
default:
- chatgpt.com
- claude.ai
- gemini.google.com
- openai.com
- anthropic.com
description: Domains associated with personal-tier AI services.
from:
kind: article
observed: '2026-10-02'
ref: https://www.huntress.com/blog/llm-security-report
type: list[domain]
personal_ai_ips:
default:
- 104.18.37.228
- 172.64.150.31
description: Observed IP addresses for personal AI services to verify data volume.
from:
kind: article
observed: '2026-10-02'
ref: https://www.huntress.com/blog/llm-security-report
type: list[ip]
scope_hosts:
default: []
description: Optional list of hosts to narrow the hunt, populated from the lead
results.
type: list[host]
provenance:
authors:
- name: Huntbase hunt generation
org: huntbase.io
generated:
by: huntbase-hunt-generation
from: https://www.huntress.com/blog/llm-security-report
gates:
- dry-run
- lint
model: hb_google/gemini-3-flash-preview
rationale: Identify exposed AI assets in the first step to establish the organization's
external attack surface. The analyst then uses findings from the shadow AI lead
to populate the time window and host parameters for the deeper file-level investigation.
references:
- name: "Huntress \u2014 Companies Push AI Use But Skip Training and Official Policy"
url: https://www.huntress.com/blog/llm-security-report
related:
- hunt: dlp-sensitive-keyword-exfiltration
reason: Broader exfiltration hunts cover more targets; this is tuned for AI platforms.
relation: sibling
scenario:
stages:
- name: Prompt Injection against Public-Facing Chatbots
observables:
- public-facing AI chatbot
- malicious instructions slipped into prompts
- manipulated chatbot responses
- credential leakage via AI interface
slug: initial-access-prompt-injection
tactic: initial-access
techniques:
- T1190
- name: Unauthorized Processing of Sensitive Documents
observables:
- copying client contracts
- client contract summary generation
- accessing sensitive information for AI processing
slug: collection-sensitive-file-access
tactic: collection
- name: Data Exfiltration via Personal AI Accounts
observables:
- personal ChatGPT account usage
- personal Claude account usage
- chatgpt.com
- claude.ai
- gemini.google.com
- uploading data to non-enterprise AI vendors
slug: exfiltration-shadow-ai-usage
tactic: exfiltration
summary: 'This analysis identifies two primary AI-related threat vectors: prompt
injection attacks against public-facing corporate chatbots used to leak credentials
or sensitive data, and the exfiltration of sensitive information by employees
using personal AI accounts without authorization. A lack of formal policy and
training leaves organizations vulnerable to both external exploitation and internal
data leakage via personal accounts.'
severity: medium
targets:
analyst:
name: Tier-2 analyst
role: analyst
endpoint:
category: endpoint
name: Endpoint telemetry (hb_ surfaces)
telemetry:
- endpoint
hunter:
agent: true
name: Hunt agent
network:
category: network
name: Network telemetry
telemetry:
- network
web:
category: siem
name: Web server / proxy logs
telemetry:
- network
tlp: clear
type: investigation
---
# Shadow AI Usage and Prompt Injection Exposure
The hunt identifies exposed AI assets and detects internal traffic to personal AI platforms. It gates the deeper investigation on initial AI traffic to focus on high-risk hosts where an analyst must confirm whether sensitive files were moved to personal accounts. By separating external prompt-injection threats from internal shadow AI usage, the hunt ensures a focused response to both risk vectors.
## external-ai-exposure
<!-- Establish external AI attack surface -->
Identify public-facing AI applications that attackers might target for prompt injection.
```sqlite target=endpoint role=scoping
~~~yaml
expected: A list of IP addresses or domains hosting AI-related software exposed to
the internet.
reads:
- domain_or_ip
- asset_type
- product
- discovered_at
silence: evidence_of_absence
source: hb_exposed_assets
verified: dry-run
verified_at: '2026-10-04'
~~~
SELECT domain_or_ip, asset_type, product, port, discovered_at FROM hb_exposed_assets WHERE (LOWER(product) LIKE '%ai%' OR LOWER(product) LIKE '%llama%' OR LOWER(product) LIKE '%gpt%' OR LOWER(product) LIKE '%chat%' OR LOWER(product) LIKE '%notebook%' OR LOWER(product) LIKE '%langchain%')
```
## personal-ai-access
<!-- Lead: Detect personal AI platform usage -->
Find internal hosts accessing personal AI domains that lack enterprise data protections.
```sqlite target=web role=baseline params=(lookback_days=lookback_days, personal_ai_domains=personal_ai_domains, scope_hosts=scope_hosts)
~~~yaml
baseline:
compare: first_seen
window: '{{lookback_days}}d'
expected: Users and hosts accessing personal AI tools. Silence indicates absence of
traffic to these domains.
prevalence:
by: device_hostname
key:
- url_hostname
rare_below: 5
reads:
- device_hostname
- actor_user_name
- url_hostname
- time
silence: evidence_of_absence
source: hb_http_activity
verified: dry-run
verified_at: '2026-10-04'
~~~
SELECT device_hostname, actor_user_name, url_hostname, COUNT(*) as request_count, MIN(time) as first_access, MAX(time) as last_access FROM hb_http_activity WHERE (instr(',' || '{{personal_ai_domains}}' || ',', ',' || LOWER(url_hostname) || ',') > 0) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, actor_user_name, url_hostname
```
## evaluate-ai-lead
<!-- Evaluate AI lead -->
```agent target=hunter
cite: required
context:
- personal-ai-access
max_iterations: 3
objective: Determine if the HTTP activity in personal-ai-access indicates probable
personal account usage.
success_criteria: A per-host verdict of personal-usage or benign.
tools:
- endpoint
- network
- web
```
## gate-on-lead
<!-- Gate: Confirm personal usage -->
if~: "the evaluate-ai-lead verdict identifies personal-usage for at least one host" (confidence: high, judge=hunter)
then: → investigate-risk-context
indeterminate: → analyst-remediation
unavailable: → analyst-remediation (blind_spot: no-tls-visibility)
else: → hunt-close-out
## investigate-risk-context
<!-- Investigate context -->
parallel:
- → sensitive-file-activity
- → network-data-volume
join: → triage-ai-risk
## sensitive-file-activity
<!-- Correlate sensitive document access -->
Identify if hosts accessing personal AI were touching sensitive files within a one-hour window of the access.
```sqlite target=endpoint role=enrichment params=(ai_access_timestamp=ai_access_timestamp, scope_hosts=scope_hosts)
~~~yaml
expected: File touch events on sensitive documents occurring near the time of AI platform
access.
reads:
- device_hostname
- actor_user_name
- file_name
- file_path
- time
silence: not_evidence_of_absence
source: hb_file_activity
verified: dry-run
verified_at: '2026-10-04'
~~~
SELECT device_hostname, actor_user_name, file_name, file_path, time FROM hb_file_activity WHERE (LOWER(file_name) LIKE '%contract%' OR LOWER(file_name) LIKE '%confidential%' OR LOWER(file_name) LIKE '%salary%') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time BETWEEN datetime('{{ai_access_timestamp}}', '-1 hour') AND datetime('{{ai_access_timestamp}}', '+1 hour')
```
## network-data-volume
<!-- Verify data exfiltration volume -->
Examine network traffic to AI IPs to find large data uploads that confirm document exfiltration.
```sqlite target=network role=enrichment params=(lookback_days=lookback_days, personal_ai_ips=personal_ai_ips, scope_hosts=scope_hosts)
~~~yaml
expected: High total_bytes counts to known AI infrastructure indicating potential
file uploads.
reads:
- device_hostname
- dst_endpoint_ip
- traffic_bytes
- time
silence: not_evidence_of_absence
source: hb_network_connection
verified: dry-run
verified_at: '2026-10-04'
~~~
SELECT device_hostname, dst_endpoint_ip, SUM(traffic_bytes) as total_bytes, MIN(time) as start_time FROM hb_network_connection WHERE (instr(',' || '{{personal_ai_ips}}' || ',', ',' || dst_endpoint_ip || ',') > 0) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, dst_endpoint_ip
```
## triage-ai-risk
<!-- Triage AI risks -->
```agent target=hunter
cite: required
context:
- evaluate-ai-lead
- sensitive-file-activity
- network-data-volume
- external-ai-exposure
max_iterations: 6
objective: Determine if any host shows sensitive file access immediately preceding
personal AI usage with significant traffic volume.
success_criteria: A verdict of malicious | suspicious | benign per host.
tools:
- endpoint
- network
- web
```
## route-on-verdict
<!-- Route on risk verdict -->
if~: "the triage-ai-risk verdict identifies malicious data exfiltration for at least one host" (confidence: high, judge=hunter)
then: → isolate-suspect-host
indeterminate: → analyst-remediation
unavailable: → analyst-remediation (blind_spot: no-clipboard-logs)
else: → analyst-remediation
## isolate-suspect-host
<!-- Isolate suspect host -->
```action target=endpoint
~~~yaml
approval: required
~~~
Isolate the host immediately. Preserve browser cache and local history to reconstruct AI prompts.
```
→ analyst-remediation
## analyst-remediation
<!-- Analyst remediation -->
```manual target=analyst
Check exposed assets from Step 1 for configuration weaknesses. Verify if users in Step 2 violated data handling policies.
```
→ hunt-close-out
## hunt-close-out
<!-- Hunt close-out -->
```manual target=analyst
Document the number of identified shadow AI users. Provide the list of exposed assets to the firewall team.
```
→ end
Run it
Take this hunt into your environment.
Open it in Huntbase to run every step against your own connections, with Scout weighing the evidence and your analysts in command. Or take the open hunt.md file anywhere that reads the format.
Machine-drafted by huntbase-hunt-generation using hb_google/gemini-3-flash-preview, gated by dry-run, lint, then reviewed by a person.