AI-Agentic Escalation and Infrastructure Hijacking
An automated AI agent loop is conducting high-speed privilege escalation via secrets managers, tampering with CI/CD configurations, and hijacking cloud AI endpoints for external orchestration.
Based on research by Unit 42 2026-09-20 10 steps · 4 queries T0010 T0016 T0043 T1078 T1555 T1578
Brief
Why Now
Our team developed this hunt following the recent Unit 42 investigation, "An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation" (https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/). The report highlights how adversaries now use agentic AI to automate complex tasks like privilege escalation and infrastructure exploitation, turning what was once a week-long effort into an afternoon of activity.
How the Hunt Flows
The hunt begins with a scoping phase using the hb_software_inventory surface. We look for hosts running Docker or other containerization software, as these environments often host the web services used as initial entry points and tunneling hubs.
Next, we execute three parallel checks to find evidence of the agentic loop. The first query monitors the hb_http_activity surface for rapid authentication state shifts. We specifically look for source IPs or user identities that experience multiple 401 Unauthorized errors followed by a 200 OK success on sensitive API endpoints within a 15-minute window. This pattern indicates an automated tool cycling through secrets or brute-forcing access.
Simultaneously, we pivot to the hb_file_activity surface to detect tampering with DevOps configurations. The hunt searches for modifications to Terraform files or CI/CD workflow directories. Agents often target these files to plant backdoors or exfiltrate environment variables, and doing so in coordination with an authentication shift provides strong evidence of an active intrusion.
Finally, the hunt baselines AI provider traffic. We monitor connections to endpoints like OpenAI, Anthropic, and AWS Bedrock. A sudden burst of traffic to these services from an internal host, especially one involved in the previous phases, suggests the adversary is hijacking your own AI infrastructure to orchestrate the next stage of their attack.
What the Hunt Cannot See
This hunt has two primary blind spots. First, we lack visibility into the specific payloads or prompts sent to AI endpoints. While we can see the volume of traffic, we cannot distinguish between a malicious agent's instructions and legitimate development activity without deep packet inspection. Second, if the agent executes scripts entirely in the memory of ephemeral CI/CD runners, our file system queries may miss the activity if those runners lack persistent monitoring agents.
In this series
Steps
-
Identify Docker-enabled infrastructure
Query · scopingFind hosts running Docker, as the initial breach targeted containerized web services used to tunnel into the network.
reads hb_software_inventorysqlSELECT DISTINCT device_hostname FROM hb_software_inventory WHERE LOWER(package_name) LIKE '{{docker_package_pattern}}'What a hit looks like. A list of hostnames providing the attack surface. Silence indicates no Docker software was found in inventory.
-
Rapid HTTP auth state shifts
Query · detection candidateDetect automated agents successfully brute-forcing or harvesting secrets by looking for a transition from 401 to 200 within a tight 15-minute window.
reads hb_http_activitysqlSELECT device_hostname, src_endpoint_ip, actor_user_name, url_hostname, COUNT(CASE WHEN status_code = 401 THEN 1 END) AS unauthorized_count, COUNT(CASE WHEN status_code = 200 THEN 1 END) AS authorized_count, MIN(time) AS first_event, MAX(time) AS last_event FROM hb_http_activity WHERE status_code IN (200, 401) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, src_endpoint_ip, actor_user_name, url_hostname HAVING unauthorized_count > 0 AND authorized_count > 0 AND (julianday(MAX(time)) - julianday(MIN(time))) * 1440 <= 15What a hit looks like. A source IP or user showing multiple failures followed by success on a sensitive API. This is the primary indicator of an automated loop.
-
DevOps configuration tampering
Query · triageIdentify attempts to plant backdoors or exfiltrate credentials via modifications to Terraform files or CI/CD workflow configurations.
reads hb_file_activitysqlSELECT device_hostname, actor_user_name, file_path, file_name, time FROM hb_file_activity WHERE (LOWER(file_name) LIKE '%.tf' OR LOWER(file_path) LIKE '%workflows%') AND activity_id IN (1, 3, 5) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. Write or modification events on infrastructure-as-code files, particularly from service accounts or in coordination with auth shifts.
-
Bursty AI endpoint usage
Query · baselineDetect AI-endpoint hijacking by finding high-volume, rare usage of frontier AI models from internal hosts.
reads hb_http_activitysqlSELECT device_hostname, url_hostname, COUNT(*) AS request_count, MIN(time) AS first_seen FROM hb_http_activity WHERE instr(',' || '{{ai_endpoints}}' || ',', ',' || LOWER(url_hostname) || ',') > 0 AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, url_hostname HAVING request_count >= 50What a hit looks like. Bursty traffic to AI providers from a specific host. A baseline count helps distinguish normal usage from hijack-driven orchestration.
-
Weigh agentic intrusion evidence
Agent triageDetermine if the auth shifts, file edits, and bursty AI requests indicate a coordinated autonomous agent breach.
-
Route on agentic threat
DecisionExecute containment if the agent confirms malicious behavior.
-
Revoke identity and isolate resources
Response actionHalt the automated agent loop by severing its access and infrastructure.
-
Terraform and CI/CD audit
Analyst taskVerify the integrity of DevOps configurations to ensure no persistent backdoors remain.
-
Hunt close-out
Analyst taskRecord findings and document any gaps in visibility identified during the hunt.
Coverage
Scenario coverage
| Stage | Covered | How, or why not |
|---|---|---|
| Secrets Manager Privilege Escalation T1555 · T0016 |
Yes | secrets-access-shifts |
| DevOps Pipeline Hijacking T1578 · T0010 |
Yes | pipeline-tamper-check |
| AI Infrastructure Post-Compromise Abuse T1078 · T0043 |
Yes | ai-endpoint-usage-burst |
| Public Web Service Breach and Automated Recon T1190 · T1046 · T0000 · T0002 |
Out of scope | Belongs to another part of the 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation' series. |
| Enterprise Repository Secrets Harvesting T1552.001 · T0014 |
Out of scope | Belongs to another part of the 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation' series. |
Blind spots
- Needs API gateway request body logging or deep packet inspection. We can see the burst of traffic but cannot distinguish between a malicious agent's plan and a legitimate large-scale AI-assisted refactor. It would answer What specifically were the agents prompting the AI endpoints to do?.
- Needs Endpoint telemetry on CI/CD runner nodes. File activity may miss in-memory script execution or environment variable scraping if the runner is ephemeral and lacks a resident agent. It would answer Did the agent execute scripts purely in memory on the CI/CD runner?.
Parameters & data
Parameters
| Parameter | Type | Default | What it is |
|---|---|---|---|
ai_endpoints | list[domain] | api.openai.com, anthropic.com, bedrock.us-east-1.amazonaws.com, sagemaker.us-east-1.amazonaws.com, api.cohere.ai | Enterprise AI service endpoints for burst detection. |
docker_package_pattern | string | %docker% | Package name pattern to identify Docker-related infrastructure. |
lookback_days | number | 14 | Days of history to examine. |
scope_hosts | list[host] | — | Hosts identified in the scoping step; leave empty to run fleet-wide. |
Telemetry
| Source | Category | Telemetry |
|---|---|---|
| Endpoint telemetry (hb_ surfaces) | endpoint | endpoint |
| Web server / proxy logs | siem | network |
Source
---
analysis: A simple rule for failed logins misses the automated adaptation of an AI
agent. This hunt correlates rapid auth shifts with downstream infrastructure changes
and AI resource usage, providing the full context of an agentic loop.
blind_spots:
- id: no-http-payload-visibility
question: What specifically were the agents prompting the AI endpoints to do?
requires: API gateway request body logging or deep packet inspection
risk: We can see the burst of traffic but cannot distinguish between a malicious
agent's plan and a legitimate large-scale AI-assisted refactor.
stage: ai-infrastructure-hijacking
- id: ephemeral-runner-memory
question: Did the agent execute scripts purely in memory on the CI/CD runner?
requires: Endpoint telemetry on CI/CD runner nodes
risk: File activity may miss in-memory script execution or environment variable
scraping if the runner is ephemeral and lacks a resident agent.
stage: cicd-pipeline-exploitation
coverage:
- stage: secrets-manager-takeover
status: covered
steps:
- secrets-access-shifts
- stage: cicd-pipeline-exploitation
status: covered
steps:
- pipeline-tamper-check
- stage: ai-infrastructure-hijacking
status: covered
steps:
- ai-endpoint-usage-burst
- reason: 'Belongs to another part of the ''An AI-Assisted Cyber Attack: Inside a
Unit 42 Investigation'' series.'
stage: web-service-breach-and-mapping
status: out_of_scope
- reason: 'Belongs to another part of the ''An AI-Assisted Cyber Attack: Inside a
Unit 42 Investigation'' series.'
stage: repository-secrets-harvesting
status: out_of_scope
guardrails:
claims: no_unsupported
evidence: citation_required
missing_data: not_benign
telemetry: untrusted
hunt:
applicability: campaign-specific
handoff: promote-to-detection
justification: AI-assisted attacks compress weeks of work into hours, making manual
triage ineffective; automated hunts for behavioral loops are necessary to catch
the intrusion before root systems are compromised.
methodology: model-assisted
trigger: intel-report
hypothesis: An automated AI agent loop is conducting high-speed privilege escalation
via secrets managers, tampering with CI/CD configurations, and hijacking cloud AI
endpoints for external orchestration.
labels:
- hunt
- attack.t1555
- attack.t1578
- attack.t1078
- attack.t0016
- attack.t0010
- attack.t0043
name: AI-Agentic Escalation and Infrastructure Hijacking
parameters:
ai_endpoints:
default:
- api.openai.com
- anthropic.com
- bedrock.us-east-1.amazonaws.com
- sagemaker.us-east-1.amazonaws.com
- api.cohere.ai
description: Enterprise AI service endpoints for burst detection.
from:
kind: article
observed: '2026-09-02'
ref: https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/
type: list[domain]
docker_package_pattern:
default: '%docker%'
description: Package name pattern to identify Docker-related infrastructure.
from:
kind: manual
observed: '2026-09-02'
ref: Article mentions Docker as a primary product in the attack chain.
type: string
lookback_days:
default: '14'
description: Days of history to examine.
from:
kind: manual
observed: '2026-09-02'
ref: Standard hunt window
type: number
scope_hosts:
default: []
description: Hosts identified in the scoping step; leave empty to run fleet-wide.
from:
kind: manual
observed: '2026-09-02'
ref: Analyst-populated from scoping result.
type: list[host]
provenance:
authors:
- name: Huntbase hunt generation
org: huntbase.io
generated:
by: huntbase-hunt-generation
from: https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/
gates:
- dry-run
- lint
- critic
model: hb_google/gemini-3-flash-preview
rationale: Focus on infrastructure hosting containerized web services and identities
with administrative access to cloud control planes and CI/CD systems.
references:
- name: "Unit 42 \u2014 An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation"
url: https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/
related:
- hunt: web-service-breach-and-mapping
reason: The initial infiltration and internal mapping stage precedes the credential
escalation and pipeline hijacking found here.
relation: follows
- hunt: automated-service-infiltration-data-harvesting
relation: follows
scenario:
stages:
- name: Public Web Service Breach and Automated Recon
observables:
- Publicly accessible web service breach
- Automated recon agent mapping internal microservices
- Service discovery tool execution
- Bursty API requests
- Structured Markdown files for inter-agent communication
slug: web-service-breach-and-mapping
tactic: initial-access
techniques:
- T1190
- T1046
- T0000
- T0002
- name: Enterprise Repository Secrets Harvesting
observables:
- Code scraping across enterprise repositories
- Extraction of hard-coded tokens and service passwords
- Presence of Python caches and paired asset folders
- Markdown files containing harvested metadata
slug: repository-secrets-harvesting
tactic: credential-access
techniques:
- T1552.001
- T0014
- name: Secrets Manager Privilege Escalation
observables:
- Infiltration of secrets management system using stolen tokens
- Harvesting of master administrative credentials
- Parallel authentications from single identities
- Rapid 401/200 HTTP state shifts during access attempts
slug: secrets-manager-takeover
tactic: privilege-escalation
techniques:
- T1555
- T0016
- name: DevOps Pipeline Hijacking
observables:
- Unauthorized CI/CD build triggers
- Execution of custom workflows in code applications
- Attempts to modify Terraform configurations
- Exfiltration of cloud access keys via pipeline actions
slug: cicd-pipeline-exploitation
tactic: persistence
techniques:
- T1578
- T0010
- name: AI Infrastructure Post-Compromise Abuse
observables:
- LLM calls to multiple frontier AI agents in parallel
- Invocation of cloud AI models via stolen API keys
- Bursty model usage from unexpected identities
- AI endpoints used as post-compromise orchestration infrastructure
slug: ai-infrastructure-hijacking
tactic: impact
techniques:
- T1078
- T0043
summary: An attacker used autonomous AI agents to compress weeks of intrusion tradecraft
into a 10-hour campaign, breaching a web service to map the internal network and
harvest secrets. The agents then escalated privileges through a secrets manager
to hijack CI/CD pipelines and repurpose enterprise AI infrastructure for post-compromise
operations.
series:
index: 2
slug: an-ai-assisted-cyber-attack-inside-a-unit-42-investigation
title: 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation'
total: 2
severity: high
targets:
analyst:
name: Tier-2 analyst
role: analyst
endpoint:
category: endpoint
name: Endpoint telemetry (hb_ surfaces)
telemetry:
- endpoint
hunter:
agent: true
name: Hunt agent
web:
category: siem
name: Web server / proxy logs
telemetry:
- network
tlp: clear
type: investigation
---
# AI-Agentic Escalation and Infrastructure Hijacking
This hunt identifies the high-speed execution characteristic of agentic AI attacks, which compress weeks of manual red-teaming into hours. It targets the transition from failed to successful access (401 to 200 state shifts) against sensitive endpoints, followed by unauthorized modifications to Terraform configurations and bursty usage of enterprise AI services. By correlating these behaviors across authentication and file surfaces, we detect the methodical orchestration of an autonomous intruder.
## docker-host-scoping
<!-- Identify Docker-enabled infrastructure -->
Find hosts running Docker, as the initial breach targeted containerized web services used to tunnel into the network.
```sqlite target=endpoint role=scoping params=(docker_package_pattern=docker_package_pattern)
~~~yaml
expected: A list of hostnames providing the attack surface. Silence indicates no Docker
software was found in inventory.
reads:
- device_hostname
- package_name
silence: not_evidence_of_absence
source: hb_software_inventory
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT DISTINCT device_hostname FROM hb_software_inventory WHERE LOWER(package_name) LIKE '{{docker_package_pattern}}'
```
## parallel-activity-check
<!-- Correlate loop behaviors -->
parallel:
- → secrets-access-shifts
- → pipeline-tamper-check
- → ai-endpoint-usage-burst
join: → triage-agent
## secrets-access-shifts
<!-- Rapid HTTP auth state shifts -->
Detect automated agents successfully brute-forcing or harvesting secrets by looking for a transition from 401 to 200 within a tight 15-minute window.
```sqlite target=web role=detection-candidate params=(lookback_days=lookback_days, scope_hosts=scope_hosts)
~~~yaml
expected: A source IP or user showing multiple failures followed by success on a sensitive
API. This is the primary indicator of an automated loop.
reads:
- device_hostname
- src_endpoint_ip
- actor_user_name
- url_hostname
- status_code
- time
silence: not_evidence_of_absence
source: hb_http_activity
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT device_hostname, src_endpoint_ip, actor_user_name, url_hostname, COUNT(CASE WHEN status_code = 401 THEN 1 END) AS unauthorized_count, COUNT(CASE WHEN status_code = 200 THEN 1 END) AS authorized_count, MIN(time) AS first_event, MAX(time) AS last_event FROM hb_http_activity WHERE status_code IN (200, 401) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, src_endpoint_ip, actor_user_name, url_hostname HAVING unauthorized_count > 0 AND authorized_count > 0 AND (julianday(MAX(time)) - julianday(MIN(time))) * 1440 <= 15
```
## pipeline-tamper-check
<!-- DevOps configuration tampering -->
Identify attempts to plant backdoors or exfiltrate credentials via modifications to Terraform files or CI/CD workflow configurations.
```sqlite target=endpoint role=triage params=(lookback_days=lookback_days, scope_hosts=scope_hosts)
~~~yaml
expected: Write or modification events on infrastructure-as-code files, particularly
from service accounts or in coordination with auth shifts.
reads:
- device_hostname
- actor_user_name
- file_path
- file_name
- time
- activity_id
silence: not_evidence_of_absence
source: hb_file_activity
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT device_hostname, actor_user_name, file_path, file_name, time FROM hb_file_activity WHERE (LOWER(file_name) LIKE '%.tf' OR LOWER(file_path) LIKE '%workflows%') AND activity_id IN (1, 3, 5) AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')
```
## ai-endpoint-usage-burst
<!-- Bursty AI endpoint usage -->
Detect AI-endpoint hijacking by finding high-volume, rare usage of frontier AI models from internal hosts.
```sqlite target=web role=baseline params=(lookback_days=lookback_days, scope_hosts=scope_hosts, ai_endpoints=ai_endpoints)
~~~yaml
baseline:
compare: first_seen
window: '{{lookback_days}}d'
expected: Bursty traffic to AI providers from a specific host. A baseline count helps
distinguish normal usage from hijack-driven orchestration.
prevalence:
by: device_hostname
key:
- url_hostname
rare_below: 2
reads:
- device_hostname
- url_hostname
- time
silence: not_evidence_of_absence
source: hb_http_activity
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT device_hostname, url_hostname, COUNT(*) AS request_count, MIN(time) AS first_seen FROM hb_http_activity WHERE instr(',' || '{{ai_endpoints}}' || ',', ',' || LOWER(url_hostname) || ',') > 0 AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY device_hostname, url_hostname HAVING request_count >= 50
```
## triage-agent
<!-- Weigh agentic intrusion evidence -->
```agent target=hunter
cite: required
context:
- docker-host-scoping
- secrets-access-shifts
- pipeline-tamper-check
- ai-endpoint-usage-burst
max_iterations: 6
objective: Decide whether the identified telemetry suggests an AI-orchestrated intrusion
based on temporal correlation and behavioral shifts.
success_criteria: A per-host verdict of malicious, suspicious, or benign citing row
evidence.
tools:
- endpoint
- web
```
## route-on-verdict
<!-- Route on agentic threat -->
if~: "the triage-agent verdict is malicious for any host or identity" (confidence: high, judge=hunter)
then: → revoke-and-isolate
indeterminate: → terraform-state-review
unavailable: → terraform-state-review (blind_spot: no-http-payload-visibility)
else: → terraform-state-review
## revoke-and-isolate
<!-- Revoke identity and isolate resources -->
```action target=endpoint
~~~yaml
approval: required
~~~
Revoke the affected identity tokens and isolate any associated cloud resources or CI/CD runners identified in the triage.
```
→ terraform-state-review
## terraform-state-review
<!-- Terraform and CI/CD audit -->
```manual target=analyst
Review all Terraform commits and CI/CD workflow changes in the window. Verify if branch protection was bypassed or if keys were exfiltrated.
```
→ close-out
## close-out
<!-- Hunt close-out -->
```manual target=analyst
Record the identities and IPs involved. If burst thresholds were too low for normal development activity, adjust them for future runs.
```
→ end
Run it
Take this hunt into your environment.
Open it in Huntbase to run every step against your own connections, with Scout weighing the evidence and your analysts in command. Or take the open hunt.md file anywhere that reads the format.
Machine-drafted by huntbase-hunt-generation using hb_google/gemini-3-flash-preview, gated by dry-run, lint, critic, then reviewed by a person.