← All hunts high TLP:CLEAR Part 1 of 2

Automated Service Infiltration and Data Harvesting

An intruder is using autonomous AI agents to breach public web services and map internal microservices while harvesting credentials, leaving behind unique filesystem artifacts and high-frequency network recon patterns.

Based on research by Unit 42 2026-09-20 9 steps · 3 queries T1046 T1078 T1190 T1552.001

Brief

Why now

Recent investigations by Unit 42 in An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation demonstrate how adversaries use autonomous agents to compress the attack lifecycle. AI agents handle reconnaissance, vulnerability research, and internal mapping at speeds that traditional detection rules often miss. This hunt targets the specific behavioral loop of an agentic breach.

How the hunt flows

The hunt begins by narrowing the attack surface. It queries software inventory data to identify hosts running common web service or API packages like Nginx, Apache, or Tomcat. These hosts represent the most likely entry points for an agent-driven exploit.

Once the scope is set, the hunt executes two parallel checks. The first looks for filesystem artifacts characteristic of AI frameworks. It searches for the creation of Markdown files like plan.md or task.md and identifies localized Python cache directories created in non-standard paths. These files represent the agent passing context between execution loops.

Simultaneously, the hunt analyzes internal network traffic. It looks for bursty HTTP activity where a single source endpoint requests a high diversity of URL paths in a short window. This pattern indicates an automated agent mapping internal microservices or API endpoints.

In the final phase, an analyst or automated agent triages the findings. The hunt correlates the presence of orchestration files with the bursty network traffic on the scoped web hosts. A positive match suggests an active, autonomous intrusion loop requiring immediate isolation and artifact analysis.

What the hunt cannot see

This hunt relies on an accurate software inventory. If an adversary deploys transient containers or uses software not tracked by package managers, the initial scoping query may miss the beachhead. Additionally, if internal HTTP logs are truncated or if the agent uses non-standard ports not monitored by existing proxies, the reconnaissance signal will be incomplete.

In this series

Steps

  1. Scope web service attack surface

    Query · scoping

    Identify hosts running web server software that represent the primary entry point for the reported infiltration.

    reads hb_software_inventorysql
    SELECT DISTINCT device_hostname FROM hb_software_inventory WHERE (LOWER(package_name) LIKE '%nginx%' OR LOWER(package_name) LIKE '%apache%' OR LOWER(package_name) LIKE '%httpd%' OR LOWER(package_name) LIKE '%tomcat%' OR LOWER(package_name) LIKE '%api%')

    What a hit looks like. A list of hostnames running web or API services. Absence means no such packages are installed via tracked package managers.

  2. Detect agent orchestration files

    Query · baseline

    Identify the creation of Markdown reports or Python caches used by agents to pass information between sessions.

    reads hb_file_activitysql
    SELECT device_hostname, file_name, file_path, time FROM hb_file_activity WHERE (instr(',' || '{{agent_indicators}}' || ',', ',' || LOWER(file_name) || ',') > 0 OR (LOWER(file_path) LIKE '%__pycache__%' AND LOWER(file_path) NOT LIKE '%\\usr\\lib\\%' AND LOWER(file_path) NOT LIKE '%\\windows\\%')) AND time >= datetime('now', '-{{lookback_days}} days') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) ORDER BY time DESC

    What a hit looks like. Specific Markdown filenames or localized Python cache directories created on web-facing hosts. Rare files across the fleet are more suspicious.

  3. Detect bursty internal reconnaissance

    Query · detection candidate

    Identify high-volume internal HTTP activity that indicates an automated microservice mapping agent.

    reads hb_http_activitysql
    SELECT device_hostname, src_endpoint_ip, COUNT(*) AS request_count, COUNT(DISTINCT url_path) AS path_diversity, MIN(time) AS first_request, MAX(time) AS last_request FROM hb_http_activity WHERE time >= datetime('now', '-{{lookback_days}} days') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) GROUP BY device_hostname, src_endpoint_ip HAVING request_count > 500 ORDER BY request_count DESC

    What a hit looks like. A high count of requests from a single source to many different paths in a short window. Silence suggests no automated web scanning occurred within the logs.

  4. Triage AI agent signals

    Agent triage

    Analyze whether the combination of exposed web services, orchestration files, and bursty traffic confirms an AI-driven intrusion.

  5. Route on triage verdict

    Decision

    Direct the hunt based on the agent's findings.

  6. Isolate compromised host

    Response action

    Stop the automated agent loop by severing its network connectivity and access to repositories.

  7. Analyze agent artifacts

    Analyst task

    Examine the content of identified Markdown files to determine what has been mapped or stolen.

  8. Hunt close out

    Analyst task

    Document findings and determine if follow-on hunts for secrets manager takeover are required.

Coverage

Scenario coverage

StageCoveredHow, or why not
Public Web Service Breach and Automated Recon
T1190 · T1046 · T0000 · T0002
Yes scope-web-services, detect-bursty-recon
Enterprise Repository Secrets Harvesting
T1552.001 · T0014
Yes detect-agent-orchestration-files, analyze-artifacts
Secrets Manager Privilege Escalation
T1555 · T0016
Out of scope Belongs to another part of the 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation' series.
DevOps Pipeline Hijacking
T1578 · T0010
Out of scope Belongs to another part of the 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation' series.
AI Infrastructure Post-Compromise Abuse
T1078 · T0043
Out of scope Belongs to another part of the 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation' series.

Blind spots

  • Needs hb_software_inventory with real-time updates. Snapshot-based software inventory may miss transient or newly spawned containers used as entry points. It would answer whether a newly deployed container is missing from the scoping query. Remediation: Implement continuous container image scanning and real-time inventory updates.
  • Needs hb_http_activity with full url_path and headers. If HTTP logs are truncated or if mapping occurs over non-standard ports not logged by proxies, the bursty recon signal will be incomplete. It would answer whether the agent successfully mapped specific internal microservices. Remediation: Ensure full URL logging is enabled for all internal API traffic.

Parameters & data

Parameters

ParameterTypeDefaultWhat it is
agent_indicatorslist[string]plan.md, task.md, agent.md, report.mdFilenames typically used by agentic AI frameworks for inter-session context passing.
lookback_daysnumber14Days of history to examine.
scope_hostslist[host]List of hostnames to focus on after the scoping step; leave empty to search the entire estate.

Telemetry

SourceCategoryTelemetry
Endpoint telemetry (hb_ surfaces)endpointendpoint
Web server / proxy logssiemnetwork

Source

Download hunt.md Definition (JSON) An open hunt.md file; it runs anywhere that reads the format.
---
analysis: A standard detection rule might flag a single exploit; this hunt pivots
  between infrastructure exposure, specific AI-orchestration filesystem artifacts,
  and bursty network patterns to confirm an autonomous intrusion loop.
blind_spots:
- id: inventory-lag
  question: whether a newly deployed container is missing from the scoping query
  remediation: Implement continuous container image scanning and real-time inventory
    updates.
  requires: hb_software_inventory with real-time updates
  risk: Snapshot-based software inventory may miss transient or newly spawned containers
    used as entry points.
  stage: web-service-breach-and-mapping
- id: log-truncation
  question: whether the agent successfully mapped specific internal microservices
  remediation: Ensure full URL logging is enabled for all internal API traffic.
  requires: hb_http_activity with full url_path and headers
  risk: If HTTP logs are truncated or if mapping occurs over non-standard ports not
    logged by proxies, the bursty recon signal will be incomplete.
  stage: web-service-breach-and-mapping
coverage:
- stage: web-service-breach-and-mapping
  status: covered
  steps:
  - scope-web-services
  - detect-bursty-recon
- stage: repository-secrets-harvesting
  status: covered
  steps:
  - detect-agent-orchestration-files
  - analyze-artifacts
- reason: 'Belongs to another part of the ''An AI-Assisted Cyber Attack: Inside a
    Unit 42 Investigation'' series.'
  stage: secrets-manager-takeover
  status: out_of_scope
- reason: 'Belongs to another part of the ''An AI-Assisted Cyber Attack: Inside a
    Unit 42 Investigation'' series.'
  stage: cicd-pipeline-exploitation
  status: out_of_scope
- reason: 'Belongs to another part of the ''An AI-Assisted Cyber Attack: Inside a
    Unit 42 Investigation'' series.'
  stage: ai-infrastructure-hijacking
  status: out_of_scope
guardrails:
  claims: no_unsupported
  evidence: citation_required
  missing_data: not_benign
  telemetry: untrusted
hunt:
  applicability: campaign-specific
  handoff: promote-to-detection
  justification: AI agents compress the attack timeline from weeks to hours; identifying
    the behavioral loop of an autonomous agent is the only way to stop a breach before
    it reaches the secrets management or pipeline layers.
  methodology: model-assisted
  trigger: intel-report
hypothesis: An intruder is using autonomous AI agents to breach public web services
  and map internal microservices while harvesting credentials, leaving behind unique
  filesystem artifacts and high-frequency network recon patterns.
labels:
- hunt
- attack.t1190
- attack.t1046
- attack.t1552.001
- attack.t1078
name: Automated Service Infiltration and Data Harvesting
parameters:
  agent_indicators:
    default:
    - plan.md
    - task.md
    - agent.md
    - report.md
    description: Filenames typically used by agentic AI frameworks for inter-session
      context passing.
    from:
      kind: article
      observed: '2026-09-02'
      ref: unit42-ai-attack
    type: list[string]
  lookback_days:
    default: '14'
    description: Days of history to examine.
    type: number
  scope_hosts:
    default: []
    description: List of hostnames to focus on after the scoping step; leave empty
      to search the entire estate.
    type: list[host]
provenance:
  authors:
  - name: Huntbase hunt generation
    org: huntbase.io
  generated:
    by: huntbase-hunt-generation
    from: https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/
    gates:
    - dry-run
    - lint
    model: hb_google/gemini-3-flash-preview
rationale: The hunt begins by identifying hosts with web service packages. Focus investigation
  on those that show bursty HTTP traffic or localized Python artifacts, which are
  typical for agentic AI intrusions.
references:
- name: "Unit 42 \u2014 An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation"
  url: https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/
related:
- hunt: secrets-manager-takeover
  reason: This hunt identifies the beachhead and credential harvesting; the next stage
    is the abuse of harvested secrets to escalate privileges.
  relation: follows
scenario:
  stages:
  - name: Public Web Service Breach and Automated Recon
    observables:
    - Publicly accessible web service breach
    - Automated recon agent mapping internal microservices
    - Service discovery tool execution
    - Bursty API requests
    - Structured Markdown files for inter-agent communication
    slug: web-service-breach-and-mapping
    tactic: initial-access
    techniques:
    - T1190
    - T1046
    - T0000
    - T0002
  - name: Enterprise Repository Secrets Harvesting
    observables:
    - Code scraping across enterprise repositories
    - Extraction of hard-coded tokens and service passwords
    - Presence of Python caches and paired asset folders
    - Markdown files containing harvested metadata
    slug: repository-secrets-harvesting
    tactic: credential-access
    techniques:
    - T1552.001
    - T0014
  - name: Secrets Manager Privilege Escalation
    observables:
    - Infiltration of secrets management system using stolen tokens
    - Harvesting of master administrative credentials
    - Parallel authentications from single identities
    - Rapid 401/200 HTTP state shifts during access attempts
    slug: secrets-manager-takeover
    tactic: privilege-escalation
    techniques:
    - T1555
    - T0016
  - name: DevOps Pipeline Hijacking
    observables:
    - Unauthorized CI/CD build triggers
    - Execution of custom workflows in code applications
    - Attempts to modify Terraform configurations
    - Exfiltration of cloud access keys via pipeline actions
    slug: cicd-pipeline-exploitation
    tactic: persistence
    techniques:
    - T1578
    - T0010
  - name: AI Infrastructure Post-Compromise Abuse
    observables:
    - LLM calls to multiple frontier AI agents in parallel
    - Invocation of cloud AI models via stolen API keys
    - Bursty model usage from unexpected identities
    - AI endpoints used as post-compromise orchestration infrastructure
    slug: ai-infrastructure-hijacking
    tactic: impact
    techniques:
    - T1078
    - T0043
  summary: An attacker used autonomous AI agents to compress weeks of intrusion tradecraft
    into a 10-hour campaign, breaching a web service to map the internal network and
    harvest secrets. The agents then escalated privileges through a secrets manager
    to hijack CI/CD pipelines and repurpose enterprise AI infrastructure for post-compromise
    operations.
series:
  index: 1
  slug: an-ai-assisted-cyber-attack-inside-a-unit-42-investigation
  title: 'An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation'
  total: 2
severity: high
targets:
  analyst:
    name: Tier-2 analyst
    role: analyst
  endpoint:
    category: endpoint
    name: Endpoint telemetry (hb_ surfaces)
    telemetry:
    - endpoint
  hunter:
    agent: true
    name: Hunt agent
  web:
    category: siem
    name: Web server / proxy logs
    telemetry:
    - network
tlp: clear
type: investigation
---


# Automated Service Infiltration and Data Harvesting

This hunt focuses on the initial infiltration and mapping stages of an AI-agentic attack. It begins by scoping the attack surface to hosts running common web service packages. It then searches for the behavioral indicators of AI orchestration: the creation of structured Markdown communication files and rapid Python cache generation. Simultaneously, it identifies automated internal reconnaissance by searching for bursty HTTP traffic patterns. An agent then weighs these findings to identify the beachhead and the extent of the internal mapping.

## scope-web-services
<!-- Scope web service attack surface -->
Identify hosts running web server software that represent the primary entry point for the reported infiltration.

```sqlite target=endpoint role=scoping
~~~yaml
expected: A list of hostnames running web or API services. Absence means no such packages
  are installed via tracked package managers.
reads:
- device_hostname
- package_name
silence: not_evidence_of_absence
source: hb_software_inventory
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT DISTINCT device_hostname FROM hb_software_inventory WHERE (LOWER(package_name) LIKE '%nginx%' OR LOWER(package_name) LIKE '%apache%' OR LOWER(package_name) LIKE '%httpd%' OR LOWER(package_name) LIKE '%tomcat%' OR LOWER(package_name) LIKE '%api%')
```

## agent-behavior-check
<!-- Check for agentic indicators and recon -->
parallel:
- → detect-agent-orchestration-files
- → detect-bursty-recon
join: → triage-agent-signals

## detect-agent-orchestration-files
<!-- Detect agent orchestration files -->
Identify the creation of Markdown reports or Python caches used by agents to pass information between sessions.

```sqlite target=endpoint role=baseline params=(lookback_days=lookback_days, scope_hosts=scope_hosts, agent_indicators=agent_indicators)
~~~yaml
baseline:
  compare: first_seen
  window: '{{lookback_days}}d'
expected: Specific Markdown filenames or localized Python cache directories created
  on web-facing hosts. Rare files across the fleet are more suspicious.
prevalence:
  by: device_hostname
  key:
  - file_name
  rare_below: 3
reads:
- device_hostname
- file_name
- file_path
- time
silence: not_evidence_of_absence
source: hb_file_activity
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT device_hostname, file_name, file_path, time FROM hb_file_activity WHERE (instr(',' || '{{agent_indicators}}' || ',', ',' || LOWER(file_name) || ',') > 0 OR (LOWER(file_path) LIKE '%__pycache__%' AND LOWER(file_path) NOT LIKE '%\\usr\\lib\\%' AND LOWER(file_path) NOT LIKE '%\\windows\\%')) AND time >= datetime('now', '-{{lookback_days}} days') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) ORDER BY time DESC
```

## detect-bursty-recon
<!-- Detect bursty internal reconnaissance -->
Identify high-volume internal HTTP activity that indicates an automated microservice mapping agent.

```sqlite target=web role=detection-candidate params=(lookback_days=lookback_days, scope_hosts=scope_hosts)
~~~yaml
expected: A high count of requests from a single source to many different paths in
  a short window. Silence suggests no automated web scanning occurred within the logs.
reads:
- device_hostname
- src_endpoint_ip
- url_path
- time
silence: not_evidence_of_absence
source: hb_http_activity
verified: dry-run
verified_at: '2026-09-20'
~~~
SELECT device_hostname, src_endpoint_ip, COUNT(*) AS request_count, COUNT(DISTINCT url_path) AS path_diversity, MIN(time) AS first_request, MAX(time) AS last_request FROM hb_http_activity WHERE time >= datetime('now', '-{{lookback_days}} days') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) GROUP BY device_hostname, src_endpoint_ip HAVING request_count > 500 ORDER BY request_count DESC
```

## triage-agent-signals
<!-- Triage AI agent signals -->
```agent target=hunter
cite: required
context:
- detect-agent-orchestration-files
- detect-bursty-recon
max_iterations: 4
objective: Determine if the bursty HTTP traffic and Markdown context files together
  represent an autonomous AI agent breach on the scoped hosts.
success_criteria: A verdict of malicious, suspicious, or benign per host with cited
  rows.
tools:
- endpoint
- web
```

## route-on-verdict
<!-- Route on triage verdict -->
if~: "the triage-agent-signals verdict is malicious for at least one host" (confidence: high, judge=hunter)
then: → isolate-host
indeterminate: → analyze-artifacts
unavailable: → analyze-artifacts (blind_spot: log-truncation)
else: → close-out

## isolate-host
<!-- Isolate compromised host -->
```action target=endpoint
~~~yaml
approval: required
~~~
Isolate the identified host immediately to stop the autonomous agent from further internal mapping or secrets extraction.
```
→ analyze-artifacts

## analyze-artifacts
<!-- Analyze agent artifacts -->
```manual target=analyst
Review the content of the .md files and Python cache directories found in the filesystem step; look for lists of internal IPs, extracted tokens, or service discovery summaries.
```
→ close-out

## close-out
<!-- Hunt close out -->
```manual target=analyst
Record the scoped hosts and findings. If credentials were found in the analyzed artifacts, trigger the 'Secrets Manager Takeover' follow-on hunt.
```
→ end

Run it

Take this hunt into your environment.

Open it in Huntbase to run every step against your own connections, with Scout weighing the evidence and your analysts in command. Or take the open hunt.md file anywhere that reads the format.

Machine-drafted by huntbase-hunt-generation using hb_google/gemini-3-flash-preview, gated by dry-run, lint, then reviewed by a person.