AI Coding Agent Sandbox Activity
An unauthorized user is running an AI coding harness to modify production codebases, using automated testing to validate the changes and communicating with external LLM APIs.
Based on research by Huntress 2026-09-29 12 steps · 5 queries T1059.006 T1071.001 T1204.002 T1610
Brief
Why this hunt matters
The recent analysis by Huntress, Fighting AI Slop in Production Codebases, highlights a new challenge for engineering teams: the rapid introduction of automated, often unvetted code via AI coding agents. While these tools promise efficiency, they can bypass traditional code review and introduce insecure patterns if deployed without oversight. This hunt provides the visibility needed to track these agents as they operate within your environment.
How the hunt flows
The hunt begins by identifying Docker-capable infrastructure. Because agentic coding harnesses often require isolated environments to execute and test code safely, they frequently provision local containers. The first query scopes the environment to hosts running Docker, providing the necessary sandbox capability. This narrows the investigation to endpoints where automated code manipulation is most likely to occur.
Once the scope is set, the hunt searches for the operational footprint of the harness itself. It looks for rare or unauthorized process names and command-line keywords associated with tools like Lemans or Claude-Code. Simultaneously, it monitors for the modification of configuration files such as CLAUDE.md or specific skill directories. These files serve as the 'brain' for the agent, containing the instructions and project rules the LLM follows during the session.
After establishing that a harness is active, the hunt correlates subsequent operational activity. It searches for automated test validation runs, specifically targeting RSpec executions that follow the modification of codebase artifacts. An automated agent typically runs tests immediately after a change to verify its work. The hunt then joins this endpoint activity with DNS lookups to known LLM providers like Anthropic or OpenAI, confirming the agent's external communication loop.
This workflow is structured as a hunt rather than a single detection because the individual actions—running Docker, updating markdown files, or executing RSpec—are common developer tasks. A single detection rule would generate excessive noise. By treating these events as a session-based sequence, an analyst can weight the cumulative evidence to identify unauthorized automation that mimics human behavior.
Blind spots and limitations
This hunt has two primary blind spots. First, it relies on snapshots of host activity. If an adversary runs a short-lived container that finishes between inventory intervals, the initial sandbox provisioning may be missed. Second, while the hunt identifies the destination of network traffic, it cannot see the content of the LLM prompts. Without TLS inspection, we cannot determine if proprietary source code or sensitive credentials were sent to the LLM provider.
Steps
-
Identify Docker-capable hosts
Query · scopingScope the hunt to hosts running Docker, which provides the required sandbox infrastructure for AI coding harnesses.
reads hb_software_inventorysqlSELECT DISTINCT device_hostname, device_uid FROM hb_software_inventory WHERE (LOWER(package_name) LIKE '%docker%' OR LOWER(vendor_name) LIKE '%docker%') AND asset_scope = 'endpoint'What a hit looks like. A list of hosts with Docker installed. Silence suggests no containerization capability is present via this package manager.
-
Harness process prevalence
Query · baselineStack-count harness processes to identify rare or unauthorized agent activity across the fleet.
reads hb_process_activitysqlSELECT process_name, process_path, user_name, COUNT(DISTINCT device_hostname) AS host_count, MIN(time) AS first_seen FROM hb_process_activity WHERE (LOWER(process_name) LIKE '%lemans%' OR LOWER(process_name) LIKE '%claude-code%' OR LOWER(process_name) LIKE '%fable-agent%' OR LOWER(process_cmd_line) LIKE '%lemans%' OR instr(',' || '{{harness_keywords}}' || ',', ',' || LOWER(process_name) || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY process_name, process_path, user_nameWhat a hit looks like. Harness names or paths seen on very few hosts. Frequent occurrences on many hosts may indicate authorized developer workstations.
-
Monitor instruction file updates
Query · detection candidateDetect modifications to the instruction files that agents use to define coding rules and API recall preferences.
reads hb_file_activitysqlSELECT device_hostname, file_name, file_path, process_name, time FROM hb_file_activity WHERE (LOWER(file_name) = 'claude.md' OR LOWER(file_path) LIKE '%/.claude/skills/%') AND activity_id IN (1, 3, 5) AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. Creation or modification of CLAUDE.md files. Silence means no agent-specific configuration was observed in this window.
-
Assess session establishment
Agent triageEstablish whether the process and file activity confirm the start of an AI-driven coding session.
-
Detect automated RSpec runs
Query · triageIdentify the automated testing phase that follows codebase modification in coding harnesses.
reads hb_process_activitysqlSELECT device_hostname, process_name, process_cmd_line, time FROM hb_process_activity WHERE (LOWER(process_name) LIKE '%rspec%' OR LOWER(process_cmd_line) LIKE '%rspec%') AND (LOWER(process_cmd_line) LIKE '%invitation.create%' OR LOWER(process_cmd_line) LIKE '%schema.rb%') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. RSpec command lines targeting artifacts mentioned in the article. Silence suggests the session did not reach the validation phase or used a different test runner.
-
DNS lookups to LLM providers
Query · enrichmentCorroborate the session by finding network traffic to the LLM providers used by coding agents.
reads hb_dns_activitysqlSELECT device_hostname, query_hostname, process_name, time FROM hb_dns_activity WHERE instr(',' || '{{llm_domains}}' || ',', ',' || LOWER(query_hostname) || ',') > 0 AND time >= datetime('now', '-{{lookback_days}} days')What a hit looks like. DNS queries for LLM domains from identified hosts. Silence means the agent may be using a different provider or a local proxy.
-
Triage agentic workflow
Agent triageDetermine if the entire sequence from sandbox start to LLM communication indicates unauthorized activity.
-
Route on triage verdict
DecisionRoute the findings based on whether the agent activity is confirmed as unauthorized.
-
Isolate endpoint
Response actionHalt further automated code manipulation by isolating the host.
-
Review code modifications
Analyst taskThe analyst manually inspects the codebase for changes introduced by the agent to verify their impact and authorization.
-
Document and close
Analyst taskFinalize the hunt and record whether the session was a legitimate engineering test.
Coverage
Scenario coverage
| Stage | Covered | How, or why not |
|---|---|---|
| Docker Sandbox Initialization T1610 |
Yes | scope-docker-hosts, harness-process-baseline |
| Automated Codebase Modification T1059.006 |
Yes | agent-instruction-writes |
| Automated Code Validation T1204.002 |
Yes | rspec-validation-runs |
| LLM API Coordination T1071.001 |
Yes | llm-api-connections |
Blind spots
- Needs Docker event logs. An agent session lasting only a few minutes might not be captured in hb_software_inventory or hb_process_activity snapshots. It would answer Did a container run and finish between inventory snapshots?.
- Needs TLS inspection for LLM domains. DNS and network connection logs confirm the destination but hide the content of the LLM interaction, which may leak proprietary code. It would answer What source code or sensitive data was included in the prompt?.
Parameters & data
Parameters
| Parameter | Type | Default | What it is |
|---|---|---|---|
harness_keywords | list[string] | lemans, claude-code, fable-agent, anthropic-agent | Process filenames or command-line keywords for coding harnesses. |
llm_domains | list[domain] | api.anthropic.com, api.openai.com, api.mistral.ai, api.groq.com | Domains of LLM providers commonly used by coding agents. |
lookback_days | number | 14 | Days of history to examine. |
scope_hosts | list[host] | — | Optional list of hostnames to focus the investigation on. |
Telemetry
| Source | Category | Telemetry |
|---|---|---|
| Endpoint telemetry (hb_ surfaces) | endpoint | endpoint |
Source
---
analysis: A single detection rule cannot correlate the sequence of Docker setup, CLAUDE.md
instruction updates, RSpec execution, and LLM API traffic. This phased hunt uses
the session context to weight the risk of subsequent automated activities that look
like standard developer behavior when viewed in isolation.
blind_spots:
- id: short-lived-containers
question: Did a container run and finish between inventory snapshots?
requires: Docker event logs
risk: An agent session lasting only a few minutes might not be captured in hb_software_inventory
or hb_process_activity snapshots.
stage: sandbox-container-provisioning
- id: encrypted-prompt-content
question: What source code or sensitive data was included in the prompt?
requires: TLS inspection for LLM domains
risk: DNS and network connection logs confirm the destination but hide the content
of the LLM interaction, which may leak proprietary code.
stage: agent-llm-communication
coverage:
- stage: sandbox-container-provisioning
status: covered
steps:
- scope-docker-hosts
- harness-process-baseline
- stage: automated-code-modification
status: covered
steps:
- agent-instruction-writes
- stage: automated-test-validation
status: covered
steps:
- rspec-validation-runs
- stage: agent-llm-communication
status: covered
steps:
- llm-api-connections
guardrails:
claims: no_unsupported
evidence: citation_required
missing_data: not_benign
telemetry: untrusted
hunt:
applicability: campaign-specific
handoff: promote-to-detection
justification: AI coding agents can rapidly introduce handrolled slop or insecure
code into production environments; identifying the operational footprint of these
agents ensures automated changes are visible and vetted.
methodology: model-assisted
trigger: intel-report
hypothesis: An unauthorized user is running an AI coding harness to modify production
codebases, using automated testing to validate the changes and communicating with
external LLM APIs.
labels:
- hunt
- attack.t1610
- attack.t1059.006
- attack.t1204.002
- attack.t1071.001
- command and control
- execution
name: AI Coding Agent Sandbox Activity
parameters:
harness_keywords:
default:
- lemans
- claude-code
- fable-agent
- anthropic-agent
description: Process filenames or command-line keywords for coding harnesses.
from:
kind: article
observed: '2026-09-22'
ref: huntress-fable-api-recall
type: list[string]
llm_domains:
default:
- api.anthropic.com
- api.openai.com
- api.mistral.ai
- api.groq.com
description: Domains of LLM providers commonly used by coding agents.
from:
kind: article
observed: '2026-09-22'
ref: huntress-fable-api-recall
type: list[domain]
lookback_days:
default: '14'
description: Days of history to examine.
from:
kind: manual
observed: '2026-09-22'
ref: hunt-standard-lookback
type: number
scope_hosts:
default: []
description: Optional list of hostnames to focus the investigation on.
from:
kind: manual
observed: '2026-09-22'
ref: analyst-defined-scope
type: list[host]
provenance:
authors:
- name: Huntbase hunt generation
org: huntbase.io
generated:
by: huntbase-hunt-generation
from: https://www.huntress.com/blog/claude-fable-api-recall
gates:
- dry-run
- lint
model: hb_google/gemini-3-flash-preview
rationale: Target hosts with Docker installed or those known for development work.
Engineering subnets are the highest priority scope.
references:
- name: "Huntress \u2014 Fighting AI Slop in Production Codebases"
url: https://www.huntress.com/blog/claude-fable-api-recall
related:
- hunt: unauthorized-llm-data-exfiltration
reason: This hunt focuses on codebase modification within a harness, not general
data theft via LLM prompts.
relation: out-of-scope-alternative
scenario:
stages:
- name: Docker Sandbox Initialization
observables:
- Docker sandbox starts from a pinned commit
- Databases running in container
- rails/lemans harness
- docker_info
slug: sandbox-container-provisioning
tactic: execution
techniques:
- T1610
- name: Automated Codebase Modification
observables:
- CLAUDE.md
- schema.rb
- has_secure_token
- generates_token_for
- normalizes
- perform_all_later
- 'comparison:'
- self.token
- invitation.create
- .claude/skills
slug: automated-code-modification
tactic: execution
techniques:
- T1059.006
- name: Automated Code Validation
observables:
- RSpec.describe
- invitation.reload.token
- first.token
- second.token
- invitation.token
- invitation.errors
- bundle exec rspec
- rubocop execution
slug: automated-test-validation
tactic: execution
techniques:
- T1204.002
- name: LLM API Coordination
observables:
- HTTPS LLM calls
- Anthropic API communication
- Restricted network access lookups
slug: agent-llm-communication
tactic: command-and-control
techniques:
- T1071.001
summary: This scenario describes a research environment where the Claude Fable 5.1
AI model is deployed within a Dockerized Ruby on Rails harness to evaluate its
API recall accuracy. The agent modifies the codebase based on natural language
tickets, followed by automated validation using RSpec tests and a secondary LLM
reviewer, with all activity confined to isolated containers and monitored LLM
API communications.
severity: medium
targets:
analyst:
name: Tier-2 analyst
role: analyst
endpoint:
category: endpoint
name: Endpoint telemetry (hb_ surfaces)
telemetry:
- endpoint
hunter:
agent: true
name: Hunt agent
tlp: clear
type: investigation
---
# AI Coding Agent Sandbox Activity
This hunt identifies the operational footprint of agentic coding harnesses such as rails/lemans. It tracks the lifecycle from Docker sandbox provisioning and the update of agent instructions in CLAUDE.md to the subsequent execution of automated RSpec tests and network calls to LLM providers like Anthropic. By correlating these endpoint and network events, the hunt distinguishes legitimate development activity from unauthorized automated code manipulation.
## scope-docker-hosts
<!-- Identify Docker-capable hosts -->
Scope the hunt to hosts running Docker, which provides the required sandbox infrastructure for AI coding harnesses.
```sqlite target=endpoint role=scoping
~~~yaml
expected: A list of hosts with Docker installed. Silence suggests no containerization
capability is present via this package manager.
reads:
- device_hostname
- device_uid
- package_name
- vendor_name
- asset_scope
silence: not_evidence_of_absence
source: hb_software_inventory
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT DISTINCT device_hostname, device_uid FROM hb_software_inventory WHERE (LOWER(package_name) LIKE '%docker%' OR LOWER(vendor_name) LIKE '%docker%') AND asset_scope = 'endpoint'
```
## early-indicators
<!-- Search for harness execution and configuration -->
parallel:
- → harness-process-baseline
- → agent-instruction-writes
join: → assess-session-start
## harness-process-baseline
<!-- Harness process prevalence -->
Stack-count harness processes to identify rare or unauthorized agent activity across the fleet.
```sqlite target=endpoint role=baseline params=(lookback_days=lookback_days, harness_keywords=harness_keywords)
~~~yaml
baseline:
compare: first_seen
window: '{{lookback_days}}d'
expected: Harness names or paths seen on very few hosts. Frequent occurrences on many
hosts may indicate authorized developer workstations.
prevalence:
by: device_hostname
key:
- process_name
rare_below: 3
reads:
- process_name
- process_path
- user_name
- device_hostname
- process_cmd_line
- time
silence: not_evidence_of_absence
source: hb_process_activity
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT process_name, process_path, user_name, COUNT(DISTINCT device_hostname) AS host_count, MIN(time) AS first_seen FROM hb_process_activity WHERE (LOWER(process_name) LIKE '%lemans%' OR LOWER(process_name) LIKE '%claude-code%' OR LOWER(process_name) LIKE '%fable-agent%' OR LOWER(process_cmd_line) LIKE '%lemans%' OR instr(',' || '{{harness_keywords}}' || ',', ',' || LOWER(process_name) || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days') GROUP BY process_name, process_path, user_name
```
## agent-instruction-writes
<!-- Monitor instruction file updates -->
Detect modifications to the instruction files that agents use to define coding rules and API recall preferences.
```sqlite target=endpoint role=detection-candidate params=(lookback_days=lookback_days)
~~~yaml
expected: Creation or modification of CLAUDE.md files. Silence means no agent-specific
configuration was observed in this window.
reads:
- device_hostname
- file_name
- file_path
- process_name
- time
silence: not_evidence_of_absence
source: hb_file_activity
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT device_hostname, file_name, file_path, process_name, time FROM hb_file_activity WHERE (LOWER(file_name) = 'claude.md' OR LOWER(file_path) LIKE '%/.claude/skills/%') AND activity_id IN (1, 3, 5) AND time >= datetime('now', '-{{lookback_days}} days')
```
## assess-session-start
<!-- Assess session establishment -->
```agent target=hunter
cite: required
context:
- harness-process-baseline
- agent-instruction-writes
max_iterations: 3
objective: Determine if the observed process execution and configuration file updates
indicate an active AI coding harness session.
success_criteria: A verdict on session presence citing specific hosts and process
names.
tools:
- endpoint
```
## agent-operation
<!-- Correlate operation activity -->
parallel:
- → rspec-validation-runs
- → llm-api-connections
join: → triage-full-lifecycle
## rspec-validation-runs
<!-- Detect automated RSpec runs -->
Identify the automated testing phase that follows codebase modification in coding harnesses.
```sqlite target=endpoint role=triage params=(lookback_days=lookback_days, scope_hosts=scope_hosts)
~~~yaml
expected: RSpec command lines targeting artifacts mentioned in the article. Silence
suggests the session did not reach the validation phase or used a different test
runner.
reads:
- device_hostname
- process_name
- process_cmd_line
- time
silence: not_evidence_of_absence
source: hb_process_activity
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT device_hostname, process_name, process_cmd_line, time FROM hb_process_activity WHERE (LOWER(process_name) LIKE '%rspec%' OR LOWER(process_cmd_line) LIKE '%rspec%') AND (LOWER(process_cmd_line) LIKE '%invitation.create%' OR LOWER(process_cmd_line) LIKE '%schema.rb%') AND ('{{scope_hosts}}' = '' OR instr(',' || '{{scope_hosts}}' || ',', ',' || device_hostname || ',') > 0) AND time >= datetime('now', '-{{lookback_days}} days')
```
## llm-api-connections
<!-- DNS lookups to LLM providers -->
Corroborate the session by finding network traffic to the LLM providers used by coding agents.
```sqlite target=endpoint role=enrichment params=(lookback_days=lookback_days, llm_domains=llm_domains)
~~~yaml
expected: DNS queries for LLM domains from identified hosts. Silence means the agent
may be using a different provider or a local proxy.
reads:
- device_hostname
- query_hostname
- process_name
- time
silence: not_evidence_of_absence
source: hb_dns_activity
verified: dry-run
verified_at: '2026-09-29'
~~~
SELECT device_hostname, query_hostname, process_name, time FROM hb_dns_activity WHERE instr(',' || '{{llm_domains}}' || ',', ',' || LOWER(query_hostname) || ',') > 0 AND time >= datetime('now', '-{{lookback_days}} days')
```
## triage-full-lifecycle
<!-- Triage agentic workflow -->
```agent target=hunter
cite: required
context:
- assess-session-start
- rspec-validation-runs
- llm-api-connections
max_iterations: 6
objective: Determine if the combined evidence of harness setup, RSpec execution, and
LLM communication indicates a suspicious or unauthorized automated code modification
session.
success_criteria: A final verdict citing rows from all phases including the initial
session establishment.
tools:
- endpoint
```
## route-verdict
<!-- Route on triage verdict -->
if~: "the triage-full-lifecycle verdict for any host is malicious or suspicious" (confidence: high, judge=hunter)
then: → isolate-endpoint
indeterminate: → review-code-modifications
unavailable: → review-code-modifications (blind_spot: short-lived-containers)
else: → document-and-close
## isolate-endpoint
<!-- Isolate endpoint -->
```action target=endpoint
~~~yaml
approval: required
~~~
Isolate the host and stop all active Docker containers associated with the coding harness.
```
→ review-code-modifications
## review-code-modifications
<!-- Review code modifications -->
```manual target=analyst
Examine the local git repository on the isolated host. Identify new files or modifications to schema.rb, CLAUDE.md, and Ruby model files. Compare these changes against recent engineering tickets.
```
→ document-and-close
## document-and-close
<!-- Document and close -->
```manual target=analyst
Record the identified session details and update the allow-list for hosts where AI agent experimentation is permitted.
```
→ end
Run it
Take this hunt into your environment.
Open it in Huntbase to run every step against your own connections, with Scout weighing the evidence and your analysts in command. Or take the open hunt.md file anywhere that reads the format.
Machine-drafted by huntbase-hunt-generation using hb_google/gemini-3-flash-preview, gated by dry-run, lint, then reviewed by a person.