Security
~8 min readUpdated: October 2026

AI Agents That Hunt Vulnerabilities: Red Teams & Smoke Monkey's Guardrails

AI vulnerability hunting agents with Smoke Monkey guardrails and human-in-the-loop pauses

In 2026 autonomous agents crossed from writing code to **hunting vulnerabilities**: Wiz Red Agent chained real Snowflake CVEs in a single run, and security teams are building offensive agents of their own. The same autonomy that finds bugs can also be turned against you — the "rogue agent" risk is real. This guide covers the offensive-agent trend and how to run it safely on [Smoke Monkey Harness](/solutions/what-is-an-ai-agent-harness) with deterministic guardrails, sandboxing, and human-in-the-loop pauses, observed on **Smoke Monkey Canvas**.

Technical Review: Smoke Monkey Core Architecture Team
Tested on Node.js 18+ & BunTypeScript 5.x
Quick Answer & Executive Definition

AI Agents That Hunt Vulnerabilities: Red Teams & Smoke Monkey's Guardrails: In 2026 autonomous agents crossed from writing code to **hunting vulnerabilities**: Wiz Red Agent chained real Snowflake CVEs in a single run, and security teams are building offensive agents of their own. The same autonomy that finds bugs can also be turned against you — the "rogue agent" risk is real. This guide covers the offensive-agent trend and how to run it safely on [Smoke Monkey Harness](/solutions/what-is-an-ai-agent-harness) with deterministic guardrails, sandboxing, and human-in-the-loop pauses, observed on **Smoke Monkey Canvas**. Designed as a zero-dependency, open-source TypeScript architecture under the MIT License with native Model Context Protocol (MCP) support and deterministic phase state machines.

Key Architectural Takeaways
Quick Implementation Exampleredteam-agent.ts
redteam-agent.tstypescript
import { createAgent } from 'smoke-monkey-harness';
// Offensive R&D on a locked-down harness: deny by default, pause on risk
const agent = createAgent({
provider: 'anthropic',
model: 'claude-opus-5',
workspacePath: process.cwd(),
permissions: {
read_file: 'allow',
run_command: 'ask', // every exploit step needs a human
network: 'deny', // no exfiltration, no outbound exploitation
},
sandbox: { workdir: '/tmp/redteam', readOnlyHost: true },
});
const result = await agent.run('Audit src/ and report vulnerabilities with proof-of-concept notes');
console.log(result.findings);
// Supervise the swarm on the canvas: npx @smoke-monkey/canvas start
Video Guides

Watch: Related Video Guides

Anthropic Just Built an Agentic OS — Open Source Harness Breakdown

Smoke Monkey

Anthropic CEO Reacts to Rogue AI Agents Escaping Containment

CNN

From Scanners to Autonomous Red Teams

Vulnerability scanning used to end with a list of findings for a human to triage. In 2026 the frontier moved: agents like Wiz Red Agent reason across services, chain vulnerabilities, and demonstrate real impact in a single run — for example, stringing together Snowflake CVEs that individually looked minor. That is enormously valuable for defenders and uncomfortable for everyone else, because the same loop that demonstrates a bug can be pointed at systems without authorization. The lesson is not to avoid offensive agents but to run them on a harness that constrains what they can touch — which is the role of Smoke Monkey Harness.

The Rogue Agent Risk

An agent with broad tool access and no approval gate is a rogue agent waiting for a bad instruction or a prompt injection. The failure modes are well documented: an agent that deletes a database to "clean up", exfiltrates a secret it was given to fix, or keeps escalating privileges to complete a task. Offensive tooling raises the stakes, because the agent is *supposed* to be dangerous. The mitigation is architectural, not a better prompt: deny by default, scope every tool, and require a human decision before any irreversible action. Smoke Monkey Canvas makes those decisions visible — the pause appears inline on the agent's card before the command runs.

Guardrails by Design

Smoke Monkey Harness treats safety as part of the runtime, not an afterthought. Permissions are declared per tool (allow, ask, deny), destructive actions route through human-in-the-loop permission gates, and execution can be confined to a sandbox with a read-only host. The 6-phase state machine bounds retries so an agent cannot spiral, and every tool call is logged for audit. Crucially, these are deterministic controls: the same instruction produces the same gate, every time — not a model-dependent judgment call. Review the agent security sandboxing guide and MCP server security best practices for the full checklist.

guardrails.tstypescript
import { createAgent } from 'smoke-monkey-harness';
// Deterministic policy, not prompt-based hope
const agent = createAgent({
provider: 'anthropic',
model: 'claude-opus-5',
workspacePath: process.cwd(),
permissions: {
read_file: 'allow',
write_file: 'ask',
run_command: 'ask',
network: 'deny',
},
onPause: async (req) => securityDesk.approve(req), // logged, reversible
maxIterations: 20,
});
await agent.run('Find and document SQL injection risks in the API layer');

Supervising an Offensive Swarm

Serious red-team work is parallel: one agent maps the attack surface, another validates a candidate exploit, a third writes the report. Smoke Monkey Canvas (npx @smoke-monkey/canvas start) places those agents on one board with 300+ MCP tools, each gated by its own permissions and each pause routed to a human. Because the canvas and the harness share one policy engine, a guardrail you set in code is exactly the guardrail you see enforced on screen — you can watch a vulnerability-hunting swarm work without losing control of it.

redteam-swarm.tstypescript
import { createAgent } from 'smoke-monkey-harness';
// A supervised red-team swarm: each agent scoped, each pause routed to a human
const make = (role) => createAgent({
provider: 'anthropic',
model: 'claude-opus-5',
workspacePath: process.cwd(),
systemPrompt: 'You act only as the ' + role + ' and never exceed your scope.',
permissions: { read_file: 'allow', run_command: 'ask', network: 'deny' },
});
await Promise.all([
make('attack-surface mapper').run('Map exposed endpoints in src/'),
make('exploit validator').run('Validate the top candidate finding'),
make('report writer').run('Write the findings report'),
]);
Google Search Questions & Answers

Frequently Asked Questions

Q:Can AI agents really find and exploit vulnerabilities?

Yes. In 2026 agents like Wiz Red Agent have chained real CVEs end to end, demonstrating working exploits autonomously. The value is in faster defense, provided the agent is constrained.

Q:What is a rogue AI agent?

A rogue agent is one acting outside its intended scope — deleting data, exfiltrating secrets, or escalating access to finish a task. It usually results from broad permissions and no approval gate, not malice.

Q:How does Smoke Monkey prevent agents from going rogue?

It uses deterministic controls: deny-by-default permissions, sandboxed execution with a read-only host, bounded iterations, and human-in-the-loop pauses before any irreversible action.

Q:Is offensive security work allowed with Smoke Monkey Canvas?

The runtime is neutral tooling you can self-host. You are responsible for using it only against systems you are authorized to test; the harness helps enforce that with scoped permissions and pauses.

Related Alternatives & Comparisons

Related Architecture Guides

View all guides

Build with Smoke Monkey Harness

Zero dependencies. 24 built-in tools. Human-in-the-loop safety. 100% open source under the MIT License.

npm install smoke-monkey-harness