Code Generation
~8 min readUpdated: October 2026

AI Security Code Review Agents: Don't Just Trust AI Code

AI security code review agents — verifier agents and review gates in Smoke Monkey

AI now writes more code than any human team can review by hand, and generated code ships its own class of vulnerabilities. Tools like **Qodo** and **Endor** turned AI code review into a category — but review is only half the loop. This guide shows how to pair review agents with automated verification in [Smoke Monkey Harness](/solutions/what-is-an-ai-agent-harness), and how **Smoke Monkey Canvas** lets you watch a review fleet work a change set before anything merges.

Technical Review: Smoke Monkey Core Architecture Team
Tested on Node.js 18+ & BunTypeScript 5.x
Quick Answer & Executive Definition

AI Security Code Review Agents: Don't Just Trust AI Code: AI now writes more code than any human team can review by hand, and generated code ships its own class of vulnerabilities. Tools like **Qodo** and **Endor** turned AI code review into a category — but review is only half the loop. This guide shows how to pair review agents with automated verification in [Smoke Monkey Harness](/solutions/what-is-an-ai-agent-harness), and how **Smoke Monkey Canvas** lets you watch a review fleet work a change set before anything merges. Designed as a zero-dependency, open-source TypeScript architecture under the MIT License with native Model Context Protocol (MCP) support and deterministic phase state machines.

Key Architectural Takeaways
Quick Implementation Examplereview-agent.ts
review-agent.tstypescript
import { createAgent } from 'smoke-monkey-harness';
// A read-only reviewer: it judges code but cannot change it
const reviewer = createAgent({
provider: 'anthropic',
model: 'claude-opus-5',
workspacePath: process.cwd(),
identity: { principal: 'agent://security-reviewer' },
permissions: {
read_file: 'allow',
run_command: 'allow', // run tests and scanners
write_file: 'deny', // cannot silently "fix" what it reviews
},
});
const report = await reviewer.run('Review the working tree for auth, injection, and secret-leak issues. List findings with severity.');
console.log(report.output);
Video Guides

Watch: Related Video Guides

Anthropic Just Built an Agentic OS — Open Source Harness Breakdown

Smoke Monkey

How to Actually Review AI Code (Don't Just Trust It)

Tech With Tim

Why AI-Generated Code Needs a Dedicated Review Pass

Language models are trained on a large body of code that includes insecure patterns, and they reproduce those patterns with confidence. That is how generated-code CVEs happen: a plausible-looking query built by string concatenation, a hardcoded secret in a "quick example", a missing authorization check on a new route. Human review was never sized for AI throughput, so the review step has to become automated. The 2026 tooling wave — Qodo, Endor, and others — proves the demand. The right architecture pairs a review agent with real verification, which is exactly what Smoke Monkey Harness does, all visible on Smoke Monkey Canvas.

Trust is not a review strategy

A review agent that can also edit the code will sometimes "resolve" a finding by hiding it. Separate the reviewer identity from the writer so the verdict stays independent.

Review Agents Belong in Their Own Identity

The security of a review depends on the reviewer being independent. Give the review agent a read-only identity: it can read files and run tests, but it cannot write. In Smoke Monkey that is a permissions block — read_file: 'allow', run_command: 'allow', write_file: 'deny'. Now a finding cannot be quietly patched away, and every reviewer action is attributable to its principal in the audit log. This is the same least-privilege discipline you would apply to any non-human identity in your stack. Assign that identity per card on Smoke Monkey Canvas so no reviewer ever runs broader than its scope.

review-gate.tstypescript
import { createAgent, defineTool } from 'smoke-monkey-harness';
// A gate that only opens when review AND tests both pass
const reviewGate = defineTool({
name: 'review_gate',
description: 'Block merges until the reviewer and the test suite both approve',
parameters: { type: 'object', properties: { diff: { type: 'string' } } },
handler: async ({ diff }) => {
const findings = await securityScanner.scan(diff);
const tests = await runTestSuite();
const ok = findings.critical === 0 && tests.passed;
return { output: ok ? 'GATE_OPEN' : 'GATE_BLOCKED', findings, tests };
},
});
const agent = createAgent({
provider: 'anthropic',
model: 'claude-opus-5',
workspacePath: process.cwd(),
tools: [reviewGate],
});

Coupling Review to the Automated Test Verification Loop

Static review catches patterns; tests catch behavior. A serious pipeline needs both, so the review agent should block on the automated test verification loop — no green tests, no merge. Smoke Monkey Harness runs the loop itself: the agent edits, the runtime executes the suite, and failure routes back through the recover phase instead of pretending success. Add a second verifier with a different model for high-risk changes and you have defense in depth without a human reading every diff. Run the reviewers side by side on Smoke Monkey Canvas and the disagreements surface visibly.

dual-verifier.tstypescript
import { createAgent } from 'smoke-monkey-harness';
// Two independent reviewers: a cheap model for breadth, a frontier model for risk
const reviewers = [
createAgent({ provider: 'anthropic', model: 'claude-haiku-5', workspacePath: process.cwd(), permissions: { write_file: 'deny' } }),
createAgent({ provider: 'anthropic', model: 'claude-opus-5', workspacePath: process.cwd(), permissions: { write_file: 'deny' } }),
];
const verdicts = await Promise.all(
reviewers.map((r) => r.run('Flag security and correctness issues in the staged diff.'))
);
// Merge only when both independent reviewers clear the change
if (verdicts.some((v) => /critical/i.test(v.output))) throw new Error('Blocked by review agent');

A Review Fleet on the Canvas

When review runs at AI speed, you need a control surface, not a terminal scroll. Smoke Monkey Canvas (npx @smoke-monkey/canvas start) lays out a whole review fleet on one infinite canvas: an explorer to map the diff, a security reviewer, a performance reviewer, and a verifier — each a card with live findings. Approve or reject gates inline, and because Canvas shares the harness runtime, 300+ MCP tools like linters, SAST scanners, and secret detectors plug straight in. It is the fastest way to see *why* a change set is blocked before it reaches main. Related: AST-aware code editing and TypeScript agent testing. Every reviewer there runs on Smoke Monkey Harness, so findings, tools, and gates stay consistent.

Google Search Questions & Answers

Frequently Asked Questions

Q:Can AI code review agents be trusted on their own?

No single reviewer should be the only gate. Use a dedicated review identity plus automated tests and, for risky changes, a second independent verifier. Smoke Monkey Harness lets you compose all three in one loop.

Q:Why give a review agent read-only permissions?

A reviewer that can also edit can "fix" a finding by hiding it. Read-only review identities keep the verdict independent and make every action attributable in the audit log.

Q:How is this different from a linter or SAST tool?

Linters and SAST scanners are tools an agent calls; they are not the whole review. Smoke Monkey plugs them in via MCP and adds reasoning, cross-file context, and a merge gate tied to passing tests.

Q:Can I watch a fleet of review agents work?

Yes. Smoke Monkey Canvas shows each reviewer as a card with live findings on one infinite canvas, and lets you approve or block gates inline before a change reaches main.

Related Alternatives & Comparisons

Related Architecture Guides

View all guides

Build with Smoke Monkey Harness

Zero dependencies. 24 built-in tools. Human-in-the-loop safety. 100% open source under the MIT License.

npm install smoke-monkey-harness