← Work
// autonomous agents · hermes agent skill

An autonomous committee that stress-tests AI code before you see it

Single-pass LLMs produce code that compiles but crashes under edge cases. Critic-Refine simulates a 3-tier adversarial committee: a generator, a polymorphic domain critic, and a judge executing mental dry-runs with explicit trade-offs.

3-tierActor-Critic-Judge
0–100Calibrated confidence
< 8 sIn-turn turnaround

The problem

Single-pass AI generators suffer from structural complacency: they write code that looks clean, passes basic happy-path unit tests, and satisfies the prompt on the surface.

In production, however, subtle failure modes remain undetected: floating-point precision imprecisions (1e-16) causing CPU micro-spins, silent float('nan') bypasses freezing thread pools, timing attacks in token comparisons, or memory leaks under concurrency. Standard chat assistants lack a built-in adversarial counterweight to catch these issues before handing code to the engineer.

What I built

A native, production-grade skill for Hermes Agent (v1.0.0) that intercepts complex engineering tasks and runs an internal 3-role committee before generating the final response.

  • The Generator (Actor): Architectures the first-pass implementation, prioritizing functional coverage and core logic.
  • The Polymorphic Critic: Dynamically adopts the domain persona of the harshest auditor (SecOps cryptographer, concurrency specialist, DBA, distributed systems engineer) to attack the solution.
  • The Mental Dry-Run: The Judge executes mental stress tests on boundary conditions (null pointers, inf, network partitions, high contention) and triggers automated revision loops (max 3).
  • Radical Trade-Off Transparency: The output documents what was intentionally sacrificed (e.g. latency vs. idempotency) alongside a 0–100 confidence score and residual production risks.
LIVE 3-ROLE AGENT COMMITTEE SIMULATOR
Test Adversarial Code Audit, Flaw Teardown & Convergence

Select a high-stakes engineering task below to see the Generator's first draft, the Polymorphic Critic's attack vectors, and how the Judge Arbitrator patches subtle failure modes before delivery.

SELECT HIGH-STAKES PROBLEM:
01 · GENERATOR FIRST-PASS DRAFT INITIAL CONFIDENCE: 52%

        
02 · POLYMORPHIC CRITIC (SECOPS / THREAD AUDITOR) CRITICAL FLAW FOUND ⚠️
03 · JUDGE ARBITRATOR & FINAL GATE FINAL SCORE: 96 / 100
// HARDENED PRODUCTION IMPLEMENTATION

        
RADICAL TRADE-OFF TRANSPARENCY

RUNTIME Hermes Agent native skill (v1.0.0)
TURNAROUND < 8 seconds (Single in-turn COT)
MAX REVISIONS 3 convergence passes
CONFIDENCE CALIBRATION Brier score calibrated (0-100)

How it works

User Prompt high-stakes task ROLE #1 01 · Generator first-pass solution ✦ ADVERSARIAL ROLE #2 Polymorphic Critic SecOps · Concurrency · DBA ROLE #3 Judge Arbitrator dry-run & trade-offs Output Report audit + calibrated fix Iteration Loop (max 3x) if critical vulnerabilities found

Adversarial audit is executed before the user ever sees a character of code.

Under the hood

PITFALL

The silent float('nan') infinite loop trap

In standard rate limiter or validation code, guards like if tokens <= 0: raise ValueError look airtight. But in Python and IEEE 754, comparisons against float('nan') always return False.

If an input payload passes NaN, it slips right past the validator into the while-loop, where self.tokens >= tokens is permanently False, triggering a 100% CPU infinite spin that hangs the worker thread.

// IEEE 754 EXPLOIT AUDIT · SILENT WORKER THREAD HANG
Naive Guard: tokens <= 0
100% CPU LOCK THREAD HANG
❌ float('nan') <= 0 evaluates to False! Bypasses guard into infinite loop.
Hardened Gate: math.isfinite()
AIRTIGHT EXCEPTION
✅ Traps NaN and Inf immediately, raising clean 400 Bad Request before thread acquires lock.
ARCHITECTURAL RULE: Never rely on inequality operators (<, >, <=, >=) to validate untrusted floats. Always enforce math.isfinite() at the network boundary.

FIX The Critic systematically checks numeric boundary conditions and enforces math.isfinite() across all arithmetic entry points.

DECISION

In-turn multi-persona reasoning vs. isolated subagent processes

Hermes Agent can spawn independent background OS subagents via delegate_task, each with isolated terminal and context. While powerful, multi-agent spawning incurs 30 to 60 seconds of process coordination latency.

TRADE-OFF For 95% of tasks, Critic-Refine executes as a sequential multi-role Chain-of-Thought in a single fast turn (< 8 seconds). Background subagents are reserved only for full repository deep-dives and terminal execution suites.

PITFALL

The illusion of "flawless" AI code

Most AI assistants present their code as universally optimal. In real software engineering, every production decision involves trade-offs: mutexes prevent race conditions but introduce lock contention; synchronous fsync ensures durability but throttles throughput.

DECISION The Judge is strictly forbidden from claiming perfection. It must explicitly articulate the trade-offs made and score the residual risks before clearing deployment.

Results

Benchmarked across complex distributed and concurrent tasks:

  • Token Bucket Rate Limiter: Caught IEEE 754 micro-spin (sub-millisecond busy-waits) and reentrant deadlocks on threading.Lock.
  • Payment Webhooks: Replaced variable-time string comparisons with hmac.compare_digest (eliminating timing attacks) and blocked memory-bomb DoS attacks via stream chunk capping.
  • Write-Ahead Log Engine: Replaced fragile JSON appending with binary frame CRC32 verification and asynchronous group commit batching.

Stack

  • Hermes Agent: Open-source agent runtime by Nous Research
  • Skill Engine: Declarative markdown-based procedural agent memory
  • Multi-Agent Protocol: Actor-Critic adversarial pipeline with calibrated scoring
  • Google AI Studio (Gemini Flash): High-throughput reasoning backbone (>180 tokens/sec)

What I took away

Adversarial tension beats model size. A lightweight, cost-effective model constrained by an aggressive critic consistently out-engineers an unchecked flagship model that is left to grade its own work.