The firm you call before your agent goes live.

We test and attack AI agents before they touch anything that matters: client money, client data, production systems. Independent audits, built for regulated finance first.

We never declare an agent safe. No serious auditor does. We attest: tested against these attack classes, against this standard, on this date. Here is what broke. Here is what held.

Agents get stuck at sign-off.

The board’s question is fair. How do we know the agent will not pay the wrong person, leak client data, or obey a hostile email? Whoever answers that question should not be the team that built the agent, and should have nothing to gain from a pass. We answer it, independently. That is what gets an agent out of pilot.

There is no independent, attestation-grade auditor of agent deployments for EU regulated finance.

Three days. Map, attack, evidence.

You cannot certify a language model. You can audit everything built around it: permissions, spend limits, approval gates, logs. And you can prove whether they hold under attack. The gate review does that in three days; the full audit takes each day deeper.

Day 1. Map

What can your agent read, write, spend, and send? We draw the complete map: every tool it can call, every permission it holds, every human approval it must clear, and what its logs would actually prove.

Day 2. Attack

We attack everything the map shows, under written authorization: hostile documents, poisoned tool outputs, attempts to exceed its authority, and every path that could move data or money out.

Day 3. Evidence

Every result becomes a finding: what broke, what held, how to reproduce it, how to fix it, which standard it maps to. Written to be forwarded to your risk committee without us in the room.

Scope is declared in writing before testing starts: which agent, which version, which tools, which attack classes. The report states what the days found and where testing stopped.

What a risk committee receives.

Gate review · Report excerpt

Fictional sample

Findings summary from the fictional sample gate-review report
IDFindingSeverityClass
GR-01Indirect prompt injection via hostile PDF drafts an attacker-directed payoutCriticalAF-01
GR-02Spend and authority limits are prompt-level only, not externally enforcedCriticalAF-04
GR-03Action log is writable by the agent’s own service account; not tamper-evidentHighAF-10
noneWhat held: payment execution was outside the agent’s reach. A manual treasury approval stopped the drafted payout. Tested and held. A single human approval is currently the only barrier.Heldnone
From the sample gate-review report, prepared as a fictional engagement. Request the sample report.

Start with a fixed-price gate review.

Every engagement is scoped to your system and quoted on a call; every price is fixed before we start.

Gate review

3 days · paid upfront

For deployers

The three days above, run on your agent in your environment. You get the findings, the fix list, and the mapping to your regulatory obligations. Small enough to skip the procurement cycle, and credited in full against a full audit signed within 90 days.

Full audit

Scoped on a call

For deployers and vendors

Each day expanded to the full attack surface, plus a review of the controls around the agent. Reported against external standards: AIUC-1, a certification standard for AI agents; ISO/IEC 42001; EU AI Act readiness. Findings are built to be the evidence an insurer will ask for.

For agents with payment authority, the policy-engine audit: is there any sequence of inputs that moves funds outside the declared limits?

Certification readiness

Scoped on a call

For AI vendors

Selling an agent into enterprises means months in security review. We build the evidence pack that shortens it, mapped to the standards the buyer’s procurement team already recognizes, so procurement stops being where deals die.

Technical due diligence

Typically two weeks

For investors

Is the AI real, and what breaks under attack? Answered in a fixed scope: architecture review, verification of the target’s technical claims, and adversarial spot-testing.

Retainer

Quarterly

After the audit

Quarterly re-testing keeps an attestation current; AIUC-1 mandates it. For clients in regulated finance, the testing follows the threat-led exercises their regulators ask for by name: the TIBER-EU and DORA TLPT tradition, extended to AI agents. Incident-response standby is quoted separately.

The failure classes.

These classes cover the ways agents fail under attack and in operation. Every finding is tagged with one and maps to the standards named above: public yardsticks, set by others.

Agent failure-mode taxonomy: ID, class, and definition
IDClassDefinition
AF-01Prompt injectionUntrusted content redirects the agent’s behavior against the operator’s intent.
AF-02Tool / MCP supply-chain compromiseA tool, plugin, or MCP server the agent trusts is malicious, compromised, or silently changed.
AF-03Excessive agency / authority escalationThe agent holds, acquires, or is manipulated into using more authority than the task requires.
AF-04Spend-limit / policy-engine bypassFinancial or action limits are advisory in the prompt rather than enforced outside the model.
AF-05Data exfiltrationThe agent is induced to move sensitive data outside its authorized boundary.
AF-06Unsafe action chainingIndividually permitted actions compose into a harmful outcome no single step would flag.
AF-07Hallucinated actionsThe agent invents tool calls, parameters, recipients, or facts and acts on them.
AF-08Memory / context poisoningPersistent memory or retrieved context is corrupted so future runs inherit hostile instructions.
AF-09Human-gate circumventionAn approval step exists but can be skipped, spoofed, fatigued, or rendered meaningless.
AF-10Logging / audit-trail gapsThe record of what the agent did is missing, incomplete, or tamperable.

Taxonomy v1 (2026), summary view. Sub-vectors, passing-control definitions, and clause-level standard mappings are engagement material, shared on request.

Before your agent goes live.

Start with a three-day gate review, or a call to scope the full engagement.