The firm you call before your agent goes live.

We test and attack AI agents before they touch anything that matters: client money, client data, production systems. Independent audits, built for regulated finance first.

We never declare an agent safe. No serious auditor does. We attest: tested against these attack classes, against this standard, on this date. Here is what broke. Here is what held.

Agents get stuck at sign-off.

The board's question is fair. How do we know it will not pay the wrong person, leak client data, or obey a hostile email? Whoever answers it should not be the team that built the agent, and should have nothing to gain from it passing. We answer it, independently. That is what gets an agent out of pilot.

There is no independent, attestation-grade auditor of agent deployments for EU regulated finance.

Three days. Map, attack, evidence.

You cannot certify a language model. You can audit everything built around it: permissions, spend limits, approval gates, logs. And you can prove whether they hold under attack. The gate review does that in three days; the full audit takes each day deeper.

Day 1. Map

What can your agent read, write, spend, and send? We draw the complete map: every tool it can call, every permission it holds, every gate a human must pass, and what its logs would actually prove.

Day 2. Attack

We attack everything the map shows, under written authorisation: hostile documents, poisoned tool outputs, attempts to exceed its authority, and every path that could move data or money out.

Day 3. Evidence

Every result becomes a finding: what broke, what held, how to reproduce it, how to fix it, which standard it maps to. Written to be forwarded to your risk committee without us in the room.

Scope is declared in writing before testing starts: which agent, which version, which tools, which attack classes. The report states what the days found and where testing stopped.

What a risk committee receives.

Fictional sample engagement

Findings summary from the fictional sample gate-review report
IDFindingSeverityClass
GR-01Indirect prompt injection via hostile PDF drafts an attacker-directed payoutCriticalAF-01
GR-02Spend and authority limits are prompt-level only, not externally enforcedCriticalAF-04
GR-03Action log is writable by the agent's own service account; not tamper-evidentHighAF-10
What held: fund execution was outside the agent's reach. A manual treasury approval stopped the drafted payout. Tested and held. A single human approval is currently the only barrier.Held
From the sample gate-review report, prepared as a fictional engagement. Request the sample report.

Start with a fixed-price gate review.

Larger engagements are scoped to your system and quoted on a call.

Gate review

€3.5k · 3 days

For deployers

The three days above, run on your agent in your environment. You get the findings, the fix list, and the mapping to your regulatory obligations. Paid upfront and credited entirely against a full audit within 90 days. No procurement cycle needed.

Full audit

Scoped per engagement

For deployers and vendors

Each day expanded to the full attack surface, plus a review of the controls around the agent. Reported against external standards: AIUC-1, a certification standard for AI agents; ISO/IEC 42001; EU AI Act readiness. Findings are built to be the evidence an insurer will ask for.

For agents with payment authority, the policy-engine audit: is there any sequence of inputs that moves funds outside the declared limits?

Certification readiness

Scoped per engagement

For AI vendors

Selling an agent into enterprises means months in security review. We build the evidence pack that shortens it, mapped to the standards the buyer's procurement team already recognises, so procurement stops being where deals die.

Technical due diligence

Scoped per engagement

For investors

Is the AI real, and what breaks under attack? Answered in a fixed scope, typically two weeks: architecture review, verification of the target's technical claims, and adversarial spot-testing.

Retainer and incident response

From €2k/quarter

After the audit

Quarterly re-testing keeps an attestation current; AIUC-1 mandates it. For clients in regulated finance, the testing follows the threat-led exercises their regulators ask for by name: the TIBER-EU and DORA TLPT tradition, extended to AI agents. Incident-response standby is priced separately: €1–2k/month, €2–3k/day on call-out.

The ten failure classes.

Ten classes cover the ways agents fail under attack and in operation. Every finding is tagged with one, and maps to the standards named above: public yardsticks, set by others.

Agent failure-mode taxonomy: ID, class, and definition
IDClassDefinition
AF-01Prompt injectionUntrusted content redirects the agent's behaviour against the operator's intent.
AF-02Tool / MCP supply-chain compromiseA tool, plugin, or MCP server the agent trusts is malicious, compromised, or silently changed.
AF-03Excessive agency / authority escalationThe agent holds, acquires, or is manipulated into using more authority than the task requires.
AF-04Spend-limit / policy-engine bypassFinancial or action limits are advisory in the prompt rather than enforced outside the model.
AF-05Data exfiltrationThe agent is induced to move sensitive data outside its authorised boundary.
AF-06Unsafe action chainingIndividually permitted actions compose into a harmful outcome no single step would flag.
AF-07Hallucinated actionsThe agent invents tool calls, parameters, recipients, or facts and acts on them.
AF-08Memory / context poisoningPersistent memory or retrieved context is corrupted so future runs inherit hostile instructions.
AF-09Human-gate circumventionAn approval step exists but can be skipped, spoofed, fatigued, or rendered meaningless.
AF-10Logging / audit-trail gapsThe record of what the agent did is missing, incomplete, or tamperable.

Taxonomy v1 (2026), summary view. Sub-vectors, passing-control definitions, and clause-level standard mappings are engagement material, shared on request.

Before your agent goes live.

Start with a three-day gate review, or a call to scope the full engagement.