AIUC announced a $40 million Series A on Tuesday, led by Ribbit Capital with First Harmonic participating, to audit and certify AI agents against a standard the company wrote itself. The founders are Kvist, one of Anthropic's first employees, and Dattani, who was chief operating officer of METR from 2024 to 2025. Their argument is not that models are too weak to deploy. It is that banks, hospitals, government agencies and militaries promise their own customers what a system will and will not do, and right now nobody can guarantee it.
That framing is worth pausing on, because it inverts the usual safety pitch. Kvist's claim is that the blocker in regulated industries is not capability but liability — the buyer needs a document to point at. AIUC sells the document.
Total funding now stands at $55 million. The earlier $15 million seed came from Nat Friedman through his NFDG fund, along with Emergence, Anthropic co-founder Ben Mann and others. Named clients include Cursor, Lovable, Harvey and ElevenLabs.
The product is a standard called AIUC-1, built on the chassis of SOC 2, plus a testing service that checks whether an agent meets it. To write the standard, AIUC assembled a consortium of roughly 250 security and risk management specialists — the people who actually sign off on buying an AI agent. The company meets with them regularly to establish what they want to see at purchase, what they ask vendors, and which risks and system characteristics have to be disclosed.
Those answers set the test content. AIUC then runs an agent through roughly 5,000 checks covering behavior under system-prompt jailbreaks, hallucination, and data leakage. The output is a report of about 100 pages describing where the agent is safe and reliable and where it fails the requirements. The tests are executed by AI agents and the data is analyzed by AI; the final audit, Kvist says, is reviewed by humans.
The choice of SOC 2 as the template is the most revealing decision here. SOC 2 is the thing enterprise buyers ask for before signing, and its function in practice is to convert an open-ended security question into a closed procurement step. Porting that to agents is a clean product insight and an honest admission of scope: this is compliance infrastructure, not an alignment program.
Which leads to the question the announcement does not raise. Every named client — Cursor, Lovable, Harvey, ElevenLabs — is a vendor of agents, not a buyer of them. The consortium of 250 that defines the tests is made of buyers; the customers who pay for the audit appear to be the sellers. That is also how SOC 2 works, and it is how credit rating agencies worked, and the structural weakness is the same one every time: the auditor's revenue comes from the party whose product is being graded. A 100-page report enumerating where an agent "does not meet requirements" is a commercially awkward document to hand the company that commissioned it. Nothing in the announcement explains what stops that pressure.
The METR comparison is more interesting than it first looks. Dattani remains on METR's board, and METR runs similar evaluations for frontier labs, though until recently it mostly measured performance — whether an agent can complete tasks reliably. METR was among the independent research organizations OpenAI brought in to investigate the Hugging Face incident. Anthropic CEO Dario Amodei recently called on the industry to slow down frontier model development, citing rapid growth in incidents of dangerous system behavior, and proposed requiring frontier labs to host embedded independent evaluators who observe systems in operation and verify their safety. He named METR as one possible candidate.
AIUC has no plans to place its people on customer sites. That is the whole distance between the two models. Amodei's version puts an outsider inside the lab, watching continuously; AIUC's version runs 5,000 tests from outside and issues a certificate. One is supervision, the other is attestation, and only the second one scales to a venture-backed business with $55 million behind it.
Between an industry-wide slowdown that no lab has agreed to and a certification any vendor can buy, the market has already chosen. The gap that leaves is not in the testing — 5,000 checks is a serious number — but in the timing: a certificate describes an agent as it was on the day it was audited, and agents change weekly.