About a-team Marketing Services
The knowledge platform for the financial technology industry

A-Team Insight Blogs

Castellum.AI’s Financial Crime Benchmark Finds Accuracy & Consistency Gaps in Standalone LLMs

Subscribe to our newsletter

Castellum.AI announced its FinCrime Agent Benchmark (FCA Bench) in September that compares the performance of nine standalone large language models (LLMs) performing financial-crime checks and when those same models operate within the Castellum.AI Harness. The benchmark measures risk identification accuracy (sanctions, politically exposed persons (PEPs) and adverse media screening), applying an institution’s standard operating procedures (SOPs) and decision consistency across successive runs.

Speaking with RegTech Insight, CEO Peter Piatetsky and VP of Growth Spencer Vuksic explained that the research was a response to prospective buyers struggling to distinguish competing AI claims. Piatetsky recalled a bank CEO telling him: “I don’t know how to tell you guys apart.”

Banks considering building their own agents also need to resource their operation after launch. Piatetsky acknowledged that some prospective clients employ more engineers than Castellum.AI, but asked: “Are you going to assign engineers to maintain it?” and “Who’s going to manage the integrations?”

Testing for Accuracy & Consistency

Accuracy and consistency are non-negotiable quality measures in managing financial crime risk. Accuracy measures whether a decision is correct; consistency measures whether the model reaches the same decision across repeated runs.

During the benchmark, each model assessed 500 cases covering sanctions, politically exposed persons (PEPs) and adverse media, completing three standalone runs and three runs within the Castellum.AI Harness. Both configurations used the same cases and baseline standard operating procedure (SOP).

Castellum.AI’s harness combines its risk database and screening tools with Arbiter agents that apply institutional policies and resolve sanctions, PEP and adverse-media alerts.

For standalone runs, each model received prompts incorporating the baseline SOP and directions to relevant online risk sources. It then used the information retrieved to clear or escalate alerts.

Castellum.AI reported end-to-end accuracy of 30.8%–71.1% for standalone models, compared with 98.1%–99.6% within the harness. Average decision consistency rose from 45.2% to 97.2%. The benchmark highlights considerable inconsistency across the standalone models with different incorrect answers across successive runs. These results reflect Castellum.AI’s test of 500 cases against one baseline SOP

Screening Data Foundations

At an earlier career point while working on Russia and Iran sanctions at the US Treasury Department, Piatetsky pressed foreign partners to adopt US designations. Comparing lists exposed a basic data problem: governments invested in identifying targets but published the results in hard-to-find / hard-to-interpret spreadsheets. He launched Castellum.AI in September 2019 to make sanctions data accessible, building the data foundation its agents now use.

The benchmark also exposed weaknesses in standalone models’ retrieval of screening data. Piatetsky said standalone models found an average of 61% of expected hits, despite receiving directions to the screening sources Castellum.AI uses. Some sanctions checks took minutes costing more than a dollar per screen as models located, analysed and cleaned information.

Arbiter supplies structured risk records and matching tools, allowing the LLM to concentrate on assessing the resulting alerts. Piatetsky argued that judgement and adjudication better suit LLM capabilities than primary screening. He also highlighted improved performance from open-weight models (e.g. Gemma 4-31B) within the harness, suggesting firms could consider that combination for in-house deployments.

Institutional Policy Tuning

Piatetsky described how one bank might welcome a high-risk customer whose business it understands, while another would refuse the same customer. The agent must apply each institution’s procedures to decide how to handle the risk.

For client deployments, Castellum.AI requests policies and procedures, alerts with completed adjudications and another set without adjudications. It tunes the agent against the completed cases using the client’s answer key before testing the unseen set. “And then we run it on basically the closed book test,” Piatetsky said.

Asked how performance compared with human reviewers, he said the process had identified analyst mistakes in every deployment. Vuksic noted that testing can also reveal problems in the written policies: “We also find inaccuracies or inconsistencies in the SOP itself,” he said. He explained that policies and procedures specify which matches or discrepancies in names, locations, dates of birth and identifiers warrant clearing or escalating an alert. However, when subjected to agent testing, contradictory rules can be revealed.

Partnerships Extend Reach

Castellum.AI’s partnerships provide routes into existing compliance systems and investment due-diligence workflows. Its LexisNexis Bridger integration places Arbiter’s decisions, explanations and audit records inside the screening platform. Piatetsky described that relationship as a route into asset management: “It’s literally putting our agent in the existing platform.”

Global risk advisory and disputes consultancy Mintz Group adds a private-markets connection. The companies announced plans earlier this year to integrate Castellum.AI’s sanctions and watchlist screening into Mintz’s Verity platform and enhanced due-diligence workflows, serving clients including private-equity firms.

The Hummingbird partnership brings Castellum.AI’s sanctions, PEP and adverse-media data into Hummingbird’s investigation and case-management platform. Compliance teams can screen subjects and record alert decisions within Hummingbird, keeping supporting data, evidence and a complete audit trail in the same system.

Subscribe to our newsletter

Related content

WEBINAR

Upcoming Webinar: Generative and Agentic AI in Financial Markets: What the Data Really Shows

Date: 15 October 2026 Time: 10:00am ET / 3:00pm London / 4:00pm CET Duration: 50 minutes Artificial intelligence is reshaping financial markets – but the reality on the ground is more nuanced, more uneven, and more instructive than the headlines suggest. A new A-Team Insight research programme, drawing on responses from senior AI decision-makers at...

BLOG

EBA Framework 4.3 Maps the Data Behind AMLA’s First Risk Assessment

The European Banking Authority’s Reporting Framework 4.3 gives financial institutions their first machine-readable view of the data that will support the EU’s selection of firms for direct anti-money laundering supervision. Published earlier this month, the framework marks an intermediate step towards the first formal risk-assessment and selection exercise by the Authority for Anti-Money Laundering and...

EVENT

Eagle Alpha Alternative Data Conference, Fall, New York, hosted by A-Team Group

Now in its 8th year, the Eagle Alpha Alternative Data Conference managed by A-Team Group, is the premier content forum and networking event for investment firms and hedge funds.

GUIDE

AI in Capital Markets Handbook 2026

AI adoption in capital markets has moved into a more disciplined phase. The priority is now controlled deployment: where AI can be used safely, where it can deliver measurable value, and how outputs can be governed, monitored and evidenced. The 2026 edition of the AI in Capital Markets Handbook examines how AI is being applied...