Anthropic's Claude now scores 90% on legal benchmarks, but hallucinations are still appearing in court filings. Benchmark scores mean nothing without governance — and most firms are deploying AI without it.
AI Governance  Trovix AriaLegal · Financial Services · Insurance

Anthropic's announcement of 20+ integrations and a Claude 4.7 model scoring 90.9% on Harvey's BigLaw Bench looks impressive on a spreadsheet. What Fortune buried in paragraph six is the real story: hallucinations are still turning up in actual court filings. For UK mid-market law firms, insurers and financial services practices bound by SRA Handbook standards, FCA Consumer Duty PS22/9, and FRC ISA UK audit requirements, this tells you something uncomfortable. A 90% accuracy tool running unsupervised on M&A due diligence, employment handbooks or regulatory filings is not 90% safe — it is a compliance liability waiting to crystallise.

We are watching an industry pattern repeat itself: vendors benchmark their models against narrow, controlled datasets, then firms deploy them into messy reality. Harvey, Luminance, Legora and now Anthropic are all racing to embed AI deeper into legal workflow. The benchmarks get better. The products get faster. And the hallucination problem — the fundamental issue that AI language models confabulate facts when they do not know answers — stays exactly the same. What has changed is the stakes. When AI was a research toy, hallucinations were academic. When AI is drafting your employment handbook or summarising acquisition targets, hallucinations become regulatory breaches.

Trovix's view is that embedding AI directly into live legal work without governance infrastructure is negligent. You cannot fix a hallucination problem with a better model; you fix it with oversight, audit trails and human verification checkpoints. The firms getting real value from AI are not the ones chasing the latest Claude release. They are the ones building verification layers — document intelligence that flags inconsistencies, knowledge systems that ground AI outputs in fact-checked sources, and audit dashboards that show exactly what the AI did and who approved it. Trovix Aria and Trovix Sift are built on this principle: AI as a productivity tool inside a governance boundary, not as a replacement for human judgment. Trovix Audit gives you the compliance record you need when a regulator asks why an AI made a particular decision.

If you are a mid-market firm considering Anthropic's new plugins or any generative AI tool in your core workflow, ask your vendor three questions before you sign: one, how do you prevent hallucinations in my live work? Two, can you show me an audit trail of every AI decision? Three, what happens when the AI gets it wrong and your client or regulator asks why you relied on it? If they cannot answer those questions clearly, do not deploy it. The productivity gain is not worth the regulatory risk. Build your AI capability on a foundation of governance, not on benchmark scores.

Source: Fortune

Related Trovix product:

Trovix Aria →Book a demo →