Anthropic's announcement of 20+ legal integrations and Claude Opus 4.7's 90.9% score on Harvey's BigLaw Bench reads like vindication for AI in law. But the headline obscures a critical reality for mid-market UK regulated firms. A system that performs well on standardised benchmarks—even very good systems—still hallucinates in live matters. The Fortune piece itself acknowledges this tension: hallucinations are appearing in actual legal filings even as vendors double down on capability claims. For UK law firms subject to SRA Code provisions on competence and client protection, this is not an edge case. It is a liability vector that benchmark scores do not measure.
What we are witnessing is a bifurcation in AI adoption. Big Law—the dozen or so mega-firms with 2,000+ fee-earners, dedicated AI teams, and redundant QA workflows—can absorb the risk of hallucination through volume and process. They can afford to deploy Claude into M&A due diligence with a secondary human review layer that catches the 9% of errors Harvey's benchmark misses. Mid-market firms cannot. A 200-person practice does not have the bandwidth for that model. Instead, they face a choice: adopt enterprise AI at Big Law's risk profile, or stay defensive. The market narrative ignores this gap entirely. Every vendor story talks about 'transformation' and 'integration' and never asks: what does a mid-market partnership actually look like when an AI system gives you a plausible but false contract interpretation that you miss in review?
Trovix's perspective is that raw benchmark performance is not the right filter for regulated professional services. What matters is verifiable lineage, explainability, and graceful failure. Trovix Sift approaches document intelligence differently: it focuses on extraction and structured data—tasks where hallucination is detectable and containable—rather than synthesis and reasoning, where it is lethal in practice. The Anthropic approach (and Harvey, and Luminance) all prioritise frontier model capability. That makes sense at scale. For mid-market firms, the safer bet is narrower, deeper, auditable AI that sits upstream of human judgment, not downstream of it. Copilot for Microsoft 365 in legal contexts falls into the same trap as Claude: it is trained for fluency, not for firms operating under FCA Consumer Duty PS22/9 or SRA post-2024 competence standards.
If you are a partner at a 150-person practice, the right move now is not to wait for Claude to be safer or to bet your reputation on 'humans-in-the-loop' review that your team will still rush. Instead, ask: what are the three most time-consuming, low-judgment tasks in our workflow? Document triage. Data extraction. Precedent assembly. Regulatory change tracking. These are the places AI can work today without generating professional negligence risk. Deploy AI there first. Use systems that show their work. Plan your own audit trail. If you are at an insurer or accountancy firm, the same logic applies: Trovix Watch and Trovix Brief sit in the intake and governance space where errors are catchable, not in the advisory core where they are catastrophic.
Source: Fortune