The question is never whether a tool is good. It is whether the work you would put it on can tolerate the errors it makes, and whether the time it saves is worth what it costs to check. Where the answer is no, we say so.
That matters more in your work than in most, because the duty does not transfer. A tool's error rate becomes the error rate of whoever signed, and no vendor carries that exposure with you.
So the discipline is the one you already apply everywhere else: understand what these systems actually do, test them against a standard before they reach client work, and measure what changes once they do. Governance and safeguards are designed in from the start, because in a firm that carries a duty they have to be.
Volume work, where a person still sees the output before it counts.
What this rests on
In a randomized controlled trial of 444 professionals, mid-level writing tasks took about 40% less time and were rated about 18% higher in quality. The weakest writers gained the most.
Preregistered randomized trial. Tasks were business writing, on an early-generation model.
Among about 5,000 customer support agents, issues resolved per hour rose about 15%. Novices gained roughly 30%; experienced agents gained little.
Population: customer support agents.
Anything carrying judgment, and anything a professional signs.
What this rests on
In a field experiment with 758 consultants, work inside the tool's competence improved: about 12% more tasks completed, about 25% faster. On work outside it, consultants using AI were 19 percentage points less likely to reach the right answer than those working without it.
Anywhere the error rate is unknown, or cannot be established.
What this rests on
In the first preregistered evaluation of AI legal research tools, the paid, purpose-built products from LexisNexis and Thomson Reuters each hallucinated between 17% and 33% of the time. These are tools sold on the promise of grounding every answer in a real citation.
Tools were tested in mid-2024 and the vendors have iterated since. Legal research, not accounting. The paper counts an answer as a hallucination if it states the law incorrectly OR cites a source that does not support it.
Asked specific, verifiable questions about random federal court cases, general-purpose chatbots hallucinated between 58% of the time (ChatGPT 4) and 88% (Llama 2).
Tested against 2023-era models, and the spread is between models rather than across tasks. Legal questions, not accounting.
Four studies, four charts, four populations. Each answers a different question, and none of them was asked about your firm specifically. That limit is worth knowing before you read them, and it is the reason the last word on any of this has to come from your own numbers rather than someone else's.
One thing is missing here, and it is missing everywhere: no published study measures whether firms use what they bought. A purchase leaves a record and use does not, so nobody counts it. It is also the question most firms cannot answer about themselves, which is usually where our work starts.
Adoption
How to read it: This is the one figure on this page that nobody was asked about. It was read out of bank transaction records rather than a survey, so it counts paid subscriptions and nothing else. Free tiers and anything on a personal card are invisible to it, which makes every bar a floor rather than a ceiling. One limit worth stating plainly: the population is Chase's own small-business banking base, so this describes the smaller end of professional services rather than a large firm. Read it as direction rather than as your number. Professional services adopts at roughly three times the rate of construction and nearly six times transportation, and the whole picture moved from almost nothing to this in six years.
Data sources
Observational, not a controlled study. It counts paid AI subscriptions in transaction data, so it undercounts free and personally-expensed tools and should be read as a floor. 'Small business' is Chase's banking population.
Implementation
How to read it: These three are stages of one study's own funnel, which is why they may be read down into each other. Nothing else on this page may. The interesting distance is the last one: getting a tool to survive contact with a real workflow is most of the work, and it is the part a demo never shows. The firms that did make it moved fast, about 90 days from pilot to production against nine months or more for everyone else, which is the more useful half of this finding. The study is interview-based and enterprise-skewed, and its authors call their own findings directional rather than precise.
Data sources
Based on 52 interviews, 153 surveys, and more than 300 public deployments. The authors describe their own findings as directionally accurate. Not peer-reviewed.
Governance
How to read it: Every number here is from one survey, one wave, one set of 351 organizations, which is why you may read them against each other. A policy is a document, and it is what a firm holds up when a client or an insurer asks. Monitoring is a practice: somebody has to look at what the tool produced and say whether it was right. The third bar is the same question asked of the small companies in that sample, which the survey defines as 500 or fewer employees, so if that is your firm it is the bar that describes you. The survey was sponsored by a company that sells AI governance software, which is worth knowing when reading a finding about how few firms govern their AI.
Data sources
Self-reported survey of 351 participants, fielded February to May 2025, 91% with US operations. 'Small' means 500 or fewer employees, which was 32% of the sample. Sponsored by Pacific AI, which sells AI governance software. Directional, not a measured rate.
Reliability
How to read it: Both bands are hallucination rates, which is why they share an axis, and both come from Stanford's RegLab. The lower band is the paid, purpose-built research tools sold by LexisNexis and Thomson Reuters, in the first preregistered evaluation of their kind: each hallucinated between 17% and 33% of the time, on tools marketed on the promise of grounding every answer in a real citation. The upper band is general-purpose chatbots on verifiable questions about federal cases, from ChatGPT 4 at the bottom to Llama 2 at the top. The two bands come from different query sets, and the tools were tested against 2023 and 2024 models that have since iterated. This is legal research rather than accounting, and it is still the most rigorous public read on what these systems do when the answer has to be right. It is also the kind of number a firm can establish for its own workflows rather than inherit from a study.
Data sources
Tools were tested in mid-2024 and the vendors have iterated since. Legal research, not accounting. The paper counts an answer as a hallucination if it states the law incorrectly OR cites a source that does not support it.
Tested against 2023-era models, and the spread is between models rather than across tasks. Legal questions, not accounting.
The work this page describes has a shape. We form a centralized AI group inside your firm, sitting under whatever already governs technology decisions there, and run it through its formation term. Grounds brings the agenda, the research, and the technical translation; the group sets its own direction.
It is a standing engagement rather than a project. Every structured session leaves a written record your people keep, and the group ends up carrying the work without us in the room.