AI Research & Reliability

AI is not reliable or unreliable. It has an error rate.

The question is never whether a tool is good. It is whether the work you would put it on can tolerate the errors it makes, and whether the time it saves is worth what it costs to check. Where the answer is no, we say so.

That matters more in your work than in most, because the duty does not transfer. A tool's error rate becomes the error rate of whoever signed, and no vendor carries that exposure with you.

So the discipline is the one you already apply everywhere else: understand what these systems actually do, test them against a standard before they reach client work, and measure what changes once they do. Governance and safeguards are designed in from the start, because in a firm that carries a duty they have to be.

Where it is dependable today

Volume work, where a person still sees the output before it counts.

For example
  • Reading and summarizing more material than a person has time for.
  • First-pass extraction: pulling the terms, dates, and figures out of a document so someone can check them.
  • Drafting from a format you already use.
  • Finding where something lives, across systems nobody ever indexed.

What this rests on

  • In a randomized controlled trial of 444 professionals, mid-level writing tasks took about 40% less time and were rated about 18% higher in quality. The weakest writers gained the most.

    Noy and Zhang, Science, 2023 · Verified July 2026

    Preregistered randomized trial. Tasks were business writing, on an early-generation model.

  • Among about 5,000 customer support agents, issues resolved per hour rose about 15%. Novices gained roughly 30%; experienced agents gained little.

    Brynjolfsson, Li and Raymond, Quarterly Journal of Economics, 2025 · Verified July 2026

    Population: customer support agents.

The evidence

What the public research measures

Four studies, four charts, four populations. Each answers a different question, and none of them was asked about your firm specifically. That limit is worth knowing before you read them, and it is the reason the last word on any of this has to come from your own numbers rather than someone else's.

One thing is missing here, and it is missing everywhere: no published study measures whether firms use what they bought. A purchase leaves a record and use does not, so nobody counts it. It is also the question most firms cannot answer about themselves, which is usually where our work starts.

Adoption

Professional services is among the fastest-adopting sectors

up from 1.7% in January 2019

How to read it: This is the one figure on this page that nobody was asked about. It was read out of bank transaction records rather than a survey, so it counts paid subscriptions and nothing else. Free tiers and anything on a personal card are invisible to it, which makes every bar a floor rather than a ceiling. One limit worth stating plainly: the population is Chase's own small-business banking base, so this describes the smaller end of professional services rather than a large firm. Read it as direction rather than as your number. Professional services adopts at roughly three times the rate of construction and nearly six times transportation, and the whole picture moved from almost nothing to this in six years.

Data sources

  • Measured, not asked · Professional services. Paid AI subscriptions identified in de-identified transaction records, 4.6 million small businesses, through the end of 2025. Read it
  • Measured, not asked · Other sectors, and the small-business average, from the same records.
  • DirectionalVerified July 2026

    Observational, not a controlled study. It counts paid AI subscriptions in transaction data, so it undercounts free and personally-expensed tools and should be read as a floor. 'Small business' is Chase's banking population.

See how a firm stands up the group that decides this

Implementation

The distance between a pilot and production

How to read it: These three are stages of one study's own funnel, which is why they may be read down into each other. Nothing else on this page may. The interesting distance is the last one: getting a tool to survive contact with a real workflow is most of the work, and it is the part a demo never shows. The firms that did make it moved fast, about 90 days from pilot to production against nine months or more for everyone else, which is the more useful half of this finding. The study is interview-based and enterprise-skewed, and its authors call their own findings directional rather than precise.

Data sources

  • Self-reported · 52 interviews, 153 surveys, and more than 300 public deployments. Not peer-reviewed; the authors describe the findings as directionally accurate.
  • DirectionalVerified July 2026

    Based on 52 interviews, 153 surveys, and more than 300 public deployments. The authors describe their own findings as directionally accurate. Not peer-reviewed.

See how a workflow gets built to survive production

Governance

Written policy and active monitoring are different commitments

one survey, 351 organizations

How to read it: Every number here is from one survey, one wave, one set of 351 organizations, which is why you may read them against each other. A policy is a document, and it is what a firm holds up when a client or an insurer asks. Monitoring is a practice: somebody has to look at what the tool produced and say whether it was right. The third bar is the same question asked of the small companies in that sample, which the survey defines as 500 or fewer employees, so if that is your firm it is the bar that describes you. The survey was sponsored by a company that sells AI governance software, which is worth knowing when reading a finding about how few firms govern their AI.

Data sources

  • Self-reported · All 351 organizations: the policy, and whether anyone monitors it.
  • Self-reported · 351 participants, fielded February to May 2025, 91% with US operations. 'Small' means 500 or fewer employees. Sponsored by a company that sells AI governance software. Read it
  • DirectionalVerified July 2026

    Self-reported survey of 351 participants, fielded February to May 2025, 91% with US operations. 'Small' means 500 or fewer employees, which was 32% of the sample. Sponsored by Pacific AI, which sells AI governance software. Directional, not a measured rate.

See what governance actually involves

Reliability

Measured hallucination rates in legal AI tools

peer-reviewed

How to read it: Both bands are hallucination rates, which is why they share an axis, and both come from Stanford's RegLab. The lower band is the paid, purpose-built research tools sold by LexisNexis and Thomson Reuters, in the first preregistered evaluation of their kind: each hallucinated between 17% and 33% of the time, on tools marketed on the promise of grounding every answer in a real citation. The upper band is general-purpose chatbots on verifiable questions about federal cases, from ChatGPT 4 at the bottom to Llama 2 at the top. The two bands come from different query sets, and the tools were tested against 2023 and 2024 models that have since iterated. This is legal research rather than accounting, and it is still the most rigorous public read on what these systems do when the answer has to be right. It is also the kind of number a firm can establish for its own workflows rather than inherit from a study.

Data sources

  • Peer-reviewed testing · Paid legal research tools (Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI): each hallucinated 17% to 33% of the time. Magesh and others, JELS 2025. Read it
  • Peer-reviewed testing · General-purpose chatbots on verifiable federal-case questions: 58% (ChatGPT 4) to 88% (Llama 2). Dahl and others, Journal of Legal Analysis 2024. Read it
  • High confidenceVerified July 2026

    Tools were tested in mid-2024 and the vendors have iterated since. Legal research, not accounting. The paper counts an answer as a hallucination if it states the law incorrectly OR cites a source that does not support it.

  • High confidenceVerified July 2026

    Tested against 2023-era models, and the spread is between models rather than across tasks. Legal questions, not accounting.

See how reliability gets tested

Building a practice you own

The work this page describes has a shape. We form a centralized AI group inside your firm, sitting under whatever already governs technology decisions there, and run it through its formation term. Grounds brings the agenda, the research, and the technical translation; the group sets its own direction.

It is a standing engagement rather than a project. Every structured session leaves a written record your people keep, and the group ends up carrying the work without us in the room.