An AI audit worth paying for does one thing: it tells you exactly where AI would change a measurable outcome in your business, what that is worth, and in what order to act — with numbers you can hold a later project to. Anything that ends in a maturity score, a technology overview, or a recommendation to buy the auditor's platform is marketing wearing an audit's clothes. This article sets out what a genuine operational audit contains, the deliverables to demand, what it should never be, and how the pricing models work so you can judge a quote without a published rate card.
What should an AI audit actually include?
Six components. A credible firm may package or name them differently, but all six should be present.
Workflow walk-throughs — as-run, not as-documented. The audit's raw material is how work actually moves: who touches a tender, an invoice, a quote, a claim; where it waits; what gets retyped; which workarounds exist because the official process does not survive contact with reality. This is interview-and-observation work, done with the people who do the job, not their managers. The gap between the documented process and the real one is usually where the value hides.
Leakage quantification. Each mapped workflow gets a number: hours consumed, loaded cost, error rates and their downstream cost, and the opportunity cost of delay where it applies — expressed as a defensible range rather than false precision. This is the discipline covered in depth in operational leakage, and it is the component most often missing from cheap assessments, because it is the hardest to fake.
Data readiness. For each candidate workflow: does the data needed actually exist, where does it live, how clean is it, and can it be reached without a year of integration work? An honest audit kills use cases here — a recommendation that assumes data you do not have is a fantasy with a Gantt chart. This is also where obligations enter: if the workflow touches personal data at scale, the audit should say plainly what India's DPDP Act or Gulf data-protection regimes imply for the design.
Use-case scoring. Candidates ranked on two axes — value at stake (from the leakage numbers, not enthusiasm) and feasibility (from the data-readiness work, plus workflow stability and exception tolerance). The scoring method matters less than its inputs being the audit's own measurements rather than industry benchmark tables. How to run this honestly is expanded in choosing AI use cases.
Build-versus-buy calls. For the top two or three use cases, a reasoned recommendation: existing product, custom build, partner-delivered system, or — the option a self-interested auditor never offers — process fix with no AI at all. What the audit adds to a standard build-versus-buy framework is applying it to your specifics: your volumes, your systems, your team.
A sequenced roadmap with denominators. Not a list of twenty initiatives — an order. First project, why first, what it costs, and critically: the baseline it will be measured against, stated before any build begins. Every roadmap item should carry its denominator — the measured current cost that a future result divides into. Without that, no later claim of ROI can be honest, which is the whole argument of building an AI business case.
What deliverables should you insist on?
| Deliverable | The test it must pass |
|---|---|
| As-run workflow maps | People who do the work agree it is accurate |
| Leakage figures per workflow | Ranges with stated assumptions you could defend to a CFO |
| Data-readiness findings | Names real systems and real gaps, not generic "data silos" |
| Scored use-case shortlist | Scores trace back to the audit's own numbers |
| Build-vs-buy recommendation | Includes at least one non-AI or do-nothing option |
| Sequenced roadmap | Each item has a cost, an owner, and a baseline metric |
The single sharpest test of the whole engagement: could you hand the roadmap's first item to any competent delivery team — including one unrelated to the auditor — and start next month? If the answer is no, the audit produced atmosphere, not direction.
What an AI audit should NOT be
A tool demo in disguise. If the firm sells a platform and the audit's conclusions converge on that platform, you did not buy an audit; you bought a pre-sales process. Separating diagnosis from delivery incentives is not always possible — many good firms do both — but the audit contract should explicitly permit, and price, a roadmap you execute with someone else.
A slideware maturity model. Being placed on a five-level maturity curve, benchmarked against your industry, and handed a pyramid diagram describes your organisation without deciding anything. Maturity models are consulting furniture: comfortable, generic, and reusable across every client. Your operations are specific; the audit should be too.
A technology survey. A tour of what large language models can do, with logos, is available free on the internet. You are paying for the intersection of that capability with your workflows and your numbers — nothing else.
Interviews with management only. An audit that never sits with the people who process the documents, chase the approvals, and maintain the workaround spreadsheets will map the org chart's idea of the process. The value is in the gap between that idea and reality, and only the people inside the workflow can show it to you.
How is an AI audit priced, and what drives the cost?
Two models dominate, and the choice between them matters more than the absolute figure.
Fixed-scope, productised assessments. A defined number of workflows, a defined number of weeks, named deliverables, one price agreed up front. This is the buyer-friendly model: risk is capped, the firm carries the estimation burden, and comparing quotes is possible because scopes are explicit. It fits the mid-market especially well, where price predictability is often a precondition for the purchase happening at all.
Open-ended discovery. Billed by time, scoped loosely, expanding as findings emerge. Occasionally justified for very large, multi-entity organisations; for most buyers it is where budgets dissolve, because the incentive runs towards more discovery rather than towards a decision.
Whatever the model, cost is driven by identifiable factors: how many workflows and sites are in scope, how many people must be interviewed, how messy and distributed the data landscape is, whether regulatory analysis is needed, and the seniority of who actually does the work — a partner-led audit and a junior-analyst audit at the same price are very different purchases. Ask any quoting firm to explain their price in terms of these drivers. A firm that can is estimating; a firm that cannot is anchoring.
One more cost question worth asking directly: what happens if the audit finds nothing worth doing? A firm willing to conclude "do not invest yet, fix these two data problems first" is demonstrating the independence you are paying for. More questions in this vein — references, incentives, who shows up — are collected in questions to ask before hiring an AI consultancy.
Where to start
Scope the audit yourself before anyone scopes it for you. Pick the two or three workflows that feel most expensive — document-heavy, deadline-bound, handoff-dense are the usual suspects — and ask candidate firms to quote a fixed-scope assessment of exactly those, with the six components and the deliverables table above written into the agreement. Six components, real numbers, a first project you could start next month: that is the entire specification. An audit that meets it typically pays for itself before any AI is built, because measured workflows get managed differently — and an audit that cannot meet it was never going to be worth its price at any price.