A 2025 MIT study on enterprise AI made the rounds for one number: of the generative-AI pilots companies actually funded, roughly 95% returned nothing measurable to the business. Around the same time, the big consultancies were booking generative-AI work by the billion — Accenture alone reported on the order of $3 billion in generative-AI bookings in a single fiscal year. Hold those two facts next to each other. An enormous amount of money is being spent telling companies how to do AI, and almost none of what gets built survives contact with a P&L. That gap is the whole story of what AI consulting is in 2026, and most of the people selling it are standing on the wrong side of it.
The thing AI consulting is sold as
The standard product is easy to describe because every firm ships the same one. A team arrives, runs a 'discovery,' interviews your stakeholders, and produces a use-case inventory scored on a two-by-two of business value against technical feasibility. You get a maturity assessment, a prioritized roadmap, a 'data readiness' section, and a deck with a slide that says 'quick wins' in the top-right quadrant. The engagement ends when the deck is delivered. That is the deliverable. The definition of AI consulting that lives in most buyers' heads — and most sales decks — is exactly this: an outside expert tells you where AI fits and hands you a plan.
None of that is fraudulent. A use-case matrix is a reasonable artifact. The problem is that the artifact is treated as the product, and the product stops precisely at the point where all the actual risk begins. You now own a roadmap that asserts, with great confidence, that a retrieval-augmented chatbot over your support docs will deflect 40% of tickets. Nobody in that room has built one over your docs. The 40% is a benchmark borrowed from someone else's data.
Why the roadmap is where projects go to die
Here is the part the conventional definition refuses to absorb: with large language models, you cannot know whether something works until it is running against real data, real users, and real edge cases. This is different from most software. A CRUD app behaves the way the spec says. An LLM feature behaves the way your specific corpus, your specific prompts, and your specific users' phrasing make it behave — and that surface is unknowable from a conference room. The retrieval that looked clean in the demo falls apart on your messy PDFs. The agent that scored well on a public benchmark hallucinates an order ID the first time a customer asks something slightly off-script.
So the deck-and-roadmap model isn't just incomplete; it sells certainty about the one thing that is genuinely uncertain. Gartner has forecast that a large share of generative-AI projects — on the order of a third — get abandoned after the proof-of-concept stage. They don't die because the strategy was wrong on paper. They die in the gap between the slide and production, which is the gap the strategy engagement quietly handed back to you. The roadmap was the easy 20% of the work priced as if it were the hard 80%.
Real AI consulting is engineering wearing a cleaner shirt
If the behavior only reveals itself in production, then the only advice worth anything comes from someone who has put these systems into production and watched them break. That collapses the comfortable distinction between 'consulting' and 'building.' The useful AI consultant is not a person who knows the landscape of models and frameworks in the abstract — that knowledge has a half-life measured in months and is free on a dozen blogs. The useful AI consultant is the person who can tell you, from having shipped it, that your document-processing idea will work but your evaluation harness is the thing that will sink you, and then go write that harness.
Phrase it as a buying test. If the firm's deliverable is a document, you are buying strategy theater. If the firm's deliverable is a working system — or an honest 'we built the smallest version of this and here's what it actually did against your data' — you are buying consulting. The label on the statement of work is noise. The question is whether the same people who form the opinion are on the hook for the opinion being right in production.
The questions that separate the two
You can sort a real AI engagement from a deck factory in about ten minutes, before you sign anything. Ask who on the team has personally deployed an LLM feature that's currently serving users, and ask what broke. Ask how they'll evaluate quality — if there's no answer involving a test set and a metric, they're guessing. Ask what they'd cut from your wish list, because anyone who says yes to all of it is selling hours, not judgment. Ask what they recommend you not do with AI, because a consultant who can't name the bad use cases hasn't seen enough of them. The strategy-only firms get visibly uncomfortable here, because their entire model depends on never being measured against a running system.
What this looks like when you actually do the work
At EltexSoft this is not a philosophical preference; it's how the studio is built. Our AI-ML practice grew out of engineers who already ran multi-year production builds, not out of a freshly minted 'AI strategy' division. When a client asks 'should we use AI for this' — RAG and AI search, an agentic workflow, a copilot, document processing — the answer doesn't arrive as a deck. It arrives as a free discovery week, then a paid pilot with no lock-in: a deliberately small, real version of the thing, built against the client's own data, so the behavior we'd otherwise be guessing about becomes something we can measure. The advice and the build are the same motion. That's the whole point.
It also means the boring engineering discipline is the consulting. Every pull request is reviewed by at least one other senior engineer before it merges; the AI tools we use — Claude Code, Cursor — are treated as reviewed accelerators, not as license to ship unread output. When we run technical due diligence or vendor evaluation under a fractional-CTO engagement, the recommendation is grounded in people who would have to live with being wrong. We reject most of the tools we evaluate, because the honest output of an evaluation is usually 'no' or 'not yet.' A firm that says yes to everything is not assessing; it's quoting.
The missing 90% is engineering, and that's the part we sell. Our AI engineering services cover the pipeline from prototype to production — evals, guardrails, and the unglamorous work that makes it real.
So, what is AI consulting?
It is two products fighting over one name. One is a strategy artifact that ends at the roadmap and hands the hard, uncertain, production-shaped part back to you — and given how many funded pilots return nothing, that product is a meaningful contributor to the failure rate, not a defense against it. The other is engineering with an advisory front end: the same people who tell you what to build are the people who build it, deploy it, and watch it meet reality. Don't buy the first one. The deck is cheap to produce and expensive to act on. Buy the second one, or buy nothing and hire engineers directly — because in a field where the only ground truth is what happens in production, advice that can't ship isn't advice. It's a slide.