Back to blog
guide

How to Choose an AI Development Company: 9 Checks Before You Pay for a Demo

How to choose an AI development company: nine checks on success metrics, evaluation harnesses, data boundaries, RAG and agents in production, cost and pricing.

Dennis Vorobyov
Dennis Vorobyov
Founder & CEO
October 9, 2026 · 7 min read

Choosing an AI development company is mostly about one risk: paying for a demo that never becomes a product. MIT's Project NANDA studied more than 300 enterprise AI deployments in 2025 and found that 95% delivered no measurable P&L impact, and Gartner expects more than 40% of agentic AI projects to be canceled by 2027 over cost and unclear value. The vendors that land in the other 5% share habits you can check in one call: they build the evaluation layer before the feature, they scope agents around a human-in-the-loop boundary, they can show you cost per request and latency on a live system, and the engineers who shipped production software before the AI cycle are the ones on your project. This guide lists nine checks, the red flags and a 30-minute call script, drawn from shipping RAG systems, agents and LLM integrations for clients in healthcare, legal, fintech and commerce.

Check 1. They ask what "working" means before they ask what model

A generative AI feature has a success metric or it has nothing: answers graded correct by a domain expert on a golden set, minutes saved per document, calls graded without a human, support tickets deflected. A vendor who opens with model choice is selling you a demo. A vendor who opens with "what does a wrong answer cost you, and how will we know" is building a product.

Ask them to write the success criteria into the proposal. On RiseMD, a healthcare marketing platform used by more than 5,000 dental practices, the AI call grading system is judged by whether grading runs with no human in the loop and whether every score carries the transcript evidence a practice manager can check. Those are testable claims, and they were the spec.

Check 2. The evaluation harness exists before the feature

The difference between a RAG demo and a RAG product is the evaluation layer. Ask the vendor what they ship with every system: a golden test dataset built with your domain experts, faithfulness and retrieval scoring, regression runs on every prompt change, and production monitoring so a wrong answer is caught before a customer reports it. Tools like Langfuse make this routine; the question is whether the vendor builds it first or promises it later.

If the answer is a demo on ten hand-picked examples, you are looking at the 95%.

Check 3. Data boundaries: what the model sees and where it runs

Before architecture, inventory the data: customer records, contracts, medical records, financial filings, source code. Each class decides which models you may use, under which agreement, and whether raw data can leave your cloud at all. Healthcare data under HIPAA needs a BAA-covered model offering, de-identification, or a custom model on compliant infrastructure; raw PHI sent to a public endpoint is a reportable problem, not a shortcut. The healthcare vendor guide covers that case in detail.

Ask the vendor to draw the data flow: what is embedded, what is sent to the model, what is logged, who can read the logs, and how a deletion request propagates through vector stores and caches.

Check 4. RAG experience with your document volume

Retrieval-augmented generation connects a model to your data so it answers from facts, not training data. The parts that break at scale are specific: chunking strategy, hybrid retrieval (keyword plus semantic, because pure semantic search misses exact-match queries), re-ranking, citation tracking so users can verify the source, and retrieval evaluation that measures whether the right chunks surface at all.

Ask for numbers from a live system: document count, query latency, and how retrieval quality is measured. EltexSoft's RAG pipelines run in production for LegalTech and FinTech clients across millions of documents with sub-second query latency; the generative AI development page lists the stack (OpenAI and Cohere embeddings, Pinecone, Qdrant or pgvector, LangGraph for stateful agents).

Check 5. Agents with a human-in-the-loop boundary

Agents that plan, call tools and self-correct are the highest-value and highest-risk category. The design decision that matters is where the agent acts on its own and where it asks a human; getting that boundary wrong is how agents cause real damage. Ask the vendor to show an agent they shipped, the checkpoints it has for high-stakes actions, what happens when a tool call fails, and how a bad run is replayed and diagnosed.

Scope is the other test. Agentic projects get canceled when the scope is "automate the department". The ones that survive automate one workflow with a measurable cost, such as an automated document review pipeline or a workflow that replaced manual processes costing hundreds of engineer-hours a month, and expand from there.

Check 6. Cost, latency and fallback are engineered, not hoped for

A production system has to handle thousands of requests an hour, stay inside a token budget, answer in under two seconds, degrade gracefully when a provider is down and cost less than the value it creates. Ask what the vendor adds on top of a raw API call: prompt management and versioning, response caching, cost controls and token budgeting, rate limiting, fallback routing across providers, and an abstraction layer that lets you swap models without touching application code.

Then ask for a dashboard screenshot from a live client: latency, cost per request, error rate. A vendor who runs AI in production has one.

Check 7. Production engineering before AI engineering

The AI layer is new; the engineering discipline is not. Ask how long each engineer on the proposed team shipped production software before touching a model, and whether the same team handles infrastructure, CI/CD for prompts, observability and on-call. On Snapwire, a photography marketplace serving Dell, Starbucks and other Fortune 500 brands, ten EltexSoft engineers spent two and a half years running ML image tagging and quality scoring across millions of images on AWS; the platform was later acquired by StudioNow. Scale problems in ML are mostly ordinary scale problems.

Check 8. Team continuity and the first call

Models change every quarter; your evaluation sets, your data pipelines and your prompt history should not walk out the door with a rotated engineer. Ask for the average tenure of engineers on client accounts, the written replacement commitment, and whether the people on the first call will write the code. At EltexSoft the first call is with an engineer, the average client engagement is 3.4 years and replacement within two weeks is in the contract; the how we work page has the terms.

Check 9. Pricing that matches the phase you are in

Published numbers to compare against. A discovery sprint that ends in a working prototype on your real data and a written go/no-go runs $25K to $60K over 4 to 8 weeks. A production MVP with evaluation harness, observability, CI/CD and a runbook runs $80K to $250K over 3 to 5 months with an AI lead, one or two AI engineers, a data engineer and QA. A retained pod of 4 to 6 engineers runs $40K to $90K a month, and staff augmentation of specific roles (RAG architect, prompt engineer, LLM evaluation specialist) is $50 to $99 per hour. For context, Clutch's April 2026 data puts the average AI development project at $120K over 10 months, and senior AI engineers in the US cost $150 to $250 or more per hour with 3 to 6 month hiring timelines.

Match the engagement to the phase: pay for a discovery sprint when the use case is unproven, not for an MVP. And own the assets from day one: prompts, evaluation datasets, fine-tuned weights, vector indexes and the accounts they live in.

Red flags

  • The first conversation is about which model, not what "correct" means.
  • A demo on hand-picked examples and no golden test dataset.
  • No answer to "what happens when the provider is down" or "what does a request cost".
  • An agent proposal with no human-in-the-loop checkpoints for high-stakes actions.
  • Customer, medical or financial data sent to public endpoints with no agreement and no de-identification.
  • Engineers whose first production system is yours.
  • A 12-month minimum before a prototype exists.
  • Prompts, evaluation sets or vector indexes held in the vendor's accounts.

A 30-minute call script

  • What would make this feature "working" in your proposal, and how will we measure it?
  • What evaluation harness ships with the system, and who builds the golden dataset?
  • Draw the data flow: what is embedded, what reaches the model, what is logged, who can read it.
  • Which RAG or agent system of yours is in production today, at what volume and latency?
  • Where does your agent act alone and where does it ask a human?
  • Show me cost per request and latency on a live client.
  • How many years did each proposed engineer ship production software before working with models?
  • What does a discovery sprint cost, what do we get, and what do we own afterwards?

Where EltexSoft fits

We are a Los Angeles software engineering studio founded in 2015, with 35 to 50 senior engineers and no juniors on client projects. Our AI work includes RiseMD's call grading and AI search positioning, RAG systems for LegalTech and FinTech clients, ML image tagging for Snapwire, recommendations and demand forecasting for Woodies Clothing, and work with Arcade.ai, a $42M-funded AI company; the generative AI development and AI and ML development pages list the scope and the published prices above. If you are shortlisting, we are one of the teams worth 30 minutes. For the general version of this framework, see how to choose a software development partner.

Frequently asked

How much does AI development cost?
Published EltexSoft numbers: a discovery sprint with a working prototype on your data and a written go/no-go is $25K to $60K over 4 to 8 weeks; a production MVP with evaluation harness, observability, CI/CD and a runbook is $80K to $250K over 3 to 5 months; a retained pod of 4 to 6 engineers is $40K to $90K a month; staff augmentation of specific AI roles is $50 to $99 per hour. For context, Clutch's April 2026 data puts the average AI development project at $120K over 10 months.
What is the difference between an AI demo and a production AI system?
A demo works on ten hand-picked examples in a notebook. A production system handles thousands of requests an hour, stays inside a token budget, answers in under two seconds, does not hallucinate when the data is ambiguous, falls back gracefully when a provider is down, and costs less than the value it creates. The evaluation harness, monitoring and cost controls are what turn the first into the second.
Should we build an AI agent or start with RAG or a copilot?
Start with the smallest system that has a measurable success metric. RAG and copilots answer questions from your data and are easier to evaluate; agents take actions and need human-in-the-loop checkpoints for high-stakes steps. Gartner expects more than 40% of agentic AI projects to be canceled by 2027 over cost and unclear value, and the survivors are scoped to one workflow with a known cost, not a whole department.
Can we send customer or patient data to ChatGPT or other public LLM APIs?
Only under the right agreement and data boundary. Regulated data such as health records needs a BAA-covered model offering, de-identification before the call, or a custom model on compliant infrastructure. For any data, ask the vendor to draw what is embedded, what reaches the model, what is logged and who can read the logs, and how a deletion request propagates through vector stores and caches.
Which models and tools should an AI development company know?
Commercial and open models (OpenAI, Anthropic Claude, Google Gemini, Llama, Mistral) behind an abstraction layer so you can swap them, vector stores such as Pinecone, Qdrant, Weaviate or pgvector, agent frameworks such as LangGraph, CrewAI or the OpenAI Agents SDK, and an evaluation and observability stack such as Langfuse. More important than any single tool is that the vendor can show cost per request, latency and evaluation scores on a system they run today.

Related posts

Need engineers who think this way?

Senior developers on retainer. Same team, month 1 and month 36+.

Talk to us