EU AI Act high risk obligations are now enforceable. Check your exposure
Insights About us Careers
Contact us
AI Strategy & Consulting

AI Vendor and Model Selection

Structured evaluation of vendors and foundation models, benchmarked against your data rather than public leaderboards.

3 to 5 weeks
Typical duration
Fixed fee
Commercial model

Public benchmarks tell you how a model performs on problems that are not yours. Vendor demos show the path the vendor rehearsed. Neither predicts what happens when your actual documents, with their scanning artefacts and inconsistent formats, hit the system on a Tuesday. Selection work exists to close that gap before you sign.

Requirements first

We build a requirements matrix before looking at any vendor, because requirements written after seeing a product tend to describe that product. Requirements are separated into three tiers so that scoring is not distorted by long lists of features nobody needs.

  • Must have: absence of this disqualifies the option, no exceptions and no roadmap promises accepted.
  • Should have: materially affects value but has a workaround.
  • Nice to have: recorded, weighted low, and explicitly excluded from disqualification.

Non-functional requirements go in the same matrix: data residency, retention and training-use terms, SSO and access control, audit logging, uptime commitments, and support response times. These frequently eliminate options that scored well on capability.

Benchmarking on your data

For model selection we build a test set from your own material, including the awkward cases rather than the clean ones. That set becomes a durable asset: you keep it, and you can rerun it every time a new model version ships.

We measureWhy it matters
Task accuracy on your dataThe only accuracy number that predicts your outcome. Public scores do not transfer.
Failure modeHow the model is wrong matters more than how often. Confidently wrong is worse than uncertain.
Latency at your payload sizeDemos use short inputs. Your documents are not short.
Cost per unit of workNot cost per token. Cost per invoice, per ticket, per document processed.
Consistency across runsVariance on identical input, which determines how much human review you will need.
Behaviour at the edgesLong inputs, mixed languages, poor scans, adversarial content.

Commercial and risk review

Capability is half the decision. We review the contractual and operational side with equal weight: whether your data can be used for training, where it is processed, what happens to it on termination, how model versions are deprecated and with what notice, what the actual support path looks like at your contract tier, and what the vendor's own dependency chain is. A vendor built entirely on one model provider carries that provider's risk, and you inherit it.

What you receive

  • A weighted requirements matrix with tiers agreed before evaluation begins.
  • Benchmark results on your data, with the test set and harness handed over to you.
  • A scored shortlist with the reasoning for each score.
  • A risk register covering data handling, lock-in, deprecation and vendor concentration.
  • Negotiation notes: which terms are commonly movable in this category and where to push.
Worth knowing

The benchmark harness is yours to keep

You leave with a reusable evaluation set and the code to run it. When the next model generation arrives you can test it in an afternoon instead of commissioning another selection exercise.

Model selection versus vendor selection

These are related but separate. Choosing between foundation models is largely a benchmarking and cost exercise, and the answer changes every few months, so the decision should be built to be reversible. Choosing a vendor is a commercial commitment with switching costs measured in quarters. We treat model choice as a configuration decision and vendor choice as an architectural one, and we design so that the first can change without disturbing the second.

What you leave with

Output

Requirements matrix, benchmark results, scored shortlist, negotiation notes. Delivered in editable formats. Full IP transfers to you on final payment.

Questions

FAQ

Do you take referral fees or commission from vendors?

No. We take no commission, referral fee or reseller margin from any vendor we evaluate. If that ever changes for a specific vendor we will disclose it before the engagement starts, and you can exclude them.

How many vendors should we evaluate?

Three to five in depth. Longer lists produce shallower evaluation and slower decisions. We use a documented screening pass to get from a long list to a shortlist, so you can see why anyone was excluded.

Can you evaluate a vendor we have already chosen?

Yes. A validation review before contract signature is quicker and often catches contractual terms around data use, deprecation notice and exit that were not examined closely during the sales process.

What if no vendor meets our must-have requirements?

That is a real and useful outcome. It usually means either the requirement is stricter than the market currently serves, in which case we say so and you can decide whether to relax it, or that building is the better route.

Is this the right engagement?

Tell us what you are trying to decide. If a different service fits better, or if you do not need us at all, we will say so.