AI Proof of Concept Development
One narrow use case built against real data in four to six weeks, judged against a threshold set before work begins.
A proof of concept has one job: turn an argument into evidence. It fails that job when the scope drifts, when the success criteria are written after the results are known, or when it quietly becomes a production system nobody hardened. We run them narrow, timeboxed, and against a number agreed in writing on day one.
The threshold comes first
Before any code, we agree what success looks like as a measurable statement. Not "the model should perform well" but something closer to "correctly extracts all eight required fields from at least 92 percent of a held-out sample of 400 invoices, with no more than 2 percent confident errors". The threshold is written into the engagement. We cannot move it later and neither can you.
Two thresholds are usually needed. One for quality and one for economics, because a system that is accurate but costs more per document than the person it replaces has still failed.
Scope discipline
One use case. One data source where possible. No integrations beyond what is needed to prove the point. The interface is deliberately minimal, because a polished interface on an unproven model wastes budget and, worse, convinces stakeholders the thing is nearly finished when the hard work has not started.
Anything raised mid-pilot that falls outside the agreed scope goes on a list and is addressed in the production estimate rather than absorbed. Scope creep is the single most common reason pilots overrun.
How the weeks run
| Phase | Work |
|---|---|
| Week 1 | Data access, held-out test set construction, baseline measurement of the current manual process. |
| Weeks 2 to 4 | Build and iterate. Evaluation runs against the held-out set at the end of each week, results shared whether they are good or not. |
| Week 5 | Hardening of the evaluation, edge case testing, cost per unit measurement at realistic volume. |
| Week 6 | Production estimate, written verdict, handover of code, prompts and evaluation harness. |
What you receive
- A working prototype you can run, with the source code and infrastructure definitions.
- The evaluation harness and held-out test set, which remain useful long after the pilot.
- Measured results against the agreed threshold, including where it failed.
- Cost per unit of work at pilot volume and projected at production volume.
- A production estimate covering engineering, integration, governance and three year run cost.
- A written verdict: proceed, proceed with named changes, or stop.
A failed pilot is a successful engagement
If the threshold is missed we report that plainly, explain whether the cause is fixable, and do not propose a phase two to chase it. A four to six week pilot that prevents a nine month build is the cheapest money you will spend on AI this year.
What a proof of concept is not
It is not a production system. It has no load handling, limited error recovery, minimal monitoring, and security appropriate to a controlled environment rather than to live traffic. We state this clearly at handover because the most common failure pattern in this industry is a successful pilot being quietly pushed into production by an enthusiastic stakeholder. If you intend to run it live, that is a separate build and we will scope it as one.
Output
Working prototype, evaluation results, production estimate, go or no-go verdict. Delivered in editable formats. Full IP transfers to you on final payment.
FAQ
Can we use synthetic or sample data instead of real data?
For a first look, sometimes. For a verdict you can act on, no. Synthetic data lacks exactly the messiness that determines whether a system works, and pilots on clean data routinely overstate accuracy by a wide margin. If real data cannot leave your environment, we work inside it.
What happens to the code if we do not proceed?
You keep it, along with the evaluation harness and documentation. IP transfers at final payment regardless of the verdict. Nothing is held back to encourage a phase two.
Six weeks seems short. Can you extend it?
We would rather narrow the scope than extend the clock. Long pilots lose stakeholder attention and start to accumulate production expectations. If the question genuinely cannot be answered in six weeks, that usually indicates it should be split into two smaller questions.
Do you build the production system afterwards?
We can, and it is quoted separately with no obligation. Some clients take the prototype and evaluation set to their internal team, which is a legitimate outcome and one the handover is designed to support.
Often paired with
AI Use Case Discovery
A structured workshop that converts a long list of AI ideas into a ranked shortlist with feasibility and value scored.
Read moreAI Readiness Assessment
A two week audit that scores whether your organisation can actually support the AI you want to build, and tells you what to fix first.
Read moreAI Vendor and Model Selection
Structured evaluation of vendors and foundation models, benchmarked against your data rather than public leaderboards.
Read moreIs this the right engagement?
Tell us what you are trying to decide. If a different service fits better, or if you do not need us at all, we will say so.