EU AI Act transparency duties apply now; high-risk duties from December 2027. Check your exposure
Insights About us Careers
Contact us
AI for Marketing and Growth

Conversion Rate Optimisation with AI

Using AI where it genuinely helps — finding friction in behavioural data no one has time to read — and keeping ordinary experimental rigour for the part where people fool themselves.

8 to 16 weeks
Typical engagement
Retainer or fixed
Commercial model
Powered
Tests sized first

Conversion optimisation has a replication problem. A large share of declared wins do not hold when re-run, because tests were stopped when they looked good, run on too little traffic, or judged on a metric that moved for unrelated reasons. AI does not fix any of that — it helps find better things to test, and the discipline still has to come from the method.

In one paragraph

Conversion rate optimisation is the systematic improvement of the proportion of visitors who complete a desired action, through research, hypothesis and controlled experiment. AI contributes by analysing behavioural data at a scale humans cannot, clustering friction patterns, and generating variant ideas — not by deciding what worked.

Where AI helps, and where it must not be trusted

TaskAI contributionThe discipline that stays
Finding friction in session dataHigh — reads volumes no team canJudging which friction matters commercially
Clustering drop-off patternsHigh — segments emerge from dataDeciding which segment is worth serving
Analysing form abandonmentHigh — field-level at scaleChoosing which field to fight for
Reading open-text feedbackHigh — themes across thousandsWeighing themes against revenue
Generating variant ideasModerate — a wide starting setRejecting the generic ones
Deciding a test wonNone. This is statisticsPower, duration, significance, no peeking

The last row is where the money is lost. A tool that declares a winner early because the numbers look good is not doing statistics; it is finding noise, and acting on it costs more than not testing at all.

Research before hypotheses

Quantify where the loss actually is

Funnel analysis by segment, device and source, so effort goes to the step losing the most revenue rather than the page someone dislikes. It is common to find the largest loss on a step nobody had considered part of the funnel.

Watch what people do, at scale

Session data clustered for rage clicks, dead ends, repeated corrections, hesitation and abandonment patterns. This is genuinely well suited to machine analysis, because no team can watch enough sessions to see a pattern that affects three per cent of visitors.

Read the field-level detail

Which form field loses people, which error message is triggered most, where validation fights the user. Form friction is consistently among the highest-return findings and among the least examined.

Listen to what they say

Support tickets, survey open text, chat transcripts and reviews, themed at scale. The friction people describe in their own words is frequently different from what the analytics suggested. See sentiment and intent analysis.

Then form hypotheses that can be wrong

A hypothesis states what will change, for whom, and by roughly how much. 'Improve the page' is not testable; 'reducing the form from nine fields to five will increase completion for mobile visitors by at least three points' is.

Experimental discipline, non-negotiable

  1. Calculate power before running. If your traffic cannot detect the effect size you care about in a reasonable period, do not run the test — make the change on judgement and say so, or pick a different test.
  2. Fix the duration in advance, covering full weekly cycles, and do not stop early because the numbers look good.
  3. Do not peek and act. Repeatedly checking and stopping at significance manufactures wins that will not replicate.
  4. Test one hypothesis at a time unless you are running a properly designed multivariate test with the traffic to support it.
  5. Measure the commercial outcome, not the click. A change that lifts add-to-cart and reduces completed orders is a loss.
  6. Re-run the important winners. Replication is cheap relative to building on a false result.
Worth knowing

Most tests do not win, and that is the normal shape of this work

A programme reporting a high win rate is usually stopping tests early or measuring the wrong outcome. The value comes from a steady cadence of properly powered tests where the losses are as informative as the wins — and from not shipping the changes that would have quietly cost you money.

Process

How the engagement runs

Tests are powered before they are run, and stopped when planned rather than when they look good.

Weeks 1 to 3

Research

Funnel loss quantified by segment; session data, form analytics and open-text feedback analysed at scale.

Week 4

Hypothesis backlog

Prioritised by expected value and feasibility, each stating what changes for whom and by how much.

Weeks 5 to 13

Testing programme

Powered tests run to fixed duration, measured on commercial outcomes, with losses documented as findings.

Week 14

Replication

Important winners re-run before further work is built on them.

Weeks 15 to 16

Handover

Method, backlog, analysis tooling and the discipline documented for the owning team.

Deliverables

What you receive

A research-backed backlog and a properly run testing programme, with losses reported as findings.

01

Funnel loss analysis

Where revenue is lost, by segment, device and source.

02

Behavioural analysis at scale

Friction patterns clustered from session, form and open-text data.

03

Prioritised hypothesis backlog

Each stating the change, the audience and the expected effect size.

04

Test results

Powered, run to fixed duration, measured on commercial outcomes — wins and losses both.

05

Replication of key wins

Important results re-run before further work depends on them.

06

Method handover

Power calculation, duration setting and the anti-peeking discipline documented.

Fit check

Is this the right engagement?

Worth being direct. Conversion Rate Optimisation with AI is the wrong spend in some situations, and those are listed rather than buried.

Good fit if

  • Traffic is sufficient for tests to reach significance in a reasonable period.
  • Conversion is a meaningful commercial lever.
  • Behavioural data exists but nobody has capacity to analyse it.
  • A previous testing programme produced wins that did not hold.
  • Friction is suspected but not located.

Choose something else if

  • Traffic is too low for experiments to conclude; improve on judgement and research instead.
  • The requirement is personalising per visitor. See personalisation.
  • The product or offer is the problem rather than the funnel.
  • You want a high reported win rate, which is a sign of bad method rather than good work.
Questions

Frequently asked questions

Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.

Can AI decide which variant won?

No. Deciding a test won is statistics, and tools that declare winners early because the numbers look good are finding noise. AI is genuinely useful for finding what to test — reading session data, clustering friction, theming open-text feedback at a scale no team can — and useless for the judgement about whether a result is real.

How much traffic do we need to test?

Enough to detect the effect size you care about within a sensible period, which we calculate before running anything. Smaller effects need much more traffic than people expect. Where the power calculation says a test cannot conclude, the honest answer is to make the change on research and judgement and label it as such.

Why did our previous test wins not hold?

Most often because tests were stopped when they looked good rather than at a pre-set duration, ran on insufficient traffic, or were judged on an intermediate metric that moved for unrelated reasons. Repeatedly checking and stopping at significance manufactures wins reliably, and they do not replicate.

What is a realistic win rate?

Lower than most programmes report. A steady cadence of properly powered tests where a minority win is the normal shape of this work, and the losses have value — they stop you shipping changes that would have cost money. A high reported win rate usually indicates early stopping or the wrong success metric.

How is this different from personalisation?

CRO finds the best single experience for everyone and proves it with an experiment. Personalisation shows different experiences to different visitors. They complement each other, and CRO usually comes first because a well-optimised generic experience is the baseline personalisation has to beat.

Is this the right engagement?

Tell us what you are trying to build. If a different service fits better, or if you do not need us at all, we will say so.