EU AI Act transparency duties apply now; high-risk duties from December 2027. Check your exposure
Insights About us Careers
Contact us
Retail & E-commerce

AI for Beauty and Cosmetics Brands

Beauty runs on claims, shades and trends: three things that are respectively regulated, hard to match remotely, and impossible to forecast from history.

Tier 2
Our depth here
Claims
A regulated surface
Skin tone
Where models fail

Two things distinguish beauty from the rest of retail: almost every product carries a claim that is legally constrained, and the single most useful thing you could offer a customer online — will this shade suit me — is genuinely hard to do well and easy to do badly in a way that excludes people.

In one paragraph

AI for beauty and cosmetics covers shade and product matching, ingredient and formulation data structuring, claims-safe content generation, review and user-generated content integrity, trend-driven demand forecasting, and replenishment timing prediction.

Shade matching, and the failure mode that matters

Shade matching from a photograph is the most valuable customer-facing application in beauty and the one with the clearest and best-documented failure pattern: systems trained on unrepresentative data perform markedly worse on darker skin tones, and the failure is invisible in an aggregate accuracy figure.

Report performance by skin tone group, always

Aggregate match accuracy conceals exactly the disparity that matters. Report by tone group as a primary output from the first evaluation, using an established scale, and treat a gap as a defect rather than a limitation to be noted.

Lighting and camera variance dominate the problem

White balance, colour temperature and camera processing vary enormously across devices and rooms, and they shift measured skin tone more than most differences between people. A reference object or a calibration step in the capture flow does more for accuracy than a better model.

Undertone is what customers actually get wrong

Depth is comparatively easy; undertone is where mismatches happen and where returns come from. It is also where your own shade data is most often inconsistent between products and ranges.

Say what you do not know

A confident wrong shade recommendation costs a return and a customer's trust. Where the system is uncertain — poor lighting, an unusual undertone, a range with sparse coverage — presenting a narrow set of options honestly beats a single confident answer.

Claims, which are regulated before they are marketing

Cosmetic claims are constrained by law in every major market, and the boundary between a cosmetic claim and a medicinal one is narrow. Generated marketing copy is a fast route across it, because a language model has no sense of which verbs are load-bearing.

Claim typePositionImplication for generated content
Cosmetic effect ("hydrates", "smooths")Permitted with substantiationSafe if generated from an approved claim set
Comparative performanceNeeds evidence you can produceNever let a model invent a comparison
Structural or medicinal ("repairs", "treats")Frequently outside cosmetic scopeThe classic generated-copy failure
Sustainability and natural claimsActively enforcedSubstantiation required; vague claims are exposure
Ingredient claimsMust match the formulationGenerate from the ingredient data, never from the name
Reviews and testimonialsFTC rule appliesGenerated or incentivised reviews are in scope
Worth knowing

Generate from an approved claim library, not from a brief

The pattern that works in regulated consumer categories is a controlled vocabulary: an approved claims library with substantiation attached, and generation constrained to it. The model assembles and adapts approved language rather than writing new claims. This is slower to set up than free-form generation and it is the difference between a content operation that scales and one that produces a regulatory problem at volume.

  • Trend-driven demand does not forecast from history. A product that goes viral has no analogue in your own data, and pretending otherwise produces confident nonsense.
  • Early signal beats prediction. Detecting a spike in its first days — search volume, UGC velocity, add-to-basket rate — and responding fast is worth more than a forecast that claims to have seen it coming.
  • Supply response is the constraint. Knowing about a spike is only useful if manufacturing lead times allow anything to be done, which in beauty they frequently do not.
  • Replenishment timing is the reliable win. Consumables have predictable exhaustion, so predicting when a customer runs out and prompting then is straightforward and effective.
  • Measure replenishment against a holdout. Prompting people who were going to reorder anyway shows excellent attributed revenue and creates nothing.
Worth knowing

Review integrity is now a specific legal surface

Beauty is heavily review-driven, which makes it a focus category for review integrity. The FTC rule on consumer reviews and testimonials covers generated reviews, incentivised positive reviews and employee-written reviews, and 2026 enforcement has moved from warning letters into broader deceptive conduct cases. If you run a seeding or gifting programme, the disclosure requirements are a design question for the programme rather than a footnote in the brief.

Process

How an engagement runs

Claims and shade data first, because they constrain what can be generated and matched.

Weeks 1 to 3

Claims and shade data audit

Approved claim library with substantiation, and shade data consistency across ranges.

Weeks 4 to 8

Product data foundation

Ingredients, shades, undertones and attributes structured; capture flow designed if shade matching is in scope.

Weeks 9 to 14

Build

Shade matching with tone-group evaluation, or claims-constrained content generation.

Weeks 15 to 18

Evaluation

Match accuracy by tone group, return rate by reason, and content reviewed by regulatory before release.

Ongoing

Operation

Tone-group performance monitored as a standing control; replenishment measured against holdout.

Deliverables

What you receive

Better matching for every customer, and content that will not create a claims problem.

01

Approved claims library

Permitted language with substantiation attached, as the constraint on generation.

02

Structured shade and ingredient data

Consistent across ranges, with undertone treated as a first-class attribute.

03

Shade matching

With capture calibration, and accuracy reported by skin tone group as a primary output.

04

Claims-constrained generation

Assembling approved language rather than writing new claims.

05

Review integrity controls

Detection and disclosure design for seeding and gifting programmes.

06

Replenishment prediction

Measured against a holdout so the effect is incremental.

Fit check

Is this the right starting point?

Worth being direct. There are situations in beauty and cosmetics where custom AI work is the wrong spend, and those are listed rather than buried.

Worth doing if

  • Shade mismatch is a leading return reason and matching is unavailable or poor.
  • Content production is a bottleneck and claims review is the slow step.
  • Your shade data is inconsistent across ranges and undertone is unstructured.
  • Review integrity or seeding programme disclosure is an open exposure.
  • You sell consumables and have never modelled replenishment timing.

Do something else if

  • You want free-form generated marketing claims. That is a regulatory problem at volume.
  • Shade matching without the data or willingness to evaluate by skin tone group.
  • You want trend prediction where manufacturing lead times prevent any response.
  • Incentivised reviews without disclosure. That is squarely within the FTC rule.
Questions

Frequently asked questions

Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.

How accurate can shade matching from a photo be?

Good enough to be genuinely useful, and only if the capture is calibrated and the evaluation is honest. Lighting, white balance and camera processing shift measured skin tone more than many differences between people, so a reference object or calibration step in the capture flow improves accuracy more than a better model. And report accuracy by skin tone group from the first evaluation — an aggregate figure conceals the disparity that matters most here.

Can we generate product copy with a language model?

From an approved claims library with substantiation attached, yes, and it scales well. From a free-form brief, no — cosmetic claims are legally constrained, the line between a cosmetic and a medicinal claim is narrow, and a model has no sense of which verbs cross it. Constrain generation to approved language and let the model assemble and adapt rather than write.

Can AI predict which products will go viral?

Not reliably, and we would rather build the thing that works: fast detection of a spike already beginning, from search volume, UGC velocity and add-to-basket rate, feeding a supply and merchandising response. The harder question is usually whether your manufacturing lead times allow any response at all — if they do not, better detection changes nothing and we will say so.

What are the rules on our gifting and seeding programme?

The FTC's consumer reviews and testimonials rule covers incentivised reviews, and 2026 enforcement has moved beyond warning letters into broader deceptive conduct cases. The practical implication is that disclosure requirements are a design question for how the programme works — what creators are told, what is required of them, and what you do when they do not comply — rather than a line in the brief.

Is replenishment prediction worth building?

For consumables, usually yes, and it is one of the more reliable applications in beauty because exhaustion is genuinely predictable from purchase size and history. The discipline required is a holdout: prompting customers who were going to reorder anyway produces excellent attributed revenue and no incremental value, which is the standard way this gets over-credited.

Tell us what the problem looks like.

Thirty minutes, no charge, no deck. We will tell you whether this is an AI problem, a data problem, or a process problem — and we will say when the honest answer is to buy something rather than build it.