← The Oyster Index Guides

How to Evaluate Skin Analysis Accuracy

8 min readGlobal

Why one accuracy number lies

Every vendor quotes an accuracy figure. Most quote a single blended number across all skin and all concerns. That number can look excellent while hiding a real gap on deeper tones or on a specific concern. To know if an engine works for your customers, you have to break accuracy apart and test it yourself.

This guide gives you a repeatable protocol you can run in a few weeks.

Decide what accurate means for you

Accuracy is not one thing. Define it for your use.

  • Concern detection. Does it flag the concerns a professional would, like acne, pigmentation, redness and texture?
  • Severity. Does the score track how mild or serious the concern is?
  • Consistency. Does the same face scanned twice return a similar result?
  • Tone fairness. Does accuracy hold across Fitzpatrick types one through six?

Write down which of these matter most, then test each on its own.

Build a fair test set

Your test set decides the answer, so build it with care.

  • Recruit twenty to fifty people who reflect your real customer skin tone mix, weighted toward the tones you serve most.
  • Capture in realistic conditions, including the lighting your customers actually have.
  • Have a trained aesthetician or dermatologist label each face independently as ground truth.
  • Keep the labeller blind to the tool's output.

A test set of only light skin will tell you nothing about deeper tones, and the reverse is also true.

Measure the right things

Run each face through the tool and compare to the expert labels.

  • Agreement by concern. How often does the tool match the expert per concern?
  • Agreement by Fitzpatrick band. Report the same figure for each skin tone group, not blended.
  • Repeatability. Scan a subset twice and check how stable the scores are.
  • Failure modes. Note where it struggles, for example uneven tone on deep melanin or redness on light skin.

The gap between the blended number and the deep tone number is the most important line in your report.

How Oyster approaches accuracy

Oyster builds and measures its engine across the full Fitzpatrick range and weights toward deeper skin, an emphasis it calls the melanin moat. Because dermatology research documents how darker tones have been underrepresented in skin datasets, Oyster treats deep tone accuracy as a first order metric, not a footnote. Ask to see accuracy reported by skin tone, then verify it on your own test set. See how the scan works.

Turn the result into a decision

Once you have per concern and per tone figures for two vendors, the choice is easier. Favour the engine that holds accuracy across the tones you actually serve, stays consistent on repeat scans, and fails gracefully rather than confidently. Then confirm the post scan action, whether matching, CRM or checkout, delivers on that accuracy. To run this test with Oyster, book a demo.

Frequently asked

Define what accurate means for you across concern detection, severity, consistency and tone fairness. Build a test set of twenty to fifty people reflecting your real customer skin tone mix, labelled independently by a trained professional. Then measure agreement per concern and per Fitzpatrick band, check repeatability, and note failure modes. Compare the blended number to the deep tone number.

A single blended figure averages across all skin tones and concerns, which can hide a real gap on deeper tones or a specific concern. An engine can report high overall accuracy while performing worse on Fitzpatrick types four, five and six. Always ask for accuracy broken out by skin tone and by concern.

Twenty to fifty people is usually enough for a practical pilot, provided they reflect your real customer skin tone mix and are weighted toward the tones you serve most. Have a trained aesthetician or dermatologist label each face independently and keep them blind to the tool's output.

Oyster builds and measures its engine across the full Fitzpatrick range and weights toward deeper skin tones, an approach it calls the melanin moat. It treats deep tone accuracy as a first order metric rather than a footnote, and buyers can request accuracy reported by skin tone and verify it on their own test set.

See what skin intelligence does for your business.

Oyster reads skin accurately on every tone and turns it into the right recommendation.