Methodology

How Verdict knows what it knows.

A skin score is a measurement, and every measurement has error. Verdict estimates cosmetic score trends and makes study assumptions inspectable. It cannot diagnose a flare or establish clinical efficacy. This page explains the methods and their current limits.

1. We use raw scores, never UI scores

YouCam's Skin Analysis returns two numbers per concern. Its documentation describes ui_score as a score that “functions primarily as a psychological motivator”, adjusted “to produce more favorable results”. That is fine for a beauty app and wrong for evidence. Every Verdict statistic uses raw_score (1–100, higher is healthier). All scans in a study use SD skin analysis v2.1, because SD and HD scores can't be compared with each other.

2. We measured the instrument first

We analysed one face (AI-generated, so no real person is involved) several times with tiny capture changes: ±10% exposure, a 3° tilt, 5% re-framing and a colour shift. Each concern moved by a different amount. Breakouts were rock-steady. Pores moved up to 10 points from framing alone, and oiliness rose about 5 points when the photo was darker.

Capture noise by concern (SD, raw points)
fixtures/youcam/noise_study.json
Pores
±3.8
Oil control
±2.7
Radiance
±2.3
Hydration
±1.8
Dark circles
±1.2
Texture
±1.2
Firmness
±1.0
Dark spots
±0.9
Fine lines
±0.7
Redness
±0.5
Breakouts
±0.4
Puffiness
±0.1

Real life adds biology on top: sleep, cycle, time since washing. Each concern's noise is √(capture² + biological²), using conservative biological priors. It is then personalised from your own scans, using the median absolute deviation of scan-to-scan changes. This robust estimate ignores the occasional big jump from a flare or a bad-light day, and it is shrunk towards the prior until you have enough data.

ConcernCapture SD (measured)Biological SD (assumed prior)Prior total SDNoise floor (MDC95)Planning score threshold
Breakouts±0.40±2.5±2.5±7.05 pts
Pores±3.83±1.5±4.1±11.44 pts
Texture±1.20±1.5±1.9±5.34 pts
Redness±0.47±1.5±1.6±4.44 pts
Oil control±2.66±3.0±4.0±11.15 pts
Dark spots±0.92±1.0±1.4±3.83 pts
Hydration±1.75±3.0±3.5±9.65 pts
Fine lines±0.73±1.0±1.2±3.43 pts
Radiance±2.34±2.0±3.1±8.54 pts
Dark circles±1.25±1.5±2.0±5.43 pts
Puffiness±0.11 (floored to 0.5)±1.0±1.1±3.13 pts
Firmness±0.95±1.0±1.4±3.83 pts

For most concerns, day-to-day biology (the assumed prior) dominates the capture noise we measured. That's why each person's own scans replace the prior as soon as there are enough of them. Next step: a test–retest study on consenting adults across Fitzpatrick III–VI.

3. One person, one product: the verdict

  • Baseline: two scans about 48 hours apart around the start date.
  • Effect: mean of the latest two scans minus the baseline mean, with SE = σ·√(1/m + 1/k).
  • Working needs two things. The 95% interval must exclude zero, and there must be at least an 80% chance the change beats the assumed score threshold.
  • Too early applies before the onset window of the product's best-evidenced active, for example 8–12 weeks for retinoids and 2–4 weeks for niacinamide on oil. These are planning assumptions, not instructions to continue despite symptoms.
  • Not working is only called after the onset window has passed with no meaningful change.
  • Irritation overrides everything. Rising redness beyond your noise floor, or reported burning, itching, swelling or other concerning symptoms, pauses the product. Irritation can leave post-inflammatory hyperpigmentation on deeper skin tones.
  • Adaptive cadence: the next check-in is 3 days away during a possible reaction, 4 during a flagged flare, 14 once a result is clear, and weekly otherwise.
Onset windows we use (best-evidenced targets)
Tretinoinbreakouts 8–12w · texture 6–12w
Adapalenebreakouts 8–12w · texture 8–12w
Retinaldehydefine lines 8–16w · texture 6–12w
Retinolfine lines 8–16w · texture 6–12w
Bakuchiolfine lines 8–12w
Salicylic acid (BHA)breakouts 4–8w · pores 4–8w
Glycolic acid (AHA)texture 2–6w · radiance 1–4w
Lactic acid (AHA)texture 2–6w · hydration 2–4w
Benzoyl peroxidebreakouts 2–6w
Azelaic acidbreakouts 4–8w · redness 4–12w
Sulfurbreakouts 4–8w
Niacinamideoil control 2–4w · pores 4–8w
Vitamin C (L-ascorbic acid)radiance 2–4w · dark spots 8–12w
Tranexamic aciddark spots 8–12w

4. Flare patterns, not a diagnosis

Verdict maps detected spots from YouCam's acne mask into face-relative zones using T-zone and U-zone masks. It compares new and previous spots to describe where a pattern changed.

The current rules consider zone overlap, timing and the ingredient list. These are unvalidated heuristics. Neither a photograph nor zone overlap can establish whether worsening is purging, irritation, an allergy or another condition. A flare prompts a pause and review; symptom reports take priority even before a baseline is complete.

5. Many people, one claim: the study

  • Design: randomised 1:1, product vs usual routine. The AI assessor never knows the arm.
  • Sample size: n per arm = 2(z₀.₉₇₅ + z₀.₈₀)²σ²/δ², where σ is the SD of a change score. It combines our measured noise with person-to-person response variability (5 points). Inflated for 15% dropout, and raised to at least 20 treated per skin-tone group when the claim names a population.
  • Analysis: change from baseline per arm, Welch's t-test, 95% CIs. Responders are people improving by at least the assumed score threshold, with Wilson intervals. Time to response is a Kaplan–Meier curve.
  • Grading: each part of a claim is graded separately: effect, timeframe, magnitude and skin tones. “Visibly” requires both statistical significance and a between-arm difference of at least half the assumed score threshold. Interim readouts never grade a claim “not supported”.
  • Skin tone: Fitzpatrick type is measured with YouCam's Fitzpatrick analyzer at sign-up, not self-reported. It drives recruitment quotas and per-group results.

6. Visuals measured against the evidence

We first tried to calibrate YouCam Skin Simulation intensity against YouCam Skin Analysis by rendering one face at several intensities and measuring each render. Two findings came out of it. First, repeated identical inputs in our recorded tests gave identical outputs. This does not guarantee the behavior of future model versions. Second, the acne slider is a step in measured score: even 0.05 removes most detected spots.

Simulation intensity00.050.10.150.20.3
Measured acne Δ+0.0+14.7+12.6+12.3+12.6+15.8

Blending the original with a full simulation is flat up to a weight of about 0.75, then steep. So a fixed calibration curve would misrepresent small effects. Instead, Verdict measures its way to the answer. It bisects the blend weight and checks each candidate with Skin Analysis until the measured change matches the study's effect within ±1.5 points. The measurement log is stored next to the image, so readers can inspect the measured approximation. A generated visual is not evidence that a person will achieve that outcome.

Blend weight00.250.50.750.850.90.951
Measured acne Δ+0.0-0.6-1.0+1.9+14.2+12.3+14.4+15.8

7. Limits we're honest about

  • Verdict measures cosmetic appearance with an AI model. It is not a diagnosis, and it does not replace a dermatologist.
  • Real-world, self-administered studies are not dermatologist-supervised clinical trials. Model-score grades are exploratory and require independent scientific, ethics and legal review before supporting commercial claims.
  • Our capture-noise study used one AI-generated face and small perturbations. Biological variability priors and score thresholds are assumptions, not human-validated measures. They are refined from personal data but do not establish clinical validity.
  • The demo's cohort panelists are simulated from this same statistical model. The demo panelist's scans are real YouCam outputs on an AI-generated face.
  • Subgroup and weekly results are exploratory. Repeated comparisons, missing data and small groups can inflate apparent findings. The planned primary endpoint is used for claim wording; independent statistical review is still required.
  • Fitzpatrick describes sun response and is an imperfect proxy for skin tone. An AI estimate does not establish fair performance across populations.
  • The claim linter is a decision aid, not legal advice.

Sources

  • YouCam API documentation: AI Skin Analysis (raw_score vs ui_score), Skin Simulation, Fitzpatrick Scale Analyzer, JS Camera Kit.
  • American Academy of Dermatology: acne treatment timelines; introducing one new product at a time; patch testing.
  • Navarrete-Solís et al., 2011 (4% niacinamide, visible change at 4–8 weeks); systematic reviews of topical vitamin C (2023) and azelaic acid (2023).
  • ASCI Annual Complaints Report 2025–26; Cosmetics Rules 2020 (Drugs & Cosmetics Act); ASCI guidelines on skin-lightening advertising.
  • India's Digital Personal Data Protection Act 2023 and Rules 2025.