PelliScope Back to PelliScope

Clinical validation · evidence record

What the current evidence does—and does not—show.

PelliScope’s published research evaluates a case-level dermatology model on deidentified, patient-submitted photographs from the public SCIN dataset. The study is an internal retrospective validation. It is not an external clinical validation and does not establish autonomous diagnostic use.

Research status: The current manuscript is a v1 medRxiv preprint, posted 1 June 2026, and has not been peer reviewed.
2,336multi-photo dermatology cases
5,041image records across those cases
0.849locked-test micro ROC-AUC
0.800locked-test macro ROC-AUC

Dataset and procedure

Case-wise evaluation, with the final test split held back.

Dataset source

The study used the public Skin Condition Image Network (SCIN), a deidentified collection of self-submitted images and symptom information from adults in the United States. The evaluated subset covered ten common inflammatory, allergic and infectious dermatology conditions.

Split and validation procedure

Cases—not individual photographs—were divided 60%/20%/20% into training, validation and locked test sets: 1,401 / 467 / 468 cases. Repeated stratified experiments were followed by one final evaluation on the held-out test set. No external hospital or prospective clinical cohort was used.

Inclusion criteria

  • SCIN submissions considered diagnosable in the source data.
  • One to three photographs associated with a single patient case.
  • A dermatologist differential containing at least one of the ten target conditions.
  • Images and metadata passing the source dataset’s deidentification and quality process.

Exclusion criteria and documentation gap

Cases outside the ten-condition target vocabulary and source entries not suitable for diagnostic review were not part of the evaluated subset. The v1 manuscript does not publish a complete row-by-row exclusion flow; this remains a documentation gap to close in a future version.

Result interpretation

Numbers need context.

What “AUC 0.86” means

The rounded 0.86 figure comes from a repeated-experiment micro ROC-AUC of 0.863. ROC-AUC measures how well the model ranks positive and negative examples across decision thresholds. It does not mean that the model made the correct diagnosis in 86% of patient cases. The final locked held-out test produced a micro ROC-AUC of 0.849 and macro ROC-AUC of 0.800.

What “80.2%” means

80.2% appeared in earlier PelliScope material as an internal balanced-accuracy summary. It is not a headline result in the current v1 preprint. PelliScope therefore does not use it here as a public clinical performance claim until the exact evaluation artifact and thresholding procedure are versioned and published.

Confidence intervals

The v1 manuscript does not report confidence intervals for the headline locked-test AUC values. We state that absence rather than implying precision the current report does not establish.

Subgroups and robustness

Important analyses are still outstanding.

AreaWhat is availableCurrent conclusion
Skin tone SCIN contains estimated Fitzpatrick and Monk Skin Tone annotations. The current study did not establish definitive subgroup performance because several strata were sparse. No parity claim is made.
Device type Patient-submitted images reflect varied real-world capture conditions. No completed device-specific performance table is reported in v1. Device-level robustness remains to be evaluated.
Image quality The source data underwent quality and safety screening. No stratified performance table by blur, lighting, resolution or framing is reported in v1.
Geography The evaluated SCIN cohort consists of United States volunteers. Results should not be assumed to transfer to Indian or other clinical populations without external validation.

Known limitations

What must happen before broader clinical claims.

Evidence limitations

  • Retrospective internal validation only.
  • No prospective workflow or patient-outcome study.
  • No external Indian, European or hospital validation cohort.
  • Ten-condition subset rather than open-world dermatology.
  • Source labels are dermatologist differentials, not encounter-confirmed final diagnoses.
  • Calibration and subgroup reliability are not fully established.

Intended interpretation

The current evidence supports continued research into case organisation, review support and triage-oriented workflows. It does not support replacing a dermatologist, diagnosing from an advertisement, or treating an AI screening output as a medical decision.

PelliScope’s public patient pathway is therefore described as secure preparation for clinician review.

Version record

Model and evaluation date

This page describes the DermAssist/PelliScope case-level gated-attention multiple-instance learning model reported in preprint v1, posted 1 June 2026. The page was reviewed on 26 July 2026. No production model version is represented as externally clinically validated.