Clinically evaluating a piece of medical software with methods designed for a physical device does not work. The clinical data relevant to a SaMD are of a different nature, the sources are specific, and the criteria for assessing data quality differ from those of a clinical study on an implant or a therapeutic device.
What makes the clinical evaluation of a SaMD different
A physical device produces a measurable effect on the body: it compresses, delivers a drug, transmits energy. Its performance can be assessed against physical criteria (durability, dosing accuracy, mechanical strength) in addition to clinical criteria.
A SaMD produces information: a suggested diagnosis, a treatment recommendation, a risk indicator. Its clinical evaluation focuses on the quality of that information — its diagnostic accuracy, its impact on clinical decisions, and its effects on patient outcomes.
Acceptable clinical data sources for SaMD
Diagnostic performance studies. For diagnostic-aid software, studies assess the sensitivity, specificity, and positive and negative predictive value of the algorithm on representative cohorts, with comparison to a reference standard (clinical gold standard or human expert). These studies are often retrospective, based on existing databases of images or signals.
Prospective clinical studies. Studies assess the impact of the software on clinical decisions and patient outcomes under real-world conditions of use. More demanding to set up, they provide data with a higher level of evidence.
Real-world data. Data arising from the actual use of the software in healthcare settings (registries, hospital databases, telemedicine data) can constitute valid post-market clinical data, provided that their collection is systematic and their analysis rigorous.
Literature data on comparable algorithms. A literature review of SaMD with similar functions in closely related indications remains a valid source, subject to a demonstration of equivalence or a justification of comparability.
What MDCG 2020-1 specifies
The MDCG 2020-1 guidance on clinical data for SaMD specifies the criteria for assessing data quality: representativeness of the test population, adequacy of the reference standard, absence of selection bias, level of evidence of the study. It points out that algorithmic performance studies alone (performance on test data) are insufficient if they are not complemented by data on the real clinical impact of the software.