KI-MED – Testing Standards and Evaluation Tools for Artificial Intelligence in Medical Diagnostic and Prognostic Systems
You are here:
About the project
Artificial intelligence (AI) is increasingly being used in medical diagnostics and prognostics, opening up new possibilities for precision, efficiency, and early intervention. At the same time, this is a high-risk domain in which incorrect decisions can have serious consequences. Despite rapid technological advances, there is still a lack of standardized, practice-oriented testing standards and tools to systematically and reproducibly assess the safety, robustness, explainability, and performance of medical AI systems.
The KI-MED project addresses this gap. Its goal is to develop a comprehensive evaluation tool that assesses AI models and the underlying datasets throughout their entire lifecycle—from data selection and training to deployment in a clinical context. In addition to performance, particular emphasis is placed on AI-specific safety aspects, robustness against disturbances and attacks, as well as explainability, interpretability, and uncertainty quantification of the models.
The project combines methodological research with application-oriented evaluation based on a medical use case. Building on systematic analyses, quality characteristics, metrics, testing requirements, and evaluation methods are developed and integrated into a modular and, as far as possible, generalizable testing tool. The results will be prepared in a scientific manner and will serve as a foundation for future certification and standardization processes. The client is the Bundesamt für Sicherheit in der Informationstechnik (BSI), which will contribute the results to national and international standardization bodies.
The project thus makes a key contribution to the trustworthiness of medical AI and establishes the foundation for the safe, transparent, and responsible use of AI-supported diagnostic and prognostic systems in healthcare.

