A new report highlights that while over 90% of health systems have integrated third-party AI tools into their operations, a significant challenge remains: less than half possess the essential infrastructure to rigorously test and validate these solutions before they are fully embedded into critical patient care workflows, leading to potential risks and inefficiencies.
A recent report, jointly published by UPMC’s Center for Connected Medicine and KLAS Research, reveals a growing paradox in healthcare: an overwhelming majority, more than 90%, of health systems have enthusiastically adopted and deployed third-party artificial intelligence tools across various clinical and administrative functions. However, this rapid embrace of AI technology has far outstripped the development of adequate governance and infrastructural capabilities. Crucially, the research indicates that less than 50% of these hospitals have a dedicated and structured environment specifically designed for thoroughly testing and validating these AI solutions. This gap means many AI tools are being implemented without a comprehensive pre-deployment evaluation, posing potential risks when they are integrated into sensitive patient care processes. The validation methods currently in use vary considerably, ranging from formal testing conducted by vendors to informal pilot programs, underscoring a lack of standardized, rigorous evaluation.
The significant shortfall in dedicated AI testing environments within health systems is primarily attributed to a combination of persistent time, capital, and talent constraints, as explained by Ken Howard, vice president of technology services engineering at UPMC Enterprises. When hospitals identify a potential area where AI could offer a solution, they often face immense pressure to deploy quickly. This urgency frequently prevents them from allocating the necessary time, securing sufficient funding, or acquiring the specialized talent required to first establish a structured and robust testing environment. Consequently, organizations tend to default to existing, often less rigorous, standard IT implementation processes. This approach, Howard warns, can lead to prolonged and costly implementation timelines, often stretching to six months or more, only for health systems to later discover that the AI solution fails to deliver the expected value or perform as anticipated, highlighting the inefficiency of foregoing proper validation.
UPMC is actively addressing these industry-wide challenges by implementing its own comprehensive AI governance model and leveraging its proprietary Ahavi platform. Ahavi serves as a real-world data platform, providing a secure environment where third-party AI tools can be rigorously validated against de-identified patient data before their formal deployment into clinical settings. Rob Bart, UPMC’s chief medical information officer, emphasizes that effective governance extends far beyond initial deployment. UPMC has maintained a formal AI governance structure for over two years, which includes continuous, post-implementation monitoring of AI tools. This ongoing oversight is particularly vital for clinical algorithms, such as those that predict hospital length of stay or readmission risk. Bart highlights that regular interval monitoring ensures the AI’s guidance remains accurate and consistent with its original validation, and custom testing against UPMC’s specific patient population helps identify and mitigate issues like algorithmic bias and model drift that might be missed by generic datasets.
The rapid acceleration of AI adoption in healthcare underscores an urgent need for the development and implementation of standardized evaluation methods. Kate Eisenberg, senior medical director of DynaMed, an AI-powered clinical decision support tool, cites an American Medical Association survey indicating that clinicians' use of AI tools nearly doubled between 2023 and 2026. This swift pace of technological integration makes it increasingly difficult for current evaluation standards to keep pace. Eisenberg emphasizes that as providers intensify their scrutiny of AI, a critical focus must be placed on equity, ensuring that AI responses are free from bias. She advocates for built-in evaluation mechanisms and specialized training for clinical teams to specifically assess AI outputs for potential biases. Until a consistent, industry-wide approach for evaluating these sophisticated tools emerges, health systems and vendors largely bear the responsibility of self-governing AI, highlighting a significant regulatory and ethical gap.