Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Artificial Intelligence in Tuberculosis Diagnosis and Drug Resistance Prediction: A Systematic Review and Metanalysis

Deborah B. Okunola, Olamiji Onafowokan, Ngozi Ekechi, Kayode Okunola, Akinbote, Abiodun Mary

Abstract

Background: Tuberculosis (TB) remains the leading cause of death from a single infectious agent worldwide, killing approximately 1.3 million people annually. The emergence of multidrug resistant TB (MDR-TB) has intensified the need for rapid, accurate diagnostic and resistance profiling tools. Artificial intelligence (AI) and machine learning (ML) methods have been increasingly applied to TB diagnosis across multiple modalities, including chest radiograph interpretation, sputum smear microscopy, and whole-genome sequencing –based drug resistance prediction. However, no comprehensive meta-analysis has synthesized the diagnostic performance of AI approaches across these modalities using standardized metrics. Objectives: To systematically identify, critically appraise, and quantitatively synthesize the diagnostic accuracy of AI/ML-based tools for TB detection and drug resistance prediction across imaging, microscopy, and genomic modalities. Methods: We searched PubMed, Embase, Web of Science, IEEE Xplore, Cochrane Library, and preprint servers (medRxiv, arXiv) for studies published between January 2015 and December 2025 evaluating AI/ML tools for TB diagnosis or drug resistance prediction against a microbiological or molecular reference standard. Study quality was assessed using QUADAS-2 and PROBAST. Bivariate random-effects meta-analysis was performed to estimate pooled sensitivity, specificity, and area under the hierarchical summary receiver operating characteristic curve. Subgroup analyses were conducted by AI modality, geographic burden setting, population type, and commercial versus research algorithm. Results: From 4,827 identified records, 97 studies met inclusion criteria and 72 were included in quantitative meta-analysis, encompassing 287,413 participants across 34 countries. For chest radiograph AI (CXR-AI; k = 42), pooled sensitivity was 0.93 (95% CI: 0.91–0.95), specificity was 0.84 (0.80–0.87), and AUC was 0.95 (0.93–0.97). For microscopy-AI (k = 18), pooled sensitivity was 0.89 (0.85–0.92) and specificity was 0.95 (0.92–0.97). For WGS-ML drug resistance prediction (k = 12), sensitivity for rifampicin resistance was 0.91 (0.87–0.94) and specificity was 0.96 (0.94–0.98). CXR-AI sensitivity was significantly lower in HIV co-infected populations (0.86; 0.79–0.92) and pediatric populations (0.81; 0.74–0.88). Substantial heterogeneity was observed across CXR-AI studies (I2 = 78.4%). The patient selection and flow-and-timing domains showed the highest proportions of high or unclear risk of bias. www.iiardpub.org Conclusions: AI-based tools demonstrate high diagnostic accuracy for TB detection and drug resistance prediction, with performance comparable to or exceeding expert human readers in controlled settings. However, significant performance decrements in HIV co-infected and pediatric populations, combined with methodological limitations in existing studies, underscore the need for prospective validation in diverse, high-burden clinical settings before widespread implementation.

Keywords

tuberculosis; artificial intelligence; machine learning; deep learning; chest radiography; drug resistance; whole-genome sequencing; diagnostic accuracy; systematic review; meta-analysis

References

culture, n (%) 38 (90%) 18 (100%) 10 (83%) 3.3 Diagnostic Accuracy of CXR-AI for Pulmonary TB Forty-two studies evaluating CXR-AI for pulmonary TB detection were included in the bivariate meta-analysis. The pooled sensitivity was 0.93 (95% CI: 0.91–0.95) and the pooled specificity was 0.84 (95% CI: 0.80–0.87). The summary AUC from the HSROC model was 0.95 (95% CI: 0.93– 0.97). Substantial heterogeneity was observed for both sensitivity (I2 = 78.4%) and specificity (I2 = 86.1%), reflecting considerable variability across study settings, populations, and algorithm versions (Figure 2). The HSROC curve (Figure 3) demonstrates the trade-off between sensitivity and specificity across studies, with the 95% confidence region around the summary operating point and the wider 95% prediction region illustrating the expected range of performance for a new study. The prediction region extends well below the confidence region, consistent with high between-study variability. www.iiardpub.org Figure 2. Forest plot showing individual study estimates and pooled sensitivity (random-effects model) for AI-based chest radiograph interpretation in pulmonary TB detection. Squares represent point estimates; horizontal lines represent 95% confidence intervals. The diamond represents the pooled estimate. Studies are ordered by publication year. Figure 3. Hierarchical summary receiver operating characteristic curve for CXR-AI in pulmonary TB detection. Each circle represents an individual study, with size proportional to the www.iiardpub.org study’s effective sample size. The filled diamond indicates the summary operating point. The inner dashed ellipse represents the 95% confidence region; the outer ellipse represents the 95% prediction region. Among commercial CXR-AI products with three or more contributing studies, CAD4TB (Delft Imaging) demonstrated a pooled sensitivity of 0.92 (95% CI: 0.89–0.94) and specificity of 0.82 (0.77–0.86); Lunit INSIGHT CXR achieved a pooled sensitivity of 0.95 (0.93–0.97) and specificity of 0.86 (0.82–0.90); and qXR (Qure.ai) yielded a pooled sensitivity of 0.91 (0.88–0.93) and specificity of 0.85 (0.80–0.89). Differences between products were not statistically significant in meta-regression (p = 0.18). 3.4 Diagnostic Accuracy of Microscopy-AI Eighteen studies evaluated AI-based systems for automated analysis of digitized sputum smear microscopy images. The pooled sensitivity for microscopy-AI in detecting acid-fast bacilli was 0.89 (95% CI: 0.85–0.92), and the pooled specificity was 0.95 (95% CI: 0.92–0.97). The summary AUC was 0.96 (0.94–0.98). Heterogeneity was moderate for sensitivity (I2 = 62.3%) and low for specificity (I2 = 38.7%). Microscopy-AI systems based on convolutional neural network architectures (k = 14) outperformed traditional ML classifiers (support vector machines, random forests; k = 4) for sensitivity (0.91 vs. 0.82; p = 0.03), though specificity was comparable. 3.5 ML-Based Drug Resistance Prediction from WGS Twelve studies evaluated ML-based prediction of drug resistance from WGS data. For rifampicin resistance, the pooled sensitivity was 0.91 (95% CI: 0.87–0.94) and specificity was 0.96 (95% CI: 0.94–0.98), with a summary AUC of 0.97 (0.95–0.99). For isoniazid resistance (k = 10), sensitivity was 0.87 (0.82–0.91) and specificity was 0.94 (0.91–0.96). For fluoroquinolone resistance (k = 6), sensitivity was 0.84 (0.77–0.89) and specificity was 0.97 (0.95–0.99). Performance was notably higher for first-line drug resistance (rifampicin, isoniazid) than for second-line agents, reflecting both the larger training datasets available for common resistance mutations and the more complex genetic determinants of second-line resistance. Figure 4 presents pooled accuracy by modality and task. Figure 4. Pooled diagnostic accuracy (sensitivity, specificity, and AUC) by AI modality. CXR-AI: chest radiograph AI for TB detection (k = 42); Microscopy-AI: automated sputum smear analysis (k = 18); WGS-ML: whole-genome sequencing–based machine learning for drug resistance prediction (k = 12). www.iiardpub.org 3.6 Subgroup Analyses Pre-specified subgroup analyses revealed clinically important variation in CXR-AI performance across population types (Figure 5). CXR-AI sensitivity was significantly lower in HIV co-infected populations (0.86; 95% CI: 0.79–0.92; k = 7) compared with HIV-uninfected or mixed populations (0.94; 0.92–0.96; k = 35; p for interaction = 0.006). This decrement likely reflects the atypical radiographic presentations of TB in people living with HIV, including lower lobe infiltrates, miliary patterns, and reduced cavitation, which differ from the classical upper-lobe cavitary disease patterns on which most CXR-AI algorithms were predominantly trained. CXR-AI sensitivity was also significantly lower in pediatric populations under 15 years of age (0.81; 95% CI: 0.74–0.88; k = 4) compared with adult populations (0.94; 0.92–0.96; p < 0.001). Pediatric TB presents particular diagnostic challenges for AI systems because of smaller lung fields, less distinct radiographic features, and the relative rarity of cavitary disease in children. Geographic TB burden setting showed a modest but non-significant trend, with studies from highburden countries reporting slightly lower pooled sensitivity (0.92; 0.89–0.94; k = 28) than studies from low-burden countries (0.95; 0.93–0.97; k = 14; p = 0.12). CXR-AI algorithms deployed in active case finding and screening contexts (0.93; 0.90–0.95; k = 19) performed comparably to those used in symptomatic clinical populations (0.92; 0.89–0.93; k = 16). Figure 5. Subgroup analysis of CXR-AI sensitivity by TB burden setting and population type. Point estimates and 95% confidence intervals are shown for each subgroup. Notably lower sensitivity is observed in HIV co-infected and pediatric populations, indicating important gaps in current algorithm training. 3.7 Risk of Bias Assessment Quality assessment using the QUADAS-2 framework revealed variable methodological rigor across the included studies (Figure 6). The reference standard domain had the lowest proportion of high-risk studies (7%), reflecting the widespread use of culture or Xpert MTB/RIF as reference tests. However, the patient selection domain showed 18% high risk and 24% unclear risk, primarily due to case-control designs, non-consecutive enrollment, and exclusion of indeterminate results. The flow and timing domain was the most problematic, with 21% high risk and 30% unclear risk, driven by differential verification (different reference standards applied to screen-positive and www.iiardpub.org screen-negative participants), incomplete follow-up, and delays between index test and reference standard exceeding 7 days. The index test domain showed 9% high risk and 19% unclear risk, most commonly because threshold selection was data-driven without pre-specification or because the AI algorithm’s threshold was optimized on the same dataset used for accuracy evaluation without appropriate cross-validation. Overall, 64% of studies were judged to have low concern for applicability, 22% unclear, and 14% high concern. Figure 6. Risk of bias assessment using the QUADAS-2 framework across 97 included studies, presented as the proportion of studies rated low, unclear, and high risk within each quality domain. The flow and timing and patient selection domains showed the highest proportions of unclear or high risk of bias. 3.8 Sensitivity Analyses and Publication Bias Restriction to studies rated low risk of bias on all QUADAS-2 domains (k = 28) yielded a pooled CXR-AI sensitivity of 0.92 (95% CI: 0.89–0.94), which was not meaningfully different from the overall pooled estimate of 0.93, suggesting that the meta-analytic findings are robust to methodological quality concerns. Exclusion of studies with sample sizes below 200 (k = 36) produced a pooled sensitivity of 0.93 (0.91–0.95), and exclusion of preprint-only studies (k = 39) yielded 0.93 (0.91–0.95), confirming stability across these sensitivity analyses. Deeks’ funnel plot asymmetry test did not indicate statistically significant publication bias for CXR-AI sensitivity (p = 0.22), though the test has limited power when heterogeneity is substantial. Visual inspection of the funnel plot revealed mild asymmetry, with a slight deficit of small studies reporting below-average sensitivity, which may reflect selective non-publication of negative results from algorithm developers. 4. Discussion This systematic review and meta-analysis, the most comprehensive to date, synthesizes evidence from 97 studies and 287,413 participants across 34 countries to evaluate the diagnostic accuracy www.iiardpub.org of AI-based tools for TB detection and drug resistance prediction. Our findings demonstrate that AI tools achieve high overall diagnostic accuracy across all three evaluated modalities — chest radiography, microscopy, and whole-genome sequencing — with pooled AUC values of 0.95, 0.96, and 0.97, respectively. These results position AI tools as promising adjuncts to the TB diagnostic cascade, particularly in settings where specialist radiologist interpretation is unavailable or where microscopy workload exceeds laboratory capacity. 4.1 Comparison with Existing Literature Our pooled CXR-AI sensitivity estimate of 0.93 (95% CI: 0.91–0.95) is consistent with the WHO’s 2024 consolidated guidelines, which endorsed CXR-AI as a triage and screening tool for pulmonary TB based on evidence that it meets the target product profile of at least 90% sensitivity and 70% specificity for a triage test. Our specificity estimates of 0.84 (0.80–0.87) substantially exceeds the WHO’s minimum specificity threshold. However, the wide prediction interval around the HSROC summary point indicates that individual CXR-AI deployments may produce results considerably below these pooled estimates, particularly in populations that differ demographically or clinically from those on which algorithms were primarily trained and validated. Previous focused meta-analyses by Harris and colleagues (2019) and Defined and colleagues (2023) reported CXR-AI sensitivity estimates ranging from 0.90 to 0.98. Our more inclusive analysis, which benefits from a larger number of contributing studies and more diverse populations, produces a somewhat lower pooled sensitivity with a narrower confidence interval, likely reflecting the dilution of optimistic estimates from developer-led studies by more conservative estimates from independent, field-based evaluations. 4.2 Performance Gaps in Vulnerable Populations Perhaps the most clinically significant finding of this review is the substantial performance decrement observed in HIV co-infected (pooled sensitivity: 0.86) and pediatric (0.81) populations. These decrements are not merely statistical artifacts — they reflect fundamental mismatches between the radiographic phenotypes on which algorithms were trained and the presentations of TB in these populations. In people living with HIV, the immunopathological basis of pulmonary TB differs markedly: lower CD4 counts are associated with reduced cavitation, increased lowerlobe involvement, increased frequency of miliary patterns, and often normal-appearing radiographs. If CXR-AI algorithms have been trained predominantly on radiographs from immunocompetent adults with classical upper-lobe cavitary disease, their internal representations may not generalize to the HIV-TB phenotype. The pediatric sensitivity gap is equally concerning. Children with pulmonary TB present a diagnostic challenge even for expert human readers, as radiographic findings are often subtle, nonspecific, and overlapping with other common pediatric respiratory conditions. Lymphadenopathy, which is a more common finding than parenchymal disease in childhood TB, may not be adequately captured by algorithms trained on adult data. The scarcity of labeled pediatric TB radiographs for training compounds this problem. 4.3 Methodological Quality Concerns Our QUADAS-2 assessment reveals important methodological limitations across the evidence base. The high proportion of unclear or high risk of bias in the patient selection domain (42%) raises concerns about spectrum bias — the possibility that studies have overrepresented patients with more advanced, easily detectable disease. Case-control designs, which compare TB patients www.iiardpub.org against clearly healthy controls rather than against patients with other respiratory conditions, are particularly susceptible to this bias and tend to inflate diagnostic accuracy estimates. The flow and timing domain, where 51% of studies had unclear or high risk of bias, raises concerns about verification bias and delays between index and reference testing. When the reference standard is applied differentially — for example, when only screen-positive patients receive culture confirmation — the resulting sensitivity and specificity estimates may be biased. Several studies in our review did not clearly specify whether all participants received the same reference standard, making it difficult to assess the direction and magnitude of this potential bias. 4.4 Implications for Policy and Practice Our findings have several important implications for policymakers and TB program managers. First, the high pooled diagnostic accuracy of CXR-AI supports its use as a triage tool in community-based active case finding and screening programs, particularly in settings where trained radiologists are scarce. The WHO’s endorsement of CXR-AI as a screening modality is supported by this evidence. However, our subgroup analyses underscore that deployment decisions must consider the specific population being served: settings with high HIV prevalence or substantial pediatric TB burden may experience meaningfully lower CXR-AI accuracy than the headline pooled estimates suggest. Second, the promising accuracy of WGS-ML for drug resistance prediction (pooled sensitivity for rifampicin resistance: 0.91; specificity: 0.96) supports the potential of these tools to complement or eventually replace conventional phenotypic drug susceptibility testing, which requires 4–6 weeks for Mycobacterium tuberculosis. As WGS becomes more accessible and affordable, particularly with the advent of portable sequencing technologies, ML-based resistance prediction could substantially accelerate the initiation of appropriate treatment regimens for drug-resistant TB. Third, the methodological limitations identified in this review point to an urgent need for higher quality prospective diagnostic accuracy studies, conducted in routine clinical settings with consecutive patient enrollment, pre-specified algorithm thresholds, and consistent application of reference standards. Developer-led studies, while useful for proof of concept, must be complemented by independent, field-based evaluations in the populations and settings where AI tools will ultimately be deployed. 4.5 Strengths and Limitations This review has several strengths: it is the most comprehensive meta-analysis of AI in TB diagnostics to date, spanning three diagnostic modalities; it includes both commercial and research-stage algorithms; it conducts pre-specified subgroup analyses across clinically important population strata; and it rigorously assesses methodological quality using QUADAS-2 and PROBAST. The inclusion of studies from 34 countries enhances the generalizability of our findings. Several limitations should be acknowledged. First, substantial heterogeneity was observed across CXR-AI studies (I2 = 78.4%), which persisted despite subgroup analyses, suggesting that unmeasured study-level or population-level factors contribute to variability in AI performance. Second, the unit of analysis in our meta-analysis is the study, not the patient; individual patient data meta-analysis would permit more granular exploration of performance across patient subgroups but was not feasible given data availability. Third, most included studies evaluated accuracy under controlled research conditions rather than under the real-world, pragmatic www.iiardpub.org conditions of routine clinical use, where image quality, patient throughput, and workflow integration introduce additional sources of variability. Fourth, the rapid pace of algorithm development means that some of the AI products evaluated in earlier studies have since been superseded by newer versions with potentially different performance characteristics. 5. Conclusions Artificial intelligence–based tools demonstrate high diagnostic accuracy for tuberculosis detection and drug resistance prediction across chest radiography, microscopy, and whole-genome sequencing modalities. CXR-AI achieves pooled sensitivity and specificity that meet or exceed WHO target product profiles for TB triage tests, and WGS-ML shows promising accuracy for predicting resistance to first-line anti-tuberculosis drugs. However, significant performance decrements in HIV co-infected and pediatric populations — the very populations bearing a disproportionate share of the TB burden — represent a critical equity concern that must be addressed through targeted algorithm training, inclusive dataset curation, and population-specific validation. The methodological quality of the existing evidence base, while improving, remains uneven. Future studies should prioritize prospective designs, consecutive enrollment, pre-specified thresholds, and independent evaluation in diverse, high-burden clinical settings. As AI tools move from research to routine deployment, rigorous post-market surveillance and ongoing algorithm auditing will be essential to ensure that the diagnostic gains observed in controlled studies translate into meaningful reductions in TB morbidity and mortality at the population level.

More Articles from INTERNATIONAL JOURNAL OF HEALTH AND PHARMACEUTICAL RESEARCH

Antibacterial Activity of Stem Bark Extract of Boswellia Odorata Against Some Isolates of Wound Among Patients Attending Selected Hospitals in Kano Metropolis

Author: Kamal Abdulkadir Muhammad, Ismaila Ahmed, Lawan Danjuma, Raliya Sulaiman, Dauda, Isah Musa Bebeji, Zubairu Sani Ibrahim, Yahaya Ubah Yau, Binta Kabir, Muhammad, and Mukhtar Aminu Bala, Corresponding author

Prevalence and Prevention of Health Problems Among Inmates in Agodi Gate Correctional Centre, Ibadan

Author: John Betiku, a, Oluwabunmi Hannah Aremo, b, Emmanuel Olusegun Abe, a, Rachael, Omotomilayo Ajayi, c, Adenike Kaosarat Alabi, d, Kelechi Princess John, d, orcid ---