Comparative Evaluation of Zero-Shot Vision-Language Models and YOLO for Image-Based Weld Defect Severity Assessment
Main Article Content
Abstract
Automated visual inspection of welded joints increasingly uses supervised deep-learning models, particularly YOLO object detectors. Although effective for defect localization, these approaches require domain-specific annotated data and provide limited engineering-oriented textual explanations. This study compared supervised YOLO-based pipelines with zero-shot Vision-Language Models (VLMs) for image-based weld-defect severity assessment. A public Welding Defect–Object Detection dataset containing 1,983 images was used. A four-level ordinal severity rubric was developed as an AWS D1.1-informed engineering heuristic rather than a direct implementation of AWS acceptance criteria because the source images lacked consistent physical scale calibration. Two supervised detectors, YOLOv8n and YOLO26n, and two zero-shot VLMs, NVIDIA Nemotron and Meta Llama 3.2 Vision-90B, were evaluated. All pipelines processed 401 validation-and-test images, while primary comparison against independently rated human references used an adjudicated subset of 100 images. Human inter-rater agreement was high (Cohen’s κ = 0.839; weighted κ = 0.929). Nemotron achieved the highest accuracy (55.0%), followed by YOLOv8n (50.0%), YOLO26n (48.0%), and Llama Vision-90B (38.0%), and the highest Level-4 recall (0.871). Both YOLO pipelines showed zero recall for Level 2. Although unadjusted McNemar testing found p = 0.030 for Nemotron versus Llama Vision-90B, no pairwise comparison remained significant after Holm correction. VLM explanation quality was moderately associated with classification correctness (r = 0.533; p = 0.0001). None of the methods was sufficiently reliable for standalone engineering acceptance decisions, but their complementary failure patterns support VLMs as an auxiliary review layer within human-supervised weld inspection.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Accepted 2026-08-16
Published 2026-08-16
Plaudit
References
American Welding Society, AWS D1.1/D1.1M:2025, Structural Welding Code—Steel, 25th ed. Miami, FL, USA: American Welding Society, 2025.
T. H. Ngo, H. M. Q. Tang, and Q. B. Diep, “Weld-CNN: Advancing non-destructive testing with a hybrid deep learning model for weld defect detection,” Advances in Mechanical Engineering, 2025, doi: 10.1177/16878132251341615. DOI: https://doi.org/10.1177/16878132251341615
M. Torres-Torres, K. Velasquez, J. Escorcia-Gutierrez, and A. Valls, “Visual weld quality inspection system for identifying defects in GMAW joints with semantic segmentation on pre-trained models,” International Journal of Computational Intelligence Systems, vol. 19, Art. no. 110, 2026, doi: 10.1007/s44196-026-01197-z. DOI: https://doi.org/10.1007/s44196-026-01197-z
R. Sapkota et al., “YOLO advances to its genesis: A decadal and comprehensive review of the You Only Look Once (YOLO) series,” Artificial Intelligence Review, vol. 58, no. 9, Art. no. 274, 2025, doi: 10.1007/s10462-025-11253-3. DOI: https://doi.org/10.1007/s10462-025-11253-3
Z. Li, Y. Yan, X. Wang, Y. Ge, and L. Meng, “A survey of deep learning for industrial visual anomaly detection,” Artificial Intelligence Review, vol. 58, Art. no. 279, 2025, doi: 10.1007/s10462-025-11287-7. DOI: https://doi.org/10.1007/s10462-025-11287-7
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” International Journal of Computer Vision, vol. 128, pp. 336–359, 2020, doi: 10.1007/s11263-019-01228-7. DOI: https://doi.org/10.1007/s11263-019-01228-7
J. Cação, J. Santos, and M. Antunes, “Explainable AI for industrial fault diagnosis: A systematic review,” Journal of Industrial Information Integration, vol. 47, Art. no. 100905, 2025, doi: 10.1016/j.jii.2025.100905. DOI: https://doi.org/10.1016/j.jii.2025.100905
M. A. Mersha, K. N. Lam, J. Wood, A. K. AlShami, and J. K. Kalita, “Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction,” Neurocomputing, vol. 599, Art. no. 128111, 2024, doi: 10.1016/j.neucom.2024.128111. DOI: https://doi.org/10.1016/j.neucom.2024.128111
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4765–4774.
M. T. Ribeiro, S. Singh, and C. Guestrin, “‘Why should I trust you?’: Explaining the predictions of any classifier,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016, pp. 1135–1144, doi: 10.1145/2939672.2939778. DOI: https://doi.org/10.1145/2939672.2939778
A. Radford et al., “Learning transferable visual models from natural language supervision,” in Proc. 38th Int. Conf. Machine Learning (ICML), vol. 139, 2021, pp. 8748–8763.
W. Wang, V. W. Zheng, H. Yu, and C. Miao, “A survey of zero-shot learning: Settings, methods, and applications,” ACM Transactions on Intelligent Systems and Technology, vol. 10, no. 2, Art. no. 13, 2019, doi: 10.1145/3293318. DOI: https://doi.org/10.1145/3293318
S. Danish, A. Sadeghi-Niaraki, S. U. Khan, L. M. Dang, L. Tightiz, and H. Moon, “A comprehensive survey of vision–language models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets,” Information Fusion, vol. 126, Art. no. 103623, 2025, doi: 10.1016/j.inffus.2025.103623. DOI: https://doi.org/10.1016/j.inffus.2025.103623
C. X. Liang et al., “A comprehensive survey and guide to multimodal large language models in vision–language tasks,” Computation, vol. 14, no. 6, Art. no. 125, 2026, doi: 10.3390/computation14060125. DOI: https://doi.org/10.3390/computation14060125
G. Khvatskii, Y. S. Lee, C. Angst, M. Gibbs, R. Landers, and N. V. Chawla, “Do multimodal large language models understand welding?,” Information Fusion, vol. 120, Art. no. 103121, 2025, doi: 10.1016/j.inffus.2025.103121. DOI: https://doi.org/10.1016/j.inffus.2025.103121
G. Yong, K. Jeon, D. Gil, and G. Lee, “Prompt engineering for zero-shot and few-shot defect detection and classification using a visual-language pretrained model,” Computer-Aided Civil and Infrastructure Engineering, vol. 38, pp. 1536–1554, 2023, doi: 10.1111/mice.12954. DOI: https://doi.org/10.1111/mice.12954
C. Sub-r-pa and R.-C. Chen, “Adapting vision–language models for few-shot industrial defect detection,” Algorithms, vol. 19, no. 4, Art. no. 259, 2026, doi: 10.3390/a19040259. DOI: https://doi.org/10.3390/a19040259
J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960, doi: 10.1177/001316446002000104. DOI: https://doi.org/10.1177/001316446002000104
J. Cohen, “Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit,” Psychological Bulletin, vol. 70, no. 4, pp. 213–220, 1968, doi: 10.1037/h0026256. DOI: https://doi.org/10.1037/h0026256
Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153–157, 1947, doi: 10.1007/BF02295996. DOI: https://doi.org/10.1007/BF02295996
Sukmaadhiwijaya, “Welding Defect – Object Detection,” Kaggle, dataset. [Online]. Available: https://www.kaggle.com/datasets/sukmaadhiwijaya/welding-defect-object-detection. [Accessed: Aug. 9, 2026].