Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

A Survey of Object Detection and Classification in Video Using Deep Learning Approaches

BELLO, Umar, Prof E J Garba, Mustapha Kassim

Abstract

Object detection and classification in videos has emerged as a core challenge in computer vision, with applications spanning autonomous driving, surveillance, healthcare, robotics, and human– computer interaction. Unlike still-image object detection, video-based detection requires handling additional complexities such as temporal coherence, motion blur, occlusion, and real-time constraints. Traditional approaches relying on handcrafted features and background subtraction have given way to deep learning-based methods, particularly convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformers, and one-stage detectors such as the You Only Look Once (YOLO) family. This paper provides a comprehensive review of object detection in videos, surveying the evolution of methods from early handcrafted approaches to recent deep architectures. We examine popular models, datasets, and benchmarks, with an emphasis on accuracy–speed trade-offs and the unique challenges of real-world deployment. A comparative analysis of state-of-the-art methods highlights trends and gaps, including the struggle with small object detection, domain adaptation, and efficient training on large-scale datasets. Finally, we outline promising future directions such as transformer-based detection, self-supervised learning, multimodal video understanding, and edge-device optimization. This survey aims to serve as both an introductory resource and a roadmap for advancing video-based object detection research.

Keywords

Object detection; Video classification; Deep learning; YOLO; Transformers

References

Bernardin, K., & Stiefelhagen, R. (2008). Evaluating multiple object tracking performance: The CLEAR MOT metrics. EURASIP Journal on Image and Video Processing, 2008(1), 1–10. https://doi.org/10.1155/2008/246309 Bochkovskiy, A., Wang, C. Y., & Liao, H. Y. M. (2020). YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 1–15. Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to- end object detection with transformers. European Conference on Computer Vision (ECCV), 213–229. Springer. https://doi.org/10.1007/978-3-030-58452-8_13 Chen, L. C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2019). Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv preprint arXiv:1412.7062. Chen, X., Li, S., & Luo, J. (2020). Video object detection via collaborative learning of motion cues. IEEE Transactions on Image Processing, 29, 10287–10299. https://doi.org/10.1109/TIP.2020.3027661 Cioppa, A., Giancola, S., & Ghanem, B. (2019). A context-aware loss function for action spotting in soccer videos. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13126–13136. https://doi.org/10.1109/CVPR.2019.01343 Dai, J., Li, Y., He, K., & Sun, J. (2017). R-FCN: Object detection via region-based fully convolutional networks. Advances in Neural Information Processing Systems, 29, 379– Dalal, N., & Triggs, B. (2005). Histograms of oriented gradients for human detection. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 886– IEEE. https://doi.org/10.1109/CVPR.2005.177 Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., & Zisserman, A. (2010). The Pascal Visual Object Classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303–338. https://doi.org/10.1007/s11263-009-0275-4 Feichtenhofer, C., Fan, H., Li, Y., & He, K. (2022). Masked autoencoders as spatiotemporal learners. Advances in Neural Information Processing Systems, 35, 35946–35958. Geiger, A., Lenz, P., & Urtasun, R. (2013). Are we ready for autonomous driving? The KITTI vision benchmark suite. 2012 IEEE Conference on Computer Vision and Pattern Recognition, 3354–3361. https://doi.org/10.1109/CVPR.2012.6248074 Girshick, R. (2015). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 1440–1448. https://doi.org/10.1109/ICCV.2015.169 Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 580–587. https://doi.org/10.1109/CVPR.2014.81 Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press. He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2961–2969. https://doi.org/10.1109/ICCV.2017.322 Horn, B. K. P., & Schunck, B. G. (1981). Determining optical flow. Artificial Intelligence, 17(1– 3), 185–203. https://doi.org/10.1016/0004-3702(81)90024-2 Idrees, H., Tayyab, M., Athrey, K., et al. (2018). Composition loss for counting, density map estimation and localization in dense crowds. Proceedings of the European Conference on Computer Vision (ECCV), 532–546. Jocher, G., Chaurasia, A., & Qiu, J. (2023). YOLO by Ultralytics. https://github.com/ultralytics/ultralytics Kang, K., Ouyang, W., Li, H., & Wang, X. (2016). Object detection from video tubelets with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 817–825. https://doi.org/10.1109/CVPR.2016.96 Kang, K., Li, H., Yan, J., Zeng, X., Yang, B., Xiao, T., … Wang, X. (2017). T-CNN: Tubelets with convolutional neural networks for object detection from videos. IEEE Transactions on Circuits and Systems for Video Technology, 28(10), 2896–2907. https://doi.org/10.1109/TCSVT.2017.2736559 Kamilaris, A., & Prenafeta-Boldú, F. X. (2018). Deep learning in agriculture: A survey. Computers and Electronics in Agriculture, 147, 70–90. https://doi.org/10.1016/j.compag.2018.02.016 Krizhevsky, A., Sutskever, I., & Hinton, G. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25, 1097–1105. Lee, J., Kim, S., & Lee, K. M. (2020). Learning action transition for fast action recognition. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10934–10943. https://doi.org/10.1109/CVPR42600.2020.01095 Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2117–2125. https://doi.org/10.1109/CVPR.2017.106 Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C. Y., & Berg, A. C. (2016). SSD: Single shot multibox detector. European Conference on Computer Vision (ECCV), 21–37. Springer. https://doi.org/10.1007/978-3-319-46448-0_2 Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., … Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10012–10022. https://doi.org/10.1109/ICCV48922.2021.00986 Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., & Hu, H. (2022). Video Swin Transformer. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3202–3211. https://doi.org/10.1109/CVPR52688.2022.00323 Lucas, B. D., & Kanade, T. (1981). An iterative image registration technique with an application to stereo vision. Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679. Ma, Y., Wang, X., Zhang, X., et al. (2021). Traffic flow prediction and intelligent transportation systems: A survey. Neurocomputing, 441, 166–180. https://doi.org/10.1016/j.neucom.2021.02.016 Ramachandra, B., & Jones, M. J. (2020). Street scene: A new dataset and evaluation protocol for video anomaly detection. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2569–2578. https://doi.org/10.1109/WACV45572.2020.9093335 Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real- time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788. https://doi.org/10.1109/CVPR.2016.91 Redmon, J., & Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv preprint arXiv:1804.02767. Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems, 28, 91–99. Sakaridis, C., Dai, D., & Van Gool, L. (2018). Semantic foggy scene understanding with synthetic data. International Journal of Computer Vision, 126, 973–992. https://doi.org/10.1007/s11263-018-1072-6 Santos, M. S., Silva, L. M., & Gomes, H. (2020). UAV-based video analytics for smart cities. Sensors, 20(23), 6818. https://doi.org/10.3390/s20236818 Shen, T., Zhou, J., Liu, X., & Guo, H. (2020). A vision-based fall detection system using deep learning. Journal of Healthcare Engineering, 2020, 1–8. https://doi.org/10.1155/2020/8890109 Stauffer, C., & Grimson, W. E. L. (1999). Adaptive background mixture models for

More Articles from INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY

Advances in Algorithmic Contract Scoring for Pre-Negotiation Yield Optimization and Risk Retention

Author: Ngozi Samuel Uzougbo, Michael Ominyi, Cyril Chimelie Anichukwueze, Blessing, Chika Jones

DevTest flow: Designing a Scalable Continuous Testing Pipeline for High-Velocity Software Delivery

Author: Lawal Ahmed Oladimeji, Achori Busayo, Akeju BusayoZainab, Saka Samson, Damilare, Mbah Demian Chidi, Runsewe Similoluwa Mayowa, Oladiti Luqman, Abiodun