References
Bernardin, K., & Stiefelhagen, R. (2008). Evaluating multiple object tracking performance: The CLEAR MOT metrics. EURASIP Journal on Image and Video Processing, 2008(1), 1–10. https://doi.org/10.1155/2008/246309 Bochkovskiy, A., Wang, C. Y., & Liao, H. Y. M. (2020). YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 1–15. Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to- end object detection with transformers. European Conference on Computer Vision (ECCV), 213–229. Springer. https://doi.org/10.1007/978-3-030-58452-8_13 Chen, L. C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2019). Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv preprint arXiv:1412.7062. Chen, X., Li, S., & Luo, J. (2020). Video object detection via collaborative learning of motion cues. IEEE Transactions on Image Processing, 29, 10287–10299. https://doi.org/10.1109/TIP.2020.3027661 Cioppa, A., Giancola, S., & Ghanem, B. (2019). A context-aware loss function for action spotting in soccer videos. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13126–13136. https://doi.org/10.1109/CVPR.2019.01343 Dai, J., Li, Y., He, K., & Sun, J. (2017). R-FCN: Object detection via region-based fully convolutional networks. Advances in Neural Information Processing Systems, 29, 379– Dalal, N., & Triggs, B. (2005). Histograms of oriented gradients for human detection. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 886– IEEE. https://doi.org/10.1109/CVPR.2005.177 Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., & Zisserman, A. (2010). The Pascal Visual Object Classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303–338. https://doi.org/10.1007/s11263-009-0275-4 Feichtenhofer, C., Fan, H., Li, Y., & He, K. (2022). Masked autoencoders as spatiotemporal learners. Advances in Neural Information Processing Systems, 35, 35946–35958. Geiger, A., Lenz, P., & Urtasun, R. (2013). Are we ready for autonomous driving? The KITTI vision benchmark suite. 2012 IEEE Conference on Computer Vision and Pattern Recognition, 3354–3361. https://doi.org/10.1109/CVPR.2012.6248074 Girshick, R. (2015). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 1440–1448. https://doi.org/10.1109/ICCV.2015.169 Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 580–587. https://doi.org/10.1109/CVPR.2014.81 Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press. He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2961–2969. https://doi.org/10.1109/ICCV.2017.322 Horn, B. K. P., & Schunck, B. G. (1981). Determining optical flow. Artificial Intelligence, 17(1– 3), 185–203. https://doi.org/10.1016/0004-3702(81)90024-2 Idrees, H., Tayyab, M., Athrey, K., et al. (2018). Composition loss for counting, density map estimation and localization in dense crowds. Proceedings of the European Conference on Computer Vision (ECCV), 532–546. Jocher, G., Chaurasia, A., & Qiu, J. (2023). YOLO by Ultralytics. https://github.com/ultralytics/ultralytics Kang, K., Ouyang, W., Li, H., & Wang, X. (2016). Object detection from video tubelets with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 817–825. https://doi.org/10.1109/CVPR.2016.96 Kang, K., Li, H., Yan, J., Zeng, X., Yang, B., Xiao, T., … Wang, X. (2017). T-CNN: Tubelets with convolutional neural networks for object detection from videos. IEEE Transactions on Circuits and Systems for Video Technology, 28(10), 2896–2907. https://doi.org/10.1109/TCSVT.2017.2736559 Kamilaris, A., & Prenafeta-Boldú, F. X. (2018). Deep learning in agriculture: A survey. Computers and Electronics in Agriculture, 147, 70–90. https://doi.org/10.1016/j.compag.2018.02.016 Krizhevsky, A., Sutskever, I., & Hinton, G. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25, 1097–1105. Lee, J., Kim, S., & Lee, K. M. (2020). Learning action transition for fast action recognition. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10934–10943. https://doi.org/10.1109/CVPR42600.2020.01095 Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2117–2125. https://doi.org/10.1109/CVPR.2017.106 Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C. Y., & Berg, A. C. (2016). SSD: Single shot multibox detector. European Conference on Computer Vision (ECCV), 21–37. Springer. https://doi.org/10.1007/978-3-319-46448-0_2 Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., … Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10012–10022. https://doi.org/10.1109/ICCV48922.2021.00986 Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., & Hu, H. (2022). Video Swin Transformer. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3202–3211. https://doi.org/10.1109/CVPR52688.2022.00323 Lucas, B. D., & Kanade, T. (1981). An iterative image registration technique with an application to stereo vision. Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679. Ma, Y., Wang, X., Zhang, X., et al. (2021). Traffic flow prediction and intelligent transportation systems: A survey. Neurocomputing, 441, 166–180. https://doi.org/10.1016/j.neucom.2021.02.016 Ramachandra, B., & Jones, M. J. (2020). Street scene: A new dataset and evaluation protocol for video anomaly detection. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2569–2578. https://doi.org/10.1109/WACV45572.2020.9093335 Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real- time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788. https://doi.org/10.1109/CVPR.2016.91 Redmon, J., & Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv preprint arXiv:1804.02767. Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems, 28, 91–99. Sakaridis, C., Dai, D., & Van Gool, L. (2018). Semantic foggy scene understanding with synthetic data. International Journal of Computer Vision, 126, 973–992. https://doi.org/10.1007/s11263-018-1072-6 Santos, M. S., Silva, L. M., & Gomes, H. (2020). UAV-based video analytics for smart cities. Sensors, 20(23), 6818. https://doi.org/10.3390/s20236818 Shen, T., Zhou, J., Liu, X., & Guo, H. (2020). A vision-based fall detection system using deep learning. Journal of Healthcare Engineering, 2020, 1–8. https://doi.org/10.1155/2020/8890109 Stauffer, C., & Grimson, W. E. L. (1999). Adaptive background mixture models for