0% Complete
صفحه اصلی
/
سی و دومین کنفرانس بین المللی مهندسی برق
Human Action Recognition in Still Images Using ConViT
نویسندگان :
Seyed Rohollah Hosseyni
1
Sanaz Seyedin
2
Hassan Taheri
3
1- Amirkabir University of Technology
2- Amirkabir University of Technology
3- Amirkabir University of Technology
کلمات کلیدی :
Human action recognition،Still images،Convolutional Neural Network،Vision Transformer
چکیده :
Understanding the relationship between different parts of an image is crucial in a variety of applications, including object recognition, scene understanding, and image classification. Despite the fact that Convolutional Neural Networks (CNNs) have demonstrated impressive results in classifying and detecting objects, they lack the capability to extract the relationship between different parts of an image, which is a crucial factor in Human Action Recognition (HAR). To address this problem, this paper proposes a new module that functions like a convolutional layer that uses Vision Transformer (ViT). In the proposed model, the Vision Transformer can complement a convolutional neural network in a variety of tasks by helping it to effectively extract the relationship among various parts of an image. It is shown that the proposed model, compared to a simple CNN, can extract meaningful parts of an image and suppress the misleading parts. The proposed model has been evaluated on the Stanford40 and PASCAL VOC 2012 action datasets and has achieved 95.5% mean Average Precision (mAP) and 91.5% mAP results, respectively, which are promising compared to other state-of-the-art methods.
لیست مقالات
لیست مقالات بایگانی شده
DRAU-Net: Double Residual Attention Mechanism for automatic MRI brain tumor segmentation
Mohammad Soltani gol - Morteza Fattahi - Hamid Soltanian zadeh - Samd Sheikhaei
Integrated strategy for segment BRATS using co-operation of FCM and TL under abnormal behavior of noises
Arman Zafaranchi - Pedram Salehpoor
Detecting Variance Changes in Alarm Systems Using Generalized Delay-timers
Zahra Sharifi - Iman Izadi - Jafar Ghaisari
A Novel Image Denoising Algorithm Based on Wavelet and Akamatsu Transforms Using Particle Swarm Optimization
Zeinab Pakdaman - Majid Amini-Valashani - Sattar Mirzakuchaki
کنترل بازوی ربات دو درجه آزادی با کنترلکننده مود لغزشی مرتبه کسری فازی-تطبیقی پایانهای
مائده نفیسی فر - متین جزءاسلامی - ابوالفضل جلیلوند - سمیرا نریمان پور - فرهاد بیات
Transfer learning using deep convolutional neural network for predicting dementia severity
Vahid Asayesh - Mehdi Dehghani - Majid Torabi Nikjeh - Sepideh Akhtari khosrowshahi
مکان یابی اهداف در محیط مختلط دید مستقیم و غیر مستقیم مبتنی بر اندازه گیری های RSS و TOA با مدل احتمالاتی
محمدرضا شمسیان - فریدون بهنیا
A High Linearity Wideband Low-Noise Amplifier Using Capacitor Cross-Coupled Common-Gate Structure
Abolfazl Rajaiyan - Fahimeh Rahimi - Mehdi Saberi
Global Finite-Time Nonlinear Observers for a Class of Nonlinear Systems Subjected to Mismatched Uncertainties
َAli Abooee - Saeed Amiri - Mohammad Hadi Rezaei
Robust Neuro-Adaptive Fuzzy Sliding Mode Control for a Remotely Operated Underwater Vehicle Manipulator
Mahdi Armoon - Marzie Lafouti - Babak Tavassoli - Hamid D. Taghirad
بیشتر
ثمین همایش، سامانه مدیریت کنفرانس ها و جشنواره ها - نگارش 43.6.0