0% Complete
صفحه اصلی
/
سی و سومین کنفرانس بین المللی مهندسی برق
Better Exploration In Single-Agent Q-Learning Using Controlled Linear Perturbation
نویسندگان :
Sadredin Hokmi
1
Mohammad Haeri
2
1- Sharif university of technology
2- Sharif university of technology
کلمات کلیدی :
Q-learning،Exploration،Controlled Linear perturbation،Convergence rate،Maze،Cart-Pole
چکیده :
Reinforcement learning algorithms, especially model-free algorithms like Q-learning, have shown reliable results in finding optimal solutions for many real-time applications. However, challenges such as exploration in real-time and the convergence rate need to be addressed, and many researches have proposed algorithms to tackle these challenges. Algorithms like speedy Q-learning, Zap Q-learning, algorithms based on adding a regularization term, noise injection, and many others have been introduced. In this paper, an algorithm based on controlled linear perturbation is presented, which, according to the numerical results, can significantly reduce unnecessary explorations that are risky in real-time. Additionally, the proposed algorithm does not depend on the learning rate \mathbit{\alpha}, \mathbit{\gamma}, or changes in coefficients. However, to be effective, the parameters of the algorithm should be chosen within the correct range. The results of applying the proposed algorithm have been compared with three reliable algorithms: standard Q-learning, speedy Q-learning, and noise injection. These comparisons were conducted in a 9x9 maze scenario and in the cart-pole environment.
لیست مقالات
لیست مقالات بایگانی شده
Design and Practical Implementation of Internal Model Controller for Temperature Regulation of Thermoelectric Cell
Parastoo Kamali - Sanaz Iman Shayan - Mahshid Mousapour - Fatemeh Abdolsamadi - Salar Zeinali - Sadra Rafatnia
خلاصه سازی ویدیوهای کپسول آندوسکوپی با رویکرد یادگیری انتقالی
محدثه امیریان چایجان - رضا آقائی زاده ظروفی - مسعود رضا سهرابی
Low Cost Implementation of Neural Networks Based on Stochastic Computing
Hadi Jahanirad - Ahmad Menbari
MAD-TI: Meta-path Aggregated-Graph Attention Network for Drug Target Interaction Prediction
Reza Shami Tanha - Maryam Sadighian - Arash Zabihian - Mohsen Hooshmand - Mohsen Afsharchi
Type-2 Fuzzy Wavelet Control for a Quadruple-Tank System based on Disturbance Rejection
Mohammadreza Esmaeilidehkordi - Alireza Nezamzadeh - Maryam Zekri - Iman Izadi - Farid Sheikholeslam
Simulation of planar organic-inorganic perovskite light-emitting diode
Morteza Yarahmadi - Elnaz Yazdani - Mohammad Kazem Moravvej-Farshi
کنترل بازوی ربات دو درجه آزادی با کنترلکننده مود لغزشی مرتبه کسری فازی-تطبیقی پایانهای
مائده نفیسی فر - متین جزءاسلامی - ابوالفضل جلیلوند - سمیرا نریمان پور - فرهاد بیات
Optimal Design of a Synchronous Reluctance Motor Using BioGeography-Based Optimization
Tohid Sharifi - Mojtaba Mirsalim
ساخت حسگر گاز بر پایه ی گرافن اکساید و سیلیکون متخلخل
سیده صفیه رضایی - مینا امیر مزلقانی
Combination of Classifiers to Detecting Grade of Gliblastoma using MRS
Roqaie Moqadam - Nazila Loghmani - Meysam Siyahmansoori - Armin Allahverdy
بیشتر
ثمین همایش، سامانه مدیریت کنفرانس ها و جشنواره ها - نگارش 43.6.0