لطفا منتظر بمانید ...
0% Complete
صفحه اصلی
/
سی و سومین کنفرانس بین المللی مهندسی برق
Better Exploration In Single-Agent Q-Learning Using Controlled Linear Perturbation
نویسندگان :
Sadredin Hokmi
1
Mohammad Haeri
2
1- Sharif university of technology
2- Sharif university of technology
کلمات کلیدی :
Q-learning،Exploration،Controlled Linear perturbation،Convergence rate،Maze،Cart-Pole
چکیده :
Reinforcement learning algorithms, especially model-free algorithms like Q-learning, have shown reliable results in finding optimal solutions for many real-time applications. However, challenges such as exploration in real-time and the convergence rate need to be addressed, and many researches have proposed algorithms to tackle these challenges. Algorithms like speedy Q-learning, Zap Q-learning, algorithms based on adding a regularization term, noise injection, and many others have been introduced. In this paper, an algorithm based on controlled linear perturbation is presented, which, according to the numerical results, can significantly reduce unnecessary explorations that are risky in real-time. Additionally, the proposed algorithm does not depend on the learning rate \mathbit{\alpha}, \mathbit{\gamma}, or changes in coefficients. However, to be effective, the parameters of the algorithm should be chosen within the correct range. The results of applying the proposed algorithm have been compared with three reliable algorithms: standard Q-learning, speedy Q-learning, and noise injection. These comparisons were conducted in a 9x9 maze scenario and in the cart-pole environment.
لیست مقالات
لیست مقالات بایگانی شده
Vibration Analysis of a High-Speed Switched Reluctance Motor Considering Fast Demagnetization Voltage
Nasrin Majlesi - Amir Rashidi - Morteza Saghaian Nejad
Wide-Band Linear to Linear Fish-Bone Polarizer using Metasurfaces
Fatemeh Ganji Arjenaki - Mohsen Maddah Ali - Abolghasem Zeidaabadi Nezhad
3D Microwave Imaging inside PEMC Cavity Using Combined-Norm Regularization Term and Modified CG Algorithm
Omid Babazadeh - Hassan Nasseri
Adaptive Loss-Drift Federated Learning for Non-Intrusive Load Monitoring
Ali Darvishi - Mohammad Reza Mansouri - Amir Reza Setayesh Matin - Reza Gharibi - Behnam Ranjbar - Rahman Dashti
Design and Practical Implementation of Internal Model Controller for Temperature Regulation of Thermoelectric Cell
Parastoo Kamali - Sanaz Iman Shayan - Mahshid Mousapour - Fatemeh Abdolsamadi - Salar Zeinali - Sadra Rafatnia
مدیریت انرژی یک شبکه هوشمند با ساختار هولاکراسی انرژی شامل مصرفکنندگان خودتولید بر اساس حق انتخاب مبتنی بر ترجیحات اقتصادی، زیستمحیطی و اجتماعی
پیمان افضلی - مسعود رشیدی نژاد - امیر عبداللهی - محمدرضا صالحی زاده - حسین فرهمند
Design and Simulation of Long Slot SIW Leaky Wave Antenna for Automotive Radar Application
Jamal Kazazi - Alireza Rahmani - Mahmoud Kamarei
DWT-Based Epileptic Seizure Detection Using Fuzzy Logic Model with Entropy and Table Lookup Scheme
Alireza Mohammadi - Arvin Esfandyari - Ali Doustmohammadi - Amir Abolfazl Suratgar - Masoud Shafiee
Sliding-mode H∞ Control of Continuous Singular Systems under Zeno-free Event-triggered Sampling Scheme
Hamidreza Ahmadzadeh - Masoud Shafiee
Structural Stability and Electron Density Analysis of Doped Antimonene: A First-Principles Study
Arash Yazdanpanah Goharrizi - Peyman Saberi Parsa
بیشتر
ثمین همایش، سامانه مدیریت کنفرانس ها و جشنواره ها - نگارش 44.7.2