0% Complete
صفحه اصلی
/
سی و سومین کنفرانس بین المللی مهندسی برق
Better Exploration In Single-Agent Q-Learning Using Controlled Linear Perturbation
نویسندگان :
Sadredin Hokmi
1
Mohammad Haeri
2
1- Sharif university of technology
2- Sharif university of technology
کلمات کلیدی :
Q-learning،Exploration،Controlled Linear perturbation،Convergence rate،Maze،Cart-Pole
چکیده :
Reinforcement learning algorithms, especially model-free algorithms like Q-learning, have shown reliable results in finding optimal solutions for many real-time applications. However, challenges such as exploration in real-time and the convergence rate need to be addressed, and many researches have proposed algorithms to tackle these challenges. Algorithms like speedy Q-learning, Zap Q-learning, algorithms based on adding a regularization term, noise injection, and many others have been introduced. In this paper, an algorithm based on controlled linear perturbation is presented, which, according to the numerical results, can significantly reduce unnecessary explorations that are risky in real-time. Additionally, the proposed algorithm does not depend on the learning rate \mathbit{\alpha}, \mathbit{\gamma}, or changes in coefficients. However, to be effective, the parameters of the algorithm should be chosen within the correct range. The results of applying the proposed algorithm have been compared with three reliable algorithms: standard Q-learning, speedy Q-learning, and noise injection. These comparisons were conducted in a 9x9 maze scenario and in the cart-pole environment.
لیست مقالات
لیست مقالات بایگانی شده
Adaptive Control of Telerehabilitation Systems in The Framework of Multi-Agent Systems
Mohammadreza Sheykh - Heidar Ali Talebi - ّIman Sharifi
برنامه ریزی مسیر حرکت ربات در بین عابران پیاده با پیشبینی حرکت عابران
ملیکا رضوانی - سمانه حسینی
An incentive compatible reward sharing approach for shard-based blockchains
Mojdeh Hemati - Mehdi Shajari
Ultra-Compact and Fast All-Optical Half-Subtractor Photonic Crystal Logic Gate
Ehsan Veisi - Mahmood Seifouri - Saeed Olyaee
Automotive radar target classification using micro-Doppler features
Amin Aghatabarroodbary - Mohammad Hassan Bastani - Fereidoon Behnia
Applying Parameter-Oriented Learning to Identify Statistical EEG Features Associated with Depression
Sara Bargi Barkouk - Melika Changizi - Mahdi Zolfagharzadeh Kermani - Ali Asadi Zeidabadi
تحلیل دینامیکی ماشین سنکرون مغناطیس دائم با آهنربای جانبی و تحلیل خطای اتصال کوتاه داخلی و ضعیف شدن آهنربا
آزیتا فتحی - پیمان نادری
A Digital Method for Offset Cancellation of Fully Dynamic Latched Comparators
Alireza Ahrar - Mohammad Yavari
Design of a Controllable and State-observable MEMS Nonlinear Resonator Based on the Awl-shaped Serpentine Spring
Ehsan Ranjbar - Amirabolfazl Suratgar
طراحی یک چارچوب غیر متمرکز تبادل انرژی برای مصرفکنندههای فعال در بازارهای همتا به همتا (P2P)
امیر زارع بخت پیما چمثقالی - مهدی مهدینژاد - مهرداد عابدی
بیشتر
ثمین همایش، سامانه مدیریت کنفرانس ها و جشنواره ها - نگارش 43.6.0