Arxiv:2010.02193V3 [Cs.LG] 3 May 2021
Total Page:16
File Type:pdf, Size:1020Kb
Load more
Recommended publications
-
Elements of DSAI: Game Tree Search, Learning Architectures
Introduction Games Game Search Evaluation Fns AlphaGo/Zero Summary Quiz References Elements of DSAI AlphaGo Part 2: Game Tree Search, Learning Architectures Search & Learn: A Recipe for AI Action Decisions J¨orgHoffmann Winter Term 2019/20 Hoffmann Elements of DSAI Game Tree Search, Learning Architectures 1/49 Introduction Games Game Search Evaluation Fns AlphaGo/Zero Summary Quiz References Competitive Agents? Quote AI Introduction: \Single agent vs. multi-agent: One agent or several? Competitive or collaborative?" ! Single agent! Several koalas, several gorillas trying to beat these up. BUT there is only a single acting entity { one player decides which moves to take (who gets into the boat). Hoffmann Elements of DSAI Game Tree Search, Learning Architectures 3/49 Introduction Games Game Search Evaluation Fns AlphaGo/Zero Summary Quiz References Competitive Agents! Quote AI Introduction: \Single agent vs. multi-agent: One agent or several? Competitive or collaborative?" ! Multi-agent competitive! TWO players deciding which moves to take. Conflicting interests. Hoffmann Elements of DSAI Game Tree Search, Learning Architectures 4/49 Introduction Games Game Search Evaluation Fns AlphaGo/Zero Summary Quiz References Agenda: Game Search, AlphaGo Architecture Games: What is that? ! Game categories, game solutions. Game Search: How to solve a game? ! Searching the game tree. Evaluation Functions: How to evaluate a game position? ! Heuristic functions for games. AlphaGo: How does it work? ! Overview of AlphaGo architecture, and changes in Alpha(Go) Zero. Hoffmann Elements of DSAI Game Tree Search, Learning Architectures 5/49 Introduction Games Game Search Evaluation Fns AlphaGo/Zero Summary Quiz References Positioning in the DSAI Phase Model Hoffmann Elements of DSAI Game Tree Search, Learning Architectures 6/49 Introduction Games Game Search Evaluation Fns AlphaGo/Zero Summary Quiz References Which Games? ! No chance element. -
Discrete Sequential Prediction of Continu
Under review as a conference paper at ICLR 2018 DISCRETE SEQUENTIAL PREDICTION OF CONTINU- OUS ACTIONS FOR DEEP RL Anonymous authors Paper under double-blind review ABSTRACT It has long been assumed that high dimensional continuous control problems cannot be solved effectively by discretizing individual dimensions of the action space due to the exponentially large number of bins over which policies would have to be learned. In this paper, we draw inspiration from the recent success of sequence-to-sequence models for structured prediction problems to develop policies over discretized spaces. Central to this method is the realization that com- plex functions over high dimensional spaces can be modeled by neural networks that predict one dimension at a time. Specifically, we show how Q-values and policies over continuous spaces can be modeled using a next step prediction model over discretized dimensions. With this parameterization, it is possible to both leverage the compositional structure of action spaces during learning, as well as compute maxima over action spaces (approximately). On a simple example task we demonstrate empirically that our method can perform global search, which effectively gets around the local optimization issues that plague DDPG. We apply the technique to off-policy (Q-learning) methods and show that our method can achieve the state-of-the-art for off-policy methods on several continuous control tasks. 1 INTRODUCTION Reinforcement learning has long been considered as a general framework applicable to a broad range of problems. However, the approaches used to tackle discrete and continuous action spaces have been fundamentally different. In discrete domains, algorithms such as Q-learning leverage backups through Bellman equations and dynamic programming to solve problems effectively. -
Rétro Gaming
ATARI - CONSOLE RETRO FLASHBACK 8 GOLD ACTIVISION – 130 JEUX Faites ressurgir vos meilleurs souvenirs ! Avec 130 classiques du jeu vidéo dont 39 jeux Activision, cette Atari Flashback 8 Gold édition Activision saura vous rappeler aux bons souvenirs du rétro-gaming. Avec les manettes sans fils ou vos anciennes manettes classiques Atari, vous n’avez qu’à brancher la console à votre télévision et vous voilà prêts pour l’action ! CARACTÉRISTIQUES : • 130 jeux classiques incluant les meilleurs hits de la console Atari 2600 et 39 titres Activision • Plug & Play • Inclut deux manettes sans fil 2.4G • Fonctions Sauvegarde, Reprise, Rembobinage • Sortie HD 720p • Port HDMI • Ecran FULL HD Inclut les jeux cultes : • Space Invaders • Centipede • Millipede • Pitfall! • River Raid • Kaboom! • Spider Fighter LISTE DES JEUX INCLUS LISTE DES JEUX ACTIVISION REF Adventure Black Jack Football Radar Lock Stellar Track™ Video Chess Beamrider Laser Blast Asteroids® Bowling Frog Pond Realsports® Baseball Street Racer Video Pinball Boxing Megamania JVCRETR0124 Centipede ® Breakout® Frogs and Flies Realsports® Basketball Submarine Commander Warlords® Bridge Oink! Kaboom! Canyon Bomber™ Fun with Numbers Realsports® Soccer Super Baseball Yars’ Return Checkers Pitfall! Missile Command® Championship Golf Realsports® Volleyball Super Breakout® Chopper Command Plaque Attack Pitfall! Soccer™ Gravitar® Return to Haunted Save Mary Cosmic Commuter Pressure Cooker EAN River Raid Circus Atari™ Hangman House Super Challenge™ Football Crackpots Private Eye Yars’ Revenge® Combat® -
Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning
Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning Guangxiang Zhu⇤ Minghao Zhang⇤ IIIS School of Software Tsinghua University Tsinghua University [email protected] [email protected] Honglak Lee Chongjie Zhang EECS IIIS University of Michigan Tsinghua University [email protected] [email protected] Abstract Sample efficiency has been one of the major challenges for deep reinforcement learning. Recently, model-based reinforcement learning has been proposed to address this challenge by performing planning on imaginary trajectories with a learned world model. However, world model learning may suffer from overfitting to training trajectories, and thus model-based value estimation and policy search will be prone to be sucked in an inferior local policy. In this paper, we propose a novel model-based reinforcement learning algorithm, called BrIdging Reality and Dream (BIRD). It maximizes the mutual information between imaginary and real trajectories so that the policy improvement learned from imaginary trajectories can be easily generalized to real trajectories. We demonstrate that our approach improves sample efficiency of model-based planning, and achieves state-of-the-art performance on challenging visual control benchmarks. 1 Introduction Reinforcement learning (RL) is proposed as a general-purpose learning framework for artificial intelligence problems, and has led to tremendous progress in a variety of domains [1, 2, 3, 4]. Model- free RL adopts a trail-and-error paradigm, which directly learns a mapping function from observations to values or actions through interactions with environments. It has achieved remarkable performance in certain video games and continuous control tasks because of its simplicity and minimal assumptions about environments. -
List of Intellivision Games
List of Intellivision Games 1) 4-Tris 25) Checkers 2) ABPA Backgammon 26) Chip Shot: Super Pro Golf 3) ADVANCED DUNGEONS & DRAGONS 27) Commando Cartridge 28) Congo Bongo 4) ADVANCED DUNGEONS & DRAGONS 29) Crazy Clones Treasure of Tarmin Cartridge 30) Deep Pockets: Super Pro Pool and Billiards 5) Adventure (AD&D - Cloudy Mountain) (1982) (Mattel) 31) Defender 6) Air Strike 32) Demon Attack 7) Armor Battle 33) Diner 8) Astrosmash 34) Donkey Kong 9) Atlantis 35) Donkey Kong Junior 10) Auto Racing 36) Dracula 11) B-17 Bomber 37) Dragonfire 12) Beamrider 38) Eggs 'n' Eyes 13) Beauty & the Beast 39) Fathom 14) Blockade Runner 40) Frog Bog 15) Body Slam! Super Pro Wrestling 41) Frogger 16) Bomb Squad 42) Game Factory 17) Boxing 43) Go for the Gold 18) Brickout! 44) Grid Shock 19) Bump 'n' Jump 45) Happy Trails 20) BurgerTime 46) Hard Hat 21) Buzz Bombers 47) Horse Racing 22) Carnival 48) Hover Force 23) Centipede 49) Hypnotic Lights 24) Championship Tennis 50) Ice Trek 51) King of the Mountain 52) Kool-Aid Man 80) Number Jumble 53) Lady Bug 81) PBA Bowling 54) Land Battle 82) PGA Golf 55) Las Vegas Poker & Blackjack 83) Pinball 56) Las Vegas Roulette 84) Pitfall! 57) League of Light 85) Pole Position 58) Learning Fun I 86) Pong 59) Learning Fun II 87) Popeye 60) Lock 'N' Chase 88) Q*bert 61) Loco-Motion 89) Reversi 62) Magic Carousel 90) River Raid 63) Major League Baseball 91) Royal Dealer 64) Masters of the Universe: The Power of He- 92) Safecracker Man 93) Scooby Doo's Maze Chase 65) Melody Blaster 94) Sea Battle 66) Microsurgeon 95) Sewer Sam 67) Mind Strike 96) Shark! Shark! 68) Minotaur (1981) (Mattel) 97) Sharp Shot 69) Mission-X 98) Slam Dunk: Super Pro Basketball 70) Motocross 99) Slap Shot: Super Pro Hockey 71) Mountain Madness: Super Pro Skiing 100) Snafu 72) Mouse Trap 101) Space Armada 73) Mr. -
Improving the Gating Mechanism of Recurrent
Under review as a conference paper at ICLR 2020 IMPROVING THE GATING MECHANISM OF RECURRENT NEURAL NETWORKS Anonymous authors Paper under double-blind review ABSTRACT Gating mechanisms are widely used in neural network models, where they allow gradients to backpropagate more easily through depth or time. However, their satura- tion property introduces problems of its own. For example, in recurrent models these gates need to have outputs near 1 to propagate information over long time-delays, which requires them to operate in their saturation regime and hinders gradient-based learning of the gate mechanism. We address this problem by deriving two synergistic modifications to the standard gating mechanism that are easy to implement, introduce no additional hyperparameters, and improve learnability of the gates when they are close to saturation. We show how these changes are related to and improve on alternative recently proposed gating mechanisms such as chrono-initialization and Ordered Neurons. Empirically, our simple gating mechanisms robustly improve the performance of recurrent models on a range of applications, including synthetic memorization tasks, sequential image classification, language modeling, and reinforcement learning, particularly when long-term dependencies are involved. 1 INTRODUCTION Recurrent neural networks (RNNs) have become a standard machine learning tool for learning from sequential data. However, RNNs are prone to the vanishing gradient problem, which occurs when the gradients of the recurrent weights become vanishingly small as they get backpropagated through time (Hochreiter et al., 2001). A common approach to alleviate the vanishing gradient problem is to use gating mechanisms, leading to models such as the long short term memory (Hochreiter & Schmidhuber, 1997, LSTM) and gated recurrent units (Chung et al., 2014, GRUs). -
Bayesian Policy Selection Using Active Inference
View metadata, citation and similar papers at core.ac.uk brought to you by CORE provided by Ghent University Academic Bibliography Published in the proceedings of the Workshop on “Structure & Priors in Reinforcement Learning” at ICLR 2019 BAYESIAN POLICY SELECTION USING ACTIVE INFERENCE Ozan C¸atal, Johannes Nauta, Tim Verbelen, Pieter Simoens, & Bart Dhoedt IDLab, Department of Information Technology Ghent University - imec Ghent, Belgium [email protected] ABSTRACT Learning to take actions based on observations is a core requirement for artificial agents to be able to be successful and robust at their task. Reinforcement Learn- ing (RL) is a well-known technique for learning such policies. However, current RL algorithms often have to deal with reward shaping, have difficulties gener- alizing to other environments and are most often sample inefficient. In this pa- per, we explore active inference and the free energy principle, a normative theory from neuroscience that explains how self-organizing biological systems operate by maintaining a model of the world and casting action selection as an inference problem. We apply this concept to a typical problem known to the RL community, the mountain car problem, and show how active inference encompasses both RL and learning from demonstrations. 1 INTRODUCTION Active inference is an emerging paradigm from neuroscience (Friston et al., 2017), which postulates that action selection in biological systems, in particular the human brain, is in effect an inference problem where agents are attracted to a preferred prior state distribution in a hidden state space. Contrary to many state-of-the art RL algorithms, active inference agents are not purely goal directed and exhibit an inherent epistemic exploration (Schwartenbeck et al., 2018). -
Newagearcade.Com 5000 in One Arcade Game List!
Newagearcade.com 5,000 In One arcade game list! 1. AAE|Armor Attack 2. AAE|Asteroids Deluxe 3. AAE|Asteroids 4. AAE|Barrier 5. AAE|Boxing Bugs 6. AAE|Black Widow 7. AAE|Battle Zone 8. AAE|Demon 9. AAE|Eliminator 10. AAE|Gravitar 11. AAE|Lunar Lander 12. AAE|Lunar Battle 13. AAE|Meteorites 14. AAE|Major Havoc 15. AAE|Omega Race 16. AAE|Quantum 17. AAE|Red Baron 18. AAE|Ripoff 19. AAE|Solar Quest 20. AAE|Space Duel 21. AAE|Space Wars 22. AAE|Space Fury 23. AAE|Speed Freak 24. AAE|Star Castle 25. AAE|Star Hawk 26. AAE|Star Trek 27. AAE|Star Wars 28. AAE|Sundance 29. AAE|Tac/Scan 30. AAE|Tailgunner 31. AAE|Tempest 32. AAE|Warrior 33. AAE|Vector Breakout 34. AAE|Vortex 35. AAE|War of the Worlds 36. AAE|Zektor 37. Classic Arcades|'88 Games 38. Classic Arcades|1 on 1 Government (Japan) 39. Classic Arcades|10-Yard Fight (World, set 1) 40. Classic Arcades|1000 Miglia: Great 1000 Miles Rally (94/07/18) 41. Classic Arcades|18 Holes Pro Golf (set 1) 42. Classic Arcades|1941: Counter Attack (World 900227) 43. Classic Arcades|1942 (Revision B) 44. Classic Arcades|1943 Kai: Midway Kaisen (Japan) 45. Classic Arcades|1943: The Battle of Midway (Euro) 46. Classic Arcades|1944: The Loop Master (USA 000620) 47. Classic Arcades|1945k III 48. Classic Arcades|19XX: The War Against Destiny (USA 951207) 49. Classic Arcades|2 On 2 Open Ice Challenge (rev 1.21) 50. Classic Arcades|2020 Super Baseball (set 1) 51. -
Using Deep Learning to Create a Universal Game Player
Using Deep Learning to create a Universal Game Player Tackling Transfer Learning and Catastrophic Forgetting Using Asynchronous Methods for Deep Reinforcement Learning Joao˜ Manuel Godinho Ribeiro Thesis to obtain the Master of Science Degree in Information Systems and Computer Engineering Supervisors: Prof. Francisco Antonio´ Chaves Saraiva de Melo Prof. Joao˜ Miguel De Sousa de Assis Dias Examination Committee Chairperson: Prof. Lu´ıs Manuel Antunes Veiga Supervisor: Prof. Francisco Antonio´ Chaves Saraiva de Melo Member of the Committee: Prof. Manuel Fernando Cabido Peres Lopes October 2018 Acknowledgments This dissertation is dedicated to my parents and family, who supported me not only throughout my academic career, but throughout all the difficulties I’ve encountered in life. To my two advisers, Francisco and Joao,˜ for not only providing all the necessary support, patience and availability, throughout the entire process, but especially for this wonderful opportunity to work in the scientific area for which I am most passionate about. To my beautiful girlfriend Margarida, for always believing in me in times of great doubt and adversity, and without whom the entire process would have been much tougher. Thank you for your strength. Finally, to my friends and colleagues which accompanied me throughout the entire process and without whom life wouldn’t have the same meaning. Thank you all for always being there for me. Abstract This dissertation introduces a general-purpose architecture that allows a learning agent to (i) transfer knowledge from a previously learned task to a new one that is now required to learned and (ii) remember how to perform the previously learned tasks as it learns the new one. -
Meta-Learning Deep Energy-Based Memory Models
Published as a conference paper at ICLR 2020 META-LEARNING DEEP ENERGY-BASED MEMORY MODELS Sergey Bartunov Jack W Rae Simon Osindero DeepMind DeepMind DeepMind London, United Kingdom London, United Kingdom London, United Kingdom [email protected] [email protected] [email protected] Timothy P Lillicrap DeepMind London, United Kingdom [email protected] ABSTRACT We study the problem of learning an associative memory model – a system which is able to retrieve a remembered pattern based on its distorted or incomplete version. Attractor networks provide a sound model of associative memory: patterns are stored as attractors of the network dynamics and associative retrieval is performed by running the dynamics starting from a query pattern until it converges to an attractor. In such models the dynamics are often implemented as an optimization procedure that minimizes an energy function, such as in the classical Hopfield network. In general it is difficult to derive a writing rule for a given dynamics and energy that is both compressive and fast. Thus, most research in energy- based memory has been limited either to tractable energy models not expressive enough to handle complex high-dimensional objects such as natural images, or to models that do not offer fast writing. We present a novel meta-learning approach to energy-based memory models (EBMM) that allows one to use an arbitrary neural architecture as an energy model and quickly store patterns in its weights. We demonstrate experimentally that our EBMM approach can build compressed memories for synthetic and natural data, and is capable of associative retrieval that outperforms existing memory systems in terms of the reconstruction error and compression rate. -
A Page 1 CART TITLE MANUFACTURER LABEL RARITY Atari Text
A CART TITLE MANUFACTURER LABEL RARITY 3D Tic-Tac Toe Atari Text 2 3D Tic-Tac Toe Sears Text 3 Action Pak Atari 6 Adventure Sears Text 3 Adventure Sears Picture 4 Adventures of Tron INTV White 3 Adventures of Tron M Network Black 3 Air Raid MenAvision 10 Air Raiders INTV White 3 Air Raiders M Network Black 2 Air Wolf Unknown Taiwan Cooper ? Air-Sea Battle Atari Text #02 3 Air-Sea Battle Atari Picture 2 Airlock Data Age Standard 3 Alien 20th Century Fox Standard 4 Alien Xante 10 Alpha Beam with Ernie Atari Children's 4 Arcade Golf Sears Text 3 Arcade Pinball Sears Text 3 Arcade Pinball Sears Picture 3 Armor Ambush INTV White 4 Armor Ambush M Network Black 3 Artillery Duel Xonox Standard 5 Artillery Duel/Chuck Norris Superkicks Xonox Double Ender 5 Artillery Duel/Ghost Master Xonox Double Ender 5 Artillery Duel/Spike's Peak Xonox Double Ender 6 Assault Bomb Standard 9 Asterix Atari 10 Asteroids Atari Silver 3 Asteroids Sears Text “66 Games” 2 Asteroids Sears Picture 2 Astro War Unknown Taiwan Cooper ? Astroblast Telegames Silver 3 Atari Video Cube Atari Silver 7 Atlantis Imagic Text 2 Atlantis Imagic Picture – Day Scene 2 Atlantis Imagic Blue 4 Atlantis II Imagic Picture – Night Scene 10 Page 1 B CART TITLE MANUFACTURER LABEL RARITY Bachelor Party Mystique Standard 5 Bachelor Party/Gigolo Playaround Standard 5 Bachelorette Party/Burning Desire Playaround Standard 5 Back to School Pak Atari 6 Backgammon Atari Text 2 Backgammon Sears Text 3 Bank Heist 20th Century Fox Standard 5 Barnstorming Activision Standard 2 Baseball Sears Text 49-75108 -
Shixiang (Shane) Gu
Shixiang (Shane) Gu Jesus College, Jesus Lane Cambridge, Cambridgeshire CB5 8BL, United Kingdoms [email protected] website: http://sg717.user.srcf.net/ RESEARCH INTERESTS Deep Learning, Reinforcement Learning, Robotics, Probabilistic Models and Approximate Inference, Causality, and other Machine Learning topics. EDUCATION PhD in Machine Learning 2014—Present University of Cambridge, UK. Advised by Prof. Richard E. Turner and Prof. Zoubin Ghahramani FRS. Max Planck Institute for Intelligent Systems, Germany. Advised by Prof. Bernhard Schoelkopf. BASc in Engineering Science, Electrical and Computer major 2009–2013 University of Toronto, Canada. CGPA: 3.93. Highest Rank 1st of 264 students (2010 Winter). Recipient of Award of Excellence 2013 (awarded to top 5 graduating students from the program). High School, Vancouver, Canada. 2006–2009 Middle School, Shanghai, China. 2002–2006 Primary School, Tokyo, Japan. 1997–2002 PROFESSIONAL EXPERIENCE Google Brain, Google Research, Mountain View, CA. PhD Intern June 2015—Jan. 2016, June 2016–Sept. 2016, June 2017–Sept. 2017 Supervisors & collaborators: Ilya Sutskever, Sergey Levine, Vincent Vanhoucke, Timothy Lillicrap. • Published 1 NIPS, 2 ICML, 3 ICLR, and 1 ICRA papers as a result of internships. • Worked part-time during Feb. 2016—May 2016 at Google DeepMind under Timothy Lillicrap. • Developed the state-of-the-art methods for deep RL in continuous control, and demonstrated the robots learning door opening under 3 hours. Panasonic Silicon Valley Lab, Panasonic R&D Company of America, Cupertino,