Discovering Symbolic Policies with Deep Reinforcement Learning
Date:
Conference Talk at The 38th International Conference on Machine Learning (ICML 2021), Virtual
Spotlight presentation on discovering interpretable symbolic policies with deep reinforcement learning. The approach, deep symbolic policy, uses an autoregressive recurrent neural network trained with a risk-seeking policy gradient to generate control policies written as short mathematical expressions, and scales to multi-dimensional action spaces through an “anchoring” procedure that distills pre-trained neural policies one action dimension at a time. Across eight benchmark control environments, the discovered symbolic policies outperformed seven state-of-the-art deep reinforcement learning algorithms in average rank and normalized reward despite their dramatically reduced complexity.
The full paper, spotlight recording, and PDF are available below: Proceedings · Spotlight · PDF
