Improving Exploration in Policy Gradient Search
Date:
Workshop Talk at 1st Mathematical Reasoning in General Artificial Intelligence Workshop, ICLR 2021, Virtual
Workshop talk on improving exploration in policy-gradient search for symbolic optimization problems. When neural networks trained with reinforcement learning are used to search combinatorial spaces of mathematical expressions, the search can suffer from early commitment and initialization bias, both of which limit exploration. The talk introduced two remedies, a hierarchical entropy regularizer and a soft length prior, and showed that they improve performance and sample efficiency and yield simpler solutions on symbolic regression tasks.
A recording of the talk, the workshop page, and the accompanying paper are available below: Video · Workshop page · Paper (ArXiv)
