Provably Efficient QQ-learning with Function Approximation via Distribution Shift Error Checking Oracle
Simon S. Du Thanks: Institute for Advanced Study, Email: Yuping Luo Thanks: Princeton University, Email: Ruosong Wang Thanks: Carnegie Mellon University, Email: Hanrui Zhang Thanks: Duke University, Email:
Abstract
$Q$ -learning with function approximation is one of the most popular methods in reinforcement learning. Though the idea of using function approximation was proposed at least $60$ years ago [28], even in the simplest setup, i.e, approximating $Q$ -functions with linear functions, it is still an open problem how to design a provably efficient algorithm that learns a near-optimal policy. The key challenges are how to efficiently explore the state space and how to decide when to stop exploring in conjunction with the function approximation scheme.
原文 arXiv:1906.06321;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1906.06321v2