For a linear-quadratic tracking problem with a target trajectory described by a linear dynamics, convergence of reinforcement learning using policy iteration method is investigated. A proof of convergence is given, and convergence rate is analyzed. Namely, it is proven that the policy iteration method converges and has a quadratic convergence rate under relaxed constraints on the initial control.
