temporary
英 ['temp(ə)rərɪ]美 [ˈtempəreri]
adj. 临时的,暂时的;短暂的
n. 临时工,临时雇
TD算法是RL的核心算法。TD是DP和MC算法的结合。Like DP, TD methods without waiting for a final outcome (they bootstrap)。
TD(0), or one-step TD
Advantages of TD Prediction Methods
TD methods update their estimates based in part on other estimates. They learn a guess from a guess,they bootstrap.
Q-learning: Off-policy TD Control