强化学习(reinforcement learning,简称RL),
agent
policy
state
action
目标
最大化累计reward
参考链接:
https://en.wikipedia.org/wiki/Reinforcement_learning
https://drive.google.com/file/d/1opPSz5AZ_kVa1uWOdOiveNiBFiEOHjkG/view