← Review lesson

Quiz dossier / Reinforcement Learning

AlphaStar Deep Dive: The Nature Paper (Vinyals et al., 2019)

Question-by-question drill on the AlphaStar paper: the two-phase training pipeline, network architecture, RL update (TD(lambda), V-trace, UPGO, KL), league training with PFSP and exploiters, fairness constraints, and transfer to complex multi-agent systems.

22 questionsEstimated 45 min0 flagged

Session status

0%

Draft

Answered
0/22
Score
0/22

Question 1

1 point(s)

AlphaStar's training pipeline has two main phases. What are they, in order?