Learning objectives
What you will be able to explain
- Why the partition function is difficult
- Positive and negative phases
- Energy-model gradient facts
- Contrastive divergence and stochastic maximum likelihood
- Short-run learning caveats
- Pseudolikelihood avoids global normalization
Section 01
Why the partition function is difficult to Energy-model gradient facts
01
Why the partition function is difficult
Guided checkpoint
For , why can maximum likelihood be intractable?
Source: Chapter 18, pp. 605-606
02
Positive and negative phases
Guided checkpoint
Match each gradient term to its interpretation.
Source: Chapter 18, section 18.1, pp. 606-607
03
Energy-model gradient facts
Guided checkpoint
Judge each statement.
Source: Chapter 18, section 18.1, pp. 606-607
Section 02
Contrastive divergence and stochastic maximum likelihood to Pseudolikelihood avoids global normalization
01
Contrastive divergence and stochastic maximum likelihood
Guided checkpoint
Match each method to its chain initialization or persistence.
Source: Chapter 18, section 18.2, pp. 607-615
02
Short-run learning caveats
Guided checkpoint
Select all correct statements.
Source: Chapter 18, section 18.2, pp. 607-615
03
Pseudolikelihood avoids global normalization
Guided checkpoint
What objective does it maximize?
Source: Chapter 18, section 18.3, pp. 615-617
Section 03
Pseudolikelihood tradeoffs to Denoising score matching
01
Pseudolikelihood tradeoffs
Guided checkpoint
Evaluate each claim.
Source: Chapter 18, section 18.3, pp. 615-617
02
Score matching and ratio matching
Guided checkpoint
Match each idea to its defining object.
Source: Chapter 18, section 18.4, pp. 617-619
03
Denoising score matching
Guided checkpoint
What practical bridge does denoising create?
Source: Chapter 18, section 18.5, pp. 619-620
Section 04
Noise-contrastive estimation to Estimating a partition function
01
Noise-contrastive estimation
Guided checkpoint
Match each component to its role.
Source: Chapter 18, section 18.6, pp. 620-623
02
Class prior in NCE
Guided checkpoint
For every data sample, NCE draws 9 noise samples. What is the prior probability that a randomly selected training example for the discriminator came from data?
Source: Chapter 18, section 18.6, pp. 620-623
03
Estimating a partition function
Guided checkpoint
Match each method to its characteristic.
Source: Chapter 18, section 18.7, pp. 623-630
Section 05
Annealed importance sampling reasoning
01
Annealed importance sampling reasoning
Guided checkpoint
Judge each statement.
Source: Chapter 18, section 18.7, pp. 623-630
Knowledge check
Turn understanding into recall.
The quiz now follows the same concepts in scored form. You can return to this lesson from the quiz whenever a gap appears.