Learning objectives
What you will be able to explain
- Regularization changes the learning problem
- Parameter norm penalties
- Gradient of a weight-decay term
- Penalty versus constraint
- Regularization in under-constrained problems
- Dataset augmentation as prior knowledge
Section 01
Regularization changes the learning problem to Gradient of a weight-decay term
01
Regularization changes the learning problem
Guided checkpoint
Which definition best matches the chapter?
Source: Chapter 7, p. 228
02
Parameter norm penalties
Guided checkpoint
Match each penalty or effect to its description.
Source: Chapter 7, section 7.1, pp. 230-237
03
Gradient of a weight-decay term
Guided checkpoint
A regularized objective adds with . For a coordinate , what additive gradient contribution comes from the penalty?
Source: Chapter 7, section 7.1.1, pp. 231-234
Section 02
Penalty versus constraint to Dataset augmentation as prior knowledge
01
Penalty versus constraint
Guided checkpoint
Under suitable regularity conditions, what constrained problem corresponds to an explicit norm penalty?
Source: Chapter 7, section 7.2, pp. 237-239
02
Regularization in under-constrained problems
Guided checkpoint
Judge each statement.
Source: Chapter 7, section 7.3, pp. 239-240
03
Dataset augmentation as prior knowledge
Guided checkpoint
Which statements are correct?
Source: Chapter 7, section 7.4, pp. 240-242
Section 03
Noise as regularization to Early stopping protocol
01
Noise as regularization
Guided checkpoint
Match each injection site to the intended effect or interpretation.
Source: Chapter 7, sections 7.5 and 7.13, pp. 242-243 and 268-270
02
Semi-supervised and multi-task regularization
Guided checkpoint
Match each setup to the source of extra constraint.
Source: Chapter 7, sections 7.6-7.7 and 7.9, pp. 243-246 and 253-254
03
Early stopping protocol
Guided checkpoint
Which checkpoint should be retained when validation loss first improves and later deteriorates while training loss continues falling?
Source: Chapter 7, section 7.8, pp. 246-253
Section 04
Why early stopping regularizes to Variance of an averaged ensemble
01
Why early stopping regularizes
Guided checkpoint
Evaluate each claim.
Source: Chapter 7, section 7.8, pp. 246-253
02
Sparse parameters versus sparse representations
Guided checkpoint
Which distinction is correct?
Source: Chapter 7, section 7.10, pp. 254-256
03
Variance of an averaged ensemble
Guided checkpoint
Each of predictors has error variance 9 and pairwise error correlation 0.25. Assuming equal variance, compute the variance of their average.
Source: Chapter 7, section 7.11, pp. 256-258
Section 05
Why ensembles help to Inverted-dropout expectation
01
Why ensembles help
Guided checkpoint
Select all correct statements.
Source: Chapter 7, section 7.11, pp. 256-258
02
Dropout as model averaging
Guided checkpoint
Which interpretation best matches the chapter?
Source: Chapter 7, section 7.12, pp. 258-268
03
Inverted-dropout expectation
Guided checkpoint
An activation is , kept with probability , and divided by the keep probability when retained. What is its expected post-mask value during training?
Source: Chapter 7, section 7.12, pp. 258-268
Section 06
Tangent-based regularization
01
Tangent-based regularization
Guided checkpoint
Judge each statement.
Source: Chapter 7, section 7.14, pp. 270-273
Knowledge check
Turn understanding into recall.
The quiz now follows the same concepts in scored form. You can return to this lesson from the quiz whenever a gap appears.