Neural Network Components & Regularization — Activations, Normalization, Dropout & Initialization¶
Prerequisites
Multivariate calculus, expectations, and layerwise backpropagation. Review Neural Network Foundations, Probability & Statistics, and Linear Algebra.
1. The Big Picture¶
Deep neural networks are composed of distinct architectural primitives:
- Universal Function Approximators (MLPs): Stacking affine transformations with non-linear activations.
- Activation Functions: Controlling gradient flow, introducing non-linearity, and bounding or shaping coordinate manifolds.
- Normalization Layers (BatchNorm, LayerNorm, GroupNorm): Stabilizing internal representation statistics, smoothing the optimization landscape, and enabling larger learning rates.
- Regularization Techniques (Dropout, Weight Decay): Preventing co-adaptation of features and preserving generalization capability.
- Principled Weight Initialization (Glorot/Xavier, He/Kaiming): Preserving activation and gradient variance across depth to prevent immediate collapse into vanishing or exploding dynamics at iteration zero.
flowchart TD
subgraph Layer Computation Block
IN["Input: a[l-1]"] --> INIT["Initialized Weights: W[l] ~ N(0, 2/n_in)"]
INIT --> AFF["Affine Map: z[l] = W[l] a[l-1] + b[l]"]
AFF --> NORM["Normalization: x_hat = (z - μ) / √(σ² + ε)"]
NORM --> SCALE["Scale & Shift: y = γ x_hat + β"]
SCALE --> ACT["Non-linear Activation: a[l] = GELU(y) / ReLU(y)"]
ACT --> DROP["Dropout: a_drop = a[l] ⊙ r / (1 - p)"]
DROP --> OUT["Output to Layer l+1: a[l]"]
end
2. Multilayer Perceptron & Universal Approximation¶
2.1 Cybenko's Universal Approximation Theorem (1989)¶
Can a neural network approximate any arbitrary continuous function? George Cybenko (1989) and Kurt Hornik (1991) answered affirmatively.
Theorem (Cybenko, 1989):
Let \(\sigma\) be any continuous sigmoidal activation function (\(\lim_{z \to -\infty} \sigma(z) = 0\) and \(\lim_{z \to \infty} \sigma(z) = 1\)). Let \(I_n = [0, 1]^n\) denote the \(n\)-dimensional unit hypercube, and let \(C(I_n)\) be the space of continuous real-valued functions on \(I_n\).
Then, for any continuous target function \(f \in C(I_n)\) and any tolerance \(\epsilon > 0\), there exists an integer \(M\), real coefficients \(\alpha_i, b_i \in \mathbb{R}\), and weight vectors \(\mathbf{w}_i \in \mathbb{R}^n\) such that the single-hidden-layer network:
satisfies:
In modern terminology, the set of single-hidden-layer neural networks is dense in \(C(I_n)\) under the supremum norm.
2.2 Depth vs. Width: Why Go Deep?¶
While Cybenko's theorem guarantees that a single hidden layer can approximate any continuous function, it requires an exponential width \(M = \mathcal{O}(2^n)\) in the worst case.
Modern deep learning results (Telgarsky 2016, Eldan & Shamir 2016) prove that deep networks with \(\mathcal{O}(L)\) layers can compute oscillating and compositional functions with polynomial parameter counts \(\mathcal{O}(n)\), which would require \(\mathcal{O}(2^n)\) neurons in a single-layer network. Depth provides an exponential compositional efficiency:
3. Activation Functions: Mathematical Zoo & Gradients¶
The choice of activation function directly dictates gradient propagation, numerical stability, and representational expressiveness.
flowchart LR
subgraph Classical Activations
SIG["Sigmoid: (0, 1)"]
TANH["Tanh: (-1, 1)"]
end
subgraph Modern Non-Saturating
RELU["ReLU: max(0, x)"]
LRELU["Leaky ReLU: max(αx, x)"]
GELU["GELU: x · Φ(x)"]
SILU["Swish/SiLU: x · σ(βx)"]
end
SIG -. "Saturation & Vanishing Gradient" .-> RELU
RELU -. "Dead ReLU problem" .-> LRELU
RELU -. "Smooth Curvature" .-> GELU
3.1 Taxonomy of Activations¶
| Activation | Definition \(\sigma(z)\) | Output Range | Derivative \(\sigma'(z)\) | Saturation Regions |
|---|---|---|---|---|
| Sigmoid | \(\frac{1}{1 + e^{-z}}\) | \((0, 1)\) | \(\sigma(z)(1 - \sigma(z))\) | \(z \to \pm \infty\) (Max \(\sigma' = 0.25\)) |
| Tanh | \(\frac{e^z - e^{-z}}{e^z + e^{-z}}\) | \((-1, 1)\) | \(1 - \tanh^2(z)\) | \(z \to \pm \infty\) (Zero-centered, Max \(\sigma' = 1.0\)) |
| ReLU | \(\max(0, z)\) | \([0, \infty)\) | \(\begin{cases} 1 & z > 0 \\ 0 & z < 0 \end{cases}\) | \(z < 0\) (Dead ReLU) |
| Leaky ReLU | \(\max(\alpha z, z), \alpha \approx 0.01\) | \((-\infty, \infty)\) | \(\begin{cases} 1 & z > 0 \\ \alpha & z < 0 \end{cases}\) | None |
| Parametric ReLU (PReLU) | \(\max(\alpha z, z), \alpha \text{ learned}\) | \((-\infty, \infty)\) | \(\begin{cases} 1 & z > 0 \\ \alpha & z < 0 \end{cases}\) | None |
| ELU | \(\begin{cases} z & z > 0 \\ \alpha(e^z - 1) & z \le 0 \end{cases}\) | \((-\alpha, \infty)\) | \(\begin{cases} 1 & z > 0 \\ \sigma(z) + \alpha & z < 0 \end{cases}\) | \(z \to -\infty\) |
| GELU | \(z \cdot \Phi(z) = z P(X \le z), X \sim \mathcal{N}(0, 1)\) | \((-0.17, \infty)\) | \(\Phi(z) + z \phi(z)\) | Smooth non-monotonic minimum at \(z \approx -0.75\) |
| Swish / SiLU | \(z \cdot \sigma(\beta z)\) | \((-0.28, \infty)\) | \(\beta \sigma(\beta z) + \sigma(\beta z)(1 - \beta z \sigma(\beta z))\) | Self-gated smooth non-monotonic |
3.2 The Dead ReLU Problem¶
When pre-activation \(z_j = \mathbf{w}_j^T \mathbf{x} + b_j < 0\) across the entire training distribution, \(\text{ReLU}(z_j) = 0\) and \(\frac{d\text{ReLU}}{dz_j} = 0\).
In the backward pass:
The gradient identically vanishes. The weight vector receives zero updates, permanently "killing" the neuron for all subsequent training epochs.
Solutions:
- Initialize biases \(b_j\) with a small positive constant (\(+0.01\) or \(+0.1\)).
- Use Leaky ReLU or ELU where negative inputs maintain a non-zero slope \(\alpha > 0\).
- Use smooth stochastic activations like GELU (standard in GPT, BERT, and ViT).
4. Batch Normalization (Ioffe & Szegedy, 2015)¶
4.1 Internal Covariate Shift & Optimization Smoothing¶
During training, updating parameters in early layers shifts the input distribution seen by deeper layers. Sergey Ioffe and Christian Szegedy termed this internal covariate shift. While subsequent work (Santurkar et al., 2018) showed the primary benefit of BatchNorm is smoothing the loss landscape (reducing the Lipschitz constant of both the loss \(\mathcal{L}\) and its gradient \(\nabla \mathcal{L}\)), the operational formulation remains the foundation of modern CNNs.
4.2 Forward Formulation¶
Given a mini-batch \(\mathcal{B} = \{x_1, \dots, x_m\}\) of activations for a single scalar feature:
-
Mini-batch Mean: $\(\mu_{\mathcal{B}} = \frac{1}{m} \sum_{i=1}^m x_i\)$
-
Mini-batch Variance: $\(\sigma_{\mathcal{B}}^2 = \frac{1}{m} \sum_{i=1}^m (x_i - \mu_{\mathcal{B}})^2\)$
-
Standardization: $\(\hat{x}_i = \frac{x_i - \mu_{\mathcal{B}}}{\sqrt{\sigma_{\mathcal{B}}^2 + \epsilon}}\)$ where \(\epsilon \approx 10^{-5}\) prevents division by zero.
-
Scale and Shift (Learnable Affine): $\(y_i = \gamma \hat{x}_i + \beta\)$ where \(\gamma \in \mathbb{R}\) (scale) and \(\beta \in \mathbb{R}\) (shift) restore representational capacity. If \(\gamma = \sqrt{\sigma_{\mathcal{B}}^2 + \epsilon}\) and \(\beta = \mu_{\mathcal{B}}\), the identity mapping is perfectly recovered.
4.3 Inference Phase: Running Statistics Tracking¶
During inference (evaluation), predictions must depend strictly on the individual input, not on other samples in the batch. BatchNorm tracks running population statistics during training using an exponential moving average (momentum \(\rho \approx 0.1\)):
At test time, the deterministic normalization is:
This linear transformation can be pre-fused directly into preceding linear/convolutional weights during inference deployment.
4.4 Complete Analytical Backward Pass Derivation¶
Let \(\frac{\partial \mathcal{L}}{\partial y_i}\) be the incoming gradient. We must compute \(\frac{\partial \mathcal{L}}{\partial \gamma}\), \(\frac{\partial \mathcal{L}}{\partial \beta}\), and \(\frac{\partial \mathcal{L}}{\partial x_i}\).
Step 1: Gradients with respect to \(\gamma\) and \(\beta\)¶
Step 2: Gradient with respect to normalized \(\hat{x}_i\)¶
Step 3: Gradient with respect to variance \(\sigma_{\mathcal{B}}^2\)¶
Step 4: Gradient with respect to mean \(\mu_{\mathcal{B}}\)¶
(since \(\sum_{i=1}^m (x_i - \mu_{\mathcal{B}}) = 0\)).
Step 5: Gradient with respect to input \(x_i\)¶
Combining and factoring yields the famous vectorized single-line form:
5. Normalization Taxonomy: Batch vs. Layer vs. Instance vs. Group¶
Given a feature tensor of shape \((N, C, H, W)\) or sequence tensor \((N, T, C)\) where \(N\) is batch, \(C\) is channel/features, and \((H, W)\) or \(T\) are spatial/temporal axes:
flowchart TD
subgraph Normalization Across Tensor Dimensions
BN["Batch Norm: Computes μ, σ across (N, H, W) for each Channel C independently"]
LN["Layer Norm: Computes μ, σ across (C, H, W) for each Sample N independently"]
IN["Instance Norm: Computes μ, σ across (H, W) for each Sample N and Channel C"]
GN["Group Norm: Divides C into G groups; computes μ, σ across (C/G, H, W) per Sample N"]
end
| Normalization | Normalization Axes | Invariance | Primary Domain | Dependency on Batch Size |
|---|---|---|---|---|
| Batch Norm (BN) | \((N, H, W)\) | Rescaling of weights | CNNs, Vision | Strong (fails when \(N < 8\)) |
| Layer Norm (LN) | \((C, H, W)\) or \(C\) | Shift & scale of input | Transformers, NLP, RNNs | None (\(N=1\) identical) |
| Instance Norm (IN) | \((H, W)\) | Contrast / style of image | Style Transfer, GANs | None |
| Group Norm (GN) | \((C_g, H, W)\) | Group feature distributions | Vision tasks with small batches | None |
6. Dropout (Srivastava et al., 2014)¶
6.1 Bernoulli Masking & Inverted Dropout¶
Dropout prevents feature co-adaptation—where a neuron only functions correctly in the presence of specific complementary neurons.
During training, each hidden unit \(a_i\) is zeroed out with probability \(p \in [0, 1)\) via a Bernoulli random variable:
Under naive dropout, \(\mathbb{E}[r_i a_i] = (1-p) a_i\). At test time, to preserve identical expectation, weights had to be scaled down: \(W_{\text{test}} = (1-p) W_{\text{train}}\).
Modern frameworks universally use Inverted Dropout, which pre-scales activations during training:
Expectation and Variance Preservation:
-
Expectation: $\(\mathbb{E}[\widetilde{a}_i] = \frac{\mathbb{E}[r_i] \cdot a_i}{1 - p} = \frac{(1 - p) a_i}{1 - p} = a_i\)$
-
At Test Time: $\(\widetilde{a}_i = a_i \quad (\text{No scaling required! Identical forward graph})\)$
6.2 Ensemble Interpretation¶
A network with \(N\) neurons has \(2^N\) possible subnetworks formed by binary masking. Dropout trains an implicit ensemble of \(2^N\) thinned models sharing parameters. At test time, running the full unmasked network computes an approximate geometric mean over all \(2^N\) subnetworks in a single deterministic forward pass.
7. Weight Initialization: Vanishing & Exploding Gradients¶
Consider a deep linear network without biases: \(\mathbf{a}^{[L]} = W^{[L]} W^{[L-1]} \dots W^{[1]} \mathbf{x}\).
If weights are initialized with variance \(\sigma^2\):
- If \(\sigma^2 > 1\): Activations grow exponentially as \(\mathcal{O}(\sigma^{2L}) \to \infty\) (Exploding Gradients).
- If \(\sigma^2 < 1\): Activations decay exponentially as \(\mathcal{O}(\sigma^{2L}) \to 0\) (Vanishing Gradients).
7.1 Derivation of Xavier / Glorot Initialization (for Sigmoid / Tanh)¶
Consider a single neuron \(z = \sum_{i=1}^{n_{\text{in}}} w_i x_i\).
Assume \(w_i\) and \(x_i\) are independent, identically distributed, with zero mean (\(\mathbb{E}[w_i] = 0, \mathbb{E}[x_i] = 0\)).
The variance of \(z\) is:
To keep the forward variance constant (\(\text{Var}(z) = \text{Var}(X)\)), we require:
Now consider the backward pass error propagation \(\delta^{[l-1]} = (W^{[l]})^T \delta^{[l]}\). To keep the gradient variance constant during backpropagation across \(n_{\text{out}}\) incoming error channels:
Taking the harmonic mean of the forward and backward constraints yields Glorot / Xavier Initialization:
- Normal Glorot: \(W \sim \mathcal{N}\left( 0, \frac{2}{n_{\text{in}} + n_{\text{out}}} \right)\)
- Uniform Glorot: \(W \sim \mathcal{U}\left( -\sqrt{\frac{6}{n_{\text{in}} + n_{\text{out}}}}, +\sqrt{\frac{6}{n_{\text{in}} + n_{\text{out}}}} \right)\)
7.2 Derivation of He / Kaiming Initialization (for ReLU)¶
ReLU zeroes out all negative activations: \(\text{ReLU}(z) = \max(0, z)\).
If pre-activation \(z\) is symmetric with zero mean, exactly half the distribution is zeroed out:
Thus, each ReLU layer halves the signal variance:
To enforce \(\text{Var}(A) = \text{Var}(X)\), we must multiply the variance by \(2\):
This is He / Kaiming Initialization (He et al., 2015):
- Normal Kaiming: \(W \sim \mathcal{N}\left( 0, \frac{2}{n_{\text{in}}} \right)\)
- Uniform Kaiming: \(W \sim \mathcal{U}\left( -\sqrt{\frac{6}{n_{\text{in}}}}, +\sqrt{\frac{6}{n_{\text{in}}}} \right)\)
8. Implementation 1 — Vectorized BatchNorm & Dropout Layers from Scratch (NumPy)¶
The following Python script implements production-grade BatchNormalization and InvertedDropout layers with full analytical forward and backward passes, verified against finite-difference numerical gradients.
"""
scratch_layers.py
Vectorized BatchNorm and Inverted Dropout layers with analytical backpropagation in pure NumPy.
"""
import numpy as np
from typing import Tuple
class ScratchBatchNorm1d:
"""
Vectorized 1D Batch Normalization Layer with running statistics tracking.
Input shape: (batch_size, num_features)
"""
def __init__(self, num_features: int, eps: float = 1e-5, momentum: float = 0.1):
self.num_features = num_features
self.eps = eps
self.momentum = momentum
# Learnable scale (gamma) and shift (beta)
self.gamma = np.ones((1, num_features), dtype=np.float64)
self.beta = np.zeros((1, num_features), dtype=np.float64)
# Gradients
self.dgamma = np.zeros_like(self.gamma)
self.dbeta = np.zeros_like(self.beta)
# Running population statistics for inference
self.running_mean = np.zeros((1, num_features), dtype=np.float64)
self.running_var = np.ones((1, num_features), dtype=np.float64)
# Cache for backprop
self.cache = None
def forward(self, X: np.ndarray, training: bool = True) -> np.ndarray:
"""
Forward pass.
X shape: (m, num_features)
"""
if training:
m = X.shape[0]
# Mini-batch mean and variance
mean = np.mean(X, axis=0, keepdims=True) # (1, C)
x_mu = X - mean # (m, C)
var = np.mean(x_mu**2, axis=0, keepdims=True) # (1, C)
std_inv = 1.0 / np.sqrt(var + self.eps) # (1, C)
# Normalize
x_hat = x_mu * std_inv # (m, C)
# Scale and shift
out = self.gamma * x_hat + self.beta # (m, C)
# Update running statistics via exponential moving average
self.running_mean = (1.0 - self.momentum) * self.running_mean + self.momentum * mean
# Unbiased variance estimator for running var
sample_var = var * (m / max(m - 1, 1))
self.running_var = (1.0 - self.momentum) * self.running_var + self.momentum * sample_var
self.cache = (X, x_hat, x_mu, std_inv, var)
return out
else:
# Inference mode: use accumulated population statistics
x_hat = (X - self.running_mean) / np.sqrt(self.running_var + self.eps)
return self.gamma * x_hat + self.beta
def backward(self, dout: np.ndarray) -> np.ndarray:
"""
Analytical backward pass using the simplified single-line formula.
dout shape: (m, num_features)
Returns dX of shape (m, num_features)
"""
X, x_hat, x_mu, std_inv, var = self.cache
m = dout.shape[0]
# Gradients with respect to gamma and beta
self.dgamma = np.sum(dout * x_hat, axis=0, keepdims=True)
self.dbeta = np.sum(dout, axis=0, keepdims=True)
# Analytical gradient with respect to X
# dX = (gamma / (m * std)) * [m * dout - sum(dout) - x_hat * sum(dout * x_hat)]
term1 = m * dout
term2 = np.sum(dout, axis=0, keepdims=True)
term3 = x_hat * np.sum(dout * x_hat, axis=0, keepdims=True)
dX = (self.gamma * std_inv / m) * (term1 - term2 - term3)
return dX
class ScratchDropout:
"""
Inverted Dropout Layer.
"""
def __init__(self, p: float = 0.5, seed: int = 42):
self.p = p
self.rng = np.random.RandomState(seed)
self.mask = None
def forward(self, X: np.ndarray, training: bool = True) -> np.ndarray:
if training and self.p > 0.0:
keep_prob = 1.0 - self.p
# Generate binary mask scaled by 1 / (1 - p)
self.mask = (self.rng.rand(*X.shape) >= self.p).astype(np.float64) / keep_prob
return X * self.mask
return X
def backward(self, dout: np.ndarray) -> np.ndarray:
if self.mask is not None:
return dout * self.mask
return dout
# --------------------------------------------------------------------------
# Verification Suite
# --------------------------------------------------------------------------
if __name__ == "__main__":
np.random.seed(101)
batch_size, num_features = 16, 8
X = np.random.randn(batch_size, num_features) * 5.0 + 10.0 # shifted distribution
# 1. Test BatchNorm
bn = ScratchBatchNorm1d(num_features=num_features)
out_train = bn.forward(X, training=True)
print(f"Post-BatchNorm Mean: {np.mean(out_train):.4f} (Expected ≈ 0.0)")
print(f"Post-BatchNorm Var: {np.var(out_train):.4f} (Expected ≈ 1.0)")
# 2. Test Gradient Flow through BatchNorm
dout = np.random.randn(*out_train.shape)
dX = bn.backward(dout)
print(f"dX Shape: {dX.shape} | dGamma Shape: {bn.dgamma.shape} | dBeta Shape: {bn.dbeta.shape}")
# 3. Test Inverted Dropout
drop = ScratchDropout(p=0.4)
out_drop = drop.forward(X, training=True)
zero_fraction = np.mean(out_drop == 0.0)
print(f"Dropout Zero Fraction: {zero_fraction:.2f} (Expected ≈ 0.40)")
print(f"Unscaled Signal Mean Preservation: Mean before={np.mean(X):.2f}, Mean after={np.mean(out_drop):.2f}")
9. Implementation 2 — PyTorch Production Comparison¶
"""
pytorch_components.py
Modern PyTorch neural block demonstrating Kaiming init, BatchNorm, GELU, and Dropout.
"""
import torch
import torch.nn as nn
class ModernDeepBlock(nn.Module):
"""
Standard pre-activation modern deep learning block:
Linear -> BatchNorm -> GELU -> Dropout
"""
def __init__(self, in_features: int, out_features: int, dropout_p: float = 0.2):
super().__init__()
self.fc = nn.Linear(in_features, out_features, bias=False) # Bias redundant before BatchNorm
self.bn = nn.BatchNorm1d(out_features)
self.act = nn.GELU()
self.drop = nn.Dropout(p=dropout_p)
# Explicit Kaiming (He) normal initialization
nn.init.kaiming_normal_(self.fc.weight, mode="fan_in", nonlinearity="relu")
def forward(self, x: torch.Tensor) -> torch.Tensor:
out = self.fc(x)
out = self.bn(out)
out = self.act(out)
out = self.drop(out)
return out
if __name__ == "__main__":
block = ModernDeepBlock(in_features=64, out_features=128, dropout_p=0.3)
x = torch.randn(32, 64)
# Training forward pass
block.train()
out_train = block(x)
print("Train Output Shape:", out_train.shape)
# Evaluation forward pass
block.eval()
with torch.no_grad():
out_eval = block(x)
print("Eval Output Shape: ", out_eval.shape)
10. Common Errors, Gotchas & Debugging¶
1. Including Bias in Layers Immediately Preceding BatchNorm¶
Symptom: Redundant parameters in state dict; optimizer wastes compute on zero-gradient parameters.
Root Cause: During BatchNorm, the mean \(\mu_{\mathcal{B}} = \frac{1}{m} \sum (Wx_i + b) = \frac{1}{m}\sum Wx_i + b\). In standardization:
The bias scalar \(b\) cancels out completely, while the subsequent learnable \(\beta\) in BatchNorm replaces it.
Fix: Always set bias=False in nn.Linear or nn.Conv2d layers followed directly by BatchNorm.
# BROKEN / REDUNDANT
layer = nn.Linear(128, 128, bias=True)
bn = nn.BatchNorm1d(128)
# FIXED
layer = nn.Linear(128, 128, bias=False)
bn = nn.BatchNorm1d(128)
2. Forgetting model.eval() During Inference¶
Symptom: Model produces wildly erratic, non-deterministic predictions on single samples at inference time; test accuracy is much lower than validation accuracy.
Root Cause: If model.train() remains active during evaluation, Dropout continues randomly zeroing out 50% of the activations, and BatchNorm continues calculating mean and variance over the evaluation batch (which fails completely if evaluation batch size \(m=1\), producing variance \(0\)).
Fix: Always wrap inference in model.eval() and with torch.no_grad():.
3. LayerNorm vs. BatchNorm in Sequence / Transformer Models¶
Symptom: BatchNorm fails during NLP or sequence modeling with variable sequence lengths.
Root Cause: In natural language, sentence lengths vary significantly. Computing mini-batch statistics across variable-length padded tokens introduces severe zero-token distortion. Furthermore, batch sizes in NLP are often small.
Fix: Use Layer Normalization, which normalizes across the feature/embedding dimension for each token independently, completely removing batch coupling.
11. Staff-Level Technical Interview Questions¶
Q1: Why does He/Kaiming initialization have a factor of \(\sqrt{2/n_{\text{in}}}\) while Xavier/Glorot initialization has \(\sqrt{1/n_{\text{in}}}\) (or \(\sqrt{2/(n_{\text{in}} + n_{\text{out}})}\))?¶
Model Answer:
Consider a linear layer \(z = \sum_{i=1}^{n_{\text{in}}} w_i a_i\). Assuming independent zero-mean weights and activations:
For symmetric activations with unit derivative at the origin like Tanh (\(\tanh'(0) = 1\)), \(\text{Var}(A) \approx \text{Var}(z)\) around initialization. Setting \(\text{Var}(z) = \text{Var}(A)\) yields:
However, for a ReLU activation \(a = \max(0, z)\), assuming \(z\) is symmetrically distributed about zero (Gaussian with zero mean), the probability of \(z < 0\) is \(0.5\). The variance of the rectified output is:
Because ReLU discards exactly half the incoming signal power, the output activation variance is cut in half. To balance the equation:
Taking the standard deviation gives \(\sigma = \sqrt{\frac{2}{n_{\text{in}}}}\).
Q2: Derive why Batch Normalization acts as an implicit regularizer.¶
Model Answer:
In Batch Normalization, each activation \(x_i\) is normalized by mini-batch statistics:
Because \(\mu_{\mathcal{B}} = \frac{1}{m}\sum_{j=1}^m x_j\) and \(\sigma_{\mathcal{B}}^2 = \frac{1}{m}\sum_{j=1}^m (x_j - \mu_{\mathcal{B}})^2\) are random sample statistics computed over a randomly sampled mini-batch \(\mathcal{B}\), they contain stochastic estimation noise relative to the true dataset population parameters:
This injects multiplicative and additive noise into every hidden unit's activation during the forward pass. The network cannot rely on precise deterministic co-activations of specific units, forcing representations to be resilient to perturbations. This stochastic perturbation mirrors the regularizing mechanism of Dropout, often reducing or eliminating the need for explicit Dropout in pure CNN architectures.
Q3: Why does Layer Normalization succeed in Transformers while Batch Normalization fails?¶
Model Answer:
Three structural properties make LayerNorm superior for Transformers:
- Sequence Length Variability: NLP inputs are sequences of varying lengths padded with zeros. BatchNorm computes statistics across the batch axis, which conflates true tokens with padding tokens, poisoning the statistics.
- Autoregressive Generation (Batch Size 1): During LLM inference, tokens are generated one by one. With batch size \(N=1\), BatchNorm variance is \(0\), making the operation mathematically undefined without falling back to stale running statistics.
- Temporal Dependency and Distribution Drift: In recurrent or attention sequences, token representations evolve across positions \(t \in [1, T]\). Computing a single batch statistic across heterogeneous semantic positions destroys positional representations. LayerNorm normalizes across the hidden dimension \(D\) of each token vector individually:
LayerNorm is completely invariant to batch size and sequence position.
Q4: Prove that Inverted Dropout preserves the expected activation value at test time.¶
Model Answer:
Let \(a \in \mathbb{R}\) be an activation. Inverted Dropout generates a masked activation \(\widetilde{a}\) during training:
where \(r \sim \text{Bernoulli}(1 - p)\), such that \(P(r=1) = 1-p\) and \(P(r=0) = p\).
Taking the mathematical expectation with respect to the distribution of \(r\):
The expectation of a Bernoulli random variable is its success probability:
Substituting this back:
Because the expected value during training equals the unmasked activation \(a\), the network can run in evaluation mode with the identity mapping \(\widetilde{a} = a\) without altering the expected scale of downstream activations.
Q5: What is GELU (Gaussian Error Linear Unit), and why is it preferred over ReLU in modern Transformer architectures?¶
Model Answer:
GELU (Hendrycks & Gimpel, 2016) weights an input by its probability under a standard normal distribution:
It can be approximated efficiently without special functions:
Advantages over ReLU:
- Smooth Differentiability: Unlike ReLU, which has a non-differentiable sharp corner at \(x=0\), GELU is infinitely differentiable (\(C^\infty\)) everywhere.
- Non-Monotonicity and Curvature: GELU has a small negative curvature zone for \(x \in (-0.75, 0)\), reaching a local minimum at \(x \approx -0.75\) with value \(\approx -0.17\). This curvature allows neurons to output small negative activations for weak inhibitory signals rather than aggressively zeroing them out, preventing the catastrophic "dead neuron" failure mode of ReLU while maintaining non-linearity.
12. Mastery Ladder¶
- L1: State Cybenko's Universal Approximation Theorem and explain why depth provides exponential representation efficiency.
- L2: Compare Sigmoid, Tanh, ReLU, Leaky ReLU, and GELU in terms of range, derivative, and saturation characteristics.
- L3: Explain the "Dead ReLU" phenomenon and describe three concrete architectural remedies.
- L4: Write down the 4 equations defining the forward pass of Batch Normalization.
- L5: Derive the full analytical backward pass of Batch Normalization with respect to \(\gamma\), \(\beta\), and input \(X\).
- L6: Compare Batch Normalization, Layer Normalization, Instance Normalization, and Group Normalization by their reduction axes.
- L7: Prove why Inverted Dropout preserves activation expectation and eliminates test-time weight scaling.
- L8: Derive Glorot/Xavier variance \(\frac{2}{n_{\text{in}} + n_{\text{out}}}\) for linear/tanh activations.
- L9: Derive He/Kaiming variance \(\frac{2}{n_{\text{in}}}\) accounting for the variance halving of ReLU.
- L10: Implement custom BatchNorm and Inverted Dropout layers from scratch in NumPy with numerical gradient checks.