Connection between KL Divergence & Cross Entropy
Why cross-entropy is the loss for probabilistic classifiers: the additive identity H(P,Q)=H(P)+KL(P||Q), the coin-flip derivation, the bridge to negative log-likelihood and maximum likelihood, binary/multiclass/soft-label forms, and stable logit-based implementation.