Vector norms: L¹, L², max
A norm measures the size of a vector: intuitively, the distance from the origin to the point x. The Lᵖ norm, for p ≥ 1, is the p-th root of Σᵢ |xᵢ|ᵖ. Every norm is zero only for the zero vector, obeys the triangle inequality f(x + y) ≤ f(x) + f(y), and scales with |α|: f(αx) = |α| f(x).
The ones deep learning uses, worked for x = [3, −4]ᵀ:
- L² (Euclidean) norm: √(3² + 4²) = 5. Used so often it is written just ‖x‖.
- Squared L² norm: xᵀx = 25. Easier to work with, since each derivative depends only on one element, but it grows very slowly near the origin.
- L¹ norm: Σᵢ |xᵢ| = 3 + 4 = 7. It grows at the same rate everywhere, so it is used when telling exactly-zero entries from small nonzero ones matters.
- L∞ (max) norm: the largest absolute value, maxᵢ |xᵢ| = 4.