Project DelphiTensors Workshop

Decomposition inside neural networks

A dense convolution kernel of shape 3 × 3 × 512 × 512 holds 2,359,296 weights, and costs about 462 million multiply-adds on a 14 × 14 feature map.

A CP factorization at rank 64 stores 64 × (3 + 3 + 512 + 512) = 65,920 weights — 35.8× fewer — and runs as four skinny convolutions in sequence: 1×1 → 3×1 → 1×3 → 1×1.

One order up, a transformer output matrix W_O of shape 4096 × 4096 stores 16,777,216 weights; a TT-matrix representation at rank 16 stores 34,816, or 481.9× fewer.

The storage ratio is the easy half.