LINEAR ALGEBRA / 8. SVD & DECOMPOSITIONS
SVD & Matrix Decompositions
The most useful matrix factorization in all of data science
EXPLANATION
Singular Value Decomposition (SVD) decomposes any matrix A into: A = U Σ Vᵀ Where: • U (m×m) → left singular vectors (orthonormal), column space • Σ (m×n) → diagonal matrix of singular values σ₁ ≥ σ₂ ≥ ... ≥ 0 • Vᵀ (n×n) → right singular vectors (orthonormal), row space Geometric interpretation: any linear transformation can be broken into: 1. Rotation (Vᵀ) 2. Scaling (Σ) 3. Another rotation (U) SVD applications in ML: • PCA: eigenvectors of XᵀX = right singular vectors of X • Low-rank approximation: keep only top-k singular values • Recommender systems: matrix factorization • Solving overdetermined systems: pseudoinverse A⁺ = V Σ⁺ Uᵀ • Measuring matrix stability: condition number = σ_max/σ_min • Compressing neural network weights: LoRA fine-tuning uses low-rank decomposition! Truncated SVD (rank-k approximation): A ≈ UₖΣₖVₖᵀ (keep only k largest singular values) This gives the best rank-k approximation in terms of Frobenius norm.
DIAGRAM
SVD: A(m×n) = U(m×m) Σ(m×n) Vᵀ(n×n)
For A(3×2):
U(3×3) · Σ(3×2) · Vᵀ(2×2)
[rotates] [scales] [rotates]
Σ = [[σ₁, 0 ],
[0, σ₂], σ₁ ≥ σ₂ ≥ 0
[0, 0 ]]
Low-rank approximation (rank-1):
A ≈ σ₁ · u₁ · v₁ᵀ (outer product)
Each additional term adds more detail
LoRA (Low-Rank Adaptation) for LLMs:
ΔW = A · B where A:(d×r), B:(r×d), r << d
Only fine-tune the low-rank factors A and B
Instead of updating the full weight matrix WCODE