Width-Independent Compressibility of Deep Neural Networks

Chronological Source Flow
Back

AI Fusion Summary

Recent research explores deep neural networks through two distinct lenses. One study proves a uniform compressibility theorem for deep multilayer perceptrons with analytic activations, demonstrating that a narrow network can approximate a wide teacher network. This compressed width depends on the error budget and input dimension rather than the original width. Simultaneously, a new finite-width geometric framework utilizes commutators to quantify incompatibility among weight-generated covariance, gates, and backward sensitivities to describe learned feature geometries.
Community Comments
Loading updates...
0