I believe model distillation does this. SparseGPT was a big one, managing to rem...

		Hugsun on May 14, 2024 \| parent \| context \| favorite \| on: Automatically Detecting Under-Trained Tokens in La... I believe model distillation does this. SparseGPT was a big one, managing to remove 50% of parameters without loosing much accuracy IIRC. I saw a more recent paper citing the SparseGPT one that managed around 70-80% sparsity, pretty impressive stuff.