Implementing sparse model loading in LM Studio: reducing VRAM usage by 40% with lottery ticket pruning for RTX 4000 series
Why I Started Looking at Sparse Loading I run LM Studio on my local RTX 4070 Ti with 12GB of VRAM. That's enough for most 7B models, but the moment I wanted to...