⬢github Python · 130 ★ +55 since we first saw it · pushed 155 d ago · NOASSERTION
SamsungLabs/LittleBit
Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)
LittleBit is Samsung Labs' official implementation of two research papers (NeurIPS 2025 and ICML 2026) that compress large language models below 1 bit per weight. It factorizes weight matrices into low-rank latent factors, binarizes them, and restores magnitudes via learned scales, achieving down to 0.1 bits per weight without changing model architecture at inference. It supports OPT, Llama, Phi-4, Qwen, Gemma, and others, with quantization-aware training.
Why now: The follow-up LittleBit-2 (ICML 2026) was recently published and its Joint-ITQ initialization was added to the repo, and the project drew attention on Hacker News for its sub-1-bit LLM compression approach.
Who it is for: ML researchers and engineers working on extreme LLM quantization who want to reproduce or build on published compression results.
Stars over our 30 snapshots: 75 to 130, since 7 h ago.
Where people talked about it
API: https://socialmediatrends-api.osmike.com/v1/repos/SamsungLabs/LittleBit