
1-bit Bonsai 1.7B Runs in Browser on WebGPU
A compact 290MB 1-bit quantized Bonsai 1.7B model now runs entirely locally in web browsers using WebGPU. The demo is hosted on Hugging Face Spaces by the webml-community. This enables lightweight LLM inference without server dependency.





