No AI summary available for this article.
Why It Matters
I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming.
Provenance
Discovered via Hacker News and published by GitHub.
Key Claims
Original description
I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift. It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next
Discovered via Hacker News
Community-ranked links and discussion from the HN front page.
Publisher: github.com
ID: 49524447 · Indexed about 3 hours ago