openmic.social is an uncensored community. You may encounter strong language, controversial opinions, and mature or NSFW material. You must be 18+ to browse. Illegal content is prohibited and removed on sight — please report it. By continuing, you accept that you may see content you personally disagree with.
Even “big” open source models like DSV4 and Ling/Ring are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.
They run surprisingly well with hybrid CPU+GPU inference on desktops. And thats not even getting into the efficient attention mechanisms.
I can run DSV4 Flash, barely quantized, with ~1M context on my Ryzen desktop at ~11 tokens/s. If you told me that two years ago, I would not have believed you.