openmic.social is an uncensored community. You may encounter strong language, controversial opinions, and mature or NSFW material. You must be 18+ to browse. Illegal content is prohibited and removed on sight — please report it. By continuing, you accept that you may see content you personally disagree with.
https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.
Upgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.
I plan to once there is a version with turboquant and MTP as that huge context window is key.
A used 3090 is like 2-3k though, IF you can find one :| A month of claude is like 20 EUR. A month of opencode go is half that, but you get less usage.