▲ 531 ▼ 2 in a single week that is crazy (lemmy.ml) submitted 2 years ago by Ascend910@lemmy.ml to c/memes@lemmy.ml 20 comments fedilink hide all child comments blob:https://phtn.app/bce94c48-9b96-4b8e-a4fd-e90166d56ed7
[–] Cort@lemmy.world 1 point 2 years ago (1 child) Would a 12g 3060 work? permalink fedilink source parent hideshow 2 child comments replies: [–] brucethemoose@lemmy.world 1 point 2 years ago* Yes! Try this model: https://huggingface.co/arcee-ai/Virtuoso-Small-v2 Or the 14B thinking model: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B But for speed and coherence, instead of ollama, I'd recommend running it through Aphrodite or TabbyAPI as a backend, depending if you prioritize speed or long inputs. They both act as generic OpenAI endpoints. I'll even step you through it and upload a quantization for your card, if you want, as it looks like there's not a good-sized exl2 on huggingface. permalink fedilink source parent
[–] brucethemoose@lemmy.world 1 point 2 years ago* Yes! Try this model: https://huggingface.co/arcee-ai/Virtuoso-Small-v2 Or the 14B thinking model: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B But for speed and coherence, instead of ollama, I'd recommend running it through Aphrodite or TabbyAPI as a backend, depending if you prioritize speed or long inputs. They both act as generic OpenAI endpoints. I'll even step you through it and upload a quantization for your card, if you want, as it looks like there's not a good-sized exl2 on huggingface. permalink fedilink source parent