My local AI model was wasting VRAM by default, and changing one setting doubled its speed
When it comes to local AI, context matters for more than one reason
We haven't written up this one. XDA Developers has the full story — the link below goes straight to it.