Urgent.News

What's breaking now, across thousands of outlets.

AI

Picking Models as a Mac User

After spending the past two weeks redoing all the models around the house, I realized it might make a good topic to chat about. I know that everyone and their brother has their own way to figure out what models they want to run on their hardware, but I figure that my own criteria might help some of the Mac users out there, so I'm tossing it into the mix as well. Picking which models to even…

I recently spent a couple of weeks replacing models around my house, which led me to think that sharing my model selection criteria could be helpful for Mac users. When a new model launches, I start by checking forums and comments to gauge real-world performance. Benchmarks are useful, but seeing how people actually use the model is crucial.

I look for issues with tokenizers, bug-bitten implementations, and any overthinking or hallucinations. I also check Artificial Analysis, which has benchmarks aligned with my needs, such as context reasoning, hallucination rate, and output tokens. On a Mac with limited VRAM, I'm particularly conscious of models generating many tokens, as it can slow things down significantly.

Qwen3.8 27B is a prime example - while it matches Opus 4.6 in benchmark performance, its high token generation needs make it less practical unless you're okay with waiting for long response times. I compared a few smaller models suitable for M2 Ultra and M5 Max Macs: Qwen3.8-27B offers impressive performance but at a high token cost; Muse Glimmer High provides good reasoning and output tokens at 48M for 30B; Gemma 4 31B excels in context and tone, running well with reasoning off; GLM-5.3-Flash has a high intelligence score but fewer active parameters than some others.

For my M3 Ultra 512GB, MiniMax M3 stands out as the top choice despite having fewer active parameters compared to some alternatives. While higher-scoring models like GLM-5.3-Flash or GLM-5.2 Max might seem appealing based solely on intelligence, MiniMax M3's balanced performance across all metrics makes it the best fit for my needs, especially considering my preference for keeping output token usage low.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

The agent that refuses to guess

I work at a B2B telecom consultancy. I'm not the one auditing the bills, but every month I watch how it's done: open the invoice PDF, check every line against the signed contract, compare it with what…

More from Monday 31 August →