I run a local LLM server on an M4 Pro Mac mini — here's my setup

My local model setup on an M4 Pro Mac Mini

I run a local LLM server on an M4 Pro Mac mini — here's my setup

A developer details how they run a local LLM server on an M4 Pro Mac mini with 48GB RAM, using Qwen3.6-35B-A3B and Gemma-4-E4B models via oMLX, with Tailscale for network access. They explain the benefits: cost predictability, latency, offline capability, no rate limits, and data privacy. The post also demystifies MoE model memory requirements and shows how to swap models easily.

Cloud APIs are rented land.
  1. bambax

    Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally?

    For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality.

    So I'm not completely convinced it's really worth it; but it's tempting!

  2. amanzi

    No mention of the performance of the models? I'm able to load a bunch of different models on my little mini-PC with 16GB RAM, but the performance is terrible. I always wonder what performance people are getting with local models that they find is acceptable?

  3. c16

    > Running a large model locally comes down to one thing: how much RAM it actually needs in memory.

    Not completely true. It's memory AND memory bandwidth. You can have 1tb of memory but if you have awful memory-bandwidth you'll also have slow tok/s. A3B helps with this, but so does MTP.

    From my experience, you'd be better off running the dense 27b-mlx with MTP than the 3.6 version with A3B. You say your model is ~20GB of ram, but the 3.8:27b-mlx is 18GB and gets me very reasonable tok/s, and greater speed if you disable thinking when not required.

  4. stub_out

    Oh man, an M4 Pro. My old M1 is really starting to show its age trying to run anything bigger than 7B.

  5. thrw93747572007

    Quite a lot of "local doesn't work" in here - unfortunately, often with not much details about what the people actually want to use their models for. Which I'd be curious about.

    I, personally, do use frontier models in the cloud for a lot of (meta-)cognitive analyses that are heavy enough to have me run against the limits of payed accounts regularly - so I'm neither a Luddite nor stingy with cash in this case.

    However: I have pretty good experiences with local models as well. My solid but hardly extreme desktop (with one RX 9070 XT 16GB) mostly serves gemma4:12b and specialized models (embedding) to my local network. This is for general use like simple queries, simple code, reformatting and the like but also for two specific tasks that are permanently running:

    a) It's connected to Home Assistant (as a second stage after very simple "turn light XY on" commands which get processed without LLM). So, I can mumble into my smartwatch "computer, how much gas do we have in the warp core and how much energy did the bussard collectors make from the cosmic dust today?" (or describe a more complex light scene or create an automation I want or whatever).

    The phone transcribes that - with a local model on device - and fires it to the desktop who has agentic access to HA, looks through the sensors and data, sees that I've tagged my solar panels and battery with nerd vocabulary. It makes the right conclusion, converts a few units and gives me back a nice overview. All hands-free while I'm s […]

  6. akg_67

    Recent performance data on my M1 Max 32GB MacBook using oMLX. I have been working on identifying suitable model and config for my use case and system. Using a refactor and suggest improvements prompt for a specific Django code block using VSCode Cline extension.

    ---

    Qwen3.8-27B-4bit, Prompt Processing (PP) 66.3 tok/s, Token Generation (TG) 11.8 tok/s

    Ornith-1.5-35B-A3B-MLX-4bit, PP 379.7, TG 45.8

    Ornith-1.5-35B-A3B-MLX-4bit, PP 381.5, TG 46.4

    Qwen3.6-35B-A3B-mxfp4, PP 389.6, TG 47.6

    Qwen3.6-35B-A3B-OptiQ-4bit, PP 342.6, TG 44.4

    ---

    Qwen3.8-27B-4bit generally runs out of output token before completing the task though excellent partial results.

    Ornith-1.5-35B-A3B-MLX-4bit seems to get in the loop often specially with tool calls.

    Qwen3.6-35B-A3B-mxfp4 seems to be optimal with speed and quality output.

    I am going to test Qwen3.6-35B-A3B-4bit soon with same code block just to check my intuition that any derivatives don't seem to perform better than the originals.

  7. amelius

    > The main reason to run local: cloud APIs are rented land. They can change their pricing, hit your usage limits, or swap the model being served behind the scenes whenever they feel like it.

    Yes but it's easy to replace them.

    The main reason should be privacy.

  8. jumploops

    My biggest problem with running local LLMs on my M4 Max/128GB RAM is the prefill latency.

    I've since acquired two DGX Sparks, and it feels so much snappier.

More from this day

2026-09-01