Academy → Practical guidesOfficial documentation · clear explanation

Run Local LLMs on Mac

Run Local LLMs on Mac

Beginner12 minutes3 questions2026-08-09
The idea in one minute

Start with meaning, then move to detail.

This lesson explains Run Local LLMs on Mac as part of getting started with Hermes correctly. You will learn what it does, when it matters, and the smallest safe test that proves it works.

If you are new

If you are new, do not memorize names. Focus on three questions: what problem does this solve, what access does it need, and how can you verify the result?

For hands-on use

For practice, inspect the first example, identify its effects, run it on test data, and compare the result with the source claim.

For specialists

For advanced readers, inspect Choosing a model, Option A: llama.cpp, Install, then verify failure modes and version compatibility.

What do you need first?

No prior experience is required; follow the steps on a safe test setup first.

What will you know?

A clear outcome before you read.

  • Understand Run Local LLMs on Mac without assumed prior knowledge.
  • Separate the source description from what still needs testing in your environment.
  • Read the first command and identify its inputs and outputs before copying it.
Lesson terms

Short definitions before the details.

Provider
The service that runs or provides access and authentication to a model.
Session & memory
A session holds conversation context, while memory keeps selected facts that should persist.
Official page description

Set up a local OpenAI-compatible LLM server on macOS with llama.cpp or MLX, including model selection, memory optimization, and real benchmarks on Apple Silicon

Topic map

What does the source say, and in what order?

  1. 01
    Choosing a model

    Start here to understand the core idea or structure.

  2. 02
    Option A: llama.cpp

    Read this after the foundation, then connect it to the previous step.

  3. 03
    Install

    Read this after the foundation, then connect it to the previous step.

  4. 04
    Download the model

    Read this after the foundation, then connect it to the previous step.

  5. 05
    Start the server

    Read this after the foundation, then connect it to the previous step.

  6. 06
    Memory optimization for constrained systems

    Read this after the foundation, then connect it to the previous step.

  7. 07
    Test it

    Read this after the foundation, then connect it to the previous step.

  8. 08
    Get the model name

    Read this after the foundation, then connect it to the previous step.

  9. 09
    Option B: MLX via omlx

    Read this after the foundation, then connect it to the previous step.

  10. 10
    List available models

    Finish here to verify the result and special cases.

Examples from the official page

Copy only after you understand the effect.

brew install llama.cpp
brew install huggingface-cli
huggingface-cli download unsloth/Qwen3.5-9B-GGUF Qwen3.5-9B-Q4_K_M.gguf --local-dir ~/models
Try it now

Read the first command and identify its inputs and outputs before copying it.

Match every command to your installed Hermes version, review the files and accounts it can reach, and use non-sensitive data for the first test. If this explanation differs from the source, the official source wins.

Knowledge check

Three decisions before completion.

1. What is the source of truth when “Run Local LLMs on Mac” changes?
2. What is the best way to apply this lesson?
3. What should happen before a step can modify files or an external account?