الأكاديمية ← أدلة تطبيقيةتوثيق رسمي · إرشاد عربي

تشغيل نماذج محلية على ماك

Run Local LLMs on Mac

متوسط7 دقائق قراءةالدرس 33 أسئلة✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

الذاكرة: ما يحتفظ به Hermes عنك بين المحادثات: تفضيلاتك، أسماء مشاريعك، قرارات سبق أن اتخذتها. من دونها تبدأ كل محادثة من الصفر وتعيد شرح نفسك. معها يكمل من حيث توقفتما. ستستعمل هنا hermes model وhermes prompt-size، والقراءة نحو 7 دقائق. انتبه: الذاكرة تكبر وتتّسخ. راجعها من وقت لآخر واحذف ما لم يعد صحيحًا، وإلا بنى الوكيل على معلومة قديمة.

6أقسام
11أمثلة برمجية
7جداول
2أوامر
1,258كلمة من المصدر
الوصف الرسمي في سطر

Set up a local OpenAI-compatible LLM server on macOS with llama.cpp or MLX, including model selection, memory optimization, and real benchmarks on Apple Silicon

ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تعرف ما الذاكرة ولماذا قد تحتاجه.
  • تنفّذ hermes model وhermes prompt-size وتفهم ما يحدث بعدها.
  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
  • تضبط HERMES_STREAM_READ_TIMEOUT في المكان الصحيح.
ما ستقابله من أسماء

كما تظهر تمامًا داخل Hermes.

الأوامر
  • hermes model
  • hermes prompt-size
متغيرات البيئة
  • HERMES_STREAM_READ_TIMEOUT
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01Choosing a model
  2. 02Option A: llama.cpp
  3. 03Option B: MLX via omlx
  4. 04Benchmarks: llama.cpp vs MLX
  5. 05Connect to Hermes
  6. 06Timeouts
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

This guide walks you through running a local LLM server on macOS with an OpenAI-compatible API. You get full privacy, zero API costs, and surprisingly good performance on Apple Silicon.

We cover two backends:

BackendInstallBest atFormat
llama.cppbrew install llama.cppFastest time-to-first-token, quantized KV cache for low memoryGGUF
omlxomlx.ai ↗Fastest token generation, native Metal optimizationMLX (safetensors)

Both expose an OpenAI-compatible /v1/chat/completions endpoint. Hermes works with either one — just point it at http://localhost:8080 or http://localhost:8000.

---

Choosing a model

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: ما يحتفظ به Hermes عنك بين المحادثات: تفضيلاتك، أسماء مشاريعك، قرارات سبق أن اتخذتها.

For getting started, we recommend Qwen3.5-9B — it's a strong reasoning model that fits comfortably in 8GB+ of unified memory with quantization.

VariantSize on diskRAM needed (128K context)Backend
Qwen3.5-9B-Q4_K_M (GGUF)5.3 GB~10 GB with quantized KV cachellama.cpp
Qwen3.5-9B-mlx-lm-mxfp4 (MLX)~5 GB~12 GBomlx

Memory rule of thumb: model size + KV cache. A 9B Q4 model is ~5 GB. The KV cache at 128K context with Q4 quantization adds ~4-5 GB. With default (f16) KV cache, that balloons to ~16 GB. The quantized KV cache flags in llama.cpp are the key trick for memory-constrained systems.

For larger models (27B, 35B), you'll need 32 GB+ of unified memory. The 9B is the sweet spot for 8-16 GB machines.

---

Option A: llama.cpp

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

llama.cpp is the most portable local LLM runtime. On macOS it uses Metal for GPU acceleration out of the box.

Install

Shellسطر واحد
brew install llama.cpp

This gives you the llama-server command globally.

Download the model

You need a GGUF-format model. The easiest source is Hugging Face via the huggingface-cli:

Shellسطر واحد
brew install huggingface-cli

Then download:

Shellسطر واحد
huggingface-cli download unsloth/Qwen3.5-9B-GGUF Qwen3.5-9B-Q4_K_M.gguf --local-dir ~/models

Start the server

Shell8 أسطر
llama-server -m ~/models/Qwen3.5-9B-Q4_K_M.gguf \
  -ngl 99 \
  -c 131072 \
  -np 1 \
  -fa on \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --host 0.0.0.0

Here's what each flag does:

FlagPurpose
-ngl 99Offload all layers to GPU (Metal). Use a high number to ensure nothing stays on CPU.
-c 131072Context window size (128K tokens). Reduce this if you're low on memory.
-np 1Number of parallel slots. Keep at 1 for single-user use — more slots split your memory budget.
-fa onFlash attention. Reduces memory usage and speeds up long-context inference.
--cache-type-k q4_0Quantize the key cache to 4-bit. This is the big memory saver.
--cache-type-v q4_0Quantize the value cache to 4-bit. Together with the above, this cuts KV cache memory by ~75% vs f16.
--host 0.0.0.0Listen on all interfaces. Use 127.0.0.1 if you don't need network access.

The server is ready when you see:

Textسطران
main: server is listening on http://0.0.0.0:8080
srv  update_slots: all slots are idle

Memory optimization for constrained systems

The --cache-type-k q4_0 --cache-type-v q4_0 flags are the most important optimization for systems with limited memory. Here's the impact at 128K context:

KV cache typeKV cache memory (128K ctx, 9B model)
f16 (default)~16 GB
q8_0~8 GB
q4_0~4 GB

On an 8 GB Mac, use q4_0 KV cache and choose a smaller model that can still fit Hermes' 64K minimum context. On 16 GB, you can comfortably do 128K context. On 32 GB+, you can run larger models or multiple parallel slots.

If you're still running out of memory, reduce context only while staying at or above Hermes' 64K minimum; otherwise switch to a smaller model or smaller quantization (Q3_K_M instead of Q4_K_M).

Test it

Shell7 أسطر
curl -s http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q4_K_M.gguf",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 50
  }' | jq .choices[0].message.content

Get the model name

If you forget the model name, query the models endpoint:

Shellسطر واحد
curl -s http://localhost:8080/v1/models | jq '.data[].id'

---

Option B: MLX via omlx

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

omlx ↗ is a macOS-native app that manages and serves MLX models. MLX is Apple's own machine learning framework, optimized specifically for Apple Silicon's unified memory architecture.

Install

Download and install from omlx.ai ↗. It provides a GUI for model management and a built-in server.

Download the model

Use the omlx app to browse and download models. Search for Qwen3.5-9B-mlx-lm-mxfp4 and download it. Models are stored locally (typically in ~/.omlx/models/).

Start the server

omlx serves models on http://127.0.0.1:8000 by default. Start serving from the app UI, or use the CLI if available.

Test it

Shell7 أسطر
curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-mlx-lm-mxfp4",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 50
  }' | jq .choices[0].message.content

List available models

omlx can serve multiple models simultaneously:

Shellسطر واحد
curl -s http://127.0.0.1:8000/v1/models | jq '.data[].id'

---

Benchmarks: llama.cpp vs MLX

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

Both backends tested on the same machine (Apple M5 Max, 128 GB unified memory) running the same model (Qwen3.5-9B) at comparable quantization levels (Q4_K_M for GGUF, mxfp4 for MLX). Five diverse prompts, three runs each, backends tested sequentially to avoid resource contention.

Results

Metricllama.cpp (Q4_K_M)MLX (mxfp4)Winner
TTFT (avg)67 ms289 msllama.cpp (4.3x faster)
TTFT (p50)66 ms286 msllama.cpp (4.3x faster)
Generation (avg)70 tok/s96 tok/sMLX (37% faster)
Generation (p50)70 tok/s96 tok/sMLX (37% faster)
Total time (512 tokens)7.3s5.5sMLX (25% faster)

What this means

  • llama.cpp excels at prompt processing — its flash attention + quantized KV cache pipeline gets you the first token in ~66ms. If you're building interactive applications where perceived responsiveness matters (chatbots, autocomplete), this is a meaningful advantage.
  • MLX generates tokens ~37% faster once it gets going. For batch workloads, long-form generation, or any task where total completion time matters more than initial latency, MLX finishes sooner.
  • Both backends are extremely consistent — variance across runs was negligible. You can rely on these numbers.

Which one should you pick?

Use caseRecommendation
Interactive chat, low-latency toolsllama.cpp
Long-form generation, bulk processingMLX (omlx)
Memory-constrained (8-16 GB)llama.cpp (quantized KV cache is unmatched)
Serving multiple models simultaneouslyomlx (built-in multi-model support)
Maximum compatibility (Linux too)llama.cpp

---

Connect to Hermes

أوامر تكتبها في الطرفية. افهم ما يفعله الأمر قبل نسخه. الأوامر هنا: hermes model.

Once your local server is running:

Shellسطر واحد
hermes model

Select Custom endpoint and follow the prompts. It will ask for the base URL and model name — use the values from whichever backend you set up above.

---

Timeouts

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير. الأوامر هنا: hermes prompt-size. تضبط HERMES_STREAM_READ_TIMEOUT خارج المحادثة، في بيئة التشغيل.

Hermes automatically detects local endpoints (localhost, LAN IPs) and relaxes its streaming timeouts. No configuration needed for most setups.

If you still hit timeout errors (e.g. very large contexts on slow hardware), you can override the streaming read timeout:

Shellسطران
# In your .env — raise from the 120s default to 30 minutes
HERMES_STREAM_READ_TIMEOUT=1800
TimeoutDefaultLocal auto-adjustmentEnv var override
Stream read (socket-level)120sRaised to 1800sHERMES_STREAM_READ_TIMEOUT
Stale stream detection180sDisabled entirelyHERMES_STREAM_STALE_TIMEOUT
API call (non-streaming)1800sNo change neededHERMES_API_TIMEOUT

The stream read timeout is the one most likely to cause issues — it's the socket-level deadline for receiving the next chunk of data. During prefill on large contexts, local models may produce no output for minutes while processing the prompt. The auto-detection handles this transparently.

اختبار الفهم

3 أسئلة إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. في جدول هذا الدرس، ما «Purpose» المقابل لـ«--cache-type-k q40»؟
2. أي متغير بيئة من التالي يظهر فعليًا في هذا الدرس؟
3. أي عنوان من التالي لا يظهر في هذا الدرس؟