Skip to content

Commit a0b89cd

Browse files
committed
docs: add AEGIS_INTEGRATION.md with complete Aegis-AI sidecar setup guide
1 parent 86eac8a commit a0b89cd

1 file changed

Lines changed: 206 additions & 0 deletions

File tree

‎AEGIS_INTEGRATION.md‎

Lines changed: 206 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,206 @@
1+
# 🛡️ Aegis-AI Integration Guide
2+
3+
`mlx-server` is designed to be a **completely transparent, drop-in replacement** for `llama-server` or any cloud VLM gateway within Aegis-AI, delivering dramatically faster zero-latency inference on Apple Silicon.
4+
5+
---
6+
7+
## 🚀 Quick Start for Aegis-AI
8+
9+
### 1. Download the Binary
10+
11+
Download the latest pre-built binary from the [Releases page](https://github.com/SharpAI/mlx-server/releases) — no Xcode required:
12+
13+
```bash
14+
# Extract and make executable
15+
tar -xzf mlx-server-*-macos-arm64.tar.gz
16+
chmod +x mlx-server
17+
```
18+
19+
### 2. Point Aegis-AI at the Server
20+
21+
In your `~/.aegis-ai/llm-config.json`, set the base URL to the mlx-server endpoint:
22+
23+
```json
24+
{
25+
"provider": "local",
26+
"baseUrl": "http://127.0.0.1:5413/v1",
27+
"model": "mlx-community/Qwen2.5-7B-Instruct-4bit"
28+
}
29+
```
30+
31+
### 3. Launch the Sidecar
32+
33+
Aegis-AI should spin up `mlx-server` as a managed subprocess:
34+
35+
```bash
36+
/path/to/mlx-server \
37+
--model mlx-community/Qwen2.5-7B-Instruct-4bit \
38+
--host 127.0.0.1 \
39+
--port 5413
40+
```
41+
42+
The server will emit a machine-readable JSON ready event on stdout when it is ready to accept connections:
43+
44+
```json
45+
{"event":"ready","port":5413,"model":"mlx-community/Qwen2.5-7B-Instruct-4bit","engine":"mlx","vision":false}
46+
```
47+
48+
Aegis-AI should **wait for this event** before routing any requests to the server.
49+
50+
---
51+
52+
## 🧠 Running 122B+ MoE Models (Critical)
53+
54+
If you are running a Mixture of Experts (MoE) model — such as `Qwen3.5-122B-A10B` — you **must** pass the `--stream-experts true` flag.
55+
56+
```bash
57+
/path/to/mlx-server \
58+
--model mlx-community/Qwen3.5-122B-A10B-4bit \
59+
--host 127.0.0.1 \
60+
--port 5413 \
61+
--stream-experts true
62+
```
63+
64+
> [!CAUTION]
65+
> **Without `--stream-experts true` on MoE models**, macOS will suffer a `Data Abort` kernel-level memory mapping fault when it attempts to load >100GB of weight tensors into Unified Memory simultaneously. The entire machine will freeze and require a hard reboot.
66+
67+
### Why `--stream-experts` Works
68+
69+
MoE models like Qwen3.5-122B have 122B *total* parameters, but only ~10B are **active** on any single forward pass. `mlx-server` exploits this sparsity:
70+
71+
- The 60GB+ of expert weight matrices are `mmap`'d directly from your NVMe SSD
72+
- Only the **2-4 specific expert shards** selected by the router for the current token (~1.5MB each) are streamed into GPU RAM via a zero-copy DMA path
73+
- The remaining experts stay on disk — never touching Unified Memory
74+
75+
The result: a 122B model running stably in ~21GB of RAM on a 64GB M5 Pro.
76+
77+
### Time-To-First-Token (TTFT) Expectations
78+
79+
Due to SSD streaming, TTFT is higher than a fully in-memory model. This is **expected and normal**:
80+
81+
| Prompt Length | Expected TTFT |
82+
|---|---|
83+
| Short (~100 tokens) | 5–15 seconds |
84+
| Medium (~500 tokens) | 30–60 seconds |
85+
| Long (1000+ tokens) | 1–3 minutes |
86+
87+
> [!TIP]
88+
> **Aegis-AI Prompt Cache**: `mlx-server` automatically caches the KV state for repeated system prompts. After the first request with a given system prompt, subsequent requests with the same system prompt will skip the expensive prefill phase and start streaming almost immediately.
89+
90+
---
91+
92+
## 📡 API Reference
93+
94+
`mlx-server` is **fully OpenAI-compatible** — any client using the OpenAI SDK works without modification.
95+
96+
### Endpoints
97+
98+
| Endpoint | Method | Description |
99+
|---|---|---|
100+
| `/health` | `GET` | Server health, GPU memory stats, active request count |
101+
| `/v1/models` | `GET` | List loaded models (OpenAI format) |
102+
| `/v1/chat/completions` | `POST` | Chat completions — streaming and non-streaming |
103+
| `/v1/completions` | `POST` | Legacy text completions |
104+
| `/metrics` | `GET` | Prometheus-compatible metrics |
105+
106+
### Health Check
107+
108+
The `/health` endpoint returns detailed telemetry useful for Aegis-AI's system monitor:
109+
110+
```bash
111+
curl http://127.0.0.1:5413/health
112+
```
113+
114+
```json
115+
{
116+
"status": "ok",
117+
"model": "mlx-community/Qwen3.5-122B-A10B-4bit",
118+
"memory": {
119+
"active_mb": 21272,
120+
"peak_mb": 23500,
121+
"cache_mb": 4096,
122+
"total_system_mb": 65536,
123+
"gpu_architecture": "Apple M5 Pro"
124+
},
125+
"stats": {
126+
"requests_total": 42,
127+
"requests_active": 1,
128+
"tokens_generated": 18500,
129+
"avg_tokens_per_sec": 3.2
130+
}
131+
}
132+
```
133+
134+
### Streaming Chat Completion
135+
136+
```bash
137+
curl http://127.0.0.1:5413/v1/chat/completions \
138+
-H "Content-Type: application/json" \
139+
-d '{
140+
"model": "mlx-community/Qwen3.5-122B-A10B-4bit",
141+
"stream": true,
142+
"messages": [
143+
{"role": "system", "content": "You are Aegis-AI, a local home security agent. Always respond in JSON."},
144+
{"role": "user", "content": "Is the person in this clip a delivery courier?"}
145+
]
146+
}'
147+
```
148+
149+
---
150+
151+
## ⚙️ Full CLI Reference
152+
153+
| Flag | Default | Description |
154+
|---|---|---|
155+
| `--model` | *(required)* | HuggingFace model ID or absolute local path |
156+
| `--port` | `5413` | Port to listen on |
157+
| `--host` | `127.0.0.1` | Host interface to bind |
158+
| `--max-tokens` | `2048` | Max generation tokens per request |
159+
| `--ctx-size` | *model default* | KV cache context window size |
160+
| `--temp` | `0.6` | Default sampling temperature (0 = greedy) |
161+
| `--top-p` | `1.0` | Nucleus sampling threshold |
162+
| `--stream-experts` | `false` | **Enable SSD streaming for MoE models** |
163+
| `--thinking` | `false` | Enable reasoning/thinking mode (Qwen3 etc.) |
164+
| `--vision` | `false` | Enable VLM mode for image inputs |
165+
| `--parallel` | `1` | Number of concurrent request slots |
166+
| `--api-key` | *none* | Enable bearer token auth |
167+
| `--cors` | *none* | Allowed CORS origin (`*` for all) |
168+
| `--gpu-layers` | `auto` | Number of layers to run on GPU |
169+
| `--mem-limit` | *system default* | Hard GPU memory cap in MB |
170+
| `--prefill-size` | `512` | Prefill chunk size (lower if GPU watchdog triggers) |
171+
| `--info` | `false` | Dry-run memory profiling report and exit |
172+
173+
---
174+
175+
## 🔍 Memory Behaviour Explained
176+
177+
On Apple Silicon, GPU and system RAM are the **same physical chips** (Unified Memory Architecture). `mlx-server` uses a layered strategy to fit the largest possible models:
178+
179+
| Model Size vs. RAM | Strategy | Notes |
180+
|---|---|---|
181+
| Fits in RAM (<85%) | `full_gpu` | All layers on GPU, maximum speed |
182+
| Slightly over RAM | `swap_assisted` | macOS swap used, 2-4× slowdown |
183+
| 2-4× over RAM | `layer_partitioned` | GPU/CPU split, use `--gpu-layers` |
184+
| MoE > 2× RAM | `ssd_stream` | Use `--stream-experts true` |
185+
186+
You can always inspect the computed memory plan before loading a model:
187+
188+
```bash
189+
mlx-server --model mlx-community/Qwen3.5-122B-A10B-4bit --info
190+
```
191+
192+
---
193+
194+
## 📋 Requirements
195+
196+
- macOS 14.0+
197+
- Apple Silicon (M1 / M2 / M3 / M4 / M5)
198+
- Xcode Command Line Tools (for source builds only)
199+
200+
---
201+
202+
## 🔗 Resources
203+
204+
- [Main README](./README.md) — general usage and benchmarks
205+
- [GitHub Releases](https://github.com/SharpAI/mlx-server/releases) — pre-built binaries
206+
- [mlx-swift](https://github.com/ml-explore/mlx-swift) — underlying MLX framework

0 commit comments

Comments
 (0)