Files
mlx-knife/docs/ELECTRON_INTEGRATION_GUIDE.md
Nikesh Parajuli 6ea9bd33f5 patches
2026-02-24 22:19:59 -05:00

137 lines
5.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# MLX Knife Electron Integration Guide (macOS arm64)
This guide covers how our Electron app embeds the upstream `mlxk2` CLI and how we download models ourselves (without `pull`).
- Build the CLI once with `scripts/build.sh` (see `docs/ELECTRON_PACKAGING.md`).
- Extract the archive into `app.getPath('userData')/msty-mlx-studio` (or whatever `BIN_NAME` you built).
- Clear quarantine (`xattr -dr com.apple.quarantine <folder>`) on first install.
## CLI usage (list/show/run/server)
Wrap the CLI in a helper so we can spawn it with consistent env vars:
```js
// electron/main/mlxk.js
import { spawn } from 'child_process';
import { join } from 'path';
import { app } from 'electron';
const binRoot = join(app.getPath('userData'), 'msty-mlx-studio');
const exePath = join(binRoot, 'msty-mlx-studio');
function baseEnv(extra = {}) {
return {
...process.env,
HF_HOME: join(app.getPath('userData'), 'hf'),
TRANSFORMERS_NO_TORCH: '1',
TRANSFORMERS_NO_TF: '1',
TRANSFORMERS_NO_FLAX: '1',
HF_HUB_DISABLE_TELEMETRY: '1',
...extra,
};
}
export function spawnMlxk(args, options = {}) {
const { env = {}, stdio = ['ignore', 'inherit', 'inherit'] } = options;
return spawn(exePath, args, { env: baseEnv(env), stdio });
}
```
You can then build helpers such as:
```js
export function listModels({ all = false, verbose = false } = {}) {
const args = ['list'];
if (all) args.push('--all');
if (verbose) args.push('--verbose');
return new Promise((resolve, reject) => {
const child = spawnMlxk(args, { stdio: ['ignore', 'pipe', 'pipe'] });
const output = [];
child.stdout.on('data', (d) => output.push(d.toString()))
child.stderr.on('data', (d) => output.push(d.toString()))
child.on('exit', (code) => code === 0 ? resolve(output.join('')) : reject(code));
});
}
```
The same `spawnMlxk` wrapper works for `show`, `rm`, `run`, or starting the HTTP server (`serve`/`server`). The packaged binary runs the server in-process (single PID), so a single `SIGINT` stops it.
## Downloads handled in Node
Rather than calling the CLI `pull`, we download models from Hugging Face directly so we have full control over progress and caching. Use `MLXKModelDownloader` (see `apps/desktop/src/mlx/MLXKModelDownloader.ts`). It mirrors the existing `MLXModelDownloader` API:
```ts
import { MLXKModelDownloader } from './MLXKModelDownloader';
const downloader = new MLXKModelDownloader(app.getPath('userData'));
await downloader.downloadModel('mlx-community/Qwen3-4B-Instruct-2507-8bit', (progress) => {
// progress.fileName, progress.percentage, progress.speed, etc.
});
```
The downloader writes into the Hugging Face cache layout:
```
HF_HOME/hub/models--<owner>--<name>/snapshots/<commit>/...
HF_HOME/hub/models--<owner>--<name>/refs/main # contains <commit>
```
Once the files are in place, the CLI can list/run the model immediately.
### Cancellation
`MLXKModelDownloader.cancelDownloadForModel(modelId)` stops in-flight downloads, removes partial files, and cleans up empty directories.
## Environment for CLI processes
If you spawn the launcher directly from Electron, set these env vars:
- `MLXK_PYTHON` absolute path to embedded Python (e.g., `…/msty-mlx-studio/python/bin/python3`).
- `RESOURCES_PATH` bundle root (folder containing the launcher and `python/`).
- `HF_HOME` shared cache location for downloads and CLI usage.
- `TRANSFORMERS_NO_TORCH=1`, `TRANSFORMERS_NO_TF=1`, `TRANSFORMERS_NO_FLAX=1` faster start, no unwanted backend checks.
- `HF_HUB_DISABLE_TELEMETRY=1` avoid telemetry from Hugging Face Hub.
Note: the native launcher compiled by our release process sets `MLXK2_SUPERVISE=0` by default so uvicorn runs inprocess (no extra python supervisor). This improves stop behavior from Electron.
## Server quick start
```bash
HF_HOME="/Users/<user>/.myapp/hf" \
~/Library/Application\ Support/MyApp/msty-mlx-studio/msty-mlx-studio \
serve --host 127.0.0.1 --port 8000 --max-tokens 4000
```
Interact via the OpenAI-compatible endpoints:
```bash
curl -X POST "http://127.0.0.1:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
-d '{"model": "Phi-3-mini-4k-instruct-4bit", "messages": [{"role":"user","content":"Hello"}]}'
```
## Recommended server lifecycle from Electron
```js
// Start server
export function startServer({ host = '127.0.0.1', port = 8000, maxTokens } = {}) {
const args = ['serve', '--host', host, '--port', String(port)];
if (maxTokens != null) args.push('--max-tokens', String(maxTokens));
const child = spawnMlxk(args, { stdio: ['ignore', 'inherit', 'inherit'] });
return child; // caller holds this
}
// Stop server
export function stopServer(child) {
if (!child || child.killed) return;
try { child.kill('SIGINT'); } catch {}
}
```
Notes:
- The server runs inprocess in the packaged binary (no uvicorn supervisor). If `SIGINT` doesnt stop it (rare), escalate to `SIGKILL`.
- The CLI also supports `server` as an alias for `serve` for backward compatibility with v1.
Process model:
- You should see `msty-mlx-studio` as the parent process and a single `python` child while the server is running.
- Stopping the parent with `SIGINT`/`SIGTERM` stops the child promptly. A second signal escalates.
Thats the entire integration: package the CLI once, manage downloads yourself with `MLXKModelDownloader`, and spawn the CLI for listing/running or serving models.