Back to All Projects
Open Source Verified Architecture
laya-mlx
laya-mlx is a native MLX runtime that executes Laya‑typed decision models directly on Apple Silicon, delivering sub‑15 ms inference for short, rule‑based decisions. It skips text generation, PyTorch, and any cloud API, keeping the stack lightweight and ideal for on‑device, low‑latency applications.
6,838 Stars 551 Forks
Core Technologies & Frameworks
Python
Technical Architecture & Specifications
If you are building low-latency decision workflows on Apple Silicon, LLM text generation is usually an expensive bottleneck. `laya-mlx` takes a fundamentally different approach: it is a native MLX runtime designed specifically for Laya typed decision models. Instead of generating autoregressive text, pulling in PyTorch/Transformers dependencies, or invoking cloud APIs, `laya-mlx` processes structured decision schemas locally on Apple Unified Memory.
### How It Works Under the Hood
The core performance gain comes from generating 0 output tokens. Standard LLM structured outputs require predicting tokens sequentially until a valid JSON structure forms. `laya-mlx` bypasses token generation entirely, evaluating typed queries in a single direct forward pass.
Running natively inside Apple's MLX framework, the runtime leverages local metal optimizations, graph compilation, and prefix reuse. Tested benchmarks on an M3 Max (FP16) reveal ultra-fast end-to-end execution:
* Laya 421M (English): 13.42 ms median (P50) end-to-end latency for short decisions.
* Laya Multilingual 322M: 7.39 ms median (P50) end-to-end latency.
When enabling graph compilation and prefix-reuse paths, tests logged up to 75.40 moves/s across 2,400 moves in the built-in Snake decision loop, outperforming eager execution modes by ~6.5%.
### Environment & Requirements
* Hardware: Apple Silicon
* OS: macOS 14+ (Tested on macOS 27.2)
* Python: Python 3.11+ (Tested on 3.12.13 with MLX 0.32.2)
### Installation and Setup
Install the standard package via `pip`:
```bash
pip install laya-mlx
```
To run the interactive CLI terminal demo (requires a terminal window of at least 104 × 35 cells), install the demo extra and pull the multilingual weights:
```bash
pip install 'laya-mlx[demo]'
hf download aac6fef/laya-multilingual-mlx
laya-snake
```
To run the optimized execution path without artificial frame-pacing:
```bash
laya-snake --optimize --max-speed
```
### Programmatic Usage
You define structured schemas using standard Python dictionaries containing field definitions, decision types, instructions, and concrete selection criteria.
```python
import laya_mlx as laya
agent = laya.load("aac6fef/laya-mlx")
result = agent.predict(
"I was billed twice. Please refund the duplicate.",
{
"department": {
"type": "choice",
"instructions": "Who should handle this?",
"criteria": ["billing", "technical", "sales"],
}
},
)
print(result["answers"]["department"])
```
### Tech Review Verdict
`laya-mlx` provides sub-15 ms decision processing by completely stripping out autoregressive generation overhead. If your local Apple Silicon workflows need deterministic, structured classification—like routing requests or executing game loops—without installing heavy PyTorch runtimes, `laya-mlx` delivers an exceptionally tight, native execution pipeline.
Reviewed by DevTechPulse Editorial Board
All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.