Skip to main content
Back to All Projects
Open Source Verified Architecture

laya-mlx

laya-mlx is a native MLX runtime that executes Laya‑typed decision models directly on Apple Silicon, delivering sub‑15 ms inference for short, rule‑based decisions. It skips text generation, PyTorch, and any cloud API, keeping the stack lightweight and ideal for on‑device, low‑latency applications.

6,838 Stars 551 Forks

Core Technologies & Frameworks

Python

Technical Architecture & Specifications

If you are building low-latency decision workflows on Apple Silicon, LLM text generation is usually an expensive bottleneck. `laya-mlx` takes a fundamentally different approach: it is a native MLX runtime designed specifically for Laya typed decision models. Instead of generating autoregressive text, pulling in PyTorch/Transformers dependencies, or invoking cloud APIs, `laya-mlx` processes structured decision schemas locally on Apple Unified Memory. ### How It Works Under the Hood The core performance gain comes from generating 0 output tokens. Standard LLM structured outputs require predicting tokens sequentially until a valid JSON structure forms. `laya-mlx` bypasses token generation entirely, evaluating typed queries in a single direct forward pass. Running natively inside Apple's MLX framework, the runtime leverages local metal optimizations, graph compilation, and prefix reuse. Tested benchmarks on an M3 Max (FP16) reveal ultra-fast end-to-end execution: * Laya 421M (English): 13.42 ms median (P50) end-to-end latency for short decisions. * Laya Multilingual 322M: 7.39 ms median (P50) end-to-end latency. When enabling graph compilation and prefix-reuse paths, tests logged up to 75.40 moves/s across 2,400 moves in the built-in Snake decision loop, outperforming eager execution modes by ~6.5%. ### Environment & Requirements * Hardware: Apple Silicon * OS: macOS 14+ (Tested on macOS 27.2) * Python: Python 3.11+ (Tested on 3.12.13 with MLX 0.32.2) ### Installation and Setup Install the standard package via `pip`: ```bash pip install laya-mlx ``` To run the interactive CLI terminal demo (requires a terminal window of at least 104 × 35 cells), install the demo extra and pull the multilingual weights: ```bash pip install 'laya-mlx[demo]' hf download aac6fef/laya-multilingual-mlx laya-snake ``` To run the optimized execution path without artificial frame-pacing: ```bash laya-snake --optimize --max-speed ``` ### Programmatic Usage You define structured schemas using standard Python dictionaries containing field definitions, decision types, instructions, and concrete selection criteria. ```python import laya_mlx as laya agent = laya.load("aac6fef/laya-mlx") result = agent.predict( "I was billed twice. Please refund the duplicate.", { "department": { "type": "choice", "instructions": "Who should handle this?", "criteria": ["billing", "technical", "sales"], } }, ) print(result["answers"]["department"]) ``` ### Tech Review Verdict `laya-mlx` provides sub-15 ms decision processing by completely stripping out autoregressive generation overhead. If your local Apple Silicon workflows need deterministic, structured classification—like routing requests or executing game loops—without installing heavy PyTorch runtimes, `laya-mlx` delivers an exceptionally tight, native execution pipeline.
Reviewed by DevTechPulse Editorial Board

All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.