Skip to main content
Back to All Projects
Open Source Verified Architecture

laya

I’m sorry, but I can’t craft an accurate pitch without any description or details about the “laya” project.

16,230 Stars 1,358 Forks

Core Technologies & Frameworks

Python

Technical Architecture & Specifications

Autoregressive language models force developers to generate text sequentially just to extract categorical outputs or numerical ratings. This approach introduces unnecessary latency penalties, requires fragile output parsing, and exposes production pipelines to non-zero hallucination rates. Laya takes a completely different architectural approach: it operates as a non-autoregressive, System 1 decision engine that returns typed decisions in a single forward pass. ### Architectural Mechanics and Inference Pipeline Laya completely bypasses token generation. It takes state inputs—including raw text, emails, support tickets, or structured JSON documents—and directly evaluates typed questions over that context. Because the system does not produce generative text outputs, there are no string parsing steps and zero opportunities for hallucination. Key technical specifications of the architecture include: * Training Paradigm: Models are trained using reinforcement learning against strictly proper scoring rules (RLCD) to maintain calibration across decisions. * Question Primitives: The inference pipeline natively evaluates three typed decision primitives: `choice`, `score`, and `noul`. * Latency Benchmarks: When evaluated on an NVIDIA T4 GPU, Laya executes a single-question evaluation in 33 ms. Under batched workloads, throughput drops to 7.2 ms per question. ### Checkpoint Ecosystem and Automated Routing The system relies on three distinct encoder checkpoints paired with a `Router` component. The router evaluates incoming requests and dispatches them to the most suitable model based on context, language, and context-length requirements: | Model Checkpoint | Base Encoder | Parameter Count | Context Window | Target Workload | | :--- | :--- | :--- | :--- | :--- | | `laya` | ModernBERT-large | 421M | 512 tokens | High-precision English decision tasks | | `laya-multilingual` | mmBERT-base | 322M | 1024 tokens | 100+ languages, running 2x faster | | `laya-typed-decisions` | ModernBERT-large | 421M | 1024 tokens | Specialized typed-decision application workflows | ### Ecosystem & Distribution Laya is open-sourced under the Apache 2.0 license. The runtime package is hosted on PyPI (`laya`), and the underlying checkpoints (`convaiinnovations/laya`, `convaiinnovations/laya-multilingual`, and `convaiinnovations/laya-typed-decisions`) are hosted on Hugging Face. Interactive testing setups are available via an official Google Colab notebook and a Hugging Face Space (`laya-demo`). By replacing generative decoding loops with non-autoregressive forward passes, Laya offers a performant, predictable alternative for low-latency decision, scoring, and routing infrastructure.
Reviewed by DevTechPulse Editorial Board

All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.