Back to All Projects
Open Source Verified Architecture
kev
kev delivers a Jev‑style suite of decision‑making models that sit on top of the Qwen 3.5/3.8 LLMs. The repo ships ready‑to‑train pipelines and inference scripts so you can fine‑tune and run the models entirely on your own hardware, keeping data in‑house.
6,727 Stars 385 Forks
Core Technologies & Frameworks
Python
Technical Architecture & Specifications
### Technical Overview: Kev Decision Models
Kev is an open-source family of specialized decision models built on top of Qwen3.5 and Qwen3.8 base architectures. Derived from the design described in *Jev's Architecture Unmasked*, Kev moves away from standard open-ended text generation, focusing specifically on structured decision tasks, probability estimation, and fine-grained content scoring.
---
### Architecture & Engine Mechanics
Under the hood, Kev adapts Qwen's underlying representations to execute multiple decision evaluations within a single request. It handles three specific decision primitives concurrently:
* `noul`: Binary yes/no decisions.
* `choice`: Discrete multiple-choice selection.
* `score`: Scalar rating evaluations.
A core technical highlight is how Kev handles context isolation during batched evaluation. Multiple questions share the same input text block, but questions remain strictly isolated from one another—they cannot read each other's prompts or leak state across queries.
To prevent overconfidence, Kev emphasizes calibrated output probabilities by default. Rather than relying on raw logit outputs, every model checkpoint ships with a pre-fitted temperature parameter calculated on held-out evaluation sets, providing well-calibrated confidence scores out of the box.
---
### Model Spectrum & Hardware Profiles
Kev provides four distinct model scales depending on your available VRAM and latency constraints:
| Model | Base Architecture | Target Hardware | Usage Profile |
| :--- | :--- | :--- | :--- |
| Kev-0.8B | Qwen3.5-0.8B-Base | Apple Silicon Mac, NVIDIA L4 | Edge inference, low latency, tight hardware limits |
| Kev-4B | Qwen3.5-4B-Base | Mid-tier GPUs | Recommended starting point for general workflows |
| Kev-9B | Qwen3.5 family | Mid-to-high VRAM GPUs | Higher accuracy on complex evaluation sources |
| Kev-27B | Qwen3.5 family | Single 80 GB GPU (e.g., A100/H100) | Maximum accuracy on heavy production workloads |
---
### SDK Compatibility, Fine-Tuning & Deployment
Kev features drop-in compatibility with TypeSafe's System One API specifications. If your application already uses the TypeSafe Python SDK, you can redirect traffic from third-party services to your self-hosted Kev endpoint without modifying application-level code.
For fine-tuning on custom datasets, Kev includes an end-to-end training loop executed on Modal via a coding-agent skill. The agent handles question discovery, model adaptation on your labelled examples, and serving.
In terms of ops, Kev supports single-command HTTPS deployments via Modal. The server infrastructure automatically scales to zero when idle, cutting hosting costs when no decision traffic is active. Live model evaluations can also be tested directly in the browser via the Hugging Face Spaces demo.
Reviewed by DevTechPulse Editorial Board
All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.