Skip to main content
Back to All Projects
Open Source Verified Architecture

SemIf-OpenJev

SemIf-OpenJev lets you generate semantic if statements using open‑source models on a single RTX 3090 at home, giving you local, privacy‑preserving conditional logic. It’s an independent effort, unaffiliated with Jev or TypeSafe. The tool runs entirely on your RTX 3090 GPU, so you avoid cloud costs and keep data in‑house.

4,514 Stars 313 Forks

Core Technologies & Frameworks

Python

Technical Architecture & Specifications

Most LLM-driven agent systems waste inference cycles generating conversational text or JSON objects simply to make binary or categorical branching decisions—like routing queries, validating evidence, or triggering retries. SemIf (formerly OpenJev) solves this overhead by treating local open models as direct, typed decision engines running on hardware like a consumer RTX 3090 or Apple Silicon. ### How It Works Under the Hood Rather than running an autoregressive decoding loop paired with fragile output parsing or JSON repair, SemIf evaluates semantic `if` statements by scoring typed option probabilities directly from the model's logits during a forward pass. Same frozen model state, zero answer generation overhead. The engine relies on modular backend implementations: * MLX Backend (macOS arm64): Designed for Apple Silicon, it supports direct scoring, serial prefix reuse (caching KV states across decisions), and parallel shared-state evaluation. * PyTorch / MPS: Provides scoring across direct, serial, and shared modes on Apple Silicon devices via `--device mps`. * llama.cpp (CPU): Runs prompt scoring directly against local GGUF checkpoints (such as `Qwen_Qwen3.5-4B-Q4_K_M.gguf`) using multi-threaded CPU execution without needing a CUDA device. * Temperature Calibration & EXL3: Includes per-workload temperature calibration for calibrated prediction outputs and integrates an EXL3 bridge for larger weights like Qwen3.8-27B. * In-Browser WebGPU: Supports browser-native models like MiniCPM5 2B and Qwen3.5 4B executing client-side via WebGPU. ### Environment Setup SemIf requires Python 3.10+, CUDA, or a supported Apple Silicon/CPU target. #### Standard CUDA Setup (4B BF16 models on GPU) ```bash python -m venv .venv . .venv/bin/activate export HF_HOME=/path/to/large-drive/huggingface pip install -e '.[test]' ``` #### macOS / Apple Silicon Native (MLX) To leverage MLX prefix reuse and direct scoring on arm64: ```bash pip install -e '.[test,mlx]' ``` Pass `--backend mlx` to your scorer command. #### CPU Offloading via llama.cpp For CUDA-less environments running local GGUF quantization: ```bash pip install -e '.[test,llamacpp]' ``` Append `--backend llamacpp --gguf /path/to/model.gguf` to your scorer execution and control core utilization with `--llama-threads`. ### Technical Review SemIf addresses a critical bottleneck in local agent workflows. By completely skipping token generation in favor of direct probability scoring, it removes string-parsing latency and output format instability. Additions like MLX serial prefix reuse and temperature calibration make it a solid runtime choice for deterministic, low-latency local control flow.
Reviewed by DevTechPulse Editorial Board

All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.