Back to All Projects
Open Source Verified Architecture
fast-jev-compaction
fast-jev-compaction is a Claude Code plugin that swaps out the default compaction summary for Jev‑based decisions. It scores each tool call and its output in a single quick request, discarding or truncating stale entries while preserving the rest verbatim. Ideal for keeping AI‑generated logs lean and accurate.
7,378 Stars 492 Forks
Core Technologies & Frameworks
TypeScript
Technical Architecture & Specifications
Standard context compaction in LLM agent frameworks usually relies on asking a model to summarize past conversation turns. In practice, this process is inherently lossy; it frequently wipes out exact stack traces, constraints, and file paths that you end up needing three turns later. `fast-jev-compaction` solves this problem by ditching summarization entirely. Distributed both as an npm package (`src/`) and a Claude Code plugin (`hooks/`, `.claude-plugin/`), it keeps user and assistant text completely verbatim, relying on Jev to surgically drop or truncate stale tool calls and outputs.
### How It Handles Context Under the Hood
Rather than rewriting raw histories, the pipeline pairs every `tool_use` with its matching `tool_result` via `tool_use_id`. To preserve essential context, the conversation's first message and the newest `preserveRecentMessages` are pinned and locked against modification.
To score older turns, the library constructs a state representation of the full conversation in chronological order, replacing full tool result bodies with short inline markers like `ok, 4213 chars (omitted)`. If this state exceeds `maxStateTokens` (25k by default), the library applies a deterministic fallback pipeline until it fits:
1. Tool Input Truncation: Drops tool inputs to 1,000, then 200, then 60 characters.
2. Text Abridgement: Cuts long non-pinned text to head and tail sections, collapsing old non-pinned messages into `[… N chars omitted …]` placeholders.
3. Line Reduction: Condenses old tool calls into single-line representations (e.g., `t12 Read file_path=src/a.ts → ok 480ch`), omits old call-less messages, and folds runs of call-only messages together.
Token counts are calculated fast without an external tokenizer using a calibrated heuristic (one word per six letters, half a token per digit, ~1 per symbol) designed to land just above Jev's reported figures.
### Evaluation & Decision Pipeline
For every unpinned tool call, the library sends two `noul` questions to Jev:
1. Should the call stay? (Is knowing the call happened with its inputs still relevant?)
2. Should the result stay verbatim? (Is the raw output still required, or is re-running unviable?)
If state size plus questions exceeds `maxRequestTokens` (30k by default, keeping under Jev's 32k hard limit), the questions are split into multiple concurrent requests sharing the exact same state payload. Upon response aggregation, decisions evaluate against `keepThreshold`:
* `keepResult ≥ threshold`: Retain both the tool call and the complete result verbatim.
* `keepCall ≥ threshold`: Retain the tool call, but truncate the result to `truncateHeadChars` plus a single-line note.
* Failed thresholds: Completely drop both the tool call and its result from the sequence.
Finally, the message list is reconstructed. Messages stripped of all contents are removed entirely, while surviving user and assistant text stays completely untouched.
### Verdict
By substituting lossy re-summarization with structured deterministic pruning, `fast-jev-compaction` stops critical code context from being degraded mid-session. It's a sharp, efficient alternative for Claude Code power users who need strict historical accuracy.
Reviewed by DevTechPulse Editorial Board
All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.