Back to All Projects
Open Source Verified Architecture
arc-task-gen
arc-task-gen is an open‑source tool that automatically creates new ARC‑AGI‑1‑style reasoning tasks. It calibrates the synthetic tasks so their difficulty and concept distribution line up with the public ARC evaluation set, giving researchers a scalable way to stress‑test and benchmark general‑purpose AI systems.
9,064 Stars 60 Forks
Core Technologies & Frameworks
Python
Technical Architecture & Specifications
Evaluating frontier reasoning models on public benchmarks like ARC-AGI-1 presents a classic AI evaluation trap: distinguishing genuine few-shot rule induction from simple memorization or prior dataset familiarity. That is where `arc-task-gen` comes in. This project provides a generator that synthesizes original, distribution-matched ARC-AGI-1-style tasks, creating a clean private evaluation set for benchmarking without data contamination.
### Under the Hood: Distribution Matching for Latent Reasoning
Public benchmarks struggle to isolate in-context learning when models may have ingested parts of the dataset during pretraining. `arc-task-gen` counters this by generating novel problems matching the structural distribution of the public ARC-AGI-1 evaluation set.
The primary motivation behind this project is evaluating models like **BDH-CQ** under controlled conditions. BDH-CQ builds on Dragon Hatchling (BDH), a post-Transformer recurrent architecture where neuron-like units communicate via low-rank interactions and hold state in an evolving associative recurrent memory. Rather than generating an intermediate text-based chain of thought, BDH-CQ solves queries using iterative computation within a structured continuous latent space.
Benchmarking shows a 150M-parameter configuration of BDH-CQ achieving **29.5% pass@2** on the public ARC-AGI-1 evaluation set at an inference cost of **$0.0007 per task** (11× cheaper per task than GPT-5.6 Luna Low). By applying controlled ARC-like interventions via `arc-task-gen`, researchers can verify whether such architectures are demonstrating true rule induction on unseen tasks or simply leveraging public benchmark familiarity.
### Output Data Schema and Harness Integration
The generator emits task sets in a standard `tasks.json` file. The output format directly mirrors the canonical ARC data structure:
```json
{
"train": [],
"test": []
}
```
Because the output adheres strictly to this format, `tasks.json` integrates directly into existing ARC evaluation harnesses without requiring custom data loaders or schema translation layers. You can drop the generated JSON directly into standard evaluation pipelines to measure how models handle fresh, un-indexed task distributions.
### How to Set Up and Run
To generate tasks, clone the repository and inspect the included setup documentation:
```bash
# View the generation guidelines and setup steps
cat instructions.md
```
Follow the detailed walkthrough inside [`instructions.md`](instructions.md) to execute the generator and produce distribution-matched task files ready for model evaluation.
### The Verdict
`arc-task-gen` is an essential tool for AI researchers working on novel architectures, recurrent latent reasoning systems, and long-context evaluation. If you need to validate whether a model possesses true in-context rule induction or is merely benefiting from benchmark leakage, this generator gives you the private, distribution-matched dataset necessary for rigorous testing.
Reviewed by DevTechPulse Editorial Board
All listed blueprints, repositories, and case studies are verified against public documentation and LTS container environments. For inquiries or updates, view our Editorial Policy.