Juicer is a data-refinement model built on Qwen3.6-35B-A3B (MoE architecture: 35B total parameters / 3B activated). It turns natural-language cleaning instructions, filtering rules, and semantic-tagging requirements into structured outputs鈥攕trict tagged text or canonical JSON.
Juicer is designed for data-refinement workflows and supports local deployment to process sensitive data in your own environment.
| Resource | Link |
|---|---|
| HuggingFace Model | datajuicer/Juicer-35B-A3B |
| ModelScope Model | Data-Juicer/Juicer-35B-A3B |
| Juicer Playground | data-juicer-hub/juicer_playground |
The juicer_playground directory in data-juicer-hub provides serve.sh, requirements.txt, and app.py. Clone the repository and enter its Playground directory:
git clone https://github.com/datajuicer/data-juicer-hub.git
cd data-juicer-hub/juicer_playgroundPrepare the inference environment and model weights using the Playground deployment instructions. Juicer can be served as an OpenAI-compatible endpoint. On a single H20 (96 GB):
export MODEL_ID=/path/to/juicer-model
bash serve.sh --model "$MODEL_ID" --port 8000The Juicer Playground provides an interactive UI to try recipes, browse showcase cases, and compare Juicer against the base model:
pip install -r requirements.txt
export JUICER_BASE_URL=http://localhost:8000/v1
python app.py
# open http://localhost:7860Juicer is evaluated on CDR-Bench, covering atomic mappers/filters, compositional workflows, order-sensitive pipelines, and semantic tasks (PII, hallucination, rubric, safety).
For full deployment options (vLLM, SGLang, Transformers), showcase cases, integration code, and AB comparison setup, visit the Juicer Playground README.