Local batch annotation pipeline for the Propella-1-4b model using llama.cpp. Built as part of the OpenEuroLLM project.
Propella-1-4b is a 4B-parameter document annotation model (Qwen3 architecture) that classifies documents across multiple properties (domain, quality, toxicity, etc.) with structured JSON output. This repository provides scripts to run it locally on Apple Silicon via llama.cpp with full JSON schema enforcement.
- Batch processing — annotate JSONL document collections with configurable parallelism
- Structured output — 100% schema-valid JSON via llama.cpp GBNF grammar enforcement
- Result inspection — HTML comparison tables with match/mismatch highlighting against expected annotations
- Dataset integration — built-in downloader for Dolci-Instruct-SFT samples
- Python 3.12+
- uv (Python package manager)
- llama.cpp (
brew install llama.cppon macOS) - ~8 GB unified memory for BF16 inference
# 1. Install Python dependencies
uv sync
# 2. Download the model (~8.9 GB)
uv run python -c "from huggingface_hub import snapshot_download; snapshot_download('ellamind/propella-1-4b', local_dir='./propella-1-4b')"
# 3. Convert to GGUF format
uv run python $(which convert_hf_to_gguf.py) ./propella-1-4b --outfile propella-1-4b.gguf --outtype bf16With 32 GB unified memory you can run full BF16. For faster inference with minimal quality loss:
llama-quantize propella-1-4b.gguf propella-1-4b-Q8_0.gguf Q8_0Q8_0 (~4.3 GB) gives ~2x speed improvement over BF16 (~8 GB).
| Option | Why not |
|---|---|
| SGLang | CUDA only, no full Metal support yet |
| MLX | mlx_lm.server has no structured output support |
| HF Transformers | No schema enforcement, slower than llama.cpp |
| Ollama | No strict JSON schema enforcement |
# Run on bundled examples
./execute_examples.sh
# Custom input/output with parallelism
./execute_examples.sh data/dolci/documents.jsonl results/dolci/annotations.jsonl 4The script starts a llama-server, processes all documents, and stops the server automatically.
./inspect.sh results/annotations.jsonl
./inspect.sh results/dolci/annotations.jsonl data/dolci/expected.jsonlOpens an HTML table comparing model output against expected annotations:
- Each cell shows the model value and the expected value (in grey) below it
- Green background = match, red = mismatch
one_sentence_descriptionis excluded (too subjective for exact match)- HTML file is written next to the input (e.g.,
results/dolci/annotations.html)
Expected annotations: examples/expected.jsonl — same format as output, fill in fields you want to compare. Empty or id-only rows are ignored.
Generating expected annotations: Use examples/annotation_prompt.md as a prompt in a Claude Code session or similar. It contains the full annotation framework and points to the input documents and output path.
Download samples from the Dolci-Instruct-SFT dataset (2.15M instruction-following conversations, 70 languages, 8 domains):
# Download 1 sample per domain (8 total)
uv run python download_dolci.py
# Different samples via seed, more per domain
uv run python download_dolci.py --seed 42 --per-domain 3
# Extract only user instructions (no assistant replies)
uv run python download_dolci.py --extraction instruction
# Extract both variants (two documents per sample: instruction + full)
uv run python download_dolci.py --extraction bothThe --extraction flag controls what gets extracted from each conversation:
full(default) — entire conversation:User: ...\n\nAssistant: ...instruction— user messages only:User: ...both— one document per variant, doubling the output count
Each output document includes an extraction field ("full" or "instruction") to identify its variant. Metadata fields (dolci_id, source_dataset, domain) are preserved as pass-through.
These files are downloaded from the Hugging Face model repo during setup:
| File | Purpose |
|---|---|
propella.py |
Bundled module — provides create_messages, AnnotationResponse, get_annotation_response_schema |
inference_example.py |
Client script — talks to llama-server via OpenAI-compatible API |
property_descriptions.md |
Loaded by propella.py at import time (must be in working directory) |
model-*.safetensors |
Original model weights (used for GGUF conversion) |
propella-local/
├── main.py # Async batch annotation client
├── download_dolci.py # Dolci-Instruct-SFT downloader
├── execute_examples.sh # One-command annotation runner
├── inspect.sh # HTML result viewer/comparator
├── examples/ # Handcrafted example documents
├── data/ # Downloaded datasets
└── results/ # Annotation outputs
Apache 2.0 — see LICENSE.