Skip to content

About

Local batch annotation pipeline for Propella-1-4b using llama.cpp

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

propella-local

Local batch annotation pipeline for the Propella-1-4b model using llama.cpp. Built as part of the OpenEuroLLM project.

Propella-1-4b is a 4B-parameter document annotation model (Qwen3 architecture) that classifies documents across multiple properties (domain, quality, toxicity, etc.) with structured JSON output. This repository provides scripts to run it locally on Apple Silicon via llama.cpp with full JSON schema enforcement.

Features

  • Batch processing — annotate JSONL document collections with configurable parallelism
  • Structured output — 100% schema-valid JSON via llama.cpp GBNF grammar enforcement
  • Result inspection — HTML comparison tables with match/mismatch highlighting against expected annotations
  • Dataset integration — built-in downloader for Dolci-Instruct-SFT samples

Prerequisites

  • Python 3.12+
  • uv (Python package manager)
  • llama.cpp (brew install llama.cpp on macOS)
  • ~8 GB unified memory for BF16 inference

Setup

# 1. Install Python dependencies
uv sync

# 2. Download the model (~8.9 GB)
uv run python -c "from huggingface_hub import snapshot_download; snapshot_download('ellamind/propella-1-4b', local_dir='./propella-1-4b')"

# 3. Convert to GGUF format
uv run python $(which convert_hf_to_gguf.py) ./propella-1-4b --outfile propella-1-4b.gguf --outtype bf16

Optional: Quantization

With 32 GB unified memory you can run full BF16. For faster inference with minimal quality loss:

llama-quantize propella-1-4b.gguf propella-1-4b-Q8_0.gguf Q8_0

Q8_0 (~4.3 GB) gives ~2x speed improvement over BF16 (~8 GB).

Why llama.cpp

Option Why not
SGLang CUDA only, no full Metal support yet
MLX mlx_lm.server has no structured output support
HF Transformers No schema enforcement, slower than llama.cpp
Ollama No strict JSON schema enforcement

Usage

Batch annotation

# Run on bundled examples
./execute_examples.sh

# Custom input/output with parallelism
./execute_examples.sh data/dolci/documents.jsonl results/dolci/annotations.jsonl 4

The script starts a llama-server, processes all documents, and stops the server automatically.

Inspecting results

./inspect.sh results/annotations.jsonl
./inspect.sh results/dolci/annotations.jsonl data/dolci/expected.jsonl

Opens an HTML table comparing model output against expected annotations:

  • Each cell shows the model value and the expected value (in grey) below it
  • Green background = match, red = mismatch
  • one_sentence_description is excluded (too subjective for exact match)
  • HTML file is written next to the input (e.g., results/dolci/annotations.html)

Expected annotations: examples/expected.jsonl — same format as output, fill in fields you want to compare. Empty or id-only rows are ignored.

Generating expected annotations: Use examples/annotation_prompt.md as a prompt in a Claude Code session or similar. It contains the full annotation framework and points to the input documents and output path.

Dolci dataset

Download samples from the Dolci-Instruct-SFT dataset (2.15M instruction-following conversations, 70 languages, 8 domains):

# Download 1 sample per domain (8 total)
uv run python download_dolci.py

# Different samples via seed, more per domain
uv run python download_dolci.py --seed 42 --per-domain 3

# Extract only user instructions (no assistant replies)
uv run python download_dolci.py --extraction instruction

# Extract both variants (two documents per sample: instruction + full)
uv run python download_dolci.py --extraction both

The --extraction flag controls what gets extracted from each conversation:

  • full (default) — entire conversation: User: ...\n\nAssistant: ...
  • instruction — user messages only: User: ...
  • both — one document per variant, doubling the output count

Each output document includes an extraction field ("full" or "instruction") to identify its variant. Metadata fields (dolci_id, source_dataset, domain) are preserved as pass-through.

Key model files

These files are downloaded from the Hugging Face model repo during setup:

File Purpose
propella.py Bundled module — provides create_messages, AnnotationResponse, get_annotation_response_schema
inference_example.py Client script — talks to llama-server via OpenAI-compatible API
property_descriptions.md Loaded by propella.py at import time (must be in working directory)
model-*.safetensors Original model weights (used for GGUF conversion)

Project structure

propella-local/
├── main.py                 # Async batch annotation client
├── download_dolci.py       # Dolci-Instruct-SFT downloader
├── execute_examples.sh     # One-command annotation runner
├── inspect.sh              # HTML result viewer/comparator
├── examples/               # Handcrafted example documents
├── data/                   # Downloaded datasets
└── results/                # Annotation outputs

License

Apache 2.0 — see LICENSE.

About

Local batch annotation pipeline for Propella-1-4b using llama.cpp

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages