Skip to content

About

persian (farsi) full stopwords dataset

Resources

Stars

4 stars

Watchers

1 watching

Forks

Repository files navigation

Persian (Farsi) Stopwords

PyPI version License: MIT Python 3.9+

A comprehensive, curated list of Persian (Farsi) stopwords for text processing, NLP, search indexing, and corpus cleaning.

Installation

pip install persian-stopwords

Quick Start

from persian_stopwords import load_stopwords, clean_text

# Load the stopwords set
stopwords = load_stopwords()
print(f"{len(stopwords)} stopwords loaded")

# Remove stopwords from text
cleaned = clean_text("این یک متن نمونه است که می‌خواهیم پاک کنیم")
print(cleaned)
# Output: متن نمونه می‌خواهیم پاک

API

load_stopwords() -> frozenset[str]

Returns the full set of Persian stopwords as a frozenset. Data is loaded from package resources, so it works from any working directory.

from persian_stopwords import load_stopwords

stopwords = load_stopwords()

remove_stopwords(tokens, stopwords=None) -> list[str]

Filters stopwords out of an iterable of tokens. If stopwords is None, uses the built-in list.

from persian_stopwords import remove_stopwords

tokens = "این یک متن نمونه است".split()
cleaned = remove_stopwords(tokens)

clean_text(text, stopwords=None) -> str

Removes Persian stopwords from a whitespace-delimited string.

from persian_stopwords import clean_text

result = clean_text("این یک متن نمونه است")

What's Included

Over 2,900 stopwords including:

  • Pronouns & demonstratives — آن، این، او، آنها، خود، etc.
  • Prepositions & postpositions — از، به، در، با، برای, etc.
  • Conjunctions — و، یا، اما، که، چون, etc.
  • Verbs & verb forms — است، بود، شد، می‌کند, etc.
  • Adverbs — خیلی، بسیار، فقط، همیشه, etc.
  • Question words — چه، کی، کجا، چرا, etc.
  • Filler & interjection words — خب، آها، آهان, etc.
  • Punctuation marks — Persian and common punctuation

Data Format

  • Encoding: UTF-8
  • Format: One word per line
  • Sorting: Persian alphabetical order (آ ا ب پ ت ث ج چ ح خ د ذ ر ز ژ س ش ص ض ط ظ ع غ ف ق ک گ ل م ن و ه ی)
  • Normalization: Zero-width characters and BOM removed; duplicates removed

File Structure

.
├── pyproject.toml                    # Build config and package metadata
├── MANIFEST.in                       # Include data files in sdist
├── README.md
├── LICENSE
├── resources.txt                     # List of source references
├── example.py                        # Standalone usage example
└── src/
    └── persian_stopwords/
        ├── __init__.py               # Package API
        └── data/
            └── stopwords.txt         # The stopwords list (UTF-8)

Sources

This dataset aggregates and refines stopwords from the following sources (see resources.txt for full links):

Contributing

Contributions are welcome! If you find missing stopwords or entries that should be removed:

  1. Fork this repository
  2. Edit src/persian_stopwords/data/stopwords.txt (keep alphabetical order and one word per line)
  3. Submit a pull request

Please ensure entries are:

  • Genuinely stopwords (high-frequency words that carry little semantic meaning)
  • In standard Persian script (avoid ASCII-only or garbled entries)
  • Not duplicates of existing entries

Publishing (for maintainers)

# Build the package
python -m build

# Upload to PyPI
twine upload dist/*

License

MIT License — Copyright (c) 2022 Masoud Kaviani


دیتاست مجموعه کلمات توقف فارسی

این مجموعه داده، شامل بیش از ۲,۹۰۰ کلمه توقف فارسی است که از منابع مختلف جمع‌آوری، پاک‌سازی، و مرتب‌سازی شده است. می‌توانید برای پاک‌سازی متون، پردازش زبان طبیعی، فهرست‌سازی جستجو و کاربردهای مشابه از آن استفاده کنید.

نصب

pip install persian-stopwords

استفاده

from persian_stopwords import load_stopwords, clean_text

stopwords = load_stopwords()
cleaned = clean_text("این یک متن نمونه است")

منابع

منابع این مجموعه داده در فایل resources.txt listing شده‌اند.

About

persian (farsi) full stopwords dataset

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages