Dataset for Training and Evaluating LLM-Based SOC Agents
-
Updated
Jun 4, 2026 - Python
Dataset for Training and Evaluating LLM-Based SOC Agents
Open-source dataset for evaluating Large Language Models (LLMs), developed as part of my graduation research project.
Grounded, fact-checked instruction-tuning dataset for cyber threat intelligence — 10k+ examples across 37 CTI categories, built from MITRE ATT&CK, CVE/KEV, CWE and live threat feeds. For LLM fine-tuning (LoRA/QLoRA).
Flipper Zero Sub-GHz RF dataset (280-1100 MHz, 9 countries) with 1500 Q&A pairs for LLM fine-tuning, fact-checked allocations, and a GPU-accelerated validation pipeline (Ollama Qwen 32B + DeBERTa NLI).
The Anti-Hallucination data layer for B2B Sourcing. Deep-verified global supply chain entities designed for RAG and LLM instruction tuning.
Open-source multi-server Discord channel scraper. Configure many servers in one TOML file, pick channels by name or ID, export to JSONL for offline search and LLM pipelines.
AI-powered Q&A system for U.S. affordable housing policy using RAG over 2,500+ HUD documents and 24 CFR
A comprehensive Python tool for extracting, processing, and analyzing RPG scenarios from the Era of the Imperial Republic (EOTIR) forums. Features automated web scraping, NLP-powered content analysis, character extraction, timeline generation, and LLM dataset preparation with an interactive HTML dashboard.
This repository aims to provide a structured and easily accessible dataset of laws in Bangladesh. The data is primarily sourced from the Bangladesh Law (BDLAW) website.
A Curated RAG Dataset of 247 Articles on Chinese Muslim Food and Culture
Gittxt is an AI-focused CLI and plugin tool for extracting, filtering, and packaging text from GitHub repos. Build LLM-compatible datasets, prep code for prompt engineering, and power AI workflows with structured .txt, .json, .md, or .zip outputs.
High-quality dataset of 201 authentic articles introducing Halal restaurants across China. RAG optimized.
Prepare the Kleister NDA dataset for LLM-based extraction. Validates labels against a Pydantic schema and delivers partitioned Parquet with co-located PDFs
ONE SYSTEM Knowledge Dataset — Structured training data from 11 AI, legal & technology brands by David Sanker. UAPK Business Compiler, Legal AI, AI Governance.
Features 232 articles covering Hui Muslim culture, travel, mosques, and halal food.
Autonomous MCP server for M2M patent intelligence. Delivers structured JSON datasets (CPC A-H) enriched with biz_value_prop, tech stacks, and importance scoring. Supports instant autonomous data purchasing via ROSE cryptocurrency.
本数据集是目前已知的纯中文原生穆斯林旅游 RAG 专属语料库。所有标题和正文均为未经过任何翻译和篡改的原始中文文本
Islamic Culture Knowledge Base & RAG Dataset
Add a description, image, and links to the llm-dataset topic page so that developers can more easily learn about it.
To associate your repository with the llm-dataset topic, visit your repo's landing page and select "manage topics."