Skip to content
View anbat3's full-sized avatar

Block or report anbat3

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
anbat3/README.md

Hi, I'm Anabel 👋

LinkedIn Email

Computational Linguist | NLP & AI Engineer

I used to break down texts for top publishers; now I build and deploy the code that helps machines understand them. After 4 years managing language quality and content data structures in digital publishing, I shifted my focus to NLP and conversational systems.

Today, my day-to-day involves training NLU pipelines, orchestrating GenAI agents with CrewAI, and building RAG architectures. Coming from a strong linguistic background means I don’t just write prompts or code models; I know exactly how syntax behaves, how context changes meaning, and how to systematically mitigate biases to ensure high-quality outputs.

📍 Madrid, Spain (Open to Hybrid / Remote roles)


🛠️ Technical Stack

Category Technologies & Tools
🤖 NLP & GenAI Advanced Prompt Engineering, Agent Orchestration (CrewAI), RAG Architectures, LLM Evaluation & Bias Mitigation, LangChain, Transformers (BERT), spaCy.
💬 Conversational AI & UX Rasa, Dialogflow, Cognigy, Dialog Flow Architecture, Intent Training, UX Writing.
⚙️ Data & Cloud Tools Python (Pandas, Regex), Google Cloud Platform (GCP - Vertex AI, BigQuery, Cloud Storage), Tensorflow, Scikit-learn, Web Scraping, REST APIs.
🗄️ Databases MongoDB, SQL (PostgreSQL, Oracle, SQL Server), End-to-End Data Traceability.
🛠️ Tools & Agile Git (GitHub/GitLab), JSON, XML/XSL, Jupyter Notebooks, Plotly, Agile (Scrum).

🚀 Highlighted Projects

🔎 Financial News Agent System

An autonomous multi-agent system designed for automated financial news analysis, entity extraction, and cost-performance benchmarking. This system was developed as the successful technical selection project for my role at Whitehole Data.

  • Tech Stack: Python, CrewAI, OpenAI, LangChain, Zephyr, spaCy, Plotly, Pandas.
  • Core Value: * Designed and built a modular GenAI architecture to benchmark open-source vs. proprietary models, directly evaluating performance, API token costs, and latency before my professional internship.
    • Developed a hybrid NER pipeline combining spaCy with LLMs and integrated interactive Plotly dashboards to monitor operational metrics and token consumption in real time.
  • 🔗 Explore Code & Dashboards ↗️

🗣️ DEL.IA – Hate Speech Detection Chatbot (Master's Thesis - 9/10)

A multilingual conversational assistant for real-time hate speech detection, classification, and documentation.

  • Tech Stack: Rasa SDK, spaCy, LangChain, OpenAI, Hugging Face Transformers.
  • Core Value: * Developed a production-ready NLU pipeline in Rasa with 30+ intents using DIETClassifier and TEDPolicy.
    • Programmed a custom Rasa Action Server to integrate LLMs, managing conversation contexts, enforcing the bot's persona/tone, and automatically generating compliance reports (.docx).

📖 NLP Text Analysis – White Noise

A corpus linguistics and computational analysis assignment developed during my Master’s program.

  • Tech Stack: Python, spaCy, Plotly.
  • Core Value: * Applied text normalization pipelines (tokenization, lemmatization, and POS tagging) using spaCy to process the complete textual dataset of Don DeLillo’s novel White Noise.
    • Extracted stylistic, syntactic, and thematic patterns through computational methods, visualizing corpus linguistics metrics with interactive Plotly graphs.
  • 🔗 View Notebook ↗️

💼 Professional Experience

Repsol GBC (LCG) | Data Analyst / Data & AI Intern Nov 2025 – May 2026

  • Data Integrity & Quality: Programmed Python scripts to clean, validate, and run automated time-series checks on renewable energy asset data (wind/solar) prior to MongoDB ingestion, ensuring high-fidelity data streams.
  • Generative AI Scoping: Worked alongside the technical lead on the conceptual design, scoping, and functional documentation for upcoming Generative AI use cases and Proof of Concepts (PoCs).
  • Data Standardization: Normalized and unified heterogeneous data structures across SQL Server, Oracle, and PostgreSQL to guarantee end-to-end data traceability.

Whitehole Data | Computational Linguist (Intern) May 2025 – Aug 2025

  • LLM Optimization & Cost Reduction: Designed a recursive prompt engineering workflow using an "agent-judge" architecture to align lightweight model performance with larger foundational models, drastically cutting inference costs.
  • Model Evaluation & Guardrails: Identified and resolved overfitting issues by isolating anomalies in training datasets and iteratively refining prompt guardrails.
  • Applied NLP: Implemented a hybrid NER pipeline (spaCy + LLMs) for semantic normalization of financial streams and built interactive quality-monitoring dashboards in Plotly.

📖 The Foundation: 4 Years in Digital Publishing (2020 – 2024)

Before fully transitioning into NLP and AI engineering, I spent four years in digital editorial production and managing cross-functional external teams for top-tier educational publishers, including Oxford University Press, Anaya Educación, and McGraw Hill.

This background is where my rigorous approach to data comes from: I specialized in designing complex educational data structures, mapping metadata, and enforcing strict language quality standards. Today, I apply that exact same structural discipline to data validation pipelines, training intents, and building deterministic context controls for LLMs.


🎓 Education & Certifications

  • Specialization in Machine Learning Engineering & AI on Google Cloud Platform (GCP) – CFTIC (In Progress, Graduation: July 2026)
    Focus: Utilizing Vertex AI, BigQuery, and Cloud Storage for NLP and ML workflows.
  • MSc in Language and Artificial Intelligence – Universidad Autónoma de Madrid (2024 – 2025)
    Focus: NLP Engineering, Computational Linguistics, Conversational Systems, and RAG architectures.
  • MA in American Studies (Critical Discourse Analysis) – Universidad Complutense de Madrid
  • BA in English Studies (Semantics, Pragmatics & Phonetics) – Universidad de Barcelona

🏅 Professional Certifications

  • Cognigy Academy: Mastery Track 1, 2 & 3 (2024-2025)
  • IBM: User Experience Design & Data Fundamentals (2025)
  • Founderz: Responsible AI & Prompt Engineering (2025)
  • Cálamo & Cran: Language technologies for linguists.

🌐 Languages

  • Spanish: Native | English: Bilingual (C1.2) | Catalan: Professional | German & French: Basic (A2)

📬 Let's Connect! Whether you want to discuss multi-agent architectures, dialog flows, or how computational linguistics is shaping the future of AI pipelines; or simply you like languages, literatura, language technologies and philosophy, feel free to reach out via LinkedIn or email me at anabelarasanz@disroot.org.

Popular repositories Loading

  1. financial-news-agent financial-news-agent Public

    AI agent for financial news analysis, with entity extraction and normalization, contextual summary generation, and sentiment classification. Implements a modular pipeline using CrewAI (OpenAI) and …

    Jupyter Notebook 1 1

  2. anbat3 anbat3 Public

    Config files for my GitHub profile.

  3. BOTanica BOTanica Public

    BOTanica is a Rasa-based virtual assistant that provides interactive, accurate information about various flowers, including their traits, blooming seasons, and care tips. Perfect for educational, g…

    Python

  4. meteorite_impact_simulator meteorite_impact_simulator Public

    Este proyecto simula la caída de meteoritos sobre una cuadrícula de 50x50 durante un periodo de 24 horas. Cada hora se genera un número aleatorio de meteoritos con posiciones y intensidades también…

    Jupyter Notebook

  5. nlp_text_analysis nlp_text_analysis Public

    Este proyecto realiza un análisis completo de texto utilizando Python y SpaCy. El objetivo es normalizar, procesar y visualizar datos textuales del libro White Noise de Don DeLillo.

    Jupyter Notebook

  6. anbat3.github.io anbat3.github.io Public

    Forked from piazzai/cvless

    CV page

    SCSS