I used to break down texts for top publishers; now I build and deploy the code that helps machines understand them. After 4 years managing language quality and content data structures in digital publishing, I shifted my focus to NLP and conversational systems.
Today, my day-to-day involves training NLU pipelines, orchestrating GenAI agents with CrewAI, and building RAG architectures. Coming from a strong linguistic background means I don’t just write prompts or code models; I know exactly how syntax behaves, how context changes meaning, and how to systematically mitigate biases to ensure high-quality outputs.
📍 Madrid, Spain (Open to Hybrid / Remote roles)
| Category | Technologies & Tools |
|---|---|
| 🤖 NLP & GenAI | Advanced Prompt Engineering, Agent Orchestration (CrewAI), RAG Architectures, LLM Evaluation & Bias Mitigation, LangChain, Transformers (BERT), spaCy. |
| 💬 Conversational AI & UX | Rasa, Dialogflow, Cognigy, Dialog Flow Architecture, Intent Training, UX Writing. |
| ⚙️ Data & Cloud Tools | Python (Pandas, Regex), Google Cloud Platform (GCP - Vertex AI, BigQuery, Cloud Storage), Tensorflow, Scikit-learn, Web Scraping, REST APIs. |
| 🗄️ Databases | MongoDB, SQL (PostgreSQL, Oracle, SQL Server), End-to-End Data Traceability. |
| 🛠️ Tools & Agile | Git (GitHub/GitLab), JSON, XML/XSL, Jupyter Notebooks, Plotly, Agile (Scrum). |
An autonomous multi-agent system designed for automated financial news analysis, entity extraction, and cost-performance benchmarking. This system was developed as the successful technical selection project for my role at Whitehole Data.
- Tech Stack: Python, CrewAI, OpenAI, LangChain, Zephyr, spaCy, Plotly, Pandas.
- Core Value: * Designed and built a modular GenAI architecture to benchmark open-source vs. proprietary models, directly evaluating performance, API token costs, and latency before my professional internship.
- Developed a hybrid NER pipeline combining spaCy with LLMs and integrated interactive Plotly dashboards to monitor operational metrics and token consumption in real time.
- 🔗 Explore Code & Dashboards
↗️
A multilingual conversational assistant for real-time hate speech detection, classification, and documentation.
- Tech Stack: Rasa SDK, spaCy, LangChain, OpenAI, Hugging Face Transformers.
- Core Value: * Developed a production-ready NLU pipeline in Rasa with 30+ intents using
DIETClassifierandTEDPolicy.- Programmed a custom Rasa Action Server to integrate LLMs, managing conversation contexts, enforcing the bot's persona/tone, and automatically generating compliance reports (
.docx).
- Programmed a custom Rasa Action Server to integrate LLMs, managing conversation contexts, enforcing the bot's persona/tone, and automatically generating compliance reports (
A corpus linguistics and computational analysis assignment developed during my Master’s program.
- Tech Stack: Python, spaCy, Plotly.
- Core Value: * Applied text normalization pipelines (tokenization, lemmatization, and POS tagging) using spaCy to process the complete textual dataset of Don DeLillo’s novel White Noise.
- Extracted stylistic, syntactic, and thematic patterns through computational methods, visualizing corpus linguistics metrics with interactive Plotly graphs.
- 🔗 View Notebook
↗️
Repsol GBC (LCG) | Data Analyst / Data & AI Intern Nov 2025 – May 2026
- Data Integrity & Quality: Programmed Python scripts to clean, validate, and run automated time-series checks on renewable energy asset data (wind/solar) prior to MongoDB ingestion, ensuring high-fidelity data streams.
- Generative AI Scoping: Worked alongside the technical lead on the conceptual design, scoping, and functional documentation for upcoming Generative AI use cases and Proof of Concepts (PoCs).
- Data Standardization: Normalized and unified heterogeneous data structures across SQL Server, Oracle, and PostgreSQL to guarantee end-to-end data traceability.
Whitehole Data | Computational Linguist (Intern) May 2025 – Aug 2025
- LLM Optimization & Cost Reduction: Designed a recursive prompt engineering workflow using an "agent-judge" architecture to align lightweight model performance with larger foundational models, drastically cutting inference costs.
- Model Evaluation & Guardrails: Identified and resolved overfitting issues by isolating anomalies in training datasets and iteratively refining prompt guardrails.
- Applied NLP: Implemented a hybrid NER pipeline (spaCy + LLMs) for semantic normalization of financial streams and built interactive quality-monitoring dashboards in Plotly.
Before fully transitioning into NLP and AI engineering, I spent four years in digital editorial production and managing cross-functional external teams for top-tier educational publishers, including Oxford University Press, Anaya Educación, and McGraw Hill.
This background is where my rigorous approach to data comes from: I specialized in designing complex educational data structures, mapping metadata, and enforcing strict language quality standards. Today, I apply that exact same structural discipline to data validation pipelines, training intents, and building deterministic context controls for LLMs.
- Specialization in Machine Learning Engineering & AI on Google Cloud Platform (GCP) – CFTIC (In Progress, Graduation: July 2026)
Focus: Utilizing Vertex AI, BigQuery, and Cloud Storage for NLP and ML workflows. - MSc in Language and Artificial Intelligence – Universidad Autónoma de Madrid (2024 – 2025)
Focus: NLP Engineering, Computational Linguistics, Conversational Systems, and RAG architectures. - MA in American Studies (Critical Discourse Analysis) – Universidad Complutense de Madrid
- BA in English Studies (Semantics, Pragmatics & Phonetics) – Universidad de Barcelona
- Cognigy Academy: Mastery Track 1, 2 & 3 (2024-2025)
- IBM: User Experience Design & Data Fundamentals (2025)
- Founderz: Responsible AI & Prompt Engineering (2025)
- Cálamo & Cran: Language technologies for linguists.
- Spanish: Native | English: Bilingual (C1.2) | Catalan: Professional | German & French: Basic (A2)
📬 Let's Connect! Whether you want to discuss multi-agent architectures, dialog flows, or how computational linguistics is shaping the future of AI pipelines; or simply you like languages, literatura, language technologies and philosophy, feel free to reach out via LinkedIn or email me at anabelarasanz@disroot.org.