YADRO · ML Engineer · Aug 2024 — Nov 2025
RAG assistant over technical documentation
Problem
Large, frequently updated documentation full of similar terms, versions and configurations. Answers must rest on the current source, not on the model's general knowledge.
What I built
- Ingestion with semantic chunking and metadata (product, version, component); chunk size and overlap tuned on an eval set.
- Hybrid retrieval: dense + BM25, RRF, a cross-encoder reranker on the top-N, product and version filters.
- Grounded generation with source references and a fallback on weak context.
- LoRA/PEFT adaptation and claim-level evaluation on a regression set.
Engineering decisions
- Latency broken down by stage; caching embeddings, retrieval and context; a cap on reranker candidates.
- Every change goes through a regression suite and shadow mode.
- Reproducible experiments in ClearML.
Result
A clear Recall@5 gain over the baseline, fewer unsupported claims on the regression set, latency within the target SLA.
What it shows → The full production RAG cycle: retrieval, generation, fine-tuning, evaluation and latency.
Python · Transformers · Sentence Transformers · PEFT/LoRA · BM25 · cross-encoder · FastAPI · ClearML · Docker · Prometheus · Grafana