DKDenis Kazhaev

ML / LLM Engineer · production RAG

Production RAG and LLM systems: retrieval, evaluation, fine-tuning and MLOps.

I build LLM services from ingestion to monitoring: hybrid retrieval and reranking, grounded generation, PEFT/LoRA, claim-level evaluation, latency optimisation, reproducible experiments and production infrastructure. I own tasks independently and agree metrics and constraints with stakeholders myself.

Serbia · remote · UTC+1  /  Russian — native, English — B1

What I bring

01

Production RAG

Ingestion and metadata, chunking tuned on an eval set, dense + BM25, RRF, cross-encoder reranking, product and version filters, fallback on weak context.

02

Evaluation

Golden sets, Recall@K, NDCG / MRR for retrieval; groundedness and unsupported-claim rate for generation; regression suites and shadow mode.

03

Fine-tuning

PEFT / LoRA on good answers, real mistakes and correct-refusal examples. Versions are compared on claim-level metrics, not by eye.

04

Latency and cost

Per-stage latency breakdown, caching of embeddings, retrieval and context, a cap on reranker candidates, async I/O, context length control. P50 / P95, not the mean.

05

MLOps

ClearML for experiments, Docker, CI/CD on GitHub Actions, AWS deployment, Prometheus and Grafana with per-stage dashboards and alerts.

06

Backend

Python, FastAPI, PostgreSQL, Redis, Pydantic schemas for structured output, retries, idempotency, health checks, structured logging.

How I run an ML task

A standard ML system design cycle: framing and metrics → data → baseline → experiments → evaluation → rollout → monitoring. Every prompt, model or retrieval change goes through a regression run.

  1. 01

    Framing and metrics

    Constraints, target offline and online metrics, latency and cost-per-request budgets.

  2. 02

    Data and golden set

    Real queries with labelled sources, segments (paraphrases, exact terms and codes), a regression set of hard cases.

  3. 03

    Baseline

    A simple interpretable solution as a reference point, to see where complex components really add value.

  4. 04

    Experiments

    Chunking, retrieval, reranker, prompts, PEFT / LoRA. Parameters, data and code versions, metrics and artifacts in ClearML.

  5. 05

    Evaluation

    Retrieval, generation and system separately; per-segment slices; a regression run before any change.

  6. 06

    Rollout

    Shadow mode on real traffic, gradual rollout, fast rollback of the index and model.

  7. 07

    Monitoring

    Per-stage latency, errors, fallback, cache hit rate, indexing lag; alerts and degradation analysis.

Projects

01

YADRO · ML Engineer · Aug 2024 — Nov 2025

RAG assistant over technical documentation

Problem

Large, frequently updated documentation full of similar terms, versions and configurations. Answers must rest on the current source, not on the model's general knowledge.

What I built
  • Ingestion with semantic chunking and metadata (product, version, component); chunk size and overlap tuned on an eval set.
  • Hybrid retrieval: dense + BM25, RRF, a cross-encoder reranker on the top-N, product and version filters.
  • Grounded generation with source references and a fallback on weak context.
  • LoRA/PEFT adaptation and claim-level evaluation on a regression set.
Engineering decisions
  • Latency broken down by stage; caching embeddings, retrieval and context; a cap on reranker candidates.
  • Every change goes through a regression suite and shadow mode.
  • Reproducible experiments in ClearML.
Result

A clear Recall@5 gain over the baseline, fewer unsupported claims on the regression set, latency within the target SLA.

What it shows → The full production RAG cycle: retrieval, generation, fine-tuning, evaluation and latency.

Python · Transformers · Sentence Transformers · PEFT/LoRA · BM25 · cross-encoder · FastAPI · ClearML · Docker · Prometheus · Grafana

02

1-lab · LLM Engineer · car dealership, production

LLM call scoring with measurable reliability

Problem

Every sales call has to be scored against a 15-criterion weighted checklist — without a human, and in a way the score can be trusted. Manually, only 20–30% of calls were reviewed.

What I built
  • Transcription with speaker diarization, utterances split into manager / customer with timestamps.
  • First LLM step — a call-type classifier based on content, not duration: technical, failed, customer coming anyway, full sales conversation.
  • Second step — a scorer: binary answers per criterion plus the objection text, refusal reason and a short summary.
  • Post-validation in code: JSON extraction, snapping to a “0 or criterion weight” grid, field truncation, deterministic final score.
Engineering decisions
  • Two calls instead of one: different tasks and strictness — fewer false zeros on short but substantive calls, and the checklist changes without touching filtering.
  • A set of reference transcripts with expected scores runs on every criteria change; a sample is regularly checked against manual scores, and discrepancies feed back into prompt rules.
  • The classifier is tuned conservatively: better to score an extra call than miss a substantive one.
Result

25,950+ calls scored, 100% coverage instead of 20–30%, scores arrive 1–2 minutes after the deal.

What it shows → The LLM as a reliable component: task decomposition, constrained output, deterministic scoring and regression evaluation.

Python · AssemblyAI · LLM API · Pydantic · PostgreSQL · S3 · Docker · Grafana

03

1-lab · LLM Engineer · industrial manufacturer, production

RAG over a B2B equipment catalog

Problem

Hundreds of SKUs; specs and price lists change regularly. The assistant must not invent specs or prices or confuse similar models.

What I built
  • Ingestion from Google Drive: change detection by hash, incremental re-indexing every 15 minutes, extraction of text and technical tables.
  • Semantic chunking by sections, models and specs rather than character count; metadata: SKU, category, source, date, version.
  • Retrieval with metadata filters and a relevance threshold; grounded generation with structured output (reply, fields, confidence, escalation flag).
  • Evaluation: a golden set of real questions with labelled sources (Recall@K, NDCG@K), groundedness and the share of unsupported answers; online — handoff rate and latency.
Engineering decisions
  • RAG rather than fine-tuning: the catalog and prices change, and the index updates without retraining.
  • An SKU filter keeps search from drifting to a similar but different model; no evidence in context means no number in the answer — the question goes to a human.
  • Document version and indexing time are stored in the index — for any answer you can see which knowledge base version it came from.
Result

The assistant runs 24/7 on an up-to-date base: an updated price list or spec reaches retrieval within 15 minutes.

What it shows → Production RAG on frequently changing data: incremental indexing, metadata, hallucination guards and traceability.

Python · embeddings · vector search · Pydantic · FastAPI · PostgreSQL · Redis · Google Drive API · Docker · Prometheus · Grafana

04

ALP GROUP · ML Engineer · Jul 2022 — Jun 2024

Meeting minutes generation

Problem

Extract decisions, action items, deadlines and owners from a long ASR transcript — without adding anything that was never said.

What I built
  • ASR preprocessing: cleanup, segmentation by time and speaker, entities, timestamp linking.
  • An interpretable baseline: TF-IDF, logistic regression, gradient boosting.
  • RAG: Sentence Transformers + FAISS, BM25, RRF, a reranker; hierarchical summarisation by topic.
  • Validation: Pydantic, evidence checks against fragments, a needs_review status.
Engineering decisions
  • The whole transcript is never sent to the model: costlier, slower and loses details in the middle.
  • Unconfirmed deadlines and owners are suppressed — an empty field for review beats an invention.
  • CI/CD via GitHub Actions, deployment to AWS, Prometheus / Grafana monitoring.
Result

Higher Recall@10 on the golden set and fewer hallucinations on the regression set — a more reliable draft for review.

What it shows → A baseline before the LLM, long-context handling and a validation layer against hallucinations.

Python · scikit-learn · Sentence Transformers · FAISS · BM25 · cross-encoder · Pydantic · FastAPI · Docker · GitHub Actions · AWS

Also at 1-lab

Multi-agent pipeline with a quality gate

Four roles with a JSON contract, a quality threshold and up to 3 iterations, different models and temperatures per role, human-in-the-loop before publishing.

LLM API · agent orchestration · Telethon · PostgreSQL

Stack

Backend
  • Python
  • FastAPI
  • Pydantic
  • PostgreSQL
  • Redis
LLM & models
  • LLM API
  • open-source LLM
  • Hugging Face Transformers
  • Sentence Transformers
  • PEFT / LoRA
RAG & retrieval
  • embeddings
  • FAISS
  • BM25
  • hybrid retrieval
  • RRF
  • cross-encoder reranking
  • metadata filtering
Evaluation
  • golden sets
  • Recall@K
  • NDCG
  • MRR
  • groundedness
  • unsupported claim rate
  • regression suites
  • P50 / P95 latency
  • cache hit rate
MLOps
  • Docker
  • GitHub Actions
  • AWS
  • Prometheus
  • Grafana
  • ClearML (трекинг экспериментов)
Integrations
  • n8n
  • CRM API
  • Google Drive API
  • Telegram API

Experience

  1. Dec 2025 — present

    1-lab · LLM Engineer

    Production LLM services: catalog RAG with incremental indexing, LLM call scoring with deterministic scoring, a multi-agent pipeline with a quality gate.

  2. Aug 2024 — Nov 2025

    YADRO · ML Engineer

    Production RAG over technical documentation: hybrid retrieval, reranking, LoRA/PEFT, claim-level evaluation, latency.

  3. Jul 2022 — Jun 2024

    ALP GROUP · ML Engineer

    NLP and LLMs for meeting minutes: classical baseline, hybrid RAG, hierarchical summarisation, validation.

Let's talk

Open to remote roles. Telegram is the fastest way to reach me.

Serbia · remote · UTC+1  /  Russian — native, English — B1