Production AI systems
About Slava Dubrov
Production AI systems, evaluation, security, and the evidence loops that improve them.
I’m Slava Dubrov—formally, Dr. Viacheslav Dubrov*—a Staff AI System Engineer working across software engineering, agent systems, evaluation, retrieval, security, and production infrastructure.
These days I work at octonomy AI on production AI and software systems. Before that, at HubSpot, I led work on the Context Layer for retrieval, grounding, and memory infrastructure, then worked on Agent Execution across LLM deployment, fine-tuning, evaluation, safety, and runtime. The connecting thread is what happens after a promising model output meets a real system.
The work behind the point of view
My work has moved from ranking and fraud systems to retrieval infrastructure and production agents.
Selected experience
- HubSpot — Context Layer and Agent Execution. I owned Staff-level architecture and technical leadership across retrieval, grounding, memory, LLM deployment, fine-tuning, evaluation, safety, and runtime.
- Wayfair — fraud, scam detection, and embeddings. I built and led systems associated with about $4M in annual savings.
- OLX — recommendation and search ranking. I worked on ranking and recommendation systems at marketplace scale.
Public evidence
- Work connects engineering claims to articles, source, results, and limitations.
- Labs collects reproducible artifacts that can be run locally.
- Edge of Context holds the long-form engineering record, including the six-part Engineering the Agentic Stack series.
Research and speaking
I hold a Kandidat of Technical Sciences degree (Russian doctoral degree, PhD-equivalent) from Southern Federal University for research in AI diagnostics, with peer-reviewed papers and patents. I also spent a doctoral research semester in Mechatronics at TU Ilmenau on a DAAD Lomonosov scholarship.
At World Agentic AI Summit Berlin 2026, I presented “Engineering the Agentic Stack”: a production architecture covering the cognitive engine, memory, tools, security, and the runtime around them.
Engineering principles
- Make behavior observable before making it autonomous.
- Prefer environment state and outcomes over plausible final answers.
- Make each added harness component earn its cost through ablation.
- Keep evaluators, security policy, and promotion authority independent from the optimizer.
- Preserve failed experiments as evidence.
- Use the cheapest repair layer that solves the real failure.
Capabilities
| System capability | Working areas |
|---|---|
| Agent systems | reasoning loops, memory, tool interfaces, runtimes, harnesses, observability |
| Evaluation and assurance | trace-based evaluation, regression suites, guardrails, permissions, security gates |
| Retrieval and context | semantic and hybrid retrieval, reranking, grounding, context compression |
| Model engineering | fine-tuning, LoRA/QLoRA, serving, quantization, structured generation |
| Production infrastructure | AWS, GCP, Kubernetes, batch, streaming, and real-time ML systems |
Elsewhere
Professional background is on LinkedIn. Source and reproducible artifacts are on GitHub. New essays also appear on Substack.
* Kandidat of Technical Sciences (Russian doctoral degree, PhD-equivalent), awarded by Southern Federal University, Rostov-on-Don, 2015.