Byte, the Edge of Context beaver

Production AI systems

About Slava Dubrov

Production AI systems, evaluation, security, and the evidence loops that improve them.

I’m Slava Dubrov—formally, Dr. Viacheslav Dubrov*—a Staff AI System Engineer working across software engineering, agent systems, evaluation, retrieval, security, and production infrastructure.

These days I work at octonomy AI on production AI and software systems. Before that, at HubSpot, I led work on the Context Layer for retrieval, grounding, and memory infrastructure, then worked on Agent Execution across LLM deployment, fine-tuning, evaluation, safety, and runtime. The connecting thread is what happens after a promising model output meets a real system.


The work behind the point of view

My work has moved from ranking and fraud systems to retrieval infrastructure and production agents.

Selected experience

  • HubSpot — Context Layer and Agent Execution. I owned Staff-level architecture and technical leadership across retrieval, grounding, memory, LLM deployment, fine-tuning, evaluation, safety, and runtime.
  • Wayfair — fraud, scam detection, and embeddings. I built and led systems associated with about $4M in annual savings.
  • OLX — recommendation and search ranking. I worked on ranking and recommendation systems at marketplace scale.

Public evidence

  • Work connects engineering claims to articles, source, results, and limitations.
  • Labs collects reproducible artifacts that can be run locally.
  • Edge of Context holds the long-form engineering record, including the six-part Engineering the Agentic Stack series.

Research and speaking

I hold a Kandidat of Technical Sciences degree (Russian doctoral degree, PhD-equivalent) from Southern Federal University for research in AI diagnostics, with peer-reviewed papers and patents. I also spent a doctoral research semester in Mechatronics at TU Ilmenau on a DAAD Lomonosov scholarship.

At World Agentic AI Summit Berlin 2026, I presented “Engineering the Agentic Stack”: a production architecture covering the cognitive engine, memory, tools, security, and the runtime around them.

Engineering principles

  • Make behavior observable before making it autonomous.
  • Prefer environment state and outcomes over plausible final answers.
  • Make each added harness component earn its cost through ablation.
  • Keep evaluators, security policy, and promotion authority independent from the optimizer.
  • Preserve failed experiments as evidence.
  • Use the cheapest repair layer that solves the real failure.

Capabilities

System capabilityWorking areas
Agent systemsreasoning loops, memory, tool interfaces, runtimes, harnesses, observability
Evaluation and assurancetrace-based evaluation, regression suites, guardrails, permissions, security gates
Retrieval and contextsemantic and hybrid retrieval, reranking, grounding, context compression
Model engineeringfine-tuning, LoRA/QLoRA, serving, quantization, structured generation
Production infrastructureAWS, GCP, Kubernetes, batch, streaming, and real-time ML systems

Elsewhere

Professional background is on LinkedIn. Source and reproducible artifacts are on GitHub. New essays also appear on Substack.


* Kandidat of Technical Sciences (Russian doctoral degree, PhD-equivalent), awarded by Southern Federal University, Rostov-on-Don, 2015.