BKDownload résumé

Lead Platform Engineer · AI Infrastructure & Reliability

Bekhruz
Kasymov

Lead Platform Engineer with 5+ years of experience building reliable infrastructure for international AI, HealthTech, TravelTech, InsurTech, and banking products. Operated 8 Kubernetes clusters and 30+ microservices across five continents, supporting products used by 333K+ active users and approximately 400 clinics while reducing release time by 80%, cloud spend by 40%, and incident response time by 50%.

333K+ active usersConfidential AI platform
~400 dental clinicsDiagnocat global platform
~6,000 customersWeGoTrip during tenure
8 Kubernetes clustersGCP and Yandex Cloud
80% faster releases4–6 hours to under 1 hour
40% lower cloud spendCleanup and spot autoscaling

Technical profile

Skills mapped to production systems

Core infrastructure

Kubernetes, GCP, Azure, Linux, Docker, PostgreSQL, multi-region and hybrid-cloud production systems

Infrastructure & delivery

Terraform, Terragrunt, Ansible, Helm, ArgoCD, GitLab CI/CD, Jenkins, GitHub Actions, GitOps

Reliability & security

Prometheus, Grafana, Loki, Alertmanager, Sentry, Trivy, SAST/DAST, RBAC, ingress and certificate management

AI platforms & data

vLLM, Ollama, ClearML, KubeRay, GPU workloads, Python, MongoDB, Redis, Kafka, RabbitMQ

Professional experience

Engineering reliable delivery

2026 — Present

Hybrid cloud · International

Confidential AI Platform

Lead Platform Engineer

  • Own a hybrid Kubernetes platform spanning self-hosted development environments, Azure AKS production, and Nebius GPU infrastructure for AI workloads.
  • Operate production infrastructure serving 333K+ active users across a large portfolio of AI generation and productivity workflows.
  • Completed the migration from Bitbucket to self-hosted GitLab, including repository moves, backup retention, protected reviews, CI/CD rewiring, and GitLab Runner deployment on Azure Kubernetes.
  • Implemented highly available PostgreSQL in Azure after a managed-database outage caused mass 5xx responses and authentication failure; strengthened database monitoring and environment maintenance workflows.
  • Built AIOps workflows with HolmesGPT, MCP-connected data analysis, and automated service-error reporting to detect API/ML anomalies and shorten investigation loops.
  • Hardened production ingress against spoofed client-IP headers that could bypass authentication rate limits or block arbitrary users; evaluated Azure WAF controls and trusted-proxy boundaries.
  • Resolved recurring GPU/NVML and AI-generation incidents, added health probes and resource controls, and operated observability across Prometheus, Grafana, Loki, Sentry, Graylog, and Fluent Bit.
  • Built a repeatable Azure proxy layer for external AI providers, solving regional routing, TLS/SNI, certificate, large-payload, WebSocket, IPv6, and timeout edge cases.

2025 — 2026

International · Multi-region production

Diagnocat

Senior DevOps Engineer

  • Managed three outsourced DevOps engineers and coordinated infrastructure strategy across frontend, backend, DevOps, and MLOps teams.
  • Architected and operated eight Kubernetes clusters on GCP and Yandex Cloud, running 30+ microservices for approximately 400 dental clinics across five continents with high availability and disaster recovery.
  • Reduced production release time from 4–6 hours to under one hour by automating tests and deployment workflows—an 80% improvement in time-to-production.
  • Cut cloud infrastructure spend by 40% through resource cleanup and autoscaled spot-based GitLab Runners.
  • Built 30+ GitLab CI/CD pipelines with Trivy, SAST/DAST, Terraform, Ansible, Helm, and ArgoCD-based GitOps controls.
  • Deployed four or more production AI models—including Whisper and open-source LLMs—with KubeRay, Ollama, vLLM, ClearML, and GPU autoscaling; reduced external AI API costs by 4×.
  • Implemented Prometheus, Grafana, Loki, and Alertmanager across the platform and operated Istio, ingress, and certificate layers, reducing incident response time by 50%.

2023 — 2025

International · Remote / Paris

WeGoTrip

Senior DevOps Engineer

  • Owned infrastructure for eight or more services across four cloud providers, supporting 10+ engineers and a platform serving approximately 6,000 customers with zero unplanned downtime.
  • Built and maintained 15+ CI/CD pipelines for frontend and backend services, improving the reliability and repeatability of releases.
  • Reduced cloud infrastructure costs by up to 70% through resource right-sizing, provider selection, and Terraform/Terragrunt automation.
  • Strengthened disaster recovery and shared production knowledge across product engineering teams.

2024

International · Paris, France

Qantev

DevOps Engineer · Contract

  • Delivered an AWS-to-VK Cloud migration in four months while maintaining delivery continuity for an international product team.
  • Supported cloud delivery for an international AI health-insurance platform trusted by 20+ insurer clients worldwide.
  • Built Azure infrastructure with Terraform and Terragrunt for an AI-driven health insurance platform.
  • Optimized delivery workflows and reduced build time by 30%.

2020 — 2022

Moscow, Russia

Sberbank

DevOps Engineer · Software Engineering Intern

  • Built and maintained eight or more Jenkins delivery pipelines using GitLab, SonarQube, Python, and Shell automation.
  • Trained two engineering teams—more than eight engineers—on delivery workflows and operational practices.
  • Reduced deployment errors by 30% through repeatable automation and improved quality controls.

Selected additional engagements

Focused work beyond core roles

Volunteer experience

Service across law, technology, and community

Sep 2017 — Jan 2018 · Science and Technology

Assistant to an Associate Director

Beeline Russia

Created, introduced, maintained, and processed technical documentation for the organization.

Sep 2021 — Present · Science and Technology

Technical Consultant

Gareni / Pamix / Tabler

Advise colleagues on programming languages, databases, automation, and optimization.

May 2023 · Social Services

Distributor

Linkee

Supported a one-month social-services distribution initiative in Paris.

Jun 2023 · Arts and Culture

Administrative Assistant

STATION F

Provided short-term administrative support at the Paris startup campus.

Jun 2023 · Science and Technology

Volunteer Staff

Viva Technology

Supported on-site operations at the Viva Technology event in Paris.

Selected complex systems work

Problems solved in production

01

Diagnocat · Scale · Reliability

Global Kubernetes platform

Architected and maintained eight Kubernetes clusters across GCP and Yandex Cloud, supporting more than 30 microservices in five continental regions with high-availability and disaster-recovery requirements.

Outcome: release cycle reduced from 4–6 hours to under one hour; incident response time reduced by 50%.

02

Confidential AI Platform · PostgreSQL · Azure

Production resilience after DB outage

Responded to an Azure database outage that caused mass 5xx responses and broken authentication, then delivered a highly available PostgreSQL production setup and improved database visibility and maintenance controls.

Closed production work: HA PostgreSQL, replica-load analysis, Grafana datasource and environment cronjob alignment.

03

MLOps · GPU · AIOps

AI platform operations

Combined model-serving operations with automated incident analysis: production LLM inference at Diagnocat; GPU/NVML recovery, readiness controls, HolmesGPT, and MCP-connected service anomaly analysis at a confidential AI platform.

Outcome at Diagnocat: four or more models in production and 4× lower external AI API costs.

04

Confidential AI Platform · Security · Delivery

Secure developer platform

Migrated delivery workflows to GitLab, hardened ingress client-IP trust, introduced managed credential sharing and database RBAC, and investigated WAF protection for the production edge.

Security work covered the full path from source control and secrets to ingress, rate limits, database access, and production networking.

05

Selected product · Full stack · AI

FailUp — AI product built end to end

Designed and built a production personal-growth platform across React, Node.js, Flutter, MongoDB, Redis/BullMQ, Qdrant, and OpenAI APIs, including semantic search, asynchronous AI workflows, encryption, and data export.

Demonstrates product ownership beyond infrastructure: application architecture, mobile and web delivery, AI pipelines, privacy controls, and self-hosted K3s operations.

Latest writing

Ideas behind dependable systems

View all writing

July 25, 2026

5 min read

Systems Thinking

Stability Starts with Discipline

Why real stability comes from personal discipline, responsibility, and the ability to keep acting when circumstances change.

March 28, 2026

5 min read

AI Infrastructure

How to Implement AI in DevOps in One Evening

A practical, low-cost architecture connecting observability, source control, an LLM, and team chat to answer infrastructure questions automatically.

Education

Systems, software engineering, and law

2022 — 2025

42 Paris

RNCP Level 7 · Expert in IT Architecture — Information Systems & Networks

Project-based computer science training centered on independent problem solving, peer learning, software engineering, and systems thinking.

2019 — 2022

School 21 by Sber

Software Engineering Program

Peer-to-peer, project-based training in software engineering, systems programming, and collaborative development.

Graduated 2018

Kutafin Moscow State Law University

Law

Legal education supporting structured analysis, risk assessment, governance, and compliance-minded engineering.

Languages

Russian · Tajik · English

Next opportunity

Ready to make high-scale, cross-layer systems understandable and dependable.

I bring a builder’s mindset, production ownership, and a bias toward clear, automated, maintainable infrastructure.