Instructor
Catalin Popescu
AI & DevOps Engineer | Building production LLM/RAG platforms
About me
I build production AI systems — and I run them with the same reliability
discipline I've spent a decade applying to cloud infrastructure. AI is where
most of my work lives now; DevOps is the backbone that keeps it dependable.
On the AI side, I design and operate LLM and RAG platforms end to end:
retrieval pipelines, fine-tuning and inference services (vLLM, MLflow,
JupyterHub), and production integrations on top of OpenAI and Anthropic
models. I treat AI workloads as first-class production systems — they get
the same SLOs, observability, cost controls, and GitOps delivery as anything
else, instead of living as fragile notebooks that break the moment they meet
real traffic.
I also apply AI back into DevOps itself: agentic automation and LLM-assisted
workflows for code review, infrastructure-as-code generation, runbook and
incident triage, and faster root-cause analysis — automation that augments
the platform team rather than adding noise.
That all stands on 10+ years of cloud and platform engineering. I've
architected infrastructure, automated CI/CD, and scaled Kubernetes workloads
across AWS and GCP for companies in aviation, insurance, manufacturing, and
software — including Allianz Technology, Aer Lingus, Molex, and Yields.
My focus is the part that decides whether a platform survives production:
reliable delivery, sane IAM, observability that catches problems before users
do, and infrastructure your own engineers can confidently maintain.
What I work with:
• AI / LLMOps — RAG pipelines, fine-tuning & inference (vLLM, MLflow),
OpenAI & Anthropic in production, agentic automation
• Kubernetes platforms — Helm, ArgoCD/Flux, GitOps delivery, multi-cluster
• Infrastructure as Code — Terraform (HashiCorp Certified), Ansible, Pulumi,
Crossplane
• Cloud — AWS (EKS, RDS, IAM, VPC) and GCP (GKE, Cloud SQL), multi-cloud
landing zones, disaster recovery
• CI/CD & Observability — GitLab CI, Jenkins, Prometheus, Grafana, Loki,
OpenTelemetry
I'm the Director of DigitalCloud Ops Ltd, a UK consultancy building
production-grade Kubernetes platforms, GitOps pipelines, and AI systems for
teams that can't afford to guess — "engineering reliability into modern
infrastructure." I also build Brevyn, a Claude Certified Architect exam-prep
platform.
If you're taking an LLM prototype to production, standing up a Kubernetes
platform, or want AI woven into your delivery and operations the right way —
let's talk.
? Bucharest / London · Open to remote consulting worldwide