Principal AI/ML Platform Engineer
I make AI systems operable.
GPU Kubernetes, model serving, and internal tooling built for the engineers and agents that run production.
14 years
Across software, cloud, SRE, and AI platforms
$1.5M / yr
Observability cost removed
3,000+ services
Observed in production
300+ environments
Operated across AWS and GCP
01 / Selected impact
What changed after the platform shipped.
Architecture matters when it changes cost, autonomy, reliability, or the speed at which other engineers can move.
- 01Model serving
One interface for every model.
Architected model serving on hybrid GPU Kubernetes with KServe, vLLM-aware KEDA scaling, dynamic GPU nodes, MLflow versioning, and reusable transformer patterns.
An 11-person team now extends the platform independently.KServe / vLLM / KEDA / MLflow - 02AI platform
Made AI a first-class platform workload.
Built and operated Spryker's internal AI platform for multi-GPU, multi-LLM, and multi-agent workloads on EKS.
One operating model for applications and AI.EKS / Karpenter / NVIDIA / LangGraph - 03Observability
Rebuilt the telemetry layer, then turned it into a product.
Moved more than 3,000 distributed services from New Relic to OpenTelemetry, Prometheus, and the Grafana stack.
$1.5M saved every year. Sold to five enterprise customers.OTel / Prometheus / Loki / Tempo / Mimir - 04Agent operations
Built tools that agents can operate safely.
Developed multi-tenant MCP servers that let agents operate AWS, GCP, and Azure accounts, with end-to-end traces and metrics for every action.
Agent-facing infrastructure with human-grade controls.Python / FastMCP / OpenTelemetry / Grafana Cloud - 05Developer platforms
Measured the platform by adoption, not feature count.
Established self-service provisioning and golden paths at Spryker, AUTO1, and ING using Terraform, Backstage, and Argo CD.
Hundreds of service teams adopted the paved roads.Terraform / Backstage / Argo CD
02 / Experience
From cloud foundations to production AI.
Each chapter expanded the operating surface: software, cloud, reliability, developer platforms, then model serving and agent infrastructure.
2012 - presentRoam AI
Principal Engineer, ML, Data and Platform Engineering
Own the ML and data platform end to end: cloud infrastructure, backend services, data contracts, and model serving on hybrid GPU Kubernetes.
Production Python and TypeScript / design reviews / mentoring and hiring
Founder
Building a code-integrity platform for agent-assisted engineering. Deterministic probes issue blocking verdicts; LLM inference never does.
CLI / VS Code extension / agent harness / in-toto attestations
Drizzle Systems
Principal Consultant, Platform Engineering and MLOps
Designed Roam AI's MLOps reference architecture and built Blocks Cloud's AWS platform, production MCP servers, and agent observability stack.
Hybrid GPU Kubernetes / AWS platform engineering / agent infrastructure
Spryker
Senior Staff Engineer and Deputy Head of Cloud and SRE
Designed the internal AI platform, led the observability rebuild, and operated 300+ customer environments and 3,000+ distributed services.
Promoted from Lead Engineer after six months
AUTO1 Group
Team Lead, Site Reliability Engineering
Led reliability for more than 5,000 microservices and shipped the SLO operating model and tooling adopted by hundreds of teams.
Multi-account AWS / SLOs / reliability tooling
ING / Lendico
DevOps and SRE Lead
Led the build and run of a regulated lending platform on Azure AKS, including custom Go operators, admission control, Vault, and observability.
Promoted from Senior DevOps Engineer after six months
Deutsche Bank / Yunar
Cloud Platform and Site Reliability Expert
Operated an Istio-backed Azure AKS platform of roughly 150 microservices and defined its SLI, SLO, and SLM-based platform offer.
AKS / Istio / reliability architecture
Smile Open Source Solutions
Cloud and DevOps Consulting Engineer
Delivered cloud engagements for Renault, Monoprix, and Svensk e-identitet across Azure, AWS, GCP, and OpenStack.
Consulting / multi-cloud / open source
Rosafi Holding + Tunisian Cloud
Senior Software Engineer and Cloud Platform Engineer
Built storage, database, and big-data platforms as self-service products, beginning a career at the seam between software and infrastructure.
Python / distributed storage / data platforms
03 / Operating range
The tools follow the system.
Deepest in platform engineering and production operations, with enough software depth to build the control plane rather than just configure it.
- 01
ML platform + serving
KServe, vLLM, KubeRay, Kubeflow, MLflow, NVIDIA device plugin, GPU scheduling, LangGraph, MCP
- 02
Kubernetes + cloud
EKS, GKE, AKS, K3S, bare metal, Karpenter, KEDA, Istio, Linkerd, AWS, GCP, Azure, Hetzner
- 03
Automation + delivery
Terraform, Terragrunt, Pulumi, Ansible, Helm, Argo CD, Atlantis, GitHub Actions, GitLab CI, Backstage
- 04
Observability
OpenTelemetry, Prometheus, Grafana, Loki, Tempo, Mimir, Datadog, distributed tracing, SLOs, error budgets
- 05
Languages
Python, FastAPI, Flask, Celery, Go, Kubebuilder, TypeScript, Bash, SQL
- 06
Data + orchestration
PostgreSQL, Redis, Kafka, Airflow, Spark, Ceph, S3-compatible storage
04 / Beyond delivery
Teaching, disclosure, and foundations.
Sharing the operating lessons, and treating security as part of the platform.
Self-hosting DeepSeek-R1 on AWS EKS
A practical walkthrough of vLLM inference, KubeRay distributed scaling, and Terraform with Karpenter for GPU infrastructure.
Microsoft Security Response Center acknowledgment
Co-disclosed an Azure Kubernetes Service privilege-escalation misconfiguration.
M.Eng., Computer Systems Networking and TelecommunicationsENET'Com Sfax, Tunisia
B.Eng., Mathematics and Computer SciencePreparatory Institute for Engineering Studies, Monastir
FRFrenchNative
ENEnglishC2
DEGermanB1
Berlin / Open to a good systems conversation