Independent consulting / Principal engineer / Berlin
Platform decisions that survive production.
Hands-on platform and ML infrastructure consulting, from architecture through the critical build and into a system the team can own.
- 01 GPU Kubernetes
- 02 Model serving
- 03 Agent infrastructure
- 04 Platform reliability
End to end
Architecture, implementation, operations, and handoff3 clouds
AWS, GCP, and Azure platform experience11
Engineers enabled to extend a delivered serving architectureHands-on
Production Python, Go, TypeScript, Terraform, and KubernetesConsulting premise
Leave a system, not a dependency.
The useful consulting outcome is not a deck or a permanent expert-shaped gap. It is a working critical path, recorded decisions, production controls, and a team that can extend the platform without the consultant in the room.
I write the architecture and the production code, review the failure modes, and transfer the operating model with the implementation.
Engagement method
From bottleneck to team ownership.
The work moves through four explicit states. Each produces an artifact the team can inspect and continue using.
- 01
Frame
Find the real bottleneck and operating boundary
- 02
Decide
Write the architecture and irreversible choices
- 03
Build
Implement the critical path in production code
- 04
Transfer
Leave runbooks, milestones, and team ownership
Capabilities
The hard seams between systems.
Drizzle is strongest where cloud infrastructure, software control planes, model runtimes, and the team operating them meet.
ML platform + serving
GPU Kubernetes, KServe, vLLM, KEDA, model registries, inference contracts, and the path from notebook to observable service.
KServe / vLLM / KEDA / MLflow / GPU schedulingPlatform foundations
Cloud and Kubernetes foundations with infrastructure as code, GitOps, self-service interfaces, and production operating models.
AWS / GCP / Azure / Terraform / EKS / Argo CDAgent infrastructure
Safe tools for agents to operate cloud systems, with tenancy, credentials, constrained interfaces, traces, and measurable actions.
Python / FastMCP / OpenTelemetry / Grafana CloudReliability + observability
Telemetry architecture, SLOs, error budgets, failure modes, and the operating practices that make a platform supportable.
OpenTelemetry / Prometheus / Grafana / SLOsSelected engagements
Architecture proven through delivery.
Two public examples of the practice: one centered on model serving, one on the cloud and agent control plane around it.
A model-serving platform the team could extend.
Designed the hybrid GPU Kubernetes architecture around KServe, KEDA scaling on vLLM metrics, dynamic GPU provisioning, MLflow model versioning, and reusable transformer patterns.
- Architecture decision records
- Runbooks
- Milestone plan
- 11-person team handoff
- 01Model
- 02KServe
- 03vLLM
- 04GPU
The AWS foundation and the interfaces agents could safely use.
Built the platform from scratch with Terraform, EKS, Helm, Argo CD, and GitHub Actions, then developed multi-tenant cloud-operation tools and end-to-end agent observability.
- AWS platform foundation
- Production MCP servers
- Agent traces + metrics
- GitOps delivery
- 01Agent
- 02MCP
- 03Policy
- 04Cloud
Engagement shapes
Bounded work with a concrete exit.
The format follows the decision or build that needs to move, not an open-ended staffing slot.
- 01
Architecture sprint
A bounded decision phase for teams that know the destination but not the operating design.
- 02
Critical-path build
Hands-on implementation of the difficult control plane, serving path, or platform foundation.
- 03
Platform assessment
A production-readiness review across failure modes, cost, interfaces, security, and team ownership.
Decisions / Production code / Runbooks / Milestones / Ownership
Drizzle Systems / Platform and ML infrastructure