Independent consulting / Principal engineer / Berlin

Platform decisions that survive production.

Hands-on platform and ML infrastructure consulting, from architecture through the critical build and into a system the team can own.

  1. 01 GPU Kubernetes
  2. 02 Model serving
  3. 03 Agent infrastructure
  4. 04 Platform reliability

End to end

Architecture, implementation, operations, and handoff

3 clouds

AWS, GCP, and Azure platform experience

11

Engineers enabled to extend a delivered serving architecture

Hands-on

Production Python, Go, TypeScript, Terraform, and Kubernetes
01

Consulting premise

Leave a system, not a dependency.

The useful consulting outcome is not a deck or a permanent expert-shaped gap. It is a working critical path, recorded decisions, production controls, and a team that can extend the platform without the consultant in the room.

Principal-level means hands-on

I write the architecture and the production code, review the failure modes, and transfer the operating model with the implementation.

02

Engagement method

From bottleneck to team ownership.

The work moves through four explicit states. Each produces an artifact the team can inspect and continue using.

  1. 01

    Frame

    Find the real bottleneck and operating boundary

  2. 02

    Decide

    Write the architecture and irreversible choices

  3. 03

    Build

    Implement the critical path in production code

  4. 04

    Transfer

    Leave runbooks, milestones, and team ownership

03

Capabilities

The hard seams between systems.

Drizzle is strongest where cloud infrastructure, software control planes, model runtimes, and the team operating them meet.

C1

ML platform + serving

GPU Kubernetes, KServe, vLLM, KEDA, model registries, inference contracts, and the path from notebook to observable service.

KServe / vLLM / KEDA / MLflow / GPU scheduling
C2

Platform foundations

Cloud and Kubernetes foundations with infrastructure as code, GitOps, self-service interfaces, and production operating models.

AWS / GCP / Azure / Terraform / EKS / Argo CD
C3

Agent infrastructure

Safe tools for agents to operate cloud systems, with tenancy, credentials, constrained interfaces, traces, and measurable actions.

Python / FastMCP / OpenTelemetry / Grafana Cloud
C4

Reliability + observability

Telemetry architecture, SLOs, error budgets, failure modes, and the operating practices that make a platform supportable.

OpenTelemetry / Prometheus / Grafana / SLOs
04

Selected engagements

Architecture proven through delivery.

Two public examples of the practice: one centered on model serving, one on the cloud and agent control plane around it.

05

Engagement shapes

Bounded work with a concrete exit.

The format follows the decision or build that needs to move, not an open-ended staffing slot.

  1. 01

    Architecture sprint

    A bounded decision phase for teams that know the destination but not the operating design.

  2. 02

    Critical-path build

    Hands-on implementation of the difficult control plane, serving path, or platform foundation.

  3. 03

    Platform assessment

    A production-readiness review across failure modes, cost, interfaces, security, and team ownership.

Every engagement leaves

Decisions / Production code / Runbooks / Milestones / Ownership

Drizzle Systems / Platform and ML infrastructure

Bring the hard platform problem.

Start a conversation 01Visit drizzle.systems 02View experience 03