Open to GPU infrastructure & AI platform architecture roles

Building the infrastructure AI inferencing actually depends on.

I work at the intersection of GPU infrastructure engineering, security architecture, and compliance — designing bare-metal platforms that serve real inferencing workloads at scale, hardened for regulatory scrutiny and sovereign-cloud operation.

18,000Concurrent inferencing users, capacity-planned
H100 / L40GPU node pools across multiple AZs
BSI C5:2026Compliance architecture, owned end to end
ZeroManual changes to managed resources
platform-status — zsh
$ kubectl get nodes -l workload=gpu
NAME GPU MIG STATUS
gpu-node-01 H100 on Ready
gpu-node-02 H100 on Ready
gpu-node-03 L40 on Ready
 
$ platform compliance --check bsi-c5
✓ audit log integrity ... immutable
✓ secrets & PKI ......... encrypted
✓ identity ............... OIDC-enforced
✓ network policy ......... zero-trust
 
$
Core expertise

Where I actually add value

Not a buzzword list — these are the specific problem spaces I own end to end, from provisioning to production.

GPU & inference platform engineering

Provisioning and operating NVIDIA GPU node pools (H100, L40), MIG partitioning to right-size capacity per workload, and autoscaling driven by real DCGM utilization metrics — not generic CPU/memory thresholds.

Zero-trust security architecture

Layered defense across policy-as-code (Kyverno), runtime threat detection (Falco), self-service WAF exceptions (Coraza), and network segmentation (Istio/NetworkPolicy) across every traffic scope — intra-namespace through external.

Compliance-as-code

Owning regulatory compliance architecture end to end — audit log immutability, encrypted secrets and PKI, OIDC identity — built so compliance posture is enforced automatically, not checked manually after the fact.

Detection engineering & SIEM

Designing correlation logic from scratch — a multi-tier confidence model that correlates signals across runtime, policy, and audit sources, including cross-cluster detection, on top of Wazuh and OpenSearch.

GitOps & supply chain security

ArgoCD-driven delivery with zero manual changes to managed resources, and a compliant container supply chain — mirroring, vulnerability scanning, SBOM generation, and cryptographic signing enforced before anything reaches production.

Bare-metal & sovereign platforms

Architecting dual-cluster OpenShift/OKD environments on sovereign infrastructure, including fully air-gapped operation — mirroring and retagging container registries so the platform functions with zero external access.

Platform work

Systems I've architected and operate

Production infrastructure supporting real regulated workloads — described here without naming the organizations involved.

01
GPU / Inference Platform

A bare-metal GPU inferencing platform sized for 18,000 concurrent users

NVIDIA GPU Operator MIG DCGM Terraform Ansible

Provisioned and operate a dedicated bare-metal GPU cluster, separated from the platform-tooling cluster, running H100 and L40 node pools across multiple availability zones. Deployed the full NVIDIA GPU Operator stack — driver, container toolkit, device plugin, DCGM and DCGM Exporter — via Helm in a fully air-gapped environment, including mirroring and retagging NVIDIA's own container images across OS variants so the stack stays functional with zero external registry access.

Configured MIG partitioning so GPU capacity is right-sized per workload instead of handing out whole GPUs by default, with scheduling that keeps tenant inferencing workloads isolated and predictable rather than contending for resources. Autoscaling is driven by actual GPU utilization (DCGM Exporter → Prometheus), and the infrastructure layer beneath model-serving frameworks — networking, storage, secrets, GPU scheduling — is built to be production-ready regardless of which serving runtime sits on top.

02
Detection Engineering

A custom SIEM correlation engine, built from scratch

Wazuh OpenSearch Falco K8s audit

Off-the-shelf SIEM correlation wasn't enough for the signal sources this platform actually produces, so I designed a multi-tier confidence model from scratch — correlating signals across Falco runtime alerts, Kyverno policy events, Kubernetes audit logs, and Wazuh, including detection logic that correlates events across cluster boundaries, not just within one.

Wazuh and OpenSearch operate as the unified SIEM and observability backbone across the platform — including on the GPU nodes themselves, so security visibility doesn't stop at the edge of the tooling cluster.

03
Security & Compliance

Zero-trust security architecture aligned to BSI C5:2026

Kyverno Istio Coraza WAF Keycloak/OIDC

Own the compliance architecture end to end for a platform operating under BSI C5:2026-aligned requirements: policy-as-code enforcement, immutable and retained audit logs, encrypted secrets and PKI, and OIDC-based identity throughout. Network segmentation via Istio and NetworkPolicy is enforced across every traffic scope — intra-namespace, intra-cluster, inter-cluster, and external — not just at the perimeter.

Playground and production environments run fully isolated — separate PKI, registries, and identity per environment — connected only through GitOps promotion, with a repeatable Terraform-based onboarding pattern so new tenants inherit the same security and compliance posture automatically rather than by manual setup.

04
GitOps & Supply Chain

A compliant delivery pipeline with zero manual production changes

ArgoCD Cosign Trivy Harbor

Drive GitOps-first delivery via ArgoCD across every environment and cluster — including GPU workloads — with no manual changes to managed resources. Built the supply-chain tooling that enforces this automatically before anything reaches production: image mirroring, vulnerability scanning, SBOM generation, and cryptographic signing via Cosign, Harbor, and Trivy.

A recurring theme across this platform: diagnosing and fixing issues at the root, including root-causing bugs upstream in third-party providers, rather than patching around symptoms — and catching infrastructure risk proactively, before it becomes an incident someone else has to escalate.

Personal infrastructure lab

Independent work, built outside the day job

Smaller in scale, same standards — a place to build and stress-test ideas end to end without a client on the other side of them.

Independent projects, run and paid for personally — not affiliated with any employer or client.
Autonomous systems

An LLM-driven ops agent with real safety engineering

Built an autonomous agent, running on a locally-hosted open model, that patches and maintains a multi-host Docker fleet unattended — every update backs up data first and auto-rolls-back on a failed health check, with a curated tool surface instead of open shell access and hard-coded boundaries no prompt can override.

Observability

A metrics platform built where the standard exporters didn't fit

Full Prometheus/Grafana stack across a home fleet; when the standard AMD/ROCm GPU exporters misreported unified-memory hardware, wrote a correct one from the actual driver interface instead. Instrumented real LLM inference traffic as histograms for proper percentile latency panels, following the same conventions production inference-serving systems use.

Security

Defense-in-depth on a self-hosted credential store

Hardened a self-hosted password manager on the assumption any single layer could fail — network isolation at the firewall, strict headers and edge rate-limiting, admin surface separated from the public login — verifying every control by trying to break it, not just reading its documentation.

Platform

A self-hosted platform run like production

20+ self-hosted services across a hypervisor and container fleet, reverse-proxied with automated certificate issuance, treated as one coherent system — DNS, TLS, network policy, and service topology managed deliberately rather than accumulated ad hoc.

How I work

The engineering principles behind all of it

01

Root cause, not symptoms

Including root-causing bugs upstream in third-party providers when that's where the real problem lives — a patched symptom just moves the failure somewhere less visible.

02

Compliance built in, not bolted on

Policy-as-code, audit immutability, and identity architecture designed so the compliant path is the only path — not a checklist run after the fact.

03

Autonomy needs boundaries

Giving a system — human or AI — the power to act means deciding, explicitly and in advance, what it's never allowed to do.

04

Tradeoffs, not just execution

Translating architecture and compliance decisions into concrete tradeoffs for stakeholders, rather than silently executing whatever was asked.

Get in touch

Let's talk infrastructure.

Open to GPU infrastructure, AI platform engineering, and security/compliance architecture roles where scale and rigor both matter.