I work at the intersection of GPU infrastructure engineering, security architecture, and compliance — designing bare-metal platforms that serve real inferencing workloads at scale, hardened for regulatory scrutiny and sovereign-cloud operation.
Not a buzzword list — these are the specific problem spaces I own end to end, from provisioning to production.
Provisioning and operating NVIDIA GPU node pools (H100, L40), MIG partitioning to right-size capacity per workload, and autoscaling driven by real DCGM utilization metrics — not generic CPU/memory thresholds.
Layered defense across policy-as-code (Kyverno), runtime threat detection (Falco), self-service WAF exceptions (Coraza), and network segmentation (Istio/NetworkPolicy) across every traffic scope — intra-namespace through external.
Owning regulatory compliance architecture end to end — audit log immutability, encrypted secrets and PKI, OIDC identity — built so compliance posture is enforced automatically, not checked manually after the fact.
Designing correlation logic from scratch — a multi-tier confidence model that correlates signals across runtime, policy, and audit sources, including cross-cluster detection, on top of Wazuh and OpenSearch.
ArgoCD-driven delivery with zero manual changes to managed resources, and a compliant container supply chain — mirroring, vulnerability scanning, SBOM generation, and cryptographic signing enforced before anything reaches production.
Architecting dual-cluster OpenShift/OKD environments on sovereign infrastructure, including fully air-gapped operation — mirroring and retagging container registries so the platform functions with zero external access.
Production infrastructure supporting real regulated workloads — described here without naming the organizations involved.
Provisioned and operate a dedicated bare-metal GPU cluster, separated from the platform-tooling cluster, running H100 and L40 node pools across multiple availability zones. Deployed the full NVIDIA GPU Operator stack — driver, container toolkit, device plugin, DCGM and DCGM Exporter — via Helm in a fully air-gapped environment, including mirroring and retagging NVIDIA's own container images across OS variants so the stack stays functional with zero external registry access.
Configured MIG partitioning so GPU capacity is right-sized per workload instead of handing out whole GPUs by default, with scheduling that keeps tenant inferencing workloads isolated and predictable rather than contending for resources. Autoscaling is driven by actual GPU utilization (DCGM Exporter → Prometheus), and the infrastructure layer beneath model-serving frameworks — networking, storage, secrets, GPU scheduling — is built to be production-ready regardless of which serving runtime sits on top.
Off-the-shelf SIEM correlation wasn't enough for the signal sources this platform actually produces, so I designed a multi-tier confidence model from scratch — correlating signals across Falco runtime alerts, Kyverno policy events, Kubernetes audit logs, and Wazuh, including detection logic that correlates events across cluster boundaries, not just within one.
Wazuh and OpenSearch operate as the unified SIEM and observability backbone across the platform — including on the GPU nodes themselves, so security visibility doesn't stop at the edge of the tooling cluster.
Own the compliance architecture end to end for a platform operating under BSI C5:2026-aligned requirements: policy-as-code enforcement, immutable and retained audit logs, encrypted secrets and PKI, and OIDC-based identity throughout. Network segmentation via Istio and NetworkPolicy is enforced across every traffic scope — intra-namespace, intra-cluster, inter-cluster, and external — not just at the perimeter.
Playground and production environments run fully isolated — separate PKI, registries, and identity per environment — connected only through GitOps promotion, with a repeatable Terraform-based onboarding pattern so new tenants inherit the same security and compliance posture automatically rather than by manual setup.
Drive GitOps-first delivery via ArgoCD across every environment and cluster — including GPU workloads — with no manual changes to managed resources. Built the supply-chain tooling that enforces this automatically before anything reaches production: image mirroring, vulnerability scanning, SBOM generation, and cryptographic signing via Cosign, Harbor, and Trivy.
A recurring theme across this platform: diagnosing and fixing issues at the root, including root-causing bugs upstream in third-party providers, rather than patching around symptoms — and catching infrastructure risk proactively, before it becomes an incident someone else has to escalate.
Smaller in scale, same standards — a place to build and stress-test ideas end to end without a client on the other side of them.
Built an autonomous agent, running on a locally-hosted open model, that patches and maintains a multi-host Docker fleet unattended — every update backs up data first and auto-rolls-back on a failed health check, with a curated tool surface instead of open shell access and hard-coded boundaries no prompt can override.
Full Prometheus/Grafana stack across a home fleet; when the standard AMD/ROCm GPU exporters misreported unified-memory hardware, wrote a correct one from the actual driver interface instead. Instrumented real LLM inference traffic as histograms for proper percentile latency panels, following the same conventions production inference-serving systems use.
Hardened a self-hosted password manager on the assumption any single layer could fail — network isolation at the firewall, strict headers and edge rate-limiting, admin surface separated from the public login — verifying every control by trying to break it, not just reading its documentation.
20+ self-hosted services across a hypervisor and container fleet, reverse-proxied with automated certificate issuance, treated as one coherent system — DNS, TLS, network policy, and service topology managed deliberately rather than accumulated ad hoc.
Including root-causing bugs upstream in third-party providers when that's where the real problem lives — a patched symptom just moves the failure somewhere less visible.
Policy-as-code, audit immutability, and identity architecture designed so the compliant path is the only path — not a checklist run after the fact.
Giving a system — human or AI — the power to act means deciding, explicitly and in advance, what it's never allowed to do.
Translating architecture and compliance decisions into concrete tradeoffs for stakeholders, rather than silently executing whatever was asked.