Resume
The complete record – where I’ve worked and what I actually did there. If you just want the shape of it, About has the three-sentence version.
NVIDIA (December 2024 – present)
Senior AI Infrastructure Engineer, DGX Cloud Operational Excellence (June 2026 – present)
- Revamping DGX Cloud’s incident management practice around on-call health and learning from incidents.
Senior AI Infrastructure Engineer, DGX Cloud Infrastructure Security (January 2026 – June 2026)
- Founding member of the team, focused on building the data pipeline for agentic security workflows.
Senior AI Infrastructure Engineer, DGX Cloud COSI – Colos, On-prem, and Systems Images (December 2024 – January 2026)
- Repaved one-off GPU clusters with thousands of H100 GPUs that weren’t deployed with infrastructure-as-code into end-to-end IaC-managed clusters.
- Identified and automated the response to Linux, Slurm, and GPU (XID) errors as well as other hardware failures.
Honeycomb (May 2022 – December 2024)
Staff Platform Engineer, Platform Core (May 2022 – December 2024)
- Moved Retriever – Honeycomb’s query and storage engine – from EC2 to Kubernetes, which first meant finding and fixing a Linux kernel bug that crashed nodes during LVM snapshots.
- Tracked a performance regression down to the Go gRPC library doing an excessive number of DNS lookups on Kubernetes; fixed it by deploying NodeLocal DNSCache.
- Added a safety check to Kafka rolling restarts: a ZooKeeper lock that won’t let a broker upgrade until every replica is back in sync. Only one broker gets to be down at a time, no matter who pushes the button.
- Cut Kafka producer latency by 20% with garbage collection and heap tuning.
Time off (November 2021 – May 2022)
- Took a break.
Netflix (February 2019 – November 2021)
Senior Software Engineer, OS/Compute (April 2021 – November 2021)
- Built the service that emits metrics and events whenever the OOM killer fires on any of 250,000+ EC2 nodes, then worked with the affected teams to right-size their instances and stop the unexpected restarts.
- One of four engineers on call for the Base OS team – the Linux layer under every instance and container at Netflix.
Senior DevOps Engineer, Personalization Infrastructure (February 2019 – April 2021)
- Ran the platform for the ~60 engineers and researchers training and deploying the ML models behind Netflix recommendations – row selection, box art selection, billboard selection.
- Operated Apache Spark clusters on Mesos across tens of thousands of EC2 instances.
- Helped develop and deploy a method of sharding Mesos behind ZooKeeper after we hit its scaling ceiling around 12,000 nodes.
- Built the system that tracked data fidelity and model quality, so regressions got caught – and root-caused – before members noticed them.
Co-VP, Trans* Employee Resource Group (February 2019 – November 2021)
- Led the Education and Talent Acquisition pods, then served as Co-VP of the ERG itself.
- Developed and delivered training with the Inclusion & Diversity team on how to be a better colleague to trans and gender non-conforming people.
- Consulted with content teams on trans representation, both in front of and behind the camera. (How that ended is its own story.)
Blizzard Entertainment (January 2015 – January 2019)
Senior Systems Engineer, Big Data (February 2018 – January 2019)
- Planned and executed a zero-downtime migration of the production Kafka clusters – 20+ billion events a day – from a commercial distribution to Apache Kafka, dropping the licensing costs on the way out.
- Deployed Kafka on Kubernetes with StatefulSets, securely mirroring production data to support moving the Global Data Platform into the public cloud.
- Replaced a commercial monitoring vendor with Puppet, Telegraf, Jolokia, and Grafana – $100K+ a year cheaper, and we could finally see what was actually going on.
- Debugged open source and in-house applications in Scala, Java, and Python to help developers speed up their jobs and stop wasting resources.
UNIX Systems Engineer (January 2015 – February 2018)
- Built and ran the telemetry pipeline behind Overwatch, World of Warcraft, and Hearthstone – service health for 40+ million monthly players.
- Designed and deployed the globally distributed ZooKeeper architecture that gave Overwatch service discovery and picked the game server with the best latency for everyone in the match.
- Migrated a multi-petabyte Hadoop data warehouse to new hardware with zero downtime by dual-writing through Kafka – 4x the compute and memory, enough to actually run Spark and Impala.
- Moved the global fan-in data pipeline from RabbitMQ to Kafka: 3x the capacity, trillions of events retained, and the end of losing data.
- Cut KPI report time from 3 hours to 10 minutes by migrating jobs from Streaming MapReduce to SQL on Hive, Spark, and Impala.
Cloudera (August 2013 – December 2014)
Solutions Consultant
- Deployed, configured, and tuned dozens of Apache Hadoop clusters running Accumulo and HBase – kernel, network, filesystem, and memory management.
- Owned the Hadoop and Active Directory/Kerberos integration, eliminating a vendor dependency.
Berico Technologies / 42six (November 2010 – August 2013)
IT Manager, then Systems Engineer
- Ran corporate IT through the spinoff to 42six.
- Deployed the Hadoop and Accumulo clusters the developers built on, kept the public-facing web servers up, and stood up the Jenkins CI that built everything.
Clayton Public Schools (July 2004 – November 2010)
IT Manager / Network Administrator
- Ran a New Jersey K-12 district’s network – 700+ computers and a redundant gigabit fiber backbone installed during a $20M renovation.
- Laptops for every teacher on a zero-growth budget, and an internship program for students considering a career in IT.