
I thrive at the intersection of technical complexity and human-centric design, finding my flow when translating ambiguity into clarity. I am driven to build scalable systems that empower teams to succeed.
At a glance
Curated signals on strengths, focus areas, and how they can help.
Led mission-critical engineering initiatives and managed large-scale platforms at industry giants including Google, Netflix, and Datadog.
Currently serves as an Engineering Manager at Pinterest, driving operational excellence and cross-functional technical strategy.
Provides high-value expertise in scaling technical workflows, optimizing team efficiency, and navigating complex organizational change.
🚀 Career trajectory
Deep Reliability Craft
SRE and TLM on Google's Network Control Plane, the layer where a mistake takes everything else down. Built incident response muscle at the sharpest end of production.
✦ SRE
✦ Incident Response
Leading Reliability Orgs
Took the same discipline into leadership at Netflix, Lyft, Datadog, and now Pinterest, running an reliability and developer experience team. Same problem, bigger surface: how teams stay ahead of production instead of chasing it.
✦ Reliability Leadership
✦ Developer Experience
💪🏻 Superpowers
Keeping systems alive when it matters
✦ TLM on Google's Network Control Plane SRE team and helped build the internal incident management platform used across the company.
✦ Brings first responder patterns to incident response: clear command structure, fast triage, post-mortems that people actually read.
On-Call That Does Not Burn People Out
Designing sustainable reliability practice for teams of any size.
✦ Has run and reformed teams across Google, Netflix, Lyft, Datadog, and Pinterest.
✦ Knows the failure modes cold: alert fatigue, rotting runbooks, skipped reviews, and the quiet attrition they cause.
Reliability Leadership Inside Product Companies
Building and defending platform teams where product ships the roadmap.
✦ Leads a large reliability and developer experience group at Pinterest.
✦ Fluent in making the business case for reliability work to leadership that measures everything in features.
I'm excited about
✦ How AI actually changes on-call, incident response & running systems that fail well at scale.
✦ How teams of 20 to 200 engineers handle reliability without a dedicated SRE org
✦ Meeting Xooglers building on the side!
✦ Homelabbing and infrastructure experiements
I can help with
✦ Incident management and SRE practice, from Google scale down to teams that cannot justify a dedicated SRE org
✦ On-call design that does not burn people out: rotation structure, alert quality, escalation hygiene
✦ Why post-mortems rot and how to build a review culture that survives past the first quarter
✦ Running reliability and platform teams inside product companies, including how to defend the roadmap
✦ Navigating infra orgs at Google, Netflix, Datadog, or Pinterest, whether you are joining, growing, or interviewing
I would love your help on
✦ Intros to engineering leaders who own reliability or on-call at Series A to C companies. I want to understand how they handle incidents without an SRE org
✦ GTM lessons from founders who have sold tooling to engineering teams, especially the first ten customers
✦ Stories from Xooglers who made the employed-to-founder jump: what you got right on timing, and what you would do differently
✦ Sharpening my thinking on where AI in production engineering is genuinely useful vs. overhyped. Skeptics especially welcome!