keyLogo

Shane Killian

Hi, I'm Shane Killian - Engineering Manager, Reliability & Dev Experience · Ex-Google, Netflix, Lyft, Datadog · SRE and incident management

I thrive at the intersection of technical complexity and human-centric design, finding my flow when translating ambiguity into clarity. I am driven to build scalable systems that empower teams to succeed.

Pinterest
Austin, TX, USA
AI insights

At a glance

Curated signals on strengths, focus areas, and how they can help.

Led mission-critical engineering initiatives and managed large-scale platforms at industry giants including Google, Netflix, and Datadog.

Currently serves as an Engineering Manager at Pinterest, driving operational excellence and cross-functional technical strategy.

Provides high-value expertise in scaling technical workflows, optimizing team efficiency, and navigating complex organizational change.

🚀 Career trajectory

Deep Reliability Craft

SRE and TLM on Google's Network Control Plane, the layer where a mistake takes everything else down. Built incident response muscle at the sharpest end of production.

SRE

Incident Response

Leading Reliability Orgs

Took the same discipline into leadership at Netflix, Lyft, Datadog, and now Pinterest, running an reliability and developer experience team. Same problem, bigger surface: how teams stay ahead of production instead of chasing it.

Reliability Leadership

Developer Experience

💪🏻 Superpowers

Keeping systems alive when it matters

TLM on Google's Network Control Plane SRE team and helped build the internal incident management platform used across the company.

Brings first responder patterns to incident response: clear command structure, fast triage, post-mortems that people actually read.

On-Call That Does Not Burn People Out

Designing sustainable reliability practice for teams of any size.

Has run and reformed teams across Google, Netflix, Lyft, Datadog, and Pinterest.

Knows the failure modes cold: alert fatigue, rotting runbooks, skipped reviews, and the quiet attrition they cause.

Reliability Leadership Inside Product Companies

Building and defending platform teams where product ships the roadmap.

Leads a large reliability and developer experience group at Pinterest.

Fluent in making the business case for reliability work to leadership that measures everything in features.

I'm excited about

How AI actually changes on-call, incident response & running systems that fail well at scale.

How teams of 20 to 200 engineers handle reliability without a dedicated SRE org

Meeting Xooglers building on the side!

Homelabbing and infrastructure experiements

I can help with

Incident management and SRE practice, from Google scale down to teams that cannot justify a dedicated SRE org

On-call design that does not burn people out: rotation structure, alert quality, escalation hygiene

Why post-mortems rot and how to build a review culture that survives past the first quarter

Running reliability and platform teams inside product companies, including how to defend the roadmap

Navigating infra orgs at Google, Netflix, Datadog, or Pinterest, whether you are joining, growing, or interviewing

I would love your help on

Intros to engineering leaders who own reliability or on-call at Series A to C companies. I want to understand how they handle incidents without an SRE org

GTM lessons from founders who have sold tooling to engineering teams, especially the first ten customers

Stories from Xooglers who made the employed-to-founder jump: what you got right on timing, and what you would do differently

Sharpening my thinking on where AI in production engineering is genuinely useful vs. overhyped. Skeptics especially welcome!