keyLogo

Roberto Lupi

Hi, I'm Roberto Lupi - Senior Site Reliability Engineer (ML Infrastructure) at Google

I am passionate about solving complex reliability challenges in machine learning infrastructure. I thrive in high-scale engineering environments where I can build, optimize, and scale robust technical systems.

Google
Zürich, Switzerland
AI insights

At a glance

Curated signals on strengths, focus areas, and how they can help.

Transitioned from general web development to specializing in global-scale systems reliability.

Currently leads high-performance machine learning infrastructure projects within large-scale distributed environments.

Offers deep expertise in site reliability engineering, system optimization, and technical infrastructure strategy.

🚀 Career trajectory

Early Web Exploration

Started in virtual worlds and web development, building the foundational understanding of distributed user environments.

Consulting

Web Development

Core SRE Evolution

Transitioned to Google as a Systems Engineer to specialize in SRE, mastering the scale of global infrastructure.

Site Reliability

Linux Administration

ML Infrastructure Specialization

Currently focused on the convergence of machine learning and infrastructure, leading high-impact ML platform initiatives.

ML Infrastructure

System Design

💪🏻 Superpowers

Machine Learning Infrastructure Architect

Designing scalable frameworks for high-stakes models.

Architecting robust infrastructure to support high-performance ML models.

Optimizing system reliability in complex, data-heavy distributed environments.

Implementing scalable solutions for mission-critical engineering pipelines.

Systems Reliability Evangelist

Championing operational excellence and stability.

Applying disciplined SRE practices to increase system uptime.

Bridging gaps between development teams and production operations.

Refining automated workflows to enhance organizational deployment speed.

Technical Capability Multiplier

Leveraging deep expertise to empower technical teams.

Mentoring engineers on complex system design and architecture.

Facilitating cross-functional collaboration for large-scale migrations.

Translating complex technical requirements into actionable team goals.

I'm excited about

Exploring advancements in high-scale machine learning and AI infrastructure.

Connecting with peers to discuss future trends in distributed systems.

Identifying opportunities to share insights on engineering reliability at scale.

I can help with

Advising on system architecture for machine learning and data pipelines.

Consulting on SRE best practices to improve production environment stability.

Mentoring junior engineers looking to specialize in reliability and infrastructure.

I would love your help on

Finding new perspectives on emerging trends in distributed computing.

Identifying cross-industry use cases for advanced ML infrastructure.

Learning effective strategies for building resilient, high-growth technical cultures.