keyLogo

Long Cheng

Hi, I'm Long Cheng - Engineering Leader @ Google Vertex AI at Google

I am obsessed with the intersection of AI research and planetary-scale production. I thrive in high-stakes environments where I can build resilient systems and lead teams to solve existential engineering challenges.

Google
Woodinville, WA, USA
AI insights

At a glance

Curated signals on strengths, focus areas, and how they can help.

Built deep expertise in distributed systems through foundational roles at Microsoft and Oracle.

Currently leads the engineering serving backbone for Google Gemini at planetary scale.

Can help teams optimize LLM inference performance and reduce operational infrastructure costs.

🚀 Career trajectory

Systems Foundation

Established core competency in large-scale distributed systems at Microsoft and Oracle.

Distributed Systems

Enterprise Software

AI Infrastructure Leadership

Pivoted into deep AI infra at Google, now driving the serving architecture for Gemini.

LLM Inference

Cloud Scale

💪🏻 Superpowers

Hyper-Scale Inference Architect

Optimizing compute for planetary AI demands

Architecting low-latency serving backbones for foundational models.

Maximizing token-per-second efficiency across global data centers.

Solving hardware scarcity through radical infra optimization.

Engineering Culture Catalyst

Translating technical ambiguity into impact

Aligning engineering teams with high-stakes business objectives.

Scaling teams through periods of intense product evolution.

Translating research prototypes into high-availability commercial engines.

Cross-Domain Research Liaison

Bridging scientific theory and production realities

Translating model architecture breakthroughs into operational realities.

Partnering with research scientists to iterate on deployment strategies.

Mitigating deployment bottlenecks for bleeding-edge generative AI.

I'm excited about

Exploring cross-industry collaborations on future AI infrastructure paradigms.

Engaging with leaders defining the next wave of generative compute.

Finding new avenues to mentor the next generation of infra architects.

I can help with

Advising on scaling LLM inference architectures for production workloads.

Providing guidance on optimizing AI unit economics and infrastructure costs.

Consulting on building engineering teams for ambiguous, fast-moving AI projects.

I would love your help on

Connecting with pioneers solving AI safety and alignment infrastructure.

Learning about emerging specialized hardware beyond traditional GPUs.

Finding communities focused on sustainable AI deployment at scale.