keyLogo

Mani Pal

Hi, I'm Mani Pal - Machine Learning Engineer at Trellions

I am driven by the pursuit of pure performance in AI systems. I thrive when rewriting the fundamental code that powers intelligence, solving the most gnarly challenges in LLM inference and hardware utilization.

Trellions
New Delhi, Delhi, India
AI insights

At a glance

Curated signals on strengths, focus areas, and how they can help.

Successfully founded and scaled infrastructure within the high-stakes Web3 payments domain.

Currently architects high-performance CUDA kernels and speculative decoding runtimes for LLM efficiency.

Offers deep technical expertise in optimizing hardware-resident AI systems and model throughput scaling.

🚀 Career trajectory

Founding Phase

Established technical foundations in the Web3 space, building robust payment and infrastructure systems at RecurX.

Founder

Full-Stack

Blockchain

Systems Engineering Pivot

Shifted focus to deep AI infrastructure, specializing in CUDA and high-performance machine learning systems.

ML Systems

Research

Optimization

💪🏻 Superpowers

Kernel-Level Optimization

Achieving maximum hardware utilization through low-level systems engineering.

Rebuilt FlashAttention-2 achieving 2.1x throughput improvement over standard baselines.

Engineered custom CUDA kernels to extract maximum performance from NVIDIA A100 hardware.

Refined KV-cache and PagedAttention logic to reduce memory overhead during inference.

Algorithm Research

Pushing the frontiers of LLM efficiency and reasoning capabilities.

Trained a 700M hybrid Mamba-2/Transformer model from scratch using custom workflows.

Integrated GRPO reasoning and DPO alignment to enhance model output reliability.

Developed high-throughput speculative decoding runtimes to optimize real-time token generation.

Entrepreneurial Engineering

Applying product-minded agility to deep-tech problem solving.

Scaled decentralized payment infrastructure as a Founding Engineer at RecurX.

Navigated the full lifecycle from system architecture to production deployment.

Balance technical depth with a clear focus on end-user performance requirements.

I'm excited about

Exploring advanced AI research communities focused on high-performance model deployment.

Collaborating with distributed teams building the next generation of scalable inference runtimes.

Finding opportunities to apply my CUDA expertise to mission-critical infrastructure projects.

I can help with

Providing technical audits for GPU kernel performance and inference optimization workflows.

Sharing insights on training hybrid architectures for specific compute-constrained environments.

Mentoring junior engineers on transitioning into low-level systems and LLM research.

I would love your help on

Identifying high-impact research organizations pushing the boundaries of efficient LLM inference.

Connecting with systems leads who value mathematical rigor in model throughput improvement.

Gaining perspective on evolving infrastructure requirements for large-scale distributed training.