Speaking
Technical talks on LLM inference, ML infrastructure, production AI systems, and agentic AI — grounded in real engineering experience, not theory.
Send Speaking InquiryWhy Invite Pawan
- 15+ years production ML and AI systems experience
- Technical depth — real production insights, not slides-only talks
- Clear, practitioner-focused delivery for engineering audiences
- Flexible formats — keynote, workshop, panel, podcast
- Custom talks tailored to your audience and conference theme
Talk Topics
Signature talks I can deliver. All can be customized — length, depth, and focus — for your specific audience.
Parallelism for Large-Scale LLM Inference
45–60 min
A deep technical walkthrough of tensor parallelism, pipeline parallelism, and sequence parallelism — what they are, when to use each, and the real tradeoffs at scale. Includes live architecture comparisons.
KV Cache: The Hidden Bottleneck in LLM Serving
30–45 min
Why KV cache management is the most underappreciated problem in LLM serving, how PagedAttention solved it, and what the next generation of solutions look like.
Production Agentic AI: What Actually Breaks
45–60 min
Moving from demos to production agentic systems is a different engineering problem than people expect. A practitioner's guide to the failure modes, architectural patterns, and operational challenges of real agentic AI.
ML Platform Design at Scale
45–60 min
How to design ML platforms that actually serve the organization — from feature stores to model registries to inference infrastructure. What works at scale, what doesn't, and why.
From Prototype to Production: GenAI Architecture Decisions
45–60 min
The most important architectural decisions you'll make when taking a GenAI application from prototype to production — and how to make them well. Covers model selection, serving strategy, evaluation, and ops.
The Real Economics of LLM Inference
30–45 min
A data-driven talk on the actual cost structure of LLM serving — compute, memory, network, and everything else. How to model costs, where the leverage is, and how to build cost-efficient LLM products.
Speaking Formats
Conference Keynote
30–60 min mainstage or breakout talk
Corporate Workshop
Half-day to full-day technical workshops for engineering teams
Podcast / Interview
Technical discussions on AI systems, ML infrastructure, career
Panel Discussion
Expert panels on AI/ML trends, production systems, industry direction
University / Research Talk
Guest lectures for CS, ML, or engineering programs
Online Summit
Virtual keynotes and live Q&A sessions
Invite Pawan to Speak
Tell me about your event, audience, and topic focus. I'll get back to you within 2–3 business days.