Architecting Intelligence LogoAI Labs

Speaking

Technical talks on LLM inference, ML infrastructure, production AI systems, and agentic AI — grounded in real engineering experience, not theory.

Send Speaking Inquiry

Why Invite Pawan

  • 15+ years production ML and AI systems experience
  • Technical depth — real production insights, not slides-only talks
  • Clear, practitioner-focused delivery for engineering audiences
  • Flexible formats — keynote, workshop, panel, podcast
  • Custom talks tailored to your audience and conference theme

Talk Topics

Signature talks I can deliver. All can be customized — length, depth, and focus — for your specific audience.

Parallelism for Large-Scale LLM Inference

45–60 min

A deep technical walkthrough of tensor parallelism, pipeline parallelism, and sequence parallelism — what they are, when to use each, and the real tradeoffs at scale. Includes live architecture comparisons.

LLM SystemsInfrastructureGPU Architecture

KV Cache: The Hidden Bottleneck in LLM Serving

30–45 min

Why KV cache management is the most underappreciated problem in LLM serving, how PagedAttention solved it, and what the next generation of solutions look like.

LLM SystemsMemory ManagementvLLM

Production Agentic AI: What Actually Breaks

45–60 min

Moving from demos to production agentic systems is a different engineering problem than people expect. A practitioner's guide to the failure modes, architectural patterns, and operational challenges of real agentic AI.

Agentic AIProduction MLSystems Design

ML Platform Design at Scale

45–60 min

How to design ML platforms that actually serve the organization — from feature stores to model registries to inference infrastructure. What works at scale, what doesn't, and why.

ML PlatformsFeature StoresInfrastructure

From Prototype to Production: GenAI Architecture Decisions

45–60 min

The most important architectural decisions you'll make when taking a GenAI application from prototype to production — and how to make them well. Covers model selection, serving strategy, evaluation, and ops.

System DesignGenAIProduction ML

The Real Economics of LLM Inference

30–45 min

A data-driven talk on the actual cost structure of LLM serving — compute, memory, network, and everything else. How to model costs, where the leverage is, and how to build cost-efficient LLM products.

LLM SystemsCost OptimizationInfrastructure

Speaking Formats

Conference Keynote

30–60 min mainstage or breakout talk

Corporate Workshop

Half-day to full-day technical workshops for engineering teams

Podcast / Interview

Technical discussions on AI systems, ML infrastructure, career

Panel Discussion

Expert panels on AI/ML trends, production systems, industry direction

University / Research Talk

Guest lectures for CS, ML, or engineering programs

Online Summit

Virtual keynotes and live Q&A sessions

Invite Pawan to Speak

Tell me about your event, audience, and topic focus. I'll get back to you within 2–3 business days.