Pawan K Jha
Sr. Principal AI/ML Scientist & Systems Architect
Founder, Architecting Intelligence Labs
I'm a Sr. Principal AI/ML Scientist and Systems Architect with over 15 years building large-scale machine learning platforms, LLM inference systems, search and ranking, forecasting, and production AI architecture.
Over the years I've designed and shipped ML systems that operate at massive scale — from feature platforms serving billions of predictions per day, to LLM serving infrastructure handling complex multi-model inference pipelines. I've worked deeply on the full stack: GPU scheduling, KV cache management, tensor parallelism, model quantization, continuous batching, and the operational layer that makes it all reliable in production.
I started Architecting Intelligence Labs because I believe the most important knowledge in AI/ML — the operational, hard-won, production-grade knowledge — isn't being documented anywhere. Academic papers cover what's possible. Engineering blogs cover what's trendy. Almost nothing covers how to actually architect, scale, and operate these systems in the real world.
That's the gap I'm filling. Through deep dives, architecture blueprints, courses, and tools, I translate years of production experience into content and products that make serious ML engineers better at their craft.
Areas of Expertise
What I Do Here
Write
Deep technical articles on LLM inference, ML infrastructure, and AI systems architecture — published on Substack and this site.
Speak
Conference talks, corporate keynotes, and podcast appearances on AI systems, production ML, and infrastructure.
Build Tools
Independent tools that solve real pain points in AI/ML — starting with inference cost calculators, evaluation harnesses, and workflow optimizers.
Teach
Live cohort courses and mentorship programs for ML engineers wanting to go deeper on LLM inference and production AI systems.
The Philosophy
I believe great ML engineers understand their systems at every layer — from the math in the paper to the CUDA kernel on the GPU to the SLA in the SLA doc. Most content teaches one layer. I try to connect all of them.
Every deep dive I write starts with a real production question: why does this system behave this way under load? why does this latency spike happen at this batch size? what's the actual tradeoff between these two architectures? The answer is always more interesting — and more useful — than the toy version you'll find in a tutorial.
If you're serious about building production AI systems, you're in the right place.