Books & Ebooks
Deep practitioner knowledge on LLM systems, AI infrastructure, and production ML — written from real engineering experience, not tutorials.
Ebooks
Focused deep dives on specific topics. Concise, dense, and immediately applicable.
Architecting LLM Inference Systems
From Runtime Engines to Multi-Node Distributed Serving
A comprehensive practitioner's guide to designing, building, and operating large-scale LLM inference infrastructure. Covers vLLM internals, KV cache management, continuous batching, tensor parallelism, speculative decoding, and multi-node serving architectures.
LLM Inference for ML Engineers
Production GenAI Systems
A hands-on guide for ML engineers shipping GenAI products to production. Covers serving architecture decisions, latency optimization, cost modeling, model selection tradeoffs, and operational best practices for teams building on top of LLMs.
AI Evaluation for LLM Applications
Reliability, Groundedness, and Quality Gates
The definitive guide to evaluating LLM-powered applications in production. Covers RAG evaluation, hallucination detection, automated regression testing, human-in-the-loop feedback loops, and building evaluation pipelines that actually catch problems.
Hardcover
Long-form reference books for engineers who want the complete picture.
Architecting Intelligence
A Systems Practitioner's Guide to Production AI
A comprehensive hardcover reference for senior ML engineers and architects. Synthesizes 15+ years of production ML experience into a definitive guide to designing, scaling, and operating AI systems at the highest level.
Get Early Access
All books launch first to Substack subscribers. Subscribe to get early access, pre-launch discounts, and free preview chapters.
Subscribe for Early Access