Architecting Intelligence LogoAI Labs

Publications & Research

Technical reports, preprints, and research papers on LLM inference systems, ML infrastructure, and production AI — grounded in real engineering experience.

Note: Most of my technical writing lives on Substack as Deep Dives. Formal technical reports and preprints are listed here as they are completed and published.
Technical Report · 2024
Draft

Parallelism Strategies for Large-Scale LLM Inference: A Practitioner's Analysis

A systematic comparison of tensor, pipeline, and sequence parallelism strategies for LLM serving at scale. Includes empirical benchmarks across model sizes, hardware configurations, and serving frameworks.

LLM InferenceParallelismvLLMGPU Infrastructure
Technical Report · 2024
In Progress

KV Cache Management in Production LLM Serving Systems

Analysis of KV cache design patterns, eviction policies, and prefix caching strategies in production deployments. Based on observations from real-world serving infrastructure.

KV CacheLLM ServingMemory Management
Technical Report · 2025
Planned

Cost Modeling for LLM Inference at Scale

A framework for accurately modeling and forecasting LLM inference costs across cloud, dedicated GPU, and hybrid deployments. Covers compute, memory bandwidth, network, and operational overhead.

Cost OptimizationLLM InferenceInfrastructure Economics

Get Notified

New technical reports and preprints are announced on Substack first.

Subscribe to Newsletter