3 posts tagged with "Inference"

vLLM Optimization Techniques: 5 Practical Methods to Improve Performance

February 6, 2026 · 26 min read

Jaydev Tonde

Data Scientist

vLLM optimization techniques cover artwork with five performance methods highlighted

Running large language models efficiently can be challenging. You want good performance without overloading your servers or exceeding your budget. That's where vLLM comes in - but even this powerful inference engine can be made faster and smarter.

In this post, we'll explore five cutting-edge optimization techniques that can dramatically improve your vLLM performance:

Prefix Caching - Stop recomputing what you've already computed
FP8 KV-Cache - Pack more memory efficiency into your cache
CPU Offloading - Make your CPU and GPU work together
Disaggregated P/D - Split processing and serving for better scaling
Zero Reload Sleep Mode - Keep your models warm without wasting resources

Each technique addresses a different bottleneck, and together they can significantly improve your inference pipeline performance. Let's explore how these optimizations work.

Disaggregated Prefill-Decode: The Architecture Behind Meta's LLM Serving

January 29, 2026 · 11 min read

Vishnu Subramanian

Founder @JarvisLabs.ai

Disaggregated Prefill-Decode Architecture

Why I'm Writing This Series

I've been deep in research mode lately, studying how to optimize LLM inference. The goal is to eventually integrate these techniques into JarvisLabs - making it easier for our users to serve models efficiently without having to become infrastructure experts themselves.

As I learn, I want to share what I find. This series is part research notes, part explainer. If you're trying to understand LLM serving optimization, hopefully my journey saves you some time.

This first post covers disaggregated prefill-decode - a pattern I discovered while reading through the vLLM router repository. Meta's team has been working closely with vLLM on this, and it solves a fundamental problem that's been on my mind.

Speculative Decoding in vLLM: Complete Guide to Faster LLM Inference

December 18, 2025 · 34 min read

Jaydev Tonde

Data Scientist

Speculative Decoding vLLM Cover

Introduction

Ever waited for an AI chatbot to finish its answer, watching the text appear word by word slow? It can feel painfully slow, especially when you need a fast response from a powerful Large Language Model (LLM).

The big problem is in how LLMs generate text. They don't just write a paragraph all at once; they follow a strict, word-by-word approach.

The model looks at the prompt and the words it has generated so far.
It calculates the best next word (token).
It adds that word to the text.
It repeats the whole process for the next word.

Each step involves complex calculations, meaning the more text you ask for, the longer the wait. For developers building real-time applications (like chatbots, code assistants, or RAG systems), this slowness (high latency) is a major problem.

Why I'm Writing This Series​

Introduction​

Why I'm Writing This Series

Introduction