If you’ve ever shaved billions of FLOPs off a model and still watched latency barely budge, this resource will resonate. At AI Tech Inspire, we spotted a free, open-source book with a systems-first blueprint for speed that pushes past the usual “just quantize it” playbook. It argues that true performance starts by asking: what is the real bottleneck—compute, bandwidth, memory, or the surrounding system?
Key facts at a glance
- A free, open-source book titled How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents is available on GitHub: make-your-model-fast.
- Core premise: reducing
FLOPsdoesn’t necessarily make a model faster; measure and optimize against the true system bottleneck. - Begins with roofline analysis and hardware fundamentals, then covers kernels, compilers, quantization, pruning, computer vision, on-device LLMs, robotics, profiling, serving, and agent systems.
- Teaches practitioners to reason about performance bounds: compute-, bandwidth-, memory-, or system-bound scenarios.
- Guides when optimizations like quantization, pruning, or kernel tuning are actually worth it for a given workload and device.
- Extends the same systems thinking to serving stacks and agent orchestration.
- Open to feedback and contributions from ML systems, inference, compiler, edge AI, and performance engineering communities; stars are welcome.
Why FLOPs don’t equal speed
It’s easy to assume fewer FLOPs means faster inference, but the book highlights a more grounded truth: models often sit idle waiting on memory or I/O. That’s why a 4-bit quantized network can underperform a higher-precision baseline if the platform is bandwidth-bound and quantization introduces extra dequantize/requantize steps that increase memory traffic.
Consider a CNN running on a mobile SoC versus a high-end GPU. On the GPU with wide memory buses, kernel fusion and CUDA occupancy tuning might unlock speed. On the phone, the same network could be limited by DRAM bandwidth or NPUs with small on-chip SRAM. Same model, different bottlenecks, radically different optimizations.
Key takeaway: Optimize for the bottleneck you actually have, not the one you assume you have.
Roofline thinking, explained
The book starts where many tutorials end: with the roofline model. That mental model draws two ceilings—one for compute throughput and another for memory bandwidth—and asks where your workload lands. If your arithmetic intensity (FLOPs per byte moved) is low, you’re memory-bound; if it’s high, you’re compute-bound. In practice, many ML ops bounce between both depending on tensor shapes, batch sizes, and kernel choices.
For developers used to API-level tweaks in PyTorch or TensorFlow, the roofline view can be a lightbulb moment. It reframes performance trade-offs and helps you predict whether techniques like operator fusion, tensor reordering, or activation recomputation will move the needle on your device.
From silicon to kernels: where time really goes
After grounding in hardware, the material steps through kernels and compilers—the engines that turn graphs into fast code. Expect discussion around:
- Kernels and fusion: Cutting memory round-trips by fusing operations can beat headline-grabbing pruning ratios for certain workloads.
- Compilers: Graph- and kernel-level compilation (think TVM, XLA, and
torch.compile) can restructure compute to better fit caches and vector units. - Runtimes and backends: Choices like ONNX runtimes or TensorRT matter. The same model can diverge by multiples in latency depending on kernel libraries and available epilog fusions.
The message is practical: if you’re bound by memory, focus on data layout, tiling, and fusion; if compute-bound, pack your tensor cores, increase occupancy, and lean on mixed precision where accuracy allows.
Quantization and pruning: not always the hero
Quantization and pruning get the spotlight in many performance stories. This book takes a measured stance: yes, they’re powerful, but only if they attack the limiting factor. For example, int8 quantization can double throughput on GPUs with specialized dot-product units, but on some CPUs without robust vectorization support, the gains may be modest or even negative once conversion overhead is counted.
For edge deployments, quantization-aware training plus a backend that natively supports low-bit ops can be a game changer. But if your service is request-fanout- or network-bound, time spent on low-bit gymnastics won’t fix tail latency. The suggested approach: profile first, set a performance target, and apply the technique that changes the roof, not just the marketing slide.
Beyond the model: serving and agent systems
A standout angle is how the same systems lens applies to serving and agents. Load balancers, batching strategies, and tokenizer throughput can dominate latency as much as matmul speed. For large models—say a GPT-style decoder—KV cache layout and streaming policies can shift you from memory-bound to compute-bound or vice versa. And in multi-tool agents, orchestration delays and API round-trips can dwarf the cost of a single forward pass.
If you’re serving diffusion models like Stable Diffusion, the optimal batch size can differ for UNet steps versus VAE encode/decode; one-size batching wastes GPUs. If you’re fine-tuning or deploying via Hugging Face stacks, pre- and post-processing overhead (tokenization, image transforms) may be the silent killer.
A practical playbook to try today
- Identify the bound: Measure arithmetic intensity and memory throughput. Quick sanity checks: nvidia-smi utilization, profiler traces, kernel-level roofline plots.
- Eliminate obvious stalls: Pin memory, prefetch batches, use
channels_lastwhere supported, and coalesce small ops via fusion. - Choose the right backend: Test an optimized runtime (e.g.,
torch.compile, XLA, TensorRT, ONNX Runtime) with the shapes you actually serve. - Quantize selectively: Target layers that dominate time and have good calibration statistics; verify that the backend has true low-bit kernels.
- Right-size batches: Balance throughput and latency. For interactive LLMs, tune tokens-per-second and KV cache footprint; for vision, watch input resolution and augment pipelines.
- Audit the system: Measure queueing delays, network hops, serialization costs, tokenization, and I/O. If agents invoke tools, instrument each API call.
None of this requires a full rewrite. Many teams see wins by sequencing these steps and validating after each change. The book’s structure—hardware to kernels to serving to agents—mirrors how a production system actually behaves under load.
Where this matters most
Three scenarios stand out for practitioners:
- On-device LLMs and vision: When deploying assistants or perception models on phones, AR headsets, or robots, memory hierarchies and bandwidth become first-class citizens. Expect techniques like operator fusion, activation checkpointing, and low-bit inference to matter—if the device supports them efficiently.
- Cost-sensitive backends: For startups optimizing unit economics, switching runtimes, improving batching, and reducing CPU preprocessing can save more than aggressive model surgery.
- Latency-critical agents: In multi-step workflows, cutting orchestration overhead (parallel tool calls, local caching, compact prompts) can be more impactful than another 5% kernel speedup.
Questions worth asking your stack
- Is the current model compute-, bandwidth-, memory-, or system-bound on the target hardware?
- Which single change would move the roofline the most: fusion, layout changes, quantization, or a new runtime?
- Are preprocessing and tokenization stealing more time than the model itself?
- For agents, what fraction of time is spent in network I/O versus actual inference?
- Can shapes be standardized or padded to enable better kernel selection without hurting tail latency?
Why this resource is worth your time
Most performance guides focus on isolated tricks. This one ties the whole stack together and, importantly, shows when not to apply a trick. That’s valuable whether you’re tuning a PyTorch model for a weekend project or operating a fleet of GPUs. The practical framing—”How fast can this possibly run? Which optimization moves that limit?”—helps teams prioritize engineering time and avoid chasing small wins that don’t dent user-facing latency or cost.
The project is free and open source on GitHub. The author invites feedback and contributions from practitioners working on ML systems, inference, compilers, edge AI, and performance engineering—and notes that a star is appreciated. If your roadmap includes squeezing more out of your hardware, this seems like a solid addition to the reading queue.
Explore the book: How to Make Your Model Fast: A Systems View of Efficient ML, from Silicon to Agents. If you try the ideas, consider instrumenting your stack end to end and sharing a trace: seeing where the time actually goes can be the most persuasive performance story inside any engineering org.
Recommended Resources
As an Amazon Associate, I earn from qualifying purchases.