4 min read

My New Book Is Live: Building LLM Systems for Platform Troubleshooting

My New Book Is Live: Building LLM Systems for Platform Troubleshooting
Photo by Ian Taylor / Unsplash

I’m excited to announce the launch of my new book, Building LLM Systems for Platform Troubleshooting, now available on Amazon.

Get the book on Amazon

This book is a practical field guide for platform engineers, SREs, DevOps teams, AI platform builders, and technical leaders who want to build LLM-powered assistants that are genuinely useful in production operations, where vague answers, fragile demos, and confident guesses are not acceptable substitutes for evidence, safety, and operational discipline.

Most AI demos aimed at engineering teams still stop at a chat interface over documentation, which can look impressive in a controlled environment but often falls apart when teams try to use it during real incidents involving noisy alerts, incomplete telemetry, recent deployments, unknown dependencies, stale runbooks, and humans working under pressure. This book takes a more serious approach by showing how to design troubleshooting assistants as evidence workflows that gather context, retrieve operational knowledge, call diagnostic tools, explain uncertainty, respect safety gates, and support engineers without removing human judgment from the incident response process.

Why I Wrote This Book

I wrote this book because the industry is moving quickly from “can we use LLMs for engineering operations?” to “how do we build these systems responsibly enough to trust them when production is degraded?” That second question is the one that matters, because an assistant used during an incident has to operate inside a very different reality from a polished demo, and it has to earn trust by being grounded, observable, measurable, and constrained by clear safety boundaries.

LLMs can be extremely useful for troubleshooting when they are surrounded by the right system design, because they can synthesize complex context, connect operational signals, generate hypotheses, summarize messy evidence, and help engineers navigate unfamiliar systems more quickly. However, those strengths only become reliable when the assistant has access to the right context, understands what evidence it is using, knows when information is missing, and is prevented from taking unsafe actions without explicit policy checks and human approvals.

From Chatbot to Production Assistant

The core idea of Building LLM Systems for Platform Troubleshooting is that teams should stop treating operational assistants as generic chatbots and start designing them as structured diagnostic systems. A dependable troubleshooting assistant should not simply answer a question from memory or produce a fluent explanation from a documentation index; it should assemble context packs from logs, metrics, traces, deployments, tickets, runbooks, service metadata, and historical incidents, then use that evidence to help humans reason through what is happening.

That design shift matters because platform troubleshooting is rarely a single-question, single-answer problem. Real incidents involve evolving hypotheses, partial data, conflicting signals, uncertain ownership, recent changes, unclear blast radius, and a constant need to distinguish what is known from what is merely suspected. A useful LLM assistant has to work within that operational loop, helping engineers ask better questions, inspect the right signals, narrow the search space, and communicate findings more clearly.

What You’ll Learn

Inside the book, I walk through how to design a reference architecture for LLM-assisted troubleshooting, including how to build context packs from operational data sources such as logs, metrics, traces, deployments, tickets, and runbooks. The book also shows how to connect retrieval to operational knowledge without creating brittle answer machines that collapse when documentation is incomplete, outdated, or written for a different failure mode.

You will also learn how to define safe tool contracts, approvals, and policy gates so that assistants can call diagnostic tools, inspect production state, and support remediation workflows without creating unacceptable operational risk. The book covers evaluation with realistic incident cases, observability for assistant behavior, memory and feedback loops, rollout strategy from prototype to production, and common failure modes such as hallucinated diagnoses, unsafe remediation, stale context, and overconfident explanations.

Why This Matters Now

Every platform organization is under pressure to make operations faster, safer, and more scalable, while also dealing with increasingly complex systems that span cloud infrastructure, Kubernetes, service meshes, managed databases, queues, CI/CD pipelines, observability platforms, security controls, and organizational boundaries. At the same time, teams are being told that AI can transform engineering work, even though many of the proposed solutions still lack the operational rigor required for production use.

This book is written for teams that want the benefits of AI assistance without pretending that reliability, safety, and accountability are optional. It gives engineers a disciplined path from prototype to production, with patterns that help them build assistants capable of explaining what they checked, why they reached a conclusion, where uncertainty remains, and what should happen before any risky action is taken.

Who This Book Is For

Building LLM Systems for Platform Troubleshooting is for platform engineers who are building internal tools, SREs who want better support during incidents, DevOps teams trying to reduce operational toil, AI platform builders looking for a serious production use case, and technical leaders who need to evaluate whether LLM-assisted operations can be adopted responsibly inside their organization.

If your team is experimenting with AI for incident response, production support, platform operations, or internal engineering enablement, this book gives you a practical framework for moving beyond demos and toward systems that can actually help under real operational pressure. It is especially relevant for teams that want AI assistance during incidents while keeping humans in control, preserving safety boundaries, and making diagnostic quality measurable rather than subjective.

Get the Book

Building LLM Systems for Platform Troubleshooting is available now on Amazon.

Get the book on Amazon

If your team wants to use LLMs during incidents without handing production to a black box, this book gives you the patterns, templates, and implementation guidance to build that capability responsibly, measure it honestly, and roll it out with the level of discipline that production operations deserve.