LLM Customization and Fine-Tuning
Adaptation, distillation, and alignment
Practical techniques for customizing and fine-tuning LLMs—LoRA, QLoRA, SFT, distillation, and DPO/RLHF alignment.
Most teams that outgrow a general-purpose API hit the same question: we have a use case and a GPU budget, so which adaptation technique do we use, and how do we keep it working after launch? Bahree and Tok answer it by building one assistant, an IT help desk for a fictional company called Contoso, and pushing it up the whole ladder: prompting and retrieval, LoRA and QLoRA, full supervised fine-tuning, distillation, and DPO alignment.
The structure is the best idea here. Plenty of fine-tuning material shows you one technique on a toy dataset and lets you assume it generalizes. This one holds the base model (Qwen3-4B-Instruct), the data (real Stack Exchange IT questions plus a small general-knowledge slice) and the task fixed. Each rung is measured against the same baseline, so the comparison means something. The framework for choosing between techniques, weighing cost, latency, privacy and return, is the part we'd photocopy. It starts from the right instinct: escalate only when a measured baseline says you must, and put the effort into a few hundred clean examples and an evaluation set you trust rather than into volume. The final chapter covers what happens after training: a model registry, drift detection with canary prompts, rollback, and a safety monitor. Most books stop at the training loop, and that's the half that decides whether a model survives its first quarter in production.
Fine-tuning on generic instruction data can make a model’s wrong answer more confident, not more correct.
— Bahree and Tok, LLM Customization and Fine-Tuning, ch. 1 summary
The honesty is real, and it's checkable. The companion repository publishes the code, the trained models and the logs, failed runs included. It admits that DPO and SFT come out roughly even on token-level F1 once you score on a properly held-out set. That's the sort of result a vendor would bury.
the question should not be “Should we fine-tune?” but “Which rung of the adaptation continuum fits this specific problem?”
— Bahree and Tok, LLM Customization and Fine-Tuning, ch. 1 overview
We have two complaints. First, the hardware claims in the marketing are looser than the repository's own measurements. The pitch talks about full fine-tuning on a single 24 GB card. The authors' own numbers say full SFT of the 4B model needs about 32 GB, which means two of those cards, and full-parameter DPO needs around 54 GB. LoRA does fit one consumer card, and the book is fine on that point. A reader who skims the blurb and buys a single 4090 will be annoyed. Second, the claim that the methods scale unchanged to frontier models is asserted, not demonstrated. One 4B model on one narrow domain is a clean teaching rig, but it can't show you what breaks at 70B or on messier data. Token-F1 is also a thin yardstick for answer quality, and a help desk assistant deserves better. Finally, the book was still a work in progress when we looked, with only part of it released, so judge the late chapters by their promise.
most fine-tuned models fail not at launch but months later
— Bahree and Tok, LLM Customization and Fine-Tuning, book description (operations chapter)
Still, this is the practical fine-tuning book we'd point an ML or platform engineer to this year. It treats adaptation as an engineering and economics decision, not a magic trick, and it gives you code to check every number. Read it for the decision framework and the operations chapter, and budget for more GPU memory than the cover implies.
Read the longer summary
The argument: escalate only when the baseline says you must
Most fine-tuning books start with the training loop. Bahree and Tok start with a question most teams skip: should you train at all? Their core argument is that “customize the model” covers a graded scale of options. At one end, you change only what goes into the context window. At the other, you change the weights and keep watching the model for months afterward. Each step up costs more engineering effort and gives you more control over behavior, and the steps aren’t evenly spaced. Going from a good prompt to retrieval is a small move. Going from retrieval to a LoRA adapter is a big one, because now you own a training pipeline, an evaluation set and a model artifact that can regress.
So the book’s main tool is a decision procedure, and the techniques come second. The authors want you to ask four things before choosing a method. Is the task commodity work a frontier API already does well? Have you measured a prompting baseline? Do you have enough clean, representative examples? Which hard constraint is actually forcing the decision: cost per request, data residency, latency, or behavior a prompt can’t produce? Our reading: if you can’t answer the second one, you aren’t ready to fine-tune. We agree, and we think this is the most useful habit in the book. Most failed fine-tuning projects we’ve heard about failed because nobody wrote down a baseline. When the adapted model shipped, nobody could say whether it beat a well-written system prompt.
The business case is concrete. The authors put the per-request saving of a self-hosted adapted model over a frontier API at roughly five to ten times, depending on which frontier tier you were paying for. They add privacy (HIPAA, FINRA, FedRAMP, GDPR residency), latency budgets, and the strategic risk of building on a single API vendor. They’re just as specific about when to say no: commodity tasks, early projects with no baseline, thin or noisy data, and cases where the frontier model is already good enough. A book that spends part of its first chapter telling you not to buy what it teaches has earned some trust.
One claim in that opening deserves more attention than it will probably get. Fine-tuning doesn’t cure confabulation. Train a model on generic instruction data and it can become more confident while staying just as wrong. Bahree’s fix is training data that explicitly shows the behavior you want, such as declining to answer or naming the uncertainty, often followed by preference optimization. Many teams believe that fine-tuning on their docs will make the model stop inventing things. This corrects that belief early.
One model, one dataset, nine chapters
The book’s structural choice is also its main strength. Every hands-on chapter uses the same base model, Qwen3-4B-Instruct-2507. Every chapter uses the same task: a help-desk assistant for a fictional company called Contoso. Every chapter uses the same data, built from real Stack Exchange IT questions (Super User, Ask Ubuntu, Server Fault) plus a small slice of Databricks Dolly to protect general capability. Because model, data and task stay fixed, comparing techniques across chapters is fair. You aren’t comparing a LoRA demo on one dataset with a DPO demo on another.
The order matches how a team would actually work:
- Why adapt at all.
- The decision framework, with a LoRA quickstart that trains in about ten minutes on a 12 GB card.
- The data: sourcing, cleaning, splitting, and generating synthetic examples from a teacher model.
- Few-shot prompting and a minimal RAG pipeline, with retrieval metrics such as Precision@k and Hit@1.
- LoRA and QLoRA.
- Full-parameter supervised fine-tuning.
- Distillation.
- Preference optimization with DPO.
- The production layer: model registry, drift detection, rollback, and safety monitoring.
The model choice holds up. A 4B instruction-tuned open-weights model is small enough to train on hardware a reader can rent cheaply, and capable enough that each technique’s effect is visible. The authors say plainly that a 4B model isn’t what most enterprises will deploy. Their argument is that LoRA, QLoRA, SFT, distillation and DPO don’t change at larger sizes; you just need more memory and compute. That’s broadly true of the mechanics. It’s less obviously true of the results, which we come back to below.
The data choice is a quiet strength too. Stack Exchange IT content is messy and specific, unlike the toy datasets that make every technique look good. Every example is attributed back to its source URL, and the dataset is rebuilt from scripts instead of shipped as an opaque file. The authors also include a small Contoso demo set with made-up internal tool names, a deliberate test a prompt can’t pass by guessing. A base model can’t know your internal tool is called something it has never seen. An adapter can learn it. That’s the clearest single example we know of where weight changes beat context engineering.
Data is the real bottleneck, and the book shows it
Chapter 3 is where Bahree and Tok go furthest past the usual tutorial. Their position is that data quality and a trusted held-out evaluation set matter more than which technique you pick. Teams routinely overestimate how much data they need. In the authors’ experience, somewhere between 60 and 500 carefully chosen examples often beats tens of thousands of noisy ones. The hard part is curation and domain expertise, not volume.
They test this. The chapter trains the same model on four versions of Financial PhraseBank and scores each on the same held-out test set. Across NVIDIA A30, AMD MI300X and NVIDIA H200 runs, the ordering stays the same: the deliberately corrupted version scores worst, in the low-to-high 80s, while the clean variants sit in the high 90s. It’s a small experiment, and that’s why it works. A reader can rerun it in an afternoon and see the effect themselves.
The chapter also has a six-step synthetic-data pipeline: load seeds, build prompts, generate with a teacher, apply quality gates, check the distribution, and mix and save. There’s also a standalone manifest module that hashes content, tracks lineage and schedules retention. That last piece is easy to dismiss as bureaucracy, and most teams skip it. But months later, when someone asks which data trained the model in production and whether a deleted customer record is still in it, the manifest is the only answer. Putting lineage in chapter 3 instead of an appendix is a good editorial call.
The training chapters: what the measured numbers say
The publisher’s pitch leans on reproducibility: the authors publish the code, the trained models on Hugging Face, and the training and evaluation logs, failed runs included. The companion repository delivers on that, and its hardware documentation is the most useful thing to read before buying a GPU.
Some figures from the authors’ own A30 runs:
| Technique | Peak memory | Wall time |
|---|---|---|
| LoRA, rank 16, 450 examples, 3 epochs | 9 GB | about 11 minutes |
| QLoRA, rank 8 | 5.1 GB | about 15 minutes |
| Full SFT | 32.5 GB across two cards | about 10 minutes |
| Full-parameter DPO | about 54 GB across three cards | a few minutes |
| DPO with a LoRA adapter | under 11 GB on one card | a few minutes |
Two practical lessons follow. First, the quantised variant buys memory headroom and pays for it in wall-clock time. On every card the authors tested, it trained slower than plain LoRA, because 4-bit quantization and dequantization add overhead you only want when memory is the constraint. Second, most of the runtime goes to generation, not training. Evaluation passes and teacher-data generation dominate the wall clock, so the way to speed up iteration is faster decoding, not a bigger training GPU. Few fine-tuning tutorials say this, and it changes how you budget a project.
The cross-hardware notes are useful field notes. Pin the model to one device instead of trusting automatic device mapping on a memory-constrained machine. The authors traced NaN gradients on an M2 Pro and a device-mismatch crash on an M4 to silent layer offloading. Hugging Face rate-limits some datacenter IP ranges no matter what token you send. QLoRA on AMD needs ROCm 6.2 or later. The most counterintuitive result: on this workload, NVIDIA’s B200 ran slowest of the three NVIDIA cards they tested, because PyTorch and bitsandbytes kernels for Blackwell weren’t mature yet. That’s the kind of detail you only get from people who actually ran the code.
Now the numbers that should temper expectations. Scored on a held-out test set of 50 questions that no training or model-selection step touched, the gains are small. On token-F1, the base model scored 0.153, the SFT model 0.168, and the DPO model 0.166. The distilled student landed at 0.165, just under its 0.168 teacher. An earlier cross-hardware pass, scored on each chapter’s own validation split, showed a much bigger jump, with the teacher around 0.56 against a base around 0.26. The authors document both and explain the change, which is honest. But a reader skimming the chapters should know that on clean held-out data, the measured lift from full fine-tuning on this task is about a point and a half of token-F1.
How you read that depends on how much you trust token-F1, and we don’t trust it much for free-form support answers. It rewards word overlap and barely notices whether an answer is correct, follows the house format, or names the right internal tool. The authors seem to know this. Chapter 7 includes a separate house-format check, and the Contoso demo exists to show behavior the metric misses. We’d still have liked the headline comparisons to use a metric that tracks what the help-desk assistant is actually for.
Distillation, DPO, and the honesty about non-wins
Chapter 8 is where the book’s commitment to reporting failures matters most. The authors align the SFT model with DPO using real, human-graded preference pairs, then report that DPO roughly ties SFT on objective accuracy. Many alignment tutorials imply that preference optimization is always the final upgrade. Bahree and Tok show a case where, on the metric they measured, it isn’t. Across all their hardware runs, DPO never ranked below SFT, and it never pulled clearly ahead on token-F1 on held-out data. That result is useful to a reader deciding whether to spend a sprint on preference data.
Safety gets real attention, not a closing paragraph. Every training chapter re-runs a safety regression suite, because adaptation can erode the base model’s guardrails, let jailbreaks transfer, and, in RAG systems, let instructions injected into retrieved documents steer the model. Treating safety as a regression test you re-run each time, not a one-time audit, is the right engineering habit. Running it across red-team categories with per-category alerts makes it something a team can actually operate.
Chapter 7 needs one caveat. The chapter is framed around moving knowledge from a strong teacher into a cheaper student, and the publisher copy talks about producing a smaller model. In the repository’s default reproducible path, though, the teacher is the chapter 6 SFT model and the student is a LoRA adapter on the same 4B base. A frontier-API teacher is optional, through OpenRouter. That’s a sensible choice for reproducibility, since nobody wants a book whose results depend on an API that changes monthly. But the student isn’t smaller to serve; it’s cheaper to train. Readers who want the real compression case, a big teacher feeding a much smaller student, will need to set it up themselves.
Where the book oversells and what it leaves out
There are a few gaps between the marketing and the measurements, and readers should know them before buying.
The publisher copy says full fine-tuning runs on one 24 GB card such as an A30 or RTX 4090. The authors’ own hardware notes contradict this. Full SFT ran out of memory on one A30 at about 23 GB during optimizer initialization and needed two 24 GB cards or one 40 GB card. Full-parameter DPO needed three. The repository explains why: bf16 weights, gradients and two Adam moments come to about 8 bytes per parameter, roughly 32 GB for a 4B model before activations. The book’s own documentation is correct, and we credit the authors for publishing the measurement that undercuts their blurb. The blurb is still wrong, and someone buying a 4090 because of it will hit that wall.
The claim that the methods carry over unchanged to frontier-scale models is true of the code and less certain for the conclusions. Whether DPO ties SFT, how much data is enough, and where LoRA stops matching full fine-tuning can all shift with model size and task. The book shows the 4B case well. It asserts the larger case.
The drift detector in chapter 9 is TF-IDF, which is cheap, transparent and easy to calibrate. That’s a reasonable starting point. In the authors’ demo, it flags a topic shift toward Kubernetes clearly. But TF-IDF catches changes in what users ask about, not changes in how well the model answers. The canary prompts and safety monitor cover some of that. A team running this in production would probably add an embedding-based or judge-based quality signal fairly quickly.
Some things are missing from the nine-chapter outline entirely. Nothing covers serving: inference engines, quantizing for deployment, batching, or how to host a LoRA adapter cheaply alongside others. Since the ROI argument depends on per-request cost, that’s a real gap. Evaluation beyond token-F1, such as model-graded rubrics or task success rates, isn’t a chapter of its own. On-policy reinforcement methods beyond DPO aren’t covered either, and neither are multi-turn or tool-calling assistants, even though a help desk is exactly where tool use matters. The book is still in early access, with five of nine chapters released at the time of writing and print publication set for late December 2026, so some of this may change.
Finally, the authorship. Bahree was until recently a senior engineering leader on Azure’s AI platform and is now CTO in G42 Americas’ office of the CEO. Tok is a Microsoft product director. It’s to their credit that the book is vendor-neutral and centered on open weights, with nothing pushing you toward Azure. It’s also worth noticing that the authors frame it as the do-it-yourself counterpart to the managed fine-tuning services their employers sell. That’s a reasonable framing. It also means the build-versus-buy chapter is written by people who have sold both.
Who should read it
This book is for an ML or platform engineer who has to make an open-weights model behave like a specialist and keep it that way, and who has, or can rent, a 24 GB GPU or better. It’s especially useful for the tech lead who has to justify the project. The decision framework, the cost argument, the data-quality experiment and the honest DPO result are what you need for that conversation. If you want to understand why LoRA works or how preference optimization relates to reward modeling, it’s thin on theory. Pair it with Nathan Lambert’s Reinforcement Learning from Human Feedback or Sebastian Raschka’s Build a Reasoning Model (From Scratch). If your team is still at “should we fine-tune?”, read chapters 1 through 4 and stop there; for many teams, the honest answer the book gives is “not yet.” If you’re past that point, the reason to buy it is the companion repository and the measured hardware notes, and the prose explains why those numbers came out the way they did.