All field notes

Glossary · 1 minute read

What Is Model Distillation?

Model distillation is a technique where a smaller "student" model is trained to mimic a larger, more capable "teacher" model, capturing much of its performance in a far more efficient package. The result runs faster and cheaper and can be deployed where the large model is impractical—like on-device or at high volume—while retaining most of the quality for the target tasks. Distillation is one of the main ways teams get the benefits of powerful models at lower inference cost and latency.

By FISTA Solutions· AI-Native Engineering Team·
What Is Model Distillation? article cover

Distillation shrinks a big model into a small, fast one that keeps most of the smarts. Here's what it is, why it cuts cost, and where it fits.

What model distillation is

Model distillation trains a smaller "student" model to mimic a larger, more capable "teacher" model—capturing much of its performance in a far more efficient package.

Why it's useful

Large modelDistilled student
High capabilityMost of the capability
Slow, expensiveFaster, cheaper
Hard to deploy everywhereRuns on-device, at scale

Distillation is a key way to get powerful-model quality at lower inference cost and latency—related to small language models.

Where it fits

Use a distilled model where you need speed, low cost, or on-device/edge deployment at high volume—and the task is well-defined enough that the smaller model performs well.

Confirm the trade-off

Distillation loses some quality—usually little for target tasks. Evaluate the student on your actual use case to confirm the trade-off is acceptable, the measurement discipline.

Related efficiency techniques

Distillation pairs with quantization and model selection as ways to control cost—see generative AI cost.

Why FISTA

FISTA Solutions builds cost-efficient AI—right-sized and distilled models where they fit—so you get quality without overpaying for inference, through AI enablement, backed by 150+ projects across 12+ countries.

Optimizing AI cost and speed? Talk to FISTA.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What is model distillation?

A technique where a smaller student model is trained to mimic a larger teacher model, capturing much of its capability in a more efficient form that runs faster and cheaper while retaining most of the quality on target tasks.

02Why use model distillation?

To get much of a large model's performance at far lower inference cost and latency, enabling deployment on-device or at high volume where the large model would be too slow or expensive.

03Does distillation lose quality?

Some, but often little for the target tasks—a well-distilled model keeps most of the teacher's performance where it matters. Evaluate the student on your actual use case to confirm the trade-off is acceptable.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project