Efficiently Distilling LLMs for Edge Applications

Achintya Kundu; Fabian Lim; Aaron Chew; Laura Wynter; Penny Chong; Rhui Dih Lee

Efficiently Distilling LLMs for Edge Applications

Achintya Kundu, Fabian Lim, Aaron Chew, Laura Wynter, Penny Chong, Rhui Dih Lee

Published: 01 Jan 2024, Last Modified: 04 Oct 2024NAACL (Industry Track) 2024EveryoneRevisionsBibTeXCC BY-SA 4.0

Abstract: Supernet training of LLMs is of great interest in industrial applications as it confers the ability to produce a palette of smaller models at constant cost, regardless of the number of models (of different size / latency) produced. We propose a new method called Multistage Low-rank Fine-tuning of Super-transformers (MLFS) for parameter-efficient supernet training. We show that it is possible to obtain high-quality encoder models that are suitable for commercial edge applications, and that while decoder-only models are resistant to a comparable degree of compression, decoders can be effectively sliced for a significant reduction in training time.

Loading