Mentored Decoding: Faster Inference meets Boosting
En palabras de los autores
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can beat the target . Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called . To get there, we connect inference to a celebrated ML training theory, , and proceed via the generalization of mentored decoding to the whole set of -divergences. We uncover key properties of mentored decoding, among which (i) the particularly appealing geometric nature of the total variation case, (ii) simple approximations for any -divergence in direct relation with boosting compliance, and (iii) a space and time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem in time and constructing optimal mentored distributions in time for any -divergence.
Apareció: lunes, 28 de septiembre. arXiv. Preprint, todavía sin revisión por pares.