SEQUOIA: Scalable and Robust Speculative Decoding

研究背景

核心思想

20260907181639

SEQUOIA 首先解决 给定节点预算, 如何构建期望产出最高的 token tree. 关键是将树形优劣转化为可计算的节点收益, 再通过动态规划分配子树预算. 在此基础上, 无放回采样进一步减少重复候选, 提高验证的鲁棒性.

为什么需要优化树形?

用节点访问概率定义树的收益

Dynamic Programming: 分配子树预算

20260907181816

辅助改进: 无放回采样

关键洞察

场景 Target / Draft 数据与温度 树配置 $(n,d)$ 每 token 延迟 相对普通解码 相对表中 SpecInfer
A100 on-device Llama2-7B / JF68M C4, $T=0$ $(128,10)$ $6.0$ ms $4.04\times$ $1.17\times$
L40 offloading Llama3-70B-Instruct / Llama3-8B-Instruct MT Bench, $T=0$ $(768,18)$ $0.60$ s $9.5\times$ $1.36\times$

第一行普通解码基线为 Hugging Face 的 $24.2$ ms/token, 第二行为 DeepSpeed-Zero-Inference 的 $5.7$ s/token. SpecInfer 对照树分别为 $5\times8$ 和 $16\times48$ 的独立链. 因此, $4.04\times$ 和 $9.5\times$ 不是相对已有投机解码方法的加速倍数.

继往开来

参考文献

  1. Chen Z., May A., Svirschevski R., Huang Y., Ryabinin M., Jia Z. and Chen B. SEQUOIA: Scalable and Robust Speculative Decoding. Advances in Neural Information Processing Systems, 2024. [Code]