Library · 08-adaptive-computation-inference-cost

Adaptive compute

TitlePeerLink
Confident Adaptive Language Modeling (CALM)◦ preprintarxiv.org/abs/2207.07061
Fast Inference from Transformers via Speculative Decoding◦ preprintarxiv.org/abs/2211.17192
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding◦ preprintarxiv.org/abs/2404.16710
SpecEE: Accelerating LLM Inference with Speculative Early Exiting◦ preprintarxiv.org/abs/2504.08850
Better & Faster Large Language Models via Multi-token Prediction◦ preprintarxiv.org/abs/2404.19737
DeepSeek-V3 Technical Report (MTP module)◦ preprintarxiv.org/abs/2412.19437
Understanding Dynamic Compute Allocation in Recurrent Transformers◦ preprintarxiv.org/abs/2602.08864
Subjective Depth and Timescale Transformers: Learning Where and When to Compute◦ preprintarxiv.org/abs/2511.21408
AdaPonderLM: Gated Pondering LMs with Token-Wise Adaptive Depth◦ preprintarxiv.org/abs/2603.01914
Look Inward to Explore Outward: Temperature Policy from LLM Internal States via Hierarchical RL◦ preprintarxiv.org/abs/2602.13035
Entropy trajectory shape predicts LLM reasoning reliability◦ preprintarxiv.org/abs/2603.18940
When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training◦ preprintarxiv.org/abs/2601.00894
What Layers When: Learning to Skip Compute in LLMs with Residual Gates◦ preprintarxiv.org/abs/2510.13876
Temporally Extended Mixture-of-Experts◦ preprintarxiv.org/abs/2604.20156
Path-Constrained Mixture-of-Experts (PathMoE)◦ preprintarxiv.org/abs/2603.18297
Omni-Router: Sharing Routing Decisions across MoE layers◦ preprintarxiv.org/abs/2507.05724
Mixture-of-Control: State-Aware Fine-Tuning◦ preprintarxiv.org/abs/2606.31397
Learning to Skip the Middle Layers of Transformers◦ preprintarxiv.org/abs/2506.21103
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation◦ preprintarxiv.org/abs/2607.06601
Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents◦ preprintarxiv.org/abs/2606.15017
Token Reduction Is Not Cost Reduction◦ preprintarxiv.org/abs/2607.12161
Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference◦ preprintarxiv.org/abs/2607.25018
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models◦ preprintarxiv.org/abs/2606.30852
When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty◦ preprintarxiv.org/abs/2608.00685
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts◦ preprintarxiv.org/abs/2607.20519
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers◦ preprintarxiv.org/abs/2608.06867
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Modelsarxiv.org/abs/2607.08665
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reusearxiv.org/abs/2608.03893
Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascadesarxiv.org/abs/2606.15308