Library · 08-adaptive-computation-inference-cost
Adaptive compute
| Title | Peer | Link |
|---|---|---|
| Confident Adaptive Language Modeling (CALM) | ◦ preprint | arxiv.org/abs/2207.07061 |
| Fast Inference from Transformers via Speculative Decoding | ◦ preprint | arxiv.org/abs/2211.17192 |
| LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding | ◦ preprint | arxiv.org/abs/2404.16710 |
| SpecEE: Accelerating LLM Inference with Speculative Early Exiting | ◦ preprint | arxiv.org/abs/2504.08850 |
| Better & Faster Large Language Models via Multi-token Prediction | ◦ preprint | arxiv.org/abs/2404.19737 |
| DeepSeek-V3 Technical Report (MTP module) | ◦ preprint | arxiv.org/abs/2412.19437 |
| Understanding Dynamic Compute Allocation in Recurrent Transformers | ◦ preprint | arxiv.org/abs/2602.08864 |
| Subjective Depth and Timescale Transformers: Learning Where and When to Compute | ◦ preprint | arxiv.org/abs/2511.21408 |
| AdaPonderLM: Gated Pondering LMs with Token-Wise Adaptive Depth | ◦ preprint | arxiv.org/abs/2603.01914 |
| Look Inward to Explore Outward: Temperature Policy from LLM Internal States via Hierarchical RL | ◦ preprint | arxiv.org/abs/2602.13035 |
| Entropy trajectory shape predicts LLM reasoning reliability | ◦ preprint | arxiv.org/abs/2603.18940 |
| When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training | ◦ preprint | arxiv.org/abs/2601.00894 |
| What Layers When: Learning to Skip Compute in LLMs with Residual Gates | ◦ preprint | arxiv.org/abs/2510.13876 |
| Temporally Extended Mixture-of-Experts | ◦ preprint | arxiv.org/abs/2604.20156 |
| Path-Constrained Mixture-of-Experts (PathMoE) | ◦ preprint | arxiv.org/abs/2603.18297 |
| Omni-Router: Sharing Routing Decisions across MoE layers | ◦ preprint | arxiv.org/abs/2507.05724 |
| Mixture-of-Control: State-Aware Fine-Tuning | ◦ preprint | arxiv.org/abs/2606.31397 |
| Learning to Skip the Middle Layers of Transformers | ◦ preprint | arxiv.org/abs/2506.21103 |
| TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation | ◦ preprint | arxiv.org/abs/2607.06601 |
| Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents | ◦ preprint | arxiv.org/abs/2606.15017 |
| Token Reduction Is Not Cost Reduction | ◦ preprint | arxiv.org/abs/2607.12161 |
| Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference | ◦ preprint | arxiv.org/abs/2607.25018 |
| When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models | ◦ preprint | arxiv.org/abs/2606.30852 |
| When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty | ◦ preprint | arxiv.org/abs/2608.00685 |
| Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts | ◦ preprint | arxiv.org/abs/2607.20519 |
| LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers | ◦ preprint | arxiv.org/abs/2608.06867 |
| Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models | arxiv.org/abs/2607.08665 | |
| Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse | arxiv.org/abs/2608.03893 | |
| Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades | arxiv.org/abs/2606.15308 |