Library · 19-retrieval-augmentation-context-management

RAG & context mgmt

TitlePeerLink
RETRO — Improving language models by retrieving from trillions of tokens◦ preprintarxiv.org/abs/2112.04426
Atlas — Few-shot Learning with Retrieval Augmented Language Models◦ preprintarxiv.org/abs/2208.03299
REPLUG — Retrieval-Augmented Black-Box Language Models◦ preprintarxiv.org/abs/2301.12652
IRCoT — Interleaving Retrieval with Chain-of-Thought for multi-step QA◦ preprintarxiv.org/abs/2212.10509
Self-RAG — Learning to Retrieve, Generate and Critique through Self-Reflection◦ preprintarxiv.org/abs/2310.11511
FLARE — Active Retrieval Augmented Generation◦ preprintarxiv.org/abs/2305.06983
Adaptive-RAG — adapting retrieval to question complexity◦ preprintarxiv.org/abs/2403.14403
Adaptive-k — Efficient Context Selection for Long-Context QA◦ preprintarxiv.org/abs/2506.08479
UAR — Unified Active Retrieval for RAG◦ preprintarxiv.org/abs/2406.12534
Self-Route — RAG or Long-Context LLMs? (hybrid routing)◦ preprintarxiv.org/abs/2407.16833
DyCP — Dynamic Context Pruning for Long-Form Dialogue◦ preprintarxiv.org/abs/2601.07994
Knowing When to Stop — Efficient Context Processing via Latent Sufficiency Signals◦ preprintarxiv.org/abs/2502.01025
Quest — Query-Aware Sparsity for Efficient Long-Context Inference◦ preprintarxiv.org/abs/2406.10774
SnapKV — LLM Knows What You Are Looking For Before Generation◦ preprintarxiv.org/abs/2404.14469
PyramidKV — Dynamic KV Cache Compression via Pyramidal Information Funneling◦ preprintarxiv.org/abs/2406.02069
Ada-KV — Adaptive Budget Allocation for KV Cache Eviction◦ preprintarxiv.org/abs/2407.11550
DynamicKV — Task-Aware Adaptive KV Cache Compression◦ preprintarxiv.org/abs/2412.14838
TokenSelect — Dynamic Token-Level KV Cache Selection◦ preprintarxiv.org/abs/2411.02886
FastGen — Model Tells You What to Discard: Adaptive KV Cache Compression◦ preprintarxiv.org/abs/2310.01801
Scissorhands — Persistence of Importance for KV Cache Compression at Test Time◦ preprintarxiv.org/abs/2305.17118
LLMLingua — Compressing Prompts for Accelerated Inference◦ preprintarxiv.org/abs/2310.05736
LongLLMLingua — Prompt Compression for Long-Context Scenarios◦ preprintarxiv.org/abs/2310.06839
AutoCompressors — Adapting Language Models to Compress Contexts◦ preprintarxiv.org/abs/2305.14788
ACON — Optimizing Context Compression for Long-horizon LLM Agents◦ preprintarxiv.org/abs/2510.00615
Infini-attention — Leave No Context Behind: Efficient Infinite Context Transformers◦ preprintarxiv.org/abs/2404.07143
Compressive Transformer — Compressive Transformers for Long-Range Sequence Modelling✓ peerarxiv.org/abs/1911.05507
MemGPT — Towards LLMs as Operating Systems◦ preprintarxiv.org/abs/2310.08560
LongMem — Augmenting Language Models with Long-Term Memory◦ preprintarxiv.org/abs/2306.07174
MemoryBank — Enhancing LLMs with Long-Term Memory◦ preprintarxiv.org/abs/2305.10250
TRIME — Training Language Models with Memory Augmentation◦ preprintarxiv.org/abs/2205.12674
MemAgent — Multi-Conv RL-based Memory Agent for long context◦ preprintarxiv.org/abs/2507.02259
GAM — Hierarchical Graph-based Agentic Memory◦ preprintarxiv.org/abs/2604.12285
Agentic Memory (AgeMem) — Unified Long-/Short-Term Memory Management◦ preprintarxiv.org/abs/2601.01885
GraphRAG — From Local to Global: A Graph RAG Approach to Query-Focused Summarization— unrefarxiv.org/abs/2404.16130
HippoRAG — Neurobiologically Inspired Long-Term Memory for LLMs◦ preprintarxiv.org/abs/2405.14831
HippoRAG 2 — From RAG to Memory: Non-Parametric Continual Learning for LLMs◦ preprintarxiv.org/abs/2502.14802
LightRAG — Simple and Fast Retrieval-Augmented Generation◦ preprintarxiv.org/abs/2410.05779
RAPTOR — Recursive Abstractive Processing for Tree-Organized Retrieval◦ preprintarxiv.org/abs/2401.18059
Graph Retrieval-Augmented Generation: A Survey◦ preprintarxiv.org/abs/2408.08921
G-Retriever — RAG for Textual Graph Understanding and Question Answering◦ preprintarxiv.org/abs/2402.07630
StructRAG — Boosting Knowledge-Intensive Reasoning via Inference-time Hybrid Information Structurization◦ preprintarxiv.org/abs/2410.08815
NodeRAG — Structuring Graph-based RAG with Heterogeneous Nodes◦ preprintarxiv.org/abs/2504.11544
HeteroRAG — A Heterogeneous RAG Framework for Medical Vision Language Tasks— unrefarxiv.org/abs/2508.12778
HeteRAG — A Heterogeneous RAG Framework with Decoupled Knowledge Representations◦ preprintarxiv.org/abs/2504.10529
HetaRAG — Hybrid Deep RAG across Heterogeneous Data Stores◦ preprintarxiv.org/abs/2509.21336
Search-R1 — Training LLMs to Reason and Leverage Search Engines with RL◦ preprintarxiv.org/abs/2503.09516
DeepRAG — Thinking to Retrieve Step by Step for LLMs◦ preprintarxiv.org/abs/2502.01142
Search-o1 — Agentic Search-Enhanced Large Reasoning Models◦ preprintarxiv.org/abs/2501.05366
R1-Searcher — Incentivizing the Search Capability in LLMs via RL◦ preprintarxiv.org/abs/2503.05592
ReSearch — Learning to Reason with Search for LLMs via RL◦ preprintarxiv.org/abs/2503.19470
IterDRAG — Inference Scaling for Long-Context Retrieval-Augmented Generation◦ preprintarxiv.org/abs/2410.04343
PlanRAG — Plan-then-Retrieval Augmented Generation for Decision Making◦ preprintarxiv.org/abs/2406.12430
CRAG — Corrective Retrieval Augmented Generation◦ preprintarxiv.org/abs/2401.15884
RankRAG — Unifying Context Ranking with RAG◦ preprintarxiv.org/abs/2407.02485
Speculative RAG — Enhancing RAG through Drafting◦ preprintarxiv.org/abs/2407.08223
Blended RAG — Semantic Search + Hybrid Query-Based Retrievers◦ preprintarxiv.org/abs/2404.07220
Contextual Retrieval — Introducing Contextual Retrievalanthropic.com/news/contextual-retrieval
Retrieval-Augmented Generation for LLMs: A Survey◦ preprintarxiv.org/abs/2312.10997
Modular RAG — Transforming RAG into LEGO-like Reconfigurable Frameworks◦ preprintarxiv.org/abs/2407.21059
Agentic Retrieval-Augmented Generation: A Survey◦ preprintarxiv.org/abs/2501.09136
Zep — A Temporal Knowledge Graph Architecture for Agent Memory◦ preprintarxiv.org/abs/2501.13956
Larimar — LLMs with Episodic Memory Control◦ preprintarxiv.org/abs/2403.11901
MemInsight — Autonomous Memory Augmentation for LLM Agents◦ preprintarxiv.org/abs/2503.21760
MIRIX — Multi-Agent Memory System for LLM-Based Agents◦ preprintarxiv.org/abs/2507.07957
Generative Agents — Interactive Simulacra of Human Behavior◦ preprintarxiv.org/abs/2304.03442
Sleep-time Compute — Beyond Inference Scaling at Test-time◦ preprintarxiv.org/abs/2504.13171
A Survey on the Memory Mechanism of LLM-based Agents◦ preprintarxiv.org/abs/2404.13501
From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs◦ preprintarxiv.org/abs/2504.15965
From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms◦ preprintarxiv.org/abs/2605.06716
LOCOMO — Evaluating Very Long-Term Conversational Memory of LLM Agents◦ preprintarxiv.org/abs/2402.17753
LongMemEval — Benchmarking Chat Assistants on Long-Term Interactive Memory◦ preprintarxiv.org/abs/2410.10813
CRAG — Comprehensive RAG Benchmark◦ preprintarxiv.org/abs/2406.04744
RAGAS — Automated Evaluation of Retrieval Augmented Generation◦ preprintarxiv.org/abs/2309.15217
ARES — An Automated Evaluation Framework for RAG Systems◦ preprintarxiv.org/abs/2311.09476
RGB — Benchmarking Large Language Models in Retrieval-Augmented Generation◦ preprintarxiv.org/abs/2309.01431
Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search◦ preprintarxiv.org/abs/2607.15253
When Does Continual Learning Require Learning◦ preprintarxiv.org/abs/2607.07847
Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs◦ preprintarxiv.org/abs/2607.25600
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulatingarxiv.org/abs/2607.24667