Library · 15-modern-llm-architectures
Modern LLM arch
| Title | Peer | Link |
|---|---|---|
| BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding | — unref | arxiv.org/abs/1810.04805 |
| Language Models are Few-Shot Learners (GPT-3) | ✓ peer | arxiv.org/abs/2005.14165 |
| Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) | ✓ peer | arxiv.org/abs/1910.10683 |
| Llama 2: Open Foundation and Fine-Tuned Chat Models | ◦ preprint | arxiv.org/abs/2307.09288 |
| FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness | ◦ preprint | arxiv.org/abs/2205.14135 |
| FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning | ◦ preprint | arxiv.org/abs/2307.08691 |
| GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints | ◦ preprint | arxiv.org/abs/2305.13245 |
| RoFormer: Enhanced Transformer with Rotary Position Embedding (RoPE) | ◦ preprint | arxiv.org/abs/2104.09864 |
| Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention | ◦ preprint | arxiv.org/abs/2502.11089 |
| Efficient Attention Mechanisms for Large Language Models: A Survey | ◦ preprint | arxiv.org/abs/2507.19595 |
| FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision | ◦ preprint | arxiv.org/abs/2407.08608 |
| Mamba: Linear-Time Sequence Modeling with Selective State Spaces | ◦ preprint | arxiv.org/abs/2312.00752 |
| Transformers are SSMs: Generalized Models and Efficient Algorithms (Mamba-2) | ◦ preprint | arxiv.org/abs/2405.21060 |
| RWKV: Reinventing RNNs for the Transformer Era | ◦ preprint | arxiv.org/abs/2305.13048 |
| Jamba: A Hybrid Transformer-Mamba Language Model | ◦ preprint | arxiv.org/abs/2403.19887 |
| GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding | ◦ preprint | arxiv.org/abs/2006.16668 |
| Mixtral of Experts | ◦ preprint | arxiv.org/abs/2401.04088 |
| DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts | ◦ preprint | arxiv.org/abs/2401.06066 |
| Efficient Memory Management for LLM Serving with PagedAttention (vLLM) | ◦ preprint | arxiv.org/abs/2309.06180 |
| LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale | ◦ preprint | arxiv.org/abs/2208.07339 |
| GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers | ◦ preprint | arxiv.org/abs/2210.17323 |
| AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration | ◦ preprint | arxiv.org/abs/2306.00978 |
| QLoRA: Efficient Finetuning of Quantized LLMs | ◦ preprint | arxiv.org/abs/2305.14314 |
| Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks | — unref | arxiv.org/abs/2005.11401 |
| Scaling Laws for Neural Language Models | ◦ preprint | arxiv.org/abs/2001.08361 |
| Training Compute-Optimal Large Language Models (Chinchilla) | ◦ preprint | arxiv.org/abs/2203.15556 |
| Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters | ◦ preprint | arxiv.org/abs/2408.03314 |
| DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning | — unref | arxiv.org/abs/2501.12948 |
| Chain-of-Thought Prompting Elicits Reasoning in Large Language Models | ◦ preprint | arxiv.org/abs/2201.11903 |
| s1: Simple test-time scaling | ◦ preprint | arxiv.org/abs/2501.19393 |
| Self-Consistency Improves Chain of Thought Reasoning in Language Models | ◦ preprint | arxiv.org/abs/2203.11171 |
| Tree of Thoughts: Deliberate Problem Solving with Large Language Models | ◦ preprint | arxiv.org/abs/2305.10601 |
| Training language models to follow instructions with human feedback (InstructGPT / RLHF) | ◦ preprint | arxiv.org/abs/2203.02155 |
| Direct Preference Optimization: Your Language Model is Secretly a Reward Model (DPO) | ◦ preprint | arxiv.org/abs/2305.18290 |
| LoRA: Low-Rank Adaptation of Large Language Models | ◦ preprint | arxiv.org/abs/2106.09685 |
| Constitutional AI: Harmlessness from AI Feedback | ◦ preprint | arxiv.org/abs/2212.08073 |
| Root Mean Square Layer Normalization (RMSNorm) | ◦ preprint | arxiv.org/abs/1910.07467 |
| Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA) | ◦ preprint | arxiv.org/abs/2301.08243 |
| Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA) | ◦ preprint | arxiv.org/abs/2404.08471 |
| LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures | ◦ preprint | arxiv.org/abs/2509.14252 |
| The Llama 3 Herd of Models | — unref | arxiv.org/abs/2407.21783 |
| Qwen2.5 Technical Report | ◦ preprint | arxiv.org/abs/2412.15115 |
| DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open LMs | ◦ preprint | arxiv.org/abs/2402.03300 |
| Let's Verify Step by Step | ◦ preprint | arxiv.org/abs/2305.20050 |
| Efficient Streaming Language Models with Attention Sinks (StreamingLLM) | ◦ preprint | arxiv.org/abs/2309.17453 |
| H₂O: Heavy-Hitter Oracle for Efficient Generative Inference | ◦ preprint | arxiv.org/abs/2306.14048 |
| YaRN: Efficient Context Window Extension of Large Language Models | ◦ preprint | arxiv.org/abs/2309.00071 |
| Train Short, Test Long: Attention with Linear Biases (ALiBi) | ◦ preprint | arxiv.org/abs/2108.12409 |
| Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints | ◦ preprint | arxiv.org/abs/2212.05055 |
| Learning Transferable Visual Models From Natural Language Supervision (CLIP) | ◦ preprint | arxiv.org/abs/2103.00020 |
| High-Resolution Image Synthesis with Latent Diffusion Models | ◦ preprint | arxiv.org/abs/2112.10752 |
| Scalable Diffusion Models with Transformers (DiT) | ◦ preprint | arxiv.org/abs/2212.09748 |
| Neural Machine Translation of Rare Words with Subword Units (BPE) | — unref | arxiv.org/abs/1508.07909 |
| Attention Is All You Need (Transformer) | — unref | arxiv.org/abs/1706.03762 |
| Neural machine translation by jointly learning to align and translate | ◦ preprint | arxiv.org/abs/1409.0473 |
| Sequence to Sequence Learning with Neural Networks | ◦ preprint | arxiv.org/abs/1409.3215 |
| Efficient Estimation of Word Representations (word2vec) | ◦ preprint | arxiv.org/abs/1301.3781 |
| GloVe: Global Vectors for Word Representation | — unref | doi.org/10.3115/v1/D14-1162 |
| Adaptive Mixtures of Local Experts | ✓ peer | doi.org/10.1162/neco.1991.3.1.79 |
| Speech and Language Processing: An Introduction to NLP, Computational Linguistics, and Speech Recognition with Language Models (3e, draft) | web.stanford.edu/~jurafsky/slp3/ | |
| Sparsely gated tiny linear experts | ◦ preprint | arxiv.org/abs/2606.07414 |
| SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization | ◦ preprint | arxiv.org/abs/2608.04084 |
| A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets | ◦ preprint | arxiv.org/abs/2607.02303 |
| Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models | ◦ preprint | arxiv.org/abs/2606.05378 |
| A Controlled Study of Attention-Only Transformers | ◦ preprint | arxiv.org/abs/2607.18363 |