Library · 15-modern-llm-architectures

Modern LLM arch

TitlePeerLink
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding— unrefarxiv.org/abs/1810.04805
Language Models are Few-Shot Learners (GPT-3)✓ peerarxiv.org/abs/2005.14165
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)✓ peerarxiv.org/abs/1910.10683
Llama 2: Open Foundation and Fine-Tuned Chat Models◦ preprintarxiv.org/abs/2307.09288
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness◦ preprintarxiv.org/abs/2205.14135
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning◦ preprintarxiv.org/abs/2307.08691
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints◦ preprintarxiv.org/abs/2305.13245
RoFormer: Enhanced Transformer with Rotary Position Embedding (RoPE)◦ preprintarxiv.org/abs/2104.09864
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention◦ preprintarxiv.org/abs/2502.11089
Efficient Attention Mechanisms for Large Language Models: A Survey◦ preprintarxiv.org/abs/2507.19595
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision◦ preprintarxiv.org/abs/2407.08608
Mamba: Linear-Time Sequence Modeling with Selective State Spaces◦ preprintarxiv.org/abs/2312.00752
Transformers are SSMs: Generalized Models and Efficient Algorithms (Mamba-2)◦ preprintarxiv.org/abs/2405.21060
RWKV: Reinventing RNNs for the Transformer Era◦ preprintarxiv.org/abs/2305.13048
Jamba: A Hybrid Transformer-Mamba Language Model◦ preprintarxiv.org/abs/2403.19887
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding◦ preprintarxiv.org/abs/2006.16668
Mixtral of Experts◦ preprintarxiv.org/abs/2401.04088
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts◦ preprintarxiv.org/abs/2401.06066
Efficient Memory Management for LLM Serving with PagedAttention (vLLM)◦ preprintarxiv.org/abs/2309.06180
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale◦ preprintarxiv.org/abs/2208.07339
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers◦ preprintarxiv.org/abs/2210.17323
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration◦ preprintarxiv.org/abs/2306.00978
QLoRA: Efficient Finetuning of Quantized LLMs◦ preprintarxiv.org/abs/2305.14314
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks— unrefarxiv.org/abs/2005.11401
Scaling Laws for Neural Language Models◦ preprintarxiv.org/abs/2001.08361
Training Compute-Optimal Large Language Models (Chinchilla)◦ preprintarxiv.org/abs/2203.15556
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters◦ preprintarxiv.org/abs/2408.03314
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning— unrefarxiv.org/abs/2501.12948
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models◦ preprintarxiv.org/abs/2201.11903
s1: Simple test-time scaling◦ preprintarxiv.org/abs/2501.19393
Self-Consistency Improves Chain of Thought Reasoning in Language Models◦ preprintarxiv.org/abs/2203.11171
Tree of Thoughts: Deliberate Problem Solving with Large Language Models◦ preprintarxiv.org/abs/2305.10601
Training language models to follow instructions with human feedback (InstructGPT / RLHF)◦ preprintarxiv.org/abs/2203.02155
Direct Preference Optimization: Your Language Model is Secretly a Reward Model (DPO)◦ preprintarxiv.org/abs/2305.18290
LoRA: Low-Rank Adaptation of Large Language Models◦ preprintarxiv.org/abs/2106.09685
Constitutional AI: Harmlessness from AI Feedback◦ preprintarxiv.org/abs/2212.08073
Root Mean Square Layer Normalization (RMSNorm)◦ preprintarxiv.org/abs/1910.07467
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA)◦ preprintarxiv.org/abs/2301.08243
Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA)◦ preprintarxiv.org/abs/2404.08471
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures◦ preprintarxiv.org/abs/2509.14252
The Llama 3 Herd of Models— unrefarxiv.org/abs/2407.21783
Qwen2.5 Technical Report◦ preprintarxiv.org/abs/2412.15115
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open LMs◦ preprintarxiv.org/abs/2402.03300
Let's Verify Step by Step◦ preprintarxiv.org/abs/2305.20050
Efficient Streaming Language Models with Attention Sinks (StreamingLLM)◦ preprintarxiv.org/abs/2309.17453
H₂O: Heavy-Hitter Oracle for Efficient Generative Inference◦ preprintarxiv.org/abs/2306.14048
YaRN: Efficient Context Window Extension of Large Language Models◦ preprintarxiv.org/abs/2309.00071
Train Short, Test Long: Attention with Linear Biases (ALiBi)◦ preprintarxiv.org/abs/2108.12409
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints◦ preprintarxiv.org/abs/2212.05055
Learning Transferable Visual Models From Natural Language Supervision (CLIP)◦ preprintarxiv.org/abs/2103.00020
High-Resolution Image Synthesis with Latent Diffusion Models◦ preprintarxiv.org/abs/2112.10752
Scalable Diffusion Models with Transformers (DiT)◦ preprintarxiv.org/abs/2212.09748
Neural Machine Translation of Rare Words with Subword Units (BPE)— unrefarxiv.org/abs/1508.07909
Attention Is All You Need (Transformer)— unrefarxiv.org/abs/1706.03762
Neural machine translation by jointly learning to align and translate◦ preprintarxiv.org/abs/1409.0473
Sequence to Sequence Learning with Neural Networks◦ preprintarxiv.org/abs/1409.3215
Efficient Estimation of Word Representations (word2vec)◦ preprintarxiv.org/abs/1301.3781
GloVe: Global Vectors for Word Representation— unrefdoi.org/10.3115/v1/D14-1162
Adaptive Mixtures of Local Experts✓ peerdoi.org/10.1162/neco.1991.3.1.79
Speech and Language Processing: An Introduction to NLP, Computational Linguistics, and Speech Recognition with Language Models (3e, draft)web.stanford.edu/~jurafsky/slp3/
Sparsely gated tiny linear experts◦ preprintarxiv.org/abs/2606.07414
SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization◦ preprintarxiv.org/abs/2608.04084
A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets◦ preprintarxiv.org/abs/2607.02303
Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models◦ preprintarxiv.org/abs/2606.05378
A Controlled Study of Attention-Only Transformers◦ preprintarxiv.org/abs/2607.18363