Essential AI Papers: Transformers, LLMs, RAG, and Agents
What This Article Covers
- Which papers shaped the modern AI ecosystem.
- What each paper is about, explained simply.
- Which papers suit beginners and which require advanced knowledge.
- Where to find papers free of charge (most are on arXiv).
Introduction
Every tool like Ollama, vLLM, or a RAG system rests on a paper. Understanding the core idea behind them helps you grasp why today’s tools work the way they do. This overview organizes the most important papers by topic and difficulty. The classification is practical: what value does the paper give you, not what is mathematically elegant.
Foundations: Transformers and Attention
Attention Is All You Need (Vaswani et al., 2017)
The paper that changed everything. It introduced the Transformer architecture. Instead of processing text word by word, like RNNs do, the model looks at all words at once and weights which ones matter to each other. This “attention” mechanism is why modern LLMs exist.
- Why read it: All modern LLMs (Llama, Qwen, GPT) build on this idea. If you want to understand quantization, context windows, or attention optimizations, you need this foundation.
- Difficulty: Medium. The concept is straightforward, the math moderate.
- Link: arXiv:1706.03762
BERT (Devlin et al., 2018)
Bidirectional Encoder Representations from Transformers. BERT reads text in both directions and revolutionized text understanding tasks. Less relevant for chatbots, important for embeddings and classification.
- Why read it: Embeddings for RAG and semantic search come from this line of work. If you use vector databases, you should know where embeddings come from.
- Difficulty: Medium.
- Link: arXiv:1810.04805
GPT-3: Language Models are Few-Shot Learners (Brown et al., 2020)
The paper that showed: if a model is large enough, it can solve tasks without being trained on them. Few-shot learning through examples in the prompt alone. The starting point of the chatbot era.
- Why read it: Explains why prompt engineering works and why model size matters.
- Difficulty: Easy to medium. The core is readable, training details are optional.
- Link: arXiv:2005.14165
Open-Weight Models
LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023)
Meta released LLaMA as an open-weight model. The paper documents how to train a strong model with fewer parameters by using more and better data. Llama 2, 3, and derivatives followed.
- Why read it: Foundation for most local models on BotServ. Llama models are the reference.
- Difficulty: Medium.
- Link: arXiv:2302.13971
Mistral 7B (Jiang et al., 2023)
A 7B model that outperformed larger ones. The paper showcases techniques like Grouped-Query Attention and Sliding-Window Attention, which are standard today.
- Why read it: If you want to understand why a “small” model can be effective and what Sliding-Window Attention means.
- Difficulty: Medium.
- Link: arXiv:2310.06825
Mixtral of Experts (Jiang et al., 2024)
Mixture of Experts: the model has many “experts,” but only some activate per token. This saves compute while maintaining quality. Today the basis for many large open models.
- Why read it: Explains why Mixtral, Qwen-MoE, and similar models exist. Relevant for MoE models.
- Difficulty: Medium to advanced.
- Link: arXiv:2401.04088
Retrieval-Augmented Generation (RAG)
RAG: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020)
The original RAG paper. A model retrieves knowledge from a document store instead of learning everything during training. This lets it use current or private data without retraining.
- Why read it: All local knowledge bots and RAG setups build on this idea.
- Difficulty: Medium. The concept is simple, details optional.
- Link: arXiv:2005.11401
Dense Passage Retrieval (Karpukhin et al., 2020)
Shows that embedding-based search beats keyword search when you have good embeddings. Foundation for semantic search and vector databases.
- Why read it: If you work with Chroma, Qdrant, or similar tools, you’re using techniques from this paper.
- Difficulty: Advanced.
- Link: arXiv:2004.04906
Agents and Tool Use
ReAct: Synergizing Reasoning and Acting (Yao et al., 2022)
The model thinks (reasoning) and acts (acting) in alternation. This is how agents emerge: they reason first, then call a tool, then continue reasoning. It’s the foundation for almost all agent frameworks.
- Why read it: If you build AI agents or tool calling, you should know ReAct.
- Difficulty: Medium.
- Link: arXiv:2210.03629
Toolformer (Schick et al., 2023)
A model learns on its own when and which tool to call. It paved the way for function calling and MCP-like interfaces.
- Why read it: Explains the logic behind tool use, relevant for MCP.
- Difficulty: Advanced.
- Link: arXiv:2302.04761
Efficiency: Why Local AI Works at All
GPTQ / AWQ / GGUF Quantization
There’s no single paper, but a series. GPTQ (2022) showed that you can quantize a model to 4 bits with minimal loss. AWQ and others followed. These papers make local AI on consumer hardware possible.
- Why read it: Explains what quantization really does and why Q4 models work so well.
- Difficulty: Advanced. The concepts are explained more simply in our quantization guide.
- Link (GPTQ): arXiv:2210.17323
FlashAttention (Dao et al., 2022)
A trick that massively reduces memory use in the attention operation. This enables long contexts and faster inference. Today it’s built into nearly every inference engine.
- Why read it: Explains why context lengths of 100k+ are possible and why VRAM matters so much.
- Difficulty: Advanced.
- Link: arXiv:2205.14135
How to Find More Papers
- arXiv.org - The main preprint server. Search for “cs.CL” (Computation and Language) or “cs.LG” (Machine Learning).
- Papers with Code - paperswithcode.com links papers with their code.
- Semantic Scholar - semanticscholar.org finds related work.
- Hugging Face Daily Papers - huggingface.co/papers curates new relevant papers daily.
Tips for Reading Papers
- Start with the abstract and conclusion. If that’s not enough, look at the diagrams.
- You don’t need to understand every formula. The idea matters more than the derivation.
- Scan the related work section. Often, related papers are written more simply.
- Use TL;DR pages. Many papers have community explanations (on Hugging Face or in blog posts).
Further Resources
- Literature - Overview of books, papers, and blogs.
- Blogs - Recommended blogs and newsletters.
- Local AI - Practical implementation of concepts from papers.
- Glossary - Terms like attention, embedding, and MoE explained clearly.
FAQ - Frequently Asked Questions
Do I need to read these papers to use local AI? No. For practical use, the articles on BotServ are enough. Papers are for those who want to understand why things work.
Which paper is easiest to start with? “Attention Is All You Need” and “ReAct” are relatively readable. GPTQ and FlashAttention are for advanced readers.
Why are papers on arXiv free? arXiv is a preprint server where authors upload work before official peer review. It makes science accessible and speeds up the exchange of ideas.
How current is this list? The list contains foundational papers that don’t become outdated. For the latest work, see the blog and newsletter recommendations on the Blogs page.

