Skip to content
BotServBotServ
AI PapersTransformersAttention is all you needRAG PaperLLM ResearcharXiv

Essential AI Papers: Transformers, LLMs, RAG & Agents

Key AI papers explained for practitioners. Transformers, LLaMA, RAG, Mixture of Experts and more with clear analysis.

S

schutzgeist

5 min read

Essential AI Papers: Transformers, LLMs, RAG, and Agents

What This Article Covers

  • Which papers shaped the modern AI ecosystem.
  • What each paper is about, explained simply.
  • Which papers suit beginners and which require advanced knowledge.
  • Where to find papers free of charge (most are on arXiv).

Introduction

Every tool like Ollama, vLLM, or a RAG system rests on a paper. Understanding the core idea behind them helps you grasp why today’s tools work the way they do. This overview organizes the most important papers by topic and difficulty. The classification is practical: what value does the paper give you, not what is mathematically elegant.

Foundations: Transformers and Attention

Attention Is All You Need (Vaswani et al., 2017)

The paper that changed everything. It introduced the Transformer architecture. Instead of processing text word by word, like RNNs do, the model looks at all words at once and weights which ones matter to each other. This “attention” mechanism is why modern LLMs exist.

  • Why read it: All modern LLMs (Llama, Qwen, GPT) build on this idea. If you want to understand quantization, context windows, or attention optimizations, you need this foundation.
  • Difficulty: Medium. The concept is straightforward, the math moderate.
  • Link: arXiv:1706.03762

BERT (Devlin et al., 2018)

Bidirectional Encoder Representations from Transformers. BERT reads text in both directions and revolutionized text understanding tasks. Less relevant for chatbots, important for embeddings and classification.

  • Why read it: Embeddings for RAG and semantic search come from this line of work. If you use vector databases, you should know where embeddings come from.
  • Difficulty: Medium.
  • Link: arXiv:1810.04805

GPT-3: Language Models are Few-Shot Learners (Brown et al., 2020)

The paper that showed: if a model is large enough, it can solve tasks without being trained on them. Few-shot learning through examples in the prompt alone. The starting point of the chatbot era.

  • Why read it: Explains why prompt engineering works and why model size matters.
  • Difficulty: Easy to medium. The core is readable, training details are optional.
  • Link: arXiv:2005.14165

Open-Weight Models

LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023)

Meta released LLaMA as an open-weight model. The paper documents how to train a strong model with fewer parameters by using more and better data. Llama 2, 3, and derivatives followed.

Mistral 7B (Jiang et al., 2023)

A 7B model that outperformed larger ones. The paper showcases techniques like Grouped-Query Attention and Sliding-Window Attention, which are standard today.

  • Why read it: If you want to understand why a “small” model can be effective and what Sliding-Window Attention means.
  • Difficulty: Medium.
  • Link: arXiv:2310.06825

Mixtral of Experts (Jiang et al., 2024)

Mixture of Experts: the model has many “experts,” but only some activate per token. This saves compute while maintaining quality. Today the basis for many large open models.

  • Why read it: Explains why Mixtral, Qwen-MoE, and similar models exist. Relevant for MoE models.
  • Difficulty: Medium to advanced.
  • Link: arXiv:2401.04088

Retrieval-Augmented Generation (RAG)

RAG: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020)

The original RAG paper. A model retrieves knowledge from a document store instead of learning everything during training. This lets it use current or private data without retraining.

  • Why read it: All local knowledge bots and RAG setups build on this idea.
  • Difficulty: Medium. The concept is simple, details optional.
  • Link: arXiv:2005.11401

Dense Passage Retrieval (Karpukhin et al., 2020)

Shows that embedding-based search beats keyword search when you have good embeddings. Foundation for semantic search and vector databases.

  • Why read it: If you work with Chroma, Qdrant, or similar tools, you’re using techniques from this paper.
  • Difficulty: Advanced.
  • Link: arXiv:2004.04906

Agents and Tool Use

ReAct: Synergizing Reasoning and Acting (Yao et al., 2022)

The model thinks (reasoning) and acts (acting) in alternation. This is how agents emerge: they reason first, then call a tool, then continue reasoning. It’s the foundation for almost all agent frameworks.

Toolformer (Schick et al., 2023)

A model learns on its own when and which tool to call. It paved the way for function calling and MCP-like interfaces.

  • Why read it: Explains the logic behind tool use, relevant for MCP.
  • Difficulty: Advanced.
  • Link: arXiv:2302.04761

Efficiency: Why Local AI Works at All

GPTQ / AWQ / GGUF Quantization

There’s no single paper, but a series. GPTQ (2022) showed that you can quantize a model to 4 bits with minimal loss. AWQ and others followed. These papers make local AI on consumer hardware possible.

FlashAttention (Dao et al., 2022)

A trick that massively reduces memory use in the attention operation. This enables long contexts and faster inference. Today it’s built into nearly every inference engine.

  • Why read it: Explains why context lengths of 100k+ are possible and why VRAM matters so much.
  • Difficulty: Advanced.
  • Link: arXiv:2205.14135

How to Find More Papers

  • arXiv.org - The main preprint server. Search for “cs.CL” (Computation and Language) or “cs.LG” (Machine Learning).
  • Papers with Code - paperswithcode.com links papers with their code.
  • Semantic Scholar - semanticscholar.org finds related work.
  • Hugging Face Daily Papers - huggingface.co/papers curates new relevant papers daily.

Tips for Reading Papers

  1. Start with the abstract and conclusion. If that’s not enough, look at the diagrams.
  2. You don’t need to understand every formula. The idea matters more than the derivation.
  3. Scan the related work section. Often, related papers are written more simply.
  4. Use TL;DR pages. Many papers have community explanations (on Hugging Face or in blog posts).

Further Resources

  • Literature - Overview of books, papers, and blogs.
  • Blogs - Recommended blogs and newsletters.
  • Local AI - Practical implementation of concepts from papers.
  • Glossary - Terms like attention, embedding, and MoE explained clearly.

FAQ - Frequently Asked Questions

Do I need to read these papers to use local AI? No. For practical use, the articles on BotServ are enough. Papers are for those who want to understand why things work.

Which paper is easiest to start with? “Attention Is All You Need” and “ReAct” are relatively readable. GPTQ and FlashAttention are for advanced readers.

Why are papers on arXiv free? arXiv is a preprint server where authors upload work before official peer review. It makes science accessible and speeds up the exchange of ideas.

How current is this list? The list contains foundational papers that don’t become outdated. For the latest work, see the blog and newsletter recommendations on the Blogs page.

Back to Blog
Share:

Related Posts