Jan: Offline AI Client for Desktop
What This Article Covers
- What Jan is and how the desktop client works for local AI
- How to install Jan on Windows, macOS, and Linux
- How to load models, chat, and use the API server
- How Jan compares to LM Studio and Ollama
- Common pitfalls and how to avoid them
Introduction: Understanding Jan
Jan is an open-source desktop client for local AI. It lets you run large language models directly on your machine without sending data to cloud services. You download a model in GGUF format, open a chat, and get started. Everything happens offline on your hardware.
Anyone new to local AI often wonders the same thing: do I really need a graphical application, or is the command line enough? Jan is built for users who want an intuitive interface. You don’t need to open a terminal to launch a model. Instead, you browse through the integrated Model Hub, download a file, and chat in the same window.
In this article, you’ll learn how Jan is structured, what features it offers, and where its limits are. If you’re new to local AI, start with our guide on what local AI actually is. For broader context, our overview of local AI software is also worth a look.
Why Do You Need Jan?
Imagine you want to use AI for private notes, code snippets, or drafts containing sensitive information. You don’t want that text landing on someone else’s server. At the same time, you’re not keen on assembling a complex setup from Python environments and command-line tools.
That’s where Jan fits in. You install an application, launch it, and see an interface that looks like a chat client. Load a model, write a message, and get a response. No cloud account, no API keys, no internet required once the model is downloaded. Everything stays on your machine.
For deeper context on the topic, our article on running LLMs locally has more background.
Jan at a Glance
Jan is a desktop application built on Electron that runs language models locally. Under the hood, it uses an engine called Cortex, which handles loading and running models. The interface gives you a chat area, a model browser, and settings for parameters like temperature and context length.
Key features of Jan:
- Open source under the AGPL license
- Runs offline once models are downloaded
- Supports GGUF format for quantized models
- Provides a local API server compatible with OpenAI’s API
- Extensible through add-ons
Who Is Jan For?
Jan serves several groups:
- Newcomers who want to explore local AI without touching the command line
- Privacy-conscious users who won’t send data to cloud services
- Developers who need a local API server for prototyping
- People who must work offline, such as while traveling or on restricted networks
If you’re already using Ollama and looking for a graphical interface, check out Open WebUI. Not sure which graphical client to pick? Comparisons are coming up later in this article.
Key Terms Around Jan
| Term | Explanation |
|---|---|
| Jan | The desktop application that provides a graphical interface for local AI |
| GGUF | A file format for quantized models that Jan uses to load models |
| Cortex | The engine powering Jan, which loads and runs models |
| Thread | A single chat conversation in Jan, similar to a chat history |
| Model Hub | The built-in catalog in Jan where you download models |
| API Server | A local server in Jan that provides an OpenAI-compatible interface |
| Extension | A plugin that adds features like additional model sources or tools |
| System Prompt | An instruction at the start of a thread that shapes the model’s behavior |
| Context | The text window the model can consider when processing a request |
| Offline | Operation without an internet connection, after models are downloaded |
For more on quantization and the GGUF format, see our article on quantization.
Installation
Jan installs on the three major operating systems. Find the official download page at jan.ai.
Windows
- Open jan.ai in your browser.
- Download the Windows installer, a file ending in
.exe. - Run the file and follow the installation dialog.
- Start Jan from the Start menu.
macOS
- Open jan.ai in your browser.
- Download the macOS version, a file ending in
.dmg. - Open the
.dmgfile and drag Jan into your Applications folder. - Start Jan from Applications or via Spotlight.
Linux
- Open jan.ai in your browser.
- Download the appropriate version, either
.AppImageor.deb. - For
.AppImage: Make the file executable withchmod +x jan.AppImageand run it. - For
.deb: Install withsudo dpkg -i jan.deb.
After the first launch, you’ll see Jan’s main interface. No model is loaded yet, but that comes next.
Finding and Loading Models
Jan comes with an integrated Model Hub. You don’t need to hunt for model files on the internet. Just browse directly in the application.
To load a model:
- Click the models section in the sidebar.
- Browse the list of available models.
- Choose a model that fits your hardware. For beginners, smaller models like Llama 3.2 with 3B parameters are a good starting point.
- Click Download and wait for the GGUF file to finish.
Pay attention to model size. A 7B parameter model with moderate quantization needs several gigabytes of storage and RAM. If you’re unsure what hardware you need, our article on CPU vs. GPU can help.
Chat Function
Once a model is loaded, you can start chatting right away.
- Create a new thread using the button in the sidebar.
- Select the loaded model from the dropdown menu.
- Optional: Add a system prompt to shape the model’s behavior.
- Type your message in the input field and send it.
Each thread stores its history separately. You can run multiple threads in parallel and switch between them. The system prompt applies per thread, so you can set different instructions for different tasks.
In a thread’s settings, you’ll find parameters like temperature, Top P, and context length. Temperature controls how creative the responses are. A low value like 0.1 produces conservative answers, while a higher value like 0.8 brings more variety.
API Server
Jan includes a local API server that mimics the OpenAI API. This means you can run applications built for OpenAI’s API using local models instead.
To enable the API server:
- Open the settings in Jan.
- Navigate to the API server section.
- Enable the server.
- Note the address, typically
http://localhost:1337/v1.
In your application, set the base URL to this address and provide any API key you like, since Jan doesn’t require real authentication. Here’s an example using Python with the OpenAI library:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:1337/v1",
api_key="not-needed"
)
response = client.chat.completions.create(
model="llama3.2-3b",
messages=[
{"role": "user", "content": "Erkläre Quantisierung in einem Satz."}
]
)
print(response.choices[0].message.content)
This call runs entirely locally. No data leaves your machine.
Extensions
Jan is extensible through extensions. They can add new model providers, integrate tools, or customize the interface.
To use extensions:
- Open settings and go to the extensions section.
- You’ll see a list of installed and available extensions.
- Enable or disable extensions with a single click.
- Some extensions require additional setup, such as API keys for external services.
Common use cases for extensions:
- Model providers: Connect to cloud services like OpenAI if you still want to use cloud-based models
- Tools: Web search or code execution within the chat
- Engine settings: Fine-tune the underlying engine
Extensions make Jan flexible. You start locally and expand as needed without switching applications.
Jan vs. LM Studio vs. Ollama
Jan isn’t the only client for local AI. The two most popular alternatives are LM Studio and Ollama. Here’s a comparison:
| Feature | Jan | LM Studio | Ollama |
|---|---|---|---|
| License | Open Source (AGPL) | Proprietary, free | Open Source (MIT) |
| Interface | GUI | GUI | Command line, optional WebUI |
| Model format | GGUF | GGUF | GGUF (custom variant) |
| API server | Yes, OpenAI-compatible | Yes, OpenAI-compatible | Yes, OpenAI-compatible |
| Model browser | Built-in | Built-in | Via CLI or library |
| Extensions | Yes | No | Limited |
| Platforms | Windows, macOS, Linux | Windows, macOS, Linux | Windows, macOS, Linux |
For more details on alternatives, check out our articles on LM Studio and Ollama.
Your choice depends on your preferences. If you favor open source and appreciate extensions, Jan is a solid pick. For a particularly polished interface, LM Studio is worth exploring. If you prefer working in the terminal, Ollama is your go-to option.
Example: Running Your First Model in Jan
Here’s a concrete walkthrough of getting your first model running in Jan.
Step 1: Install Jan
Download Jan from jan.ai and install it as described above.
Step 2: Download a model
Open the model section in Jan. Search for a small model like Llama 3.2 with 3B parameters. Select a moderate quantization like Q4_K_M, which offers a good balance between file size and quality. Click download and wait for the file to finish loading.
Step 3: Create a thread
Click the button to create a new thread. Give it a name, such as “First Tests”. Select the model you just downloaded from the dropdown menu.
Step 4: Set the system prompt
Enter an instruction in the system prompt field, for example: “Du bist ein hilfreicher Assistent und antwortest auf Deutsch.”
Step 5: Start chatting
Type your first message in the input field, such as “Was kannst Du für mich tun?” and send it. The model generates a response. Depending on your hardware, this may take a moment.
Step 6: Adjust parameters
If responses are too creative or too predictable, adjust the temperature value in the thread. Save the setting and test again.
Common Pitfalls with Jan
Jan is straightforward, but a few points can trip up newcomers.
-
Insufficient RAM: Models demand a lot of memory. A 7B model needs at least 8 GB of free RAM, preferably more. If your machine doesn’t have enough, the process crashes or slows to a crawl.
-
Wrong model size: A large model like 70B parameters won’t run reasonably on a typical machine. Start with small models and work your way up.
-
GPU acceleration disabled: Jan supports GPU acceleration, but you need to enable it. Without a GPU, the model runs on the CPU, which is significantly slower.
-
Confusion about quantization: Different Q values like Q4, Q5, and Q8 confuse beginners. Q4 is smaller and faster; Q8 is more accurate but larger. For getting started, Q4_K_M is a safe choice.
-
API server not active: If you want to use the API server, you must enable it in settings. Many users expect it to run by default, but it doesn’t.
-
Dependency on the model hub: The model hub requires an internet connection. When offline, you can’t download new models, only use ones you’ve already loaded.
-
Long load times: When you first run a model, Jan loads the file into RAM. This can take several seconds to minutes depending on model size and storage type. This is normal, not an error.
-
Outdated Jan version: Jan receives regular updates. An old version may have problems with newer models or engines. Keep Jan up to date.
Hardware, Costs, and Security with Jan
Hardware
Hardware requirements depend on the model you want to run. As rough guidelines:
- 3B models: 8 GB RAM, CPU sufficient for moderate speed
- 7B models: 16 GB RAM, a GPU with 8 GB VRAM significantly accelerates performance
- 13B models and larger: 32 GB RAM, a GPU with 16 GB VRAM or more
A fast SSD helps when loading model files. For more on hardware, see our article on CPU vs. GPU.
Costs
Jan itself is free and open source. You pay nothing for the software. Costs come solely from the hardware running Jan. There are no cloud fees because everything runs locally.
Security
Because Jan works offline, your data stays on your machine. This is a major advantage for sensitive content. Keep these points in mind:
- Models can produce incorrect or invented answers; verify important information.
- Extensions that connect to cloud services transmit data externally. Check which extensions you enable.
- Download Jan only from the official site jan.ai to avoid tampered versions.
Further Resources and Information about Jan
- Official website: jan.ai
- Source code on GitHub: github.com/menloresearch/jan
- Documentation: jan.ai/docs
- Community chat: Links available on the GitHub page
For more topics around local AI:
FAQ: Jan - Common Questions
Is Jan free?
Yes, Jan is open source under the AGPL license and costs nothing. You don’t pay for the software itself.
Do I need internet to use Jan?
You need it to download models. Once a model sits on your machine, you can use Jan completely offline.
Which model should I download first?
For beginners, a small model like Llama 3.2 with 3B parameters in Q4_K_M quantization works well. It’s compact, runs on most machines, and produces usable results.
Can I use Jan instead of the OpenAI API?
Yes, Jan offers a local API server that mimics the OpenAI API. You can redirect programs written for OpenAI to run on local models.
Is Jan better than LM Studio?
That depends on your priorities. Jan is open source and offers extensions, while LM Studio has a somewhat more polished interface. Both are solid choices for local AI.
Do I need a GPU for Jan?
No, Jan runs on CPU alone. A GPU significantly speeds up generation, though. For small models, CPU suffices; for larger ones, a GPU is recommended.
Can I load multiple models at once?
Yes, you can download multiple models in Jan and switch between them. You should run only one at a time, since each model consumes RAM when loaded.
Where does Jan store models?
Jan stores models in a local directory on your machine. You’ll find the exact path in settings under the storage location section.
Can I load my own GGUF files in Jan?
Yes, you can import GGUF files from other sources into Jan. Place the file in the model directory or use the import function in the interface.
How do I update Jan?
Jan checks for updates at startup and notifies you. You can also manually download the latest version from jan.ai and install it over your existing installation. Your models and threads remain intact.
Is Jan safe for sensitive data?
Yes, as long as you work offline and don’t enable cloud extensions, all data stays on your machine. Still verify model responses for accuracy, since models can make mistakes.
Sources and Further Reading
- Jan Documentation, jan.ai/docs
- Jan GitHub Repository, github.com/menloresearch/jan
- Overview of GGUF and Quantization, in our article on Quantization
- Comparison of local AI tools, in our guide to Local AI Software


