Skip to content
BotServBotServ
JanDesktopLocal AISoftwareGGUFOpen SourceOfflineAPI

Jan: Offline AI Client for Desktop

Jan: Open-source desktop client for local AI. Load models, chat, and use API. Installation, features, and LM Studio comparison.

S

schutzgeist

11 min read
Jan: Offline AI Client for Desktop

Jan: Offline AI Client for Desktop

What This Article Covers

  • What Jan is and how the desktop client works for local AI
  • How to install Jan on Windows, macOS, and Linux
  • How to load models, chat, and use the API server
  • How Jan compares to LM Studio and Ollama
  • Common pitfalls and how to avoid them

Introduction: Understanding Jan

Jan is an open-source desktop client for local AI. It lets you run large language models directly on your machine without sending data to cloud services. You download a model in GGUF format, open a chat, and get started. Everything happens offline on your hardware.

Anyone new to local AI often wonders the same thing: do I really need a graphical application, or is the command line enough? Jan is built for users who want an intuitive interface. You don’t need to open a terminal to launch a model. Instead, you browse through the integrated Model Hub, download a file, and chat in the same window.

In this article, you’ll learn how Jan is structured, what features it offers, and where its limits are. If you’re new to local AI, start with our guide on what local AI actually is. For broader context, our overview of local AI software is also worth a look.

Why Do You Need Jan?

Imagine you want to use AI for private notes, code snippets, or drafts containing sensitive information. You don’t want that text landing on someone else’s server. At the same time, you’re not keen on assembling a complex setup from Python environments and command-line tools.

That’s where Jan fits in. You install an application, launch it, and see an interface that looks like a chat client. Load a model, write a message, and get a response. No cloud account, no API keys, no internet required once the model is downloaded. Everything stays on your machine.

For deeper context on the topic, our article on running LLMs locally has more background.

Jan at a Glance

Jan is a desktop application built on Electron that runs language models locally. Under the hood, it uses an engine called Cortex, which handles loading and running models. The interface gives you a chat area, a model browser, and settings for parameters like temperature and context length.

Key features of Jan:

  • Open source under the AGPL license
  • Runs offline once models are downloaded
  • Supports GGUF format for quantized models
  • Provides a local API server compatible with OpenAI’s API
  • Extensible through add-ons

Who Is Jan For?

Jan serves several groups:

  • Newcomers who want to explore local AI without touching the command line
  • Privacy-conscious users who won’t send data to cloud services
  • Developers who need a local API server for prototyping
  • People who must work offline, such as while traveling or on restricted networks

If you’re already using Ollama and looking for a graphical interface, check out Open WebUI. Not sure which graphical client to pick? Comparisons are coming up later in this article.

Key Terms Around Jan

TermExplanation
JanThe desktop application that provides a graphical interface for local AI
GGUFA file format for quantized models that Jan uses to load models
CortexThe engine powering Jan, which loads and runs models
ThreadA single chat conversation in Jan, similar to a chat history
Model HubThe built-in catalog in Jan where you download models
API ServerA local server in Jan that provides an OpenAI-compatible interface
ExtensionA plugin that adds features like additional model sources or tools
System PromptAn instruction at the start of a thread that shapes the model’s behavior
ContextThe text window the model can consider when processing a request
OfflineOperation without an internet connection, after models are downloaded

For more on quantization and the GGUF format, see our article on quantization.

Installation

Jan installs on the three major operating systems. Find the official download page at jan.ai.

Windows

  1. Open jan.ai in your browser.
  2. Download the Windows installer, a file ending in .exe.
  3. Run the file and follow the installation dialog.
  4. Start Jan from the Start menu.

macOS

  1. Open jan.ai in your browser.
  2. Download the macOS version, a file ending in .dmg.
  3. Open the .dmg file and drag Jan into your Applications folder.
  4. Start Jan from Applications or via Spotlight.

Linux

  1. Open jan.ai in your browser.
  2. Download the appropriate version, either .AppImage or .deb.
  3. For .AppImage: Make the file executable with chmod +x jan.AppImage and run it.
  4. For .deb: Install with sudo dpkg -i jan.deb.

After the first launch, you’ll see Jan’s main interface. No model is loaded yet, but that comes next.

Finding and Loading Models

Jan comes with an integrated Model Hub. You don’t need to hunt for model files on the internet. Just browse directly in the application.

To load a model:

  1. Click the models section in the sidebar.
  2. Browse the list of available models.
  3. Choose a model that fits your hardware. For beginners, smaller models like Llama 3.2 with 3B parameters are a good starting point.
  4. Click Download and wait for the GGUF file to finish.

Pay attention to model size. A 7B parameter model with moderate quantization needs several gigabytes of storage and RAM. If you’re unsure what hardware you need, our article on CPU vs. GPU can help.

Chat Function

Once a model is loaded, you can start chatting right away.

  1. Create a new thread using the button in the sidebar.
  2. Select the loaded model from the dropdown menu.
  3. Optional: Add a system prompt to shape the model’s behavior.
  4. Type your message in the input field and send it.

Each thread stores its history separately. You can run multiple threads in parallel and switch between them. The system prompt applies per thread, so you can set different instructions for different tasks.

In a thread’s settings, you’ll find parameters like temperature, Top P, and context length. Temperature controls how creative the responses are. A low value like 0.1 produces conservative answers, while a higher value like 0.8 brings more variety.

API Server

Jan includes a local API server that mimics the OpenAI API. This means you can run applications built for OpenAI’s API using local models instead.

To enable the API server:

  1. Open the settings in Jan.
  2. Navigate to the API server section.
  3. Enable the server.
  4. Note the address, typically http://localhost:1337/v1.

In your application, set the base URL to this address and provide any API key you like, since Jan doesn’t require real authentication. Here’s an example using Python with the OpenAI library:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1337/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="llama3.2-3b",
    messages=[
        {"role": "user", "content": "Erkläre Quantisierung in einem Satz."}
    ]
)

print(response.choices[0].message.content)

This call runs entirely locally. No data leaves your machine.

Extensions

Jan is extensible through extensions. They can add new model providers, integrate tools, or customize the interface.

To use extensions:

  1. Open settings and go to the extensions section.
  2. You’ll see a list of installed and available extensions.
  3. Enable or disable extensions with a single click.
  4. Some extensions require additional setup, such as API keys for external services.

Common use cases for extensions:

  • Model providers: Connect to cloud services like OpenAI if you still want to use cloud-based models
  • Tools: Web search or code execution within the chat
  • Engine settings: Fine-tune the underlying engine

Extensions make Jan flexible. You start locally and expand as needed without switching applications.

Jan vs. LM Studio vs. Ollama

Jan isn’t the only client for local AI. The two most popular alternatives are LM Studio and Ollama. Here’s a comparison:

FeatureJanLM StudioOllama
LicenseOpen Source (AGPL)Proprietary, freeOpen Source (MIT)
InterfaceGUIGUICommand line, optional WebUI
Model formatGGUFGGUFGGUF (custom variant)
API serverYes, OpenAI-compatibleYes, OpenAI-compatibleYes, OpenAI-compatible
Model browserBuilt-inBuilt-inVia CLI or library
ExtensionsYesNoLimited
PlatformsWindows, macOS, LinuxWindows, macOS, LinuxWindows, macOS, Linux

For more details on alternatives, check out our articles on LM Studio and Ollama.

Your choice depends on your preferences. If you favor open source and appreciate extensions, Jan is a solid pick. For a particularly polished interface, LM Studio is worth exploring. If you prefer working in the terminal, Ollama is your go-to option.

Example: Running Your First Model in Jan

Here’s a concrete walkthrough of getting your first model running in Jan.

Step 1: Install Jan

Download Jan from jan.ai and install it as described above.

Step 2: Download a model

Open the model section in Jan. Search for a small model like Llama 3.2 with 3B parameters. Select a moderate quantization like Q4_K_M, which offers a good balance between file size and quality. Click download and wait for the file to finish loading.

Step 3: Create a thread

Click the button to create a new thread. Give it a name, such as “First Tests”. Select the model you just downloaded from the dropdown menu.

Step 4: Set the system prompt

Enter an instruction in the system prompt field, for example: “Du bist ein hilfreicher Assistent und antwortest auf Deutsch.”

Step 5: Start chatting

Type your first message in the input field, such as “Was kannst Du für mich tun?” and send it. The model generates a response. Depending on your hardware, this may take a moment.

Step 6: Adjust parameters

If responses are too creative or too predictable, adjust the temperature value in the thread. Save the setting and test again.

Common Pitfalls with Jan

Jan is straightforward, but a few points can trip up newcomers.

  1. Insufficient RAM: Models demand a lot of memory. A 7B model needs at least 8 GB of free RAM, preferably more. If your machine doesn’t have enough, the process crashes or slows to a crawl.

  2. Wrong model size: A large model like 70B parameters won’t run reasonably on a typical machine. Start with small models and work your way up.

  3. GPU acceleration disabled: Jan supports GPU acceleration, but you need to enable it. Without a GPU, the model runs on the CPU, which is significantly slower.

  4. Confusion about quantization: Different Q values like Q4, Q5, and Q8 confuse beginners. Q4 is smaller and faster; Q8 is more accurate but larger. For getting started, Q4_K_M is a safe choice.

  5. API server not active: If you want to use the API server, you must enable it in settings. Many users expect it to run by default, but it doesn’t.

  6. Dependency on the model hub: The model hub requires an internet connection. When offline, you can’t download new models, only use ones you’ve already loaded.

  7. Long load times: When you first run a model, Jan loads the file into RAM. This can take several seconds to minutes depending on model size and storage type. This is normal, not an error.

  8. Outdated Jan version: Jan receives regular updates. An old version may have problems with newer models or engines. Keep Jan up to date.

Hardware, Costs, and Security with Jan

Hardware

Hardware requirements depend on the model you want to run. As rough guidelines:

  • 3B models: 8 GB RAM, CPU sufficient for moderate speed
  • 7B models: 16 GB RAM, a GPU with 8 GB VRAM significantly accelerates performance
  • 13B models and larger: 32 GB RAM, a GPU with 16 GB VRAM or more

A fast SSD helps when loading model files. For more on hardware, see our article on CPU vs. GPU.

Costs

Jan itself is free and open source. You pay nothing for the software. Costs come solely from the hardware running Jan. There are no cloud fees because everything runs locally.

Security

Because Jan works offline, your data stays on your machine. This is a major advantage for sensitive content. Keep these points in mind:

  • Models can produce incorrect or invented answers; verify important information.
  • Extensions that connect to cloud services transmit data externally. Check which extensions you enable.
  • Download Jan only from the official site jan.ai to avoid tampered versions.

Further Resources and Information about Jan

  • Official website: jan.ai
  • Source code on GitHub: github.com/menloresearch/jan
  • Documentation: jan.ai/docs
  • Community chat: Links available on the GitHub page

For more topics around local AI:

FAQ: Jan - Common Questions

Is Jan free?

Yes, Jan is open source under the AGPL license and costs nothing. You don’t pay for the software itself.

Do I need internet to use Jan?

You need it to download models. Once a model sits on your machine, you can use Jan completely offline.

Which model should I download first?

For beginners, a small model like Llama 3.2 with 3B parameters in Q4_K_M quantization works well. It’s compact, runs on most machines, and produces usable results.

Can I use Jan instead of the OpenAI API?

Yes, Jan offers a local API server that mimics the OpenAI API. You can redirect programs written for OpenAI to run on local models.

Is Jan better than LM Studio?

That depends on your priorities. Jan is open source and offers extensions, while LM Studio has a somewhat more polished interface. Both are solid choices for local AI.

Do I need a GPU for Jan?

No, Jan runs on CPU alone. A GPU significantly speeds up generation, though. For small models, CPU suffices; for larger ones, a GPU is recommended.

Can I load multiple models at once?

Yes, you can download multiple models in Jan and switch between them. You should run only one at a time, since each model consumes RAM when loaded.

Where does Jan store models?

Jan stores models in a local directory on your machine. You’ll find the exact path in settings under the storage location section.

Can I load my own GGUF files in Jan?

Yes, you can import GGUF files from other sources into Jan. Place the file in the model directory or use the import function in the interface.

How do I update Jan?

Jan checks for updates at startup and notifies you. You can also manually download the latest version from jan.ai and install it over your existing installation. Your models and threads remain intact.

Is Jan safe for sensitive data?

Yes, as long as you work offline and don’t enable cloud extensions, all data stays on your machine. Still verify model responses for accuracy, since models can make mistakes.

Sources and Further Reading

  • Jan Documentation, jan.ai/docs
  • Jan GitHub Repository, github.com/menloresearch/jan
  • Overview of GGUF and Quantization, in our article on Quantization
  • Comparison of local AI tools, in our guide to Local AI Software
Back to Blog
Share:

Related Posts