Model Finder for Local AI
What this article covers
- How to find the right local model for your needs.
- Which criteria matter when selecting a model.
- Model categories: text, coding, vision, audio, and more.
- How to filter models by hardware, license, and task.
- Popular model lists and tools for discovery.
Introduction: Finding the Right Local AI Model
The number of open-source AI models grows daily. For newcomers, it’s challenging to pick the right model for a given use case. A model finder helps you filter by task, hardware, license, and language. You save time and avoid downloading models that won’t run on your machine.
This article walks through what to consider when choosing a model and which tools and lists are available.
Key Concepts
- Parameter size: Number of model weights, for example 7B or 70B.
- Quantization: Reducing the bit depth to make models smaller.
- Context length: Maximum number of tokens the model can process.
- Tool calling: Ability to produce structured function calls.
- License: Permitted usage, for example Apache 2.0 or Llama 3 Community License.
- VRAM: GPU video memory.
- MoE: Mixture-of-Experts, a memory-efficient architecture.
Selection Criteria
1. Use Case
Start by thinking about what you’ll use the model for:
- General conversation
- Coding
- Math and reasoning
- Vision and image analysis
- Audio and speech
- Embeddings and reranking
2. Hardware
Check your RAM and VRAM:
- 7B models in 4-bit: approximately 4 to 6 GB.
- 13B models in 4-bit: approximately 8 to 10 GB.
- 70B models in 4-bit: 40 GB and higher.
- MoE models require more memory than dense models of equivalent size.
3. Language
- German models or multilingual models
- English for coding and technical documentation
- Specialized models for specific languages
4. License
- Apache 2.0: very permissive.
- MIT: also open.
- Llama 3 Community License: can be used commercially with restrictions.
- Always verify commercial usage rights.
5. Context Length
For long documents or coding projects, you need longer context windows. 8K suffices for chat; 32K to 128K works for document analysis.
Model Categories
Text Models
General-purpose models for conversation, summarization, and text generation. Examples: Llama 3.1, Qwen 2.5, Mistral 7B, Gemma.
Coding Models
Optimized for code generation and debugging. Examples: Qwen 2.5 Coder, CodeLlama, Codestral, DeepSeek-Coder.
Reasoning Models
For mathematical and logical tasks. Examples: DeepSeek-R1, Mathstral, Qwen QWQ.
Vision Models
Understand images and combine them with text. Examples: LLaVA, Qwen-VL, BakLLaVA.
Speech-to-Text and Text-to-Speech
Convert speech to text and vice versa. Examples: Whisper, faster-whisper, Piper, MeloTTS.
Embedding and Reranking Models
For RAG and semantic search. Examples: BGE, nomic-embed-text, BAAI.
Popular Model Finders and Lists
- Ollama Library: https://ollama.com/library
- Hugging Face: https://huggingface.co/models
- LMSYS Arena: Benchmarks and leaderboards.
- Artificial Analysis: Model comparison.
- LocalLLaMA Subreddit: Community recommendations.
- TheBloke on Hugging Face: Quantized versions in GGUF format.
Decision Tree
1. Choose your use case
2. Find the right category
3. Check hardware (RAM/VRAM)
4. Determine required context length
5. Verify the license
6. Test the model in 4-bit
7. Measure quality on your own dataset
Common Pitfalls
- Relying only on benchmarks: Real-world performance may differ.
- Choosing a model that’s too large: Won’t run on your hardware.
- Ignoring licensing: Commercial use may be prohibited.
- Underestimating context length: Long documents need more context.
- Overlooking tool calling: Not every model supports function calls.
Further Resources and Information
- BotServ.de Local AI Models
- BotServ.de Hardware Requirements
- BotServ.de Quantization Basics
- BotServ.de Setting Up Ollama
FAQ: Model Finder
How do I quickly find a suitable model? Start with Ollama Library, filter by category, and test 7B or 8B variants.
What size is best for getting started? 7B to 8B in 4-bit runs on many consumer systems and delivers usable results.
Should I always pick the largest model? No. Larger models are slower and consume more memory. For many tasks, a smaller model suffices.
Are MoE models better? They can be more efficient but often require more disk space. Dense models are simpler for beginners.
How do I test a model? Download it with Ollama and try typical prompts from your use case.
Sources and Further Reading
- Ollama Library: https://ollama.com/library
- Hugging Face Models: https://huggingface.co/models
- LMSYS Chatbot Arena: https://chat.lmsys.org/
Summary: Model Finder for Local AI
The right model finder saves time and hardware resources. The most important criteria are use case, hardware, language, license, and context length. For newcomers, 7B to 8B models in 4-bit quantization are a good starting point. Specialized coding, reasoning, and vision models exist for every task. Testing against your actual dataset consistently beats relying on benchmarks alone to find the right model.


