Data Analysis with Ollama
What this article covers
- Which analysis tasks Ollama can handle.
- How to evaluate texts, tables, and logs.
- How to get structured output.
- Examples of classification, extraction, and reports.
- Tips for quality, performance, and data privacy.
Introduction: Data analysis with Ollama
Language models do more than generate text. They can analyze it too. Ollama is particularly well-suited for this because data stays local and sensitive content never leaves your machine. Emails, support tickets, log files, contracts, and product descriptions can be automatically classified, summarized, or converted into structured data.
This article shows what data analysis tasks Ollama can handle and how to implement them in practice.
Key concepts
- Classification: Assigning text to a category.
- Extraction: Pulling out specific information.
- Summarization: Compressing text into fewer words.
- Sentiment analysis: Detecting emotional tone.
- Structured output: Returning data in a defined format.
- Prompt template: Repeatable input structure.
- Named Entity Recognition: Identifying names, locations, or concepts.
- Batch: Processing multiple items at once.
Use cases
- Sort emails or tickets by category.
- Rate customer feedback by sentiment.
- Search contracts for specific keywords.
- Check log files for error patterns.
- Structure product descriptions.
- Summarize documents.
- Extract entities like people or companies.
- Derive tables from text.
Classification
A straightforward prompt for categorization:
import requests
def classify(text, model="llama3.1"):
url = "http://localhost:11434/api/generate"
prompt = f"""Classify the following text as one of these categories:
- Technical
- Financial
- Personal
- Spam
Text: {text}
Reply with only the category."""
response = requests.post(url, json={
"model": model,
"prompt": prompt,
"stream": False
})
return response.json()["response"].strip()
print(classify("My invoice is incorrect."))
Summarization
def summarize(text, model="llama3.1"):
prompt = f"Summarize the following text in three sentences:\n\n{text}"
response = requests.post(
"http://localhost:11434/api/generate",
json={"model": model, "prompt": prompt, "stream": False}
)
return response.json()["response"]
Extraction
def extract_entities(text, model="llama3.1"):
prompt = f"""Extract person names, locations, and organizations from the text below.
Return the answer as JSON:
{{"persons": [], "locations": [], "organizations": []}}
Text: {text}"""
response = requests.post(
"http://localhost:11434/api/generate",
json={"model": model, "prompt": prompt, "stream": False}
)
return response.json()["response"]
Table analysis
Ollama can also evaluate CSV or Markdown tables:
def analyze_table(table_text, model="llama3.1"):
prompt = f"""Analyze the following table and name the three key insights:
{table_text}"""
response = requests.post(
"http://localhost:11434/api/generate",
json={"model": model, "prompt": prompt, "stream": False}
)
return response.json()["response"]
Analyzing log files
def analyze_log(log_text, model="llama3.1"):
prompt = f"""Analyze the following log entries and name the most common error types:
{log_text}"""
response = requests.post(
"http://localhost:11434/api/generate",
json={"model": model, "prompt": prompt, "stream": False}
)
return response.json()["response"]
Structured output
Consistent results require a clear format:
Return the answer only as JSON:
{
"category": "Technical",
"confidence": 0.9,
"keywords": ["network", "error"]
}
Models don’t always produce valid JSON, so you should validate the output.
Batch processing
import json
texts = ["Text one", "Text two", "Text three"]
results = []
for text in texts:
result = classify(text)
results.append({"text": text, "category": result})
with open("results.json", "w") as f:
json.dump(results, f, indent=2, ensure_ascii=False)
Tips
- Define categories clearly and distinctly.
- Include examples in the prompt when needed.
- Keep temperature low for consistent classification.
- Parse and validate outputs.
- Use models with strong text understanding.
- Chunk long documents first.
- Keep sensitive data local, never send it to external services.
Common pitfalls
- Inconsistent output: Model returns JSON one moment, free text the next.
- Too many categories: Model misclassifies into wrong categories.
- Long documents: Context window runs out.
- Wrong format: Parsing fails.
- Model hallucination: Extracts data that isn’t there.
- Temperature too high: Unstable results.
- Prompt too vague: Classification lacks precision.
Further reading
- BotServ.de Ollama commands
- BotServ.de Ollama API libraries
- BotServ.de Ollama automation
- BotServ.de Ollama prompt engineering
FAQ: Data analysis with Ollama
Can Ollama understand tables? Yes, especially if they’re in Markdown or CSV format.
Is classification reliable? For simple categories, usually yes. Test first with complex domains.
How many texts can I analyze? As many as your hardware and time allow. Batch processing works.
Do I need JSON output? Not required, but it makes downstream processing easier.
Which model works best for analysis?
llama3.1, qwen2.5, or mistral are solid all-rounders.
Sources and further reading
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
- NER Overview: https://en.wikipedia.org/wiki/Named-entity_recognition
- JSON validation in Python: https://docs.python.org/3/library/json.html
Summary: Data analysis with Ollama
Ollama works well for local data analysis. You can classify, summarize, evaluate, and structure text. With clear prompts, simple scripts, and batch processing, repetitive analysis tasks become automated. The keys are well-defined categories, validated output, and an appropriate model. Analyzing data locally keeps you compliant with privacy rules and lets you safely process sensitive content.


