Skip to content
BotServBotServ
OllamaPrompt EngineeringPromptsLLMChat

Prompt Engineering for Ollama

Write better prompts for local models. Roles, context, examples, chain-of-thought, and error avoidance.

S

schutzgeist

4 min read
Prompt Engineering for Ollama

Prompt Engineering for Ollama

What this article covers

  • What prompt engineering is and why it matters.
  • Core techniques like roles, context, and examples.
  • How to structure complex tasks.
  • Methods such as Chain-of-Thought and Few-Shot.
  • Common mistakes and how to avoid them.

Introduction: Prompt Engineering for Ollama

Prompt engineering is the art of guiding models toward better answers through careful task formulation. Unlike larger models, small language models running on local hardware benefit particularly from clear, precise prompts. A well-crafted prompt can coax more useful output from a 7B model than a careless approach to a 13B model, without needing more powerful hardware.

This article covers the essential techniques and patterns for working effectively with Ollama models.

Key terminology

  • Prompt: Input sent to a language model.
  • System prompt: Background instruction that steers behavior.
  • User prompt: The actual question or task.
  • Context: Additional information within the prompt.
  • Few-Shot: Examples of desired input and output.
  • Zero-Shot: No examples, just the task.
  • Chain-of-Thought: The model reasons through steps aloud.
  • Output format: The desired structure of the response.

Ground rules

  • Prompts should be clear and precise.
  • Describe the role of the model.
  • Specify the desired tone and style.
  • Give concrete output requirements.
  • Break complex tasks into smaller steps.
  • Provide examples when format matters.
  • Avoid mixing too many different instructions.

Assigning a role

A well-defined role improves answer quality.

You are an experienced Python developer. Explain concepts concisely with practical examples.

Other examples:

  • Expert AI consultant.
  • Creative copywriter.
  • Technical translator.
  • Security analyst.
  • Teacher for beginners.

Adding context

The more relevant information in the context, the more precisely the model can answer.

Context: I run Ollama on a Ryzen 9 with RTX 4070 Ti and 64 GB RAM.
Question: Which 13B models run smoothly on this hardware?

Specifying output format

Give the answer in three bullet points, followed by a one-sentence summary at the end.

Providing examples

Few-Shot

Write brief product descriptions.

Example 1:
Product: Mechanical keyboard
Description: Robust mechanical keyboard with precise switches and durable construction.

Example 2:
Product: 4K monitor
Description: Large 4K monitor with sharp image and high color accuracy for productive work.

Product: Local AI workstation
Description:

The model recognizes the pattern and continues it.

Chain-of-Thought

For complex tasks:

Solve the task step by step and briefly explain each intermediate step.

This forces the model to articulate intermediate steps rather than jumping straight to an answer. It reduces errors.

Negative examples

When the model should avoid something, exclude it explicitly:

Explain quantization without technical jargon. Avoid mathematical formulas.

Iterating and testing

A good prompt emerges through experimentation:

  1. Write an initial version.
  2. Review the response.
  3. Refine the instructions.
  4. Add or remove examples.
  5. Adjust temperature.
  6. Refine the system prompt.

Advanced techniques

Self-Consistency

Ask the same question multiple times and select the most common answer. Improves reliability.

ReAct

Reasoning and acting in turns: the model thinks, acts, observes, and draws conclusions. Useful for agents and tool use.

Tree-of-Thoughts

Develop and evaluate multiple solution paths instead of following a single trajectory.

Persona switching

For creative or controversial topics, ask the model to adopt different perspectives.

Prompt for RAG

Answer the question using only the following context. If the answer is not in the context, state that no relevant information is available.

Context: {context}
Question: {question}

Prompt for coding

You are an experienced software developer. Write clean, commented Python code for the following function. Briefly explain what the code does.

Task: A function that sorts a list of numbers and returns the largest n values.

Prompt for summaries

Summarize the following text in three to five bullet points. Keep the most important facts.

Text: {text}

Common mistakes

  • Too vague: The model has room for interpretation.
  • Overly long introduction: Important information gets buried.
  • Contradictory instructions: The model cannot do both at once.
  • Missing examples: The model delivers the wrong format.
  • Temperature too high: Creative tasks become erratic.
  • Context overloaded: Not all information is relevant.
  • No iteration: The prompt never improves.

Tips for local models

  • Local models are often smaller, so precision in prompting pays off.
  • System prompts have a large impact on output.
  • Keep temperature low when consistency matters.
  • Make formatting requirements very explicit.
  • Split complex tasks across multiple prompts.

Further reading

FAQ: Prompt Engineering

Do I need long prompts? Not necessarily. Clear and precise beats long.

Which is better: system or user prompt? Both serve different roles. System prompts govern behavior, user prompts deliver the task.

How do I find good prompts? Test, iterate, and collect examples.

What temperature works best? 0.1 to 0.3 for precise tasks, 0.7 to 0.9 for creative ones.

Do examples actually help? Yes, especially for specific output formats.

Sources and further reading

Summary: Prompt Engineering for Ollama

Prompt engineering is the fastest way to extract better results from Ollama models. Clear roles, precise context, output specifications, and examples lead to noticeably better answers. Techniques like Chain-of-Thought, Few-Shot, and negative examples help with complex tasks. Iteratively testing and refining prompts quickly reveals patterns that work well even with smaller local models.

Back to Blog
Share:

Related Posts