Stirling-PDF self-hosted: PDF toolkit in your browser
What this article covers
- What Stirling-PDF is and why it’s the essential self-hosted PDF solution.
- Docker setup in 5 minutes.
- Overview of 50+ tools, focusing on the ones you’ll actually use.
- Integration into workflows (n8n, OCR for Paperless) and local AI pipelines.
- Configuration, security, and practical examples.
Introduction
Merging, splitting, signing, OCRing, and compressing PDFs: most people rely on online services like iLovePDF or SmallPDF. Problem is, your documents go to a third party. Stirling-PDF is the self-hosted answer: a complete PDF toolkit as a web app, open source (GPL), one Docker container, zero data leaving your server.
For BotServ users, Stirling-PDF is especially relevant because it bridges document workflows and local AI: OCR for RAG pipelines, PDF manipulation in n8n workflows, automated document processing.
Typical use cases
- Prepare documents for RAG: OCR scans, then feed them into your embedding pipeline → Local RAG.
- Complement Paperless-ngx: Stirling-PDF as a flexible PDF workbench alongside your DMS, see Paperless-ngx.
- Automate workflows: n8n calls the Stirling-PDF API to compress incoming email PDFs, OCR them, and pass them downstream.
- Handle company documents: Edit invoices and contracts without uploading to iLovePDF.
- Work with PDF forms: Fill, flatten, and sign them locally.
Installation: Docker in 5 minutes
# docker-compose.yml
services:
stirling-pdf:
image: stirlingtools/stirling-pdf:latest
container_name: stirling-pdf
ports:
- "127.0.0.1:8080:8080"
volumes:
- ./configs:/configs
- ./logs:/logs
environment:
- DOCKER_ENABLE_SECURITY=false # for internal networks
- LANGS=deu,eng # Tesseract languages for OCR
- SYSTEM_DEFAULTLOCALE=de-DE
restart: always
docker compose up -d
Done. The web UI runs at http://localhost:8080. For external access, add a reverse proxy plus authentication (see Security section).
Core tools at a glance
Over 50 functions; the takeaway is simple: everything iLovePDF does, you can do locally.
| Category | Tools | What you’ll use them for |
|---|---|---|
| Organization | Merge, Split, Rotate, Reorder, Remove Pages | Assembling documents |
| Conversion | PDF↔Word/Images/HTML/Markdown | Import/export pipelines |
| OCR | Tesseract-based | Make scans searchable (input for RAG!) |
| Security | Sign, Watermark, Password, Sanitize | Protect company documents |
| Optimization | Compress, Flatten | Email delivery, archiving |
| Images | Image↔PDF, Color adjustment | Post-processing scans |
Practical example: OCR pipeline for RAG
The workflow that turns Stirling-PDF into an AI building block:
Image-only PDF (scan)
│
▼
Stirling-PDF /api/v1/convert/ocr/pdf
│ (Tesseract deu)
▼
Searchable PDF
│
▼
Extract text → Chunk → Create embeddings → Qdrant
│
▼
RAG chatbot answers questions about your documents
API call from Python:
import requests
with open("scan.pdf", "rb") as f:
r = requests.post(
"http://stirling-pdf:8080/api/v1/convert/ocr/pdf",
files={"fileInput": f},
data={"languages": "deu", "sidecar": "false", "ocrRenderType": "hocr"},
)
with open("scan_ocred.pdf", "wb") as f:
f.write(r.content)
n8n integration
Stirling-PDF exposes a REST API that n8n can call via HTTP Request nodes:
Trigger (email with PDF)
│
▼
HTTP Request → POST /api/v1/convert/ocr/pdf
│
▼
HTTP Request → POST /api/v1/convert/pdf/text
│
▼
Code node: chunk the text
│
▼
Ollama embeddings → Qdrant
See n8n for n8n basics.
Configuration
The settings.yml file lives in your configs mount. Key options:
security:
enableLogin: true # Turn on for external access
system:
defaultLocale: de-DE
maxFileSize: 100MB
enableUrlToPDF: false # Security: no URL fetches
ui:
appName: "My PDF Tool"
homeDescription: "Local PDF workbench"
enterprise:
enabled: false # Pro features require a license; community edition is plenty
Security
- Enable login (
enableLogin: true) unless the service only runs on localhost. Otherwise anyone on your network can process your PDFs. - Localhost only: If Stirling-PDF runs on the same host as Paperless or n8n, binding to
127.0.0.1:8080is enough. - Sanitize PDFs: Stirling can strip scripts and metadata from incoming documents.
- Set file size limits: Protects against DoS on shared servers.
Stirling-PDF vs. alternatives
| Tool | Strength | Weakness |
|---|---|---|
| Stirling-PDF | Complete PDF toolkit, API, actively maintained | Java (heavier than necessary) |
| Paperless-ngx | Document management with AI tagging | No PDF editor |
| PDF-Toolbox/CUPS | Lightweight | No web UI |
| iLovePDF etc. | Zero setup | Data goes to third parties |
Further reading
- IRC-Coding.de: In-depth programming tutorials on REST APIs, PDF processing, and OCR pipelines.
- Paperless-ngx: Document management.
- OCR for documents: OCR fundamentals.
- Local RAG: Turn documents into AI knowledge.
- n8n: Workflow automation.
- Docker: Container basics.
Key takeaways:
- Stirling-PDF is self-hosted iLovePDF: 50+ PDF tools in your browser.
- One Docker container, REST API, no data leaves your server.
- Killer use case: OCR scans → input for RAG or Paperless.
- n8n integration makes it an automation building block.
- Enable login if you’re not just running on localhost.
FAQ
Is Stirling-PDF free?
Can Stirling-PDF handle German OCR?
Does it have an API?
Stirling-PDF or Paperless-ngx?
How much CPU and RAM does it need?
Multiple users?
Sources and further reading
- Stirling-PDF: GitHub.
- Stirling-PDF Docs: Documentation.
- IRC-Coding.de: Programming tutorials.


