Skip to content
BotServBotServ
Stirling-PDFPDFOCRSelf-HostingDockerOpen SourceDocuments

Stirling-PDF self-hosted: PDF toolkit in browser

Self-host Stirling-PDF: merge, split, OCR, convert PDFs locally and GDPR-compliant. Docker setup, n8n integration and examples.

S

schutzgeist

4 min read
Stirling-PDF self-hosted: PDF toolkit in browser

Stirling-PDF self-hosted: PDF toolkit in your browser

What this article covers

  • What Stirling-PDF is and why it’s the essential self-hosted PDF solution.
  • Docker setup in 5 minutes.
  • Overview of 50+ tools, focusing on the ones you’ll actually use.
  • Integration into workflows (n8n, OCR for Paperless) and local AI pipelines.
  • Configuration, security, and practical examples.

Introduction

Merging, splitting, signing, OCRing, and compressing PDFs: most people rely on online services like iLovePDF or SmallPDF. Problem is, your documents go to a third party. Stirling-PDF is the self-hosted answer: a complete PDF toolkit as a web app, open source (GPL), one Docker container, zero data leaving your server.

For BotServ users, Stirling-PDF is especially relevant because it bridges document workflows and local AI: OCR for RAG pipelines, PDF manipulation in n8n workflows, automated document processing.

Typical use cases

  • Prepare documents for RAG: OCR scans, then feed them into your embedding pipeline → Local RAG.
  • Complement Paperless-ngx: Stirling-PDF as a flexible PDF workbench alongside your DMS, see Paperless-ngx.
  • Automate workflows: n8n calls the Stirling-PDF API to compress incoming email PDFs, OCR them, and pass them downstream.
  • Handle company documents: Edit invoices and contracts without uploading to iLovePDF.
  • Work with PDF forms: Fill, flatten, and sign them locally.

Installation: Docker in 5 minutes

# docker-compose.yml
services:
  stirling-pdf:
    image: stirlingtools/stirling-pdf:latest
    container_name: stirling-pdf
    ports:
      - "127.0.0.1:8080:8080"
    volumes:
      - ./configs:/configs
      - ./logs:/logs
    environment:
      - DOCKER_ENABLE_SECURITY=false      # for internal networks
      - LANGS=deu,eng                     # Tesseract languages for OCR
      - SYSTEM_DEFAULTLOCALE=de-DE
    restart: always
docker compose up -d

Done. The web UI runs at http://localhost:8080. For external access, add a reverse proxy plus authentication (see Security section).

Core tools at a glance

Over 50 functions; the takeaway is simple: everything iLovePDF does, you can do locally.

CategoryToolsWhat you’ll use them for
OrganizationMerge, Split, Rotate, Reorder, Remove PagesAssembling documents
ConversionPDF↔Word/Images/HTML/MarkdownImport/export pipelines
OCRTesseract-basedMake scans searchable (input for RAG!)
SecuritySign, Watermark, Password, SanitizeProtect company documents
OptimizationCompress, FlattenEmail delivery, archiving
ImagesImage↔PDF, Color adjustmentPost-processing scans

Practical example: OCR pipeline for RAG

The workflow that turns Stirling-PDF into an AI building block:

Image-only PDF (scan)
    │
    ▼
Stirling-PDF /api/v1/convert/ocr/pdf
    │ (Tesseract deu)
    ▼
Searchable PDF
    │
    ▼
Extract text → Chunk → Create embeddings → Qdrant
    │
    ▼
RAG chatbot answers questions about your documents

API call from Python:

import requests

with open("scan.pdf", "rb") as f:
    r = requests.post(
        "http://stirling-pdf:8080/api/v1/convert/ocr/pdf",
        files={"fileInput": f},
        data={"languages": "deu", "sidecar": "false", "ocrRenderType": "hocr"},
    )
with open("scan_ocred.pdf", "wb") as f:
    f.write(r.content)

n8n integration

Stirling-PDF exposes a REST API that n8n can call via HTTP Request nodes:

Trigger (email with PDF)
    │
    ▼
HTTP Request → POST /api/v1/convert/ocr/pdf
    │
    ▼
HTTP Request → POST /api/v1/convert/pdf/text
    │
    ▼
Code node: chunk the text
    │
    ▼
Ollama embeddings → Qdrant

See n8n for n8n basics.

Configuration

The settings.yml file lives in your configs mount. Key options:

security:
  enableLogin: true        # Turn on for external access

system:
  defaultLocale: de-DE
  maxFileSize: 100MB
  enableUrlToPDF: false    # Security: no URL fetches

ui:
  appName: "My PDF Tool"
  homeDescription: "Local PDF workbench"

enterprise:
  enabled: false           # Pro features require a license; community edition is plenty

Security

  • Enable login (enableLogin: true) unless the service only runs on localhost. Otherwise anyone on your network can process your PDFs.
  • Localhost only: If Stirling-PDF runs on the same host as Paperless or n8n, binding to 127.0.0.1:8080 is enough.
  • Sanitize PDFs: Stirling can strip scripts and metadata from incoming documents.
  • Set file size limits: Protects against DoS on shared servers.

Stirling-PDF vs. alternatives

ToolStrengthWeakness
Stirling-PDFComplete PDF toolkit, API, actively maintainedJava (heavier than necessary)
Paperless-ngxDocument management with AI taggingNo PDF editor
PDF-Toolbox/CUPSLightweightNo web UI
iLovePDF etc.Zero setupData goes to third parties

Further reading

Key takeaways:

  • Stirling-PDF is self-hosted iLovePDF: 50+ PDF tools in your browser.
  • One Docker container, REST API, no data leaves your server.
  • Killer use case: OCR scans → input for RAG or Paperless.
  • n8n integration makes it an automation building block.
  • Enable login if you’re not just running on localhost.

FAQ

Is Stirling-PDF free?

Yes, it’s open source (GPL). There’s an Enterprise version with SSO and advanced features, but the community edition is sufficient for most use cases.

Can Stirling-PDF handle German OCR?

Yes. Tesseract supports German (deu). Set LANGS=deu,eng in docker-compose, then OCR with German language recognition.

Does it have an API?

Yes, a full REST API at /api/v1/, making Stirling-PDF a building block for n8n workflows and custom scripts.

Stirling-PDF or Paperless-ngx?

Different purposes: Paperless is a document management system (storage, tagging, search). Stirling-PDF is a PDF toolkit (editing, OCR, conversion). Using both together is powerful: Paperless archives, Stirling processes.

How much CPU and RAM does it need?

It’s a Java app: around 500 MB to 1 GB RAM at idle, more during OCR jobs. Runs without problems alongside other services on a mini-PC.

Multiple users?

Yes, with login enabled you get user accounts and roles. Without login, it’s a single-user tool.

Sources and further reading

Back to Blog
Share:

Related Posts