pypdf: Merge, Split and Edit PDFs with Python (2026)
pypdf: Merge, Split and Edit PDFs with Python (2026) shows you how to do the PDF jobs that usually end up as boring manual work: merging reports, pulling out a few pages, splitting a file into one PDF per page, stamping a CONFIDENTIAL watermark, rotating a sideways page, setting metadata, locking a file with an AES-256 password, extracting text and shrinking a bloated merge. Everything uses pypdf, the pure-Python library with no system dependencies, and every output below comes from a real run.
PDF handling is a natural next step after automating Excel files with openpyxl: build the numbers in a spreadsheet, then ship them as a merged, watermarked, password-protected PDF pack. For the bigger picture of scripting repetitive work, see Python automation mastery in 2026. If your goal is feeding PDFs to an LLM or a RAG pipeline (layout, tables, Markdown), use Docling to turn PDFs into LLM-ready Markdown instead; pypdf is the tool for manipulating the PDF files themselves.
TL;DR
pip install "pypdf[crypto]". Pure Python, works on Windows, macOS, Linux, Docker and serverless. The[crypto]extra adds AES encryption support.- Merge:
writer = PdfWriter(); writer.append("a.pdf"); writer.append("b.pdf"); writer.write("merged.pdf"). - Pick pages:
writer.append("report.pdf", pages=[0, 2]). Split: onePdfWriterper page withadd_page(). - Edit:
page.merge_page(stamp)for watermarks,page.rotate(90),writer.add_metadata(),writer.encrypt(..., algorithm="AES-256"). - Read:
PdfReader("file.pdf").pages[n].extract_text(). Scanned PDFs have no text layer and need OCR. PyPDF2is the old name. New code should usepypdf; most class names are the same.- Tested with pypdf
6.19.0, cryptography50.0.2, Python3.13.5on 2026-10-08.
pypdf, PyPDF2, PyMuPDF or pdfplumber?
| Library | Best for | Install | Notes |
|---|---|---|---|
pypdf | Merge, split, rotate, watermark, encrypt, metadata, forms, basic text extraction | Pure Python | Actively maintained; the successor of PyPDF2 |
PyPDF2 | Legacy scripts only | Pure Python | Deprecated; its development moved back into pypdf |
PyMuPDF | Fast rendering, images, high-quality text extraction | Compiled wheels | AGPL or commercial licence |
pdfplumber | Tables and exact character positions | Pure Python (pdfminer.six) | Read-only; does not write PDFs |
Rule of thumb: if you are changing PDF files (combining, cutting, protecting, stamping), reach for pypdf. If you mainly need to read complex layouts or tables, use pdfplumber, PyMuPDF or Docling.
Versions tested (2026-10-08)
- Python
3.13.5 pypdf6.19.0(current release on PyPI; requires Python 3.9 or newer)cryptography50.0.2(installed by thepypdf[crypto]extra, used for AES-256)reportlab5.0.1(only to generate the sample PDFs)
python -m venv .venv && source .venv/bin/activate
pip install "pypdf[crypto]==6.19.0" reportlab
python make_samples.py
python merge_split.py
python extract_text.py
python edit_secure.py
python shrink.py
0. Create sample PDFs to practise on
To keep the tutorial reproducible, we generate three small PDFs with reportlab: a four-page Q3 sales report, a one-page invoice, and a transparent CONFIDENTIAL stamp that we will use as a watermark. You can skip this step and point the scripts at your own files.
# Creates the sample PDFs used in this tutorial (reportlab is only needed here).
from reportlab.lib.pagesizes import A4
from reportlab.pdfgen import canvas
def report(path="report.pdf"):
c = canvas.Canvas(path, pagesize=A4)
c.setTitle("Q3 Sales Report")
sections = ["Overview", "Regions", "Revenue", "Next steps"]
for n, name in enumerate(sections, start=1):
c.setFont("Helvetica-Bold", 22)
c.drawString(72, 770, f"Q3 Sales Report - {name}")
c.setFont("Helvetica", 12)
c.drawString(72, 740, f"Page {n} of {len(sections)}")
if name == "Revenue":
c.drawString(72, 700, "North 17,700.00 | South 9,872.50 | East 4,970.00")
c.drawString(72, 680, "Total revenue: 32,542.50")
else:
c.drawString(72, 700, f"Notes for the {name.lower()} section.")
c.showPage()
c.save()
def invoice(path="invoice.pdf"):
c = canvas.Canvas(path, pagesize=A4)
c.setFont("Helvetica-Bold", 22)
c.drawString(72, 770, "Invoice INV-1042")
c.setFont("Helvetica", 12)
c.drawString(72, 740, "Amount due: 1,250.00 | Due date: 2026-10-31")
c.showPage()
c.save()
def stamp(path="stamp.pdf"):
c = canvas.Canvas(path, pagesize=A4)
c.setFillColorRGB(0.85, 0.1, 0.1, alpha=0.25)
c.setFont("Helvetica-Bold", 80)
c.translate(300, 420)
c.rotate(45)
c.drawCentredString(0, 0, "CONFIDENTIAL")
c.showPage()
c.save()
if __name__ == "__main__":
report(); invoice(); stamp()
print("created report.pdf (4 pages), invoice.pdf (1 page), stamp.pdf")
$ python make_samples.py
created report.pdf (4 pages), invoice.pdf (1 page), stamp.pdf
1. Merge, pick pages and split PDFs
PdfWriter.append() is the modern way to merge: it copies all pages from a file (or just the pages you list), keeps links and bookmarks, and can add an outline entry so readers can jump to each part. Splitting is the reverse: create a new writer for each page and save it.
from pathlib import Path
from pypdf import PdfReader, PdfWriter
# 1) Merge whole files, with a bookmark for each one
writer = PdfWriter()
writer.append("report.pdf", outline_item="Q3 report")
writer.append("invoice.pdf", outline_item="Invoice")
writer.write("merged.pdf")
merged = PdfReader("merged.pdf")
print("merged.pdf pages:", len(merged.pages))
print("bookmarks:", [item.title for item in merged.outline])
# 2) Extract selected pages (0-based): report pages 1 and 3
picked = PdfWriter()
picked.append("report.pdf", pages=[0, 2])
picked.write("summary.pdf")
print("summary.pdf pages:", len(PdfReader("summary.pdf").pages))
# 3) Split: one PDF per page
out = Path("pages")
out.mkdir(exist_ok=True)
for i, page in enumerate(merged.pages, start=1):
single = PdfWriter()
single.add_page(page)
single.write(out / f"page_{i:02}.pdf")
print("split into:", sorted(p.name for p in out.glob("*.pdf")))
$ python merge_split.py
merged.pdf pages: 5
bookmarks: ['Q3 report', 'Invoice']
summary.pdf pages: 2
split into: ['page_01.pdf', 'page_02.pdf', 'page_03.pdf', 'page_04.pdf', 'page_05.pdf']
The four report pages plus the invoice give a five-page merged.pdf with two bookmarks, Q3 report and Invoice. The pages=[0, 2] argument uses 0-based indices, so it picked report pages 1 and 3 into summary.pdf. The split loop produced page_01.pdf to page_05.pdf; the zero-padded names keep them in order when sorted.
2. Extract and search text
extract_text() returns the text layer of a page. It is ideal for quick checks, routing files by keyword, or finding which page holds a total.
from pypdf import PdfReader
reader = PdfReader("merged.pdf")
print("page 3 text:")
print(reader.pages[2].extract_text())
# Find which pages mention a keyword
hits = [n for n, page in enumerate(reader.pages, start=1)
if "Total revenue" in page.extract_text()]
print("'Total revenue' found on page(s):", hits)
$ python extract_text.py
page 3 text:
Q3 Sales Report - Revenue
Page 3 of 4
North 17,700.00 | South 9,872.50 | East 4,970.00
Total revenue: 32,542.50
'Total revenue' found on page(s): [3]
The Revenue page came back line by line, and the keyword search correctly reports page 3. Remember that this only works for PDFs that contain real text. A scanned document is just an image inside a PDF, so extract_text() returns an empty string; run OCR (for example Tesseract) or use Docling for those.
3. Watermark, rotate, add metadata and encrypt
PdfWriter(clone_from=...) loads an existing PDF for editing. We stamp every page with the transparent CONFIDENTIAL page using merge_page(), rotate the invoice page by 90 degrees, set the document title and author, then encrypt with AES-256. The second half reopens the file and proves each change.
from pypdf import PdfReader, PdfWriter
stamp = PdfReader("stamp.pdf").pages[0]
writer = PdfWriter(clone_from="merged.pdf")
for page in writer.pages: # watermark every page
page.merge_page(stamp, over=True)
writer.pages[-1].rotate(90) # rotate the invoice page
writer.add_metadata({"/Title": "Q3 Sales Pack", "/Author": "Finance Bot"})
writer.encrypt(user_password="s3cret", owner_password="owner-pw",
algorithm="AES-256")
writer.write("secure.pdf")
reader = PdfReader("secure.pdf")
print("encrypted:", reader.is_encrypted)
print("wrong password:", reader.decrypt("guess").name)
print("right password:", reader.decrypt("s3cret").name)
print("title:", reader.metadata.title, "| author:", reader.metadata.author)
print("rotation of last page:", reader.pages[-1].rotation)
print("page 1 watermarked:", "CONFIDENTIAL" in reader.pages[0].extract_text())
$ python edit_secure.py
encrypted: True
wrong password: NOT_DECRYPTED
right password: USER_PASSWORD
title: Q3 Sales Pack | author: Finance Bot
rotation of last page: 90
page 1 watermarked: True
The wrong password returns NOT_DECRYPTED and the correct one returns USER_PASSWORD, so the file is really locked. The title and author survived encryption, the invoice page now has /Rotate 90, and the stamp text is part of page 1. As an independent check, poppler's pdfinfo reported Encrypted: yes (... algorithm:AES-256) for secure.pdf. The user password is what people type to open the file; the owner password controls permissions such as printing and copying.
4. Shrink a bloated merge
Merging many copies of similar files duplicates fonts, images and other resources. compress_identical_objects() removes the duplicates before you write the final file.
from pathlib import Path
from pypdf import PdfWriter
# Merging the same invoice 50 times duplicates fonts and resources
writer = PdfWriter()
for _ in range(50):
writer.append("invoice.pdf")
writer.write("bloated.pdf")
writer.compress_identical_objects(remove_duplicates=True, remove_unreferenced=True)
writer.write("compact.pdf")
for name in ("bloated.pdf", "compact.pdf"):
print(f"{name}: {Path(name).stat().st_size:,} bytes")
$ python shrink.py
bloated.pdf: 42,775 bytes
compact.pdf: 18,481 bytes
Fifty copies of the invoice dropped from 42,775 to 18,481 bytes, a 57% reduction, because the repeated font and resource objects are now stored once. In pypdf 6 the arguments are remove_duplicates and remove_unreferenced; the older remove_identicals and remove_orphans names still work but print a deprecation warning and will be removed in pypdf 7.
Common mistakes
- Still importing PyPDF2. Change
from PyPDF2 import PdfReader, PdfWritertofrom pypdf import PdfReader, PdfWriter. Very old names such asPdfFileReader,PdfFileMergerandgetPage()were removed; usePdfReader,PdfWriter.append()andreader.pages[i]. - Off-by-one page numbers.
reader.pagesand thepages=argument are 0-based. Page 1 is index 0. - Expecting text from scans.
extract_text()reads the text layer only. Empty output usually means the PDF is an image; use OCR. - Encrypting without the crypto extra. AES needs
pip install "pypdf[crypto]"(which brings incryptography). Without an AES backend (cryptography or pycryptodome), pypdf raises aDependencyError. - Treating owner restrictions as security. A file with an empty user password opens for anyone, and permission flags are honoured only by well-behaved viewers. Set a real user password for confidential files.
- Overwriting the input file. Write to a new file name, then replace the original once the output is verified.
A simple PDF automation workflow
- Generate the report data with pandas or openpyxl and export each part to PDF.
- Merge the parts with
PdfWriter.append()and add a bookmark per section. - Stamp a watermark and set the title, author and subject metadata.
- Run
compress_identical_objects(), encrypt with AES-256, and write a new file. - Reopen the result, check page count and keywords with
extract_text(), then email or upload it on a schedule.
FAQ
Is pypdf free for commercial use? Yes. pypdf is released under the permissive BSD-3-Clause licence.
Can pypdf create a PDF from scratch? It can create blank pages and combine or edit existing ones, but for drawing text and charts use a generator such as reportlab, then post-process with pypdf.
How do I merge all PDFs in a folder? Loop over sorted(Path("in").glob("*.pdf")), call writer.append(path) for each, then writer.write("all.pdf").
Can pypdf open password-protected PDFs? Yes. Call reader.decrypt("password") right after creating the reader, as section 3 shows.