When digitizing physical books, the raw output is usually a large PDF containing double-page spreads with black borders, uneven margins, and misaligned text. Cropping these scans into standardized, single-page formats like A4 or US Letter is crucial for readability and printing.

However, processing scanned book PDFs introduces friction: scanned books can run to hundreds of MB.

This guide breaks down the technical reality of PDF bounding boxes, provides a Python script for splitting spreads, and explains how to trim borders.

The Technical Reality: PDF Bounding Boxes

A scanned page is a raster image inside a PDF. Cropping a PDF typically does not delete the hidden areas of the page; it merely redefines the visible viewport using internal coordinate systems.

A page box is defined by four coordinates [left bottom right top] in points.

The PDF specification defines several page boundaries. The two most important for cropping scans are the MediaBox and CropBox. When you “crop” a scanned PDF, you are modifying these boundaries in the PDF dictionary.

Because PDF uses a standard unit of 1/72 of an inch (a point), standardizing your cropped pages requires targeting specific dimensions.

Standard PDF Page Dimensions (in Points)

For reference, common page sizes in points:

FormatDimensions (Inches)Dimensions (Millimeters)PDF Points (W x H)
US Letter8.5 × 11 in215.9 × 279.4 mm612 × 792
A48.27 × 11.69 in210 × 297 mm≈595.28 × 841.89
US Legal8.5 × 14 in215.9 × 355.6 mm612 × 1008
A5 (Half Book)5.83 × 8.27 in148 × 210 mm≈419.53 × 595.28

Method 1: The Developer Route (Python & pypdf)

If you are comfortable with the terminal, you can programmatically crop double-page spreads into standardized single pages using Python. You can install the library using pypdf on PyPI.

The following Python script uses pypdf to split a double-page scan directly down the middle, separating it into distinct left and right pages. It sets each copy’s MediaBox and CropBox using the pypdf PageObject (mediabox) and pypdf PdfWriter.add_page.

from pypdf import PdfReader, PdfWriter

reader = PdfReader("scanned_book.pdf")
writer = PdfWriter()

for page in reader.pages:
    left_edge = float(page.mediabox.left)
    bottom = float(page.mediabox.bottom)
    top = float(page.mediabox.top)
    midpoint = left_edge + float(page.mediabox.width) / 2

    left = writer.add_page(page)    # add_page returns the writer's copy
    left.mediabox.upper_right = (midpoint, top)
    left.cropbox = left.mediabox

    right = writer.add_page(page)   # a second, independent copy
    right.mediabox.lower_left = (midpoint, bottom)
    right.cropbox = right.mediabox

with open("standardized_book.pdf", "wb") as f:
    writer.write(f)

Trim borders with the PDF Cropper

To trim black scanner borders after splitting, use the PDF Cropper. The tool runs in your browser, so the file is not uploaded.

You can draw one rectangle from a page-1 preview and it is applied to all pages. The crop is shown in percent. Note that it cannot split spreads or set A4/Letter sizes, and it may misplace the crop on rotated pages.

4-Point Pre-Flight Verification

Before permanently archiving your newly cropped PDF, run through this verification checklist:

  1. Check the Logical Pagination: Did splitting the double-page spreads align correctly with the book’s logical page numbers (e.g., odd pages on the right)?
  2. Verify Page Sizes: Open the document properties in your PDF viewer and confirm each page is half the original spread width.
  3. Confirm File Size Has Not Bloated: Poorly optimized tools can accidentally duplicate image streams when saving cropped pages. The final file size should remain similar to the original scan.
  4. Expect Metadata Changes: Saving with the browser cropper changes the Producer field and the modification date.

For more techniques, see our guide on cropping PDF margins for e-readers.