Text versus pictures of text
When a program like Word saves a PDF, each letter is stored as a character code plus a reference to a font. The letter "a" costs a byte or two, however large it is printed. Lines and tables are stored as drawing instructions. That is why text PDFs are small and stay sharp at any zoom.
A scanner does not know there are letters on the page. It measures the colour of millions of tiny points and stores the result as an image. The PDF it produces is really a container holding one photograph per page. Every point of white paper costs as much to store as a point of ink.
Where the size comes from
Three settings multiply together to decide how big a scanned page is.
Resolution. Measured in dots per inch. An A4 page at 300 dpi is about 2480 × 3508 pixels — 8.7 million pixels. At 600 dpi it is four times as many. Many scanners and office copiers default to 300 dpi, and some "high quality" presets go to 600 dpi, which is far more than any screen needs.
Colour depth. Colour images store three values per pixel (red, green, blue). Greyscale stores one. Pure black-and-white stores a single bit. A colour scan of a black-and-white letter records the paper's faint tint and shadows in full colour — data that nobody wants.
Compression. Scanners usually save pages as JPG inside the PDF. Some use very light compression, or even lossless formats, to be safe. The difference between a light and a moderate JPG setting can be three or four times the size, with little visible difference for text.
Multiply these together and a single colour page at 300 dpi with light compression lands between 1 and 3 MB. Ten pages is 10 to 30 MB — over most email limits and far over most upload forms.
Phone scans have their own issues
Phone cameras capture 12 to 48 megapixels per shot, more than a 300 dpi scan of an A4 page, and they save photos in colour at high quality. Scanning apps often keep that full resolution. They also capture background — the table, your hand, the shadow of the phone — which adds noise that compresses poorly. Cropping to the page edge and choosing a "document" or black-and-white mode in the app helps a lot.
Fix it at the source if you can
The best compression is not creating the data in the first place. If you are about to scan something, check the settings:
- Resolution: 150 dpi is enough for documents that will be read on screen. Use 200–300 dpi only for small print, or if the file will be OCR'd, where 300 dpi gives the best recognition.
- Colour mode: greyscale for ordinary documents; black-and-white for plain typed text; colour only for photos, coloured stamps or where the receiver requires colour.
- File format: PDF, with "compact" or "small size" selected if the scanner offers it.
With 150 dpi greyscale, a typical text page comes out at around 100–250 KB instead of 1–3 MB.
Fix it after the fact
When you already have a large scan, re-encoding the pages is the answer. There are three routes on PDF300, and they suit different needs:
- Compress (browser) redraws every page as a JPG at a width and quality you choose. It runs on your computer, so the scan is never uploaded. Because a scan is already a picture, you lose nothing by redrawing it as a smaller picture.
- Compress to target size does the same kind of re-encoding on our server, but searches for the best resolution that fits a size you type. Use it when a form has a hard limit.
- High-quality compress uses Ghostscript on the server to resample images to about 150 dpi while leaving any text and vector content alone. It is the better choice for mixed documents — a text report with some scanned pages inside.
For scans, the first two routes usually shrink files by 70 to 95 percent. Start gently, check readability, and go further only if you need to. The guide How to reduce a PDF to under 300 KB walks through a strict limit step by step.
What about making the scan searchable?
Because a scan is a picture, you cannot search it or copy text out of it. OCR recognises the letters and can add an invisible text layer behind each page, so the scan looks the same but becomes searchable. The added text layer is small — usually a few kilobytes per page — so OCR does not make a scan meaningfully larger. Do OCR before heavy compression: recognition accuracy drops as resolution drops, and 300 dpi input gives the best results.
Things that do not help much
- Zipping the PDF. The images inside are already compressed; a ZIP typically saves only a few percent.
- "Save as reduced size" on a text PDF. If the file is large because of images, removing fonts or metadata saves little.
- Converting to PDF/A. Archival formats may embed more, not less.
In short
A scanned PDF is large because it stores every page as a photograph, usually in colour and usually at a higher resolution than the reader needs. Scan at 150 dpi greyscale when you can. When you cannot, re-encode the pages at a lower resolution — in your browser if privacy matters, or with a target size when a form sets a hard limit.