本文へスキップ
OrbitexPDF
Two stacked sheets of paper under a large magnifying glass, standing for a scanner looking at a page as a picture
PDF を圧縮

Why scans are large

Why Scanned PDFs Are So Large, and How to Shrink Them

Ten typed pages are 200 KB. Ten scanned pages are 40 MB, and they say the same thing. The difference is that a scan is a photograph of a page rather than the words on it. Here is the arithmetic, and what to do about it.

Ten pages come out of the scanner as a 40 MB PDF. The same ten pages typed in Word are 200 KB. The words are identical. If your scanned PDF is too large, the reason is not the page count: a scanner does not read your page, it photographs it. Every page becomes one image, and an image's size is decided by two numbers you chose before pressing the button, the resolution and the colour mode.

Both can be changed. Scanning the next batch at 200 to 300 dpi in greyscale rather than 600 dpi in colour takes that 40 MB file under 5 MB with nothing lost an eye can find. For the scan already in your downloads folder, redrawing those images at a lower resolution is what a compressor does, and a scan is the one file where compression makes a dramatic difference.

Below: what is inside a scanned PDF, the arithmetic that turns dots per inch into megabytes, and what each level of Compress PDF does to a scan. Also what we cannot do for you.

Why a scanned PDF is too large: it is a picture of the page

Open a typed PDF and drag your cursor across a line. The text highlights, because the file holds characters: glyph codes with coordinates, plus a font embedded once and reused on every page. The two-hundredth page costs the same handful of kilobytes as the second.

Do the same on a scan and nothing highlights, because there is no text to highlight. The page contains one image, the width and height of your sheet of paper, and a one-line instruction to draw it across the whole page. The letters you can see are dark pixels in a grid of light ones, and the file has no idea they are letters.

That is the whole explanation for scanned document file size. A typed PDF's weight follows its content; a scan's follows its pixel count, and the pixel count follows the scanner. A blank page at 600 dpi in colour is heavier than a dense page of text at 200 dpi in greyscale, which feels wrong until you remember the file is not storing words.

Two panels comparing a scanned page made of one picture per page against a typed page made of characters and coordinates
The same ten pages, two hundred times apart in size. One stores what the page says, the other what it looks like.

The arithmetic: what dots per inch actually costs

Dots per inch is how many samples the scanner takes along each inch of the page, in both directions. An A4 sheet is 8.27 by 11.69 inches. At 300 dpi that is 2,480 by 3,508 samples, or 8.7 million pixels. In colour each pixel is three numbers, a red, a green and a blue, so the page is roughly 26 million values before anything compresses it: about 26 MB of raw data for one sheet of A4.

Now raise it to 600 dpi. Because dpi applies to both directions, doubling it quadruples the pixel count: 4,961 by 7,016 is 34.8 million pixels, over 100 million values in colour. Four times the data for the same sheet. That is why 600 dpi receipts exist: the setting sounds like twice as much and is four times as much.

The arithmetic runs the useful way too. Going from 300 dpi to 150 quarters the pixels; from 600 to 150 leaves a sixteenth. The images are compressed afterwards, so the file does not shrink by exactly sixteen times, but the relationship holds: the pixel count is the budget, and dpi spends it in squares.

Three bars drawn to scale showing one A4 page holding 2.2 million pixels at 150 dpi, 8.7 million at 300 dpi and 34.8 million at 600 dpi
Drawn to scale. The 150 dpi bar is not a rounding error against the 600 dpi one; it is one sixteenth of it.

Colour mode is the other half of the multiplication

Every pixel needs a value, and how many numbers that takes depends on what the scanner was told to record. Full colour is three per pixel. Greyscale is one, so the same page at the same dpi is a third of the data before compression starts. Black and white, sometimes called bilevel or line art, is a single bit: ink or paper, nothing between.

For black text on white paper, greyscale loses nothing you want. The colour channels are recording that your paper is slightly cream and your toner slightly warm, which nobody needs from a contract. What greyscale keeps is the soft edge of each letter, which makes text look like text.

Bilevel goes further and is excellent for clean typed originals: at 300 dpi a page of text can land under 50 KB, because the codecs built for that mode compress runs of identical pixels enormously well. It is the wrong choice for a photograph, a highlighter mark, a pencil note or a faint stamp, each of which becomes solid black or nothing at all.

What one page costs at each setting

One A4 sheet of typed black text, saved as JPEG inside the PDF. Real numbers move with the paper and the scanner; the ratios hold.

A photograph on the page pushes every row up. Bilevel at 300 dpi drops a text page under 50 KB.
Scanner settingPixels on the pageTypical pageTen pages
150 dpi, greyscale2.2 millionabout 120 KBabout 1.2 MB
200 dpi, greyscale3.9 millionabout 200 KBabout 2 MB
300 dpi, greyscale8.7 millionabout 500 KBabout 5 MB
300 dpi, colour8.7 million, three channelsabout 1.5 MBabout 15 MB
600 dpi, colour34.8 million, three channelsabout 4 MBabout 40 MB

JPEG and lossless inside the same scan

The PDF is only a wrapper. The image inside each page is encoded in some format, and which one the scanner chose changes the size by a factor of five or more for identical pixels.

Most scanners write JPEG, which is lossy: it discards detail the eye is poor at noticing. On a photograph that is close to invisible. On text it shows as faint haloes around the letters, because sharp black-to-white edges are what JPEG handles worst.

Some scanners, and most phone scanning apps set to their best mode, write lossless images instead. It sounds like the safe choice and it is how a ten-page scan reaches 40 MB. Lossless stores every pixel exactly, including the sensor noise that made your white paper very slightly speckled, and noise is the one thing lossless compression can do nothing with. A lossless colour scan of a blank sheet can be several megabytes of nothing. For text, greyscale at 300 dpi with ordinary JPEG quality is the better trade.

How to reduce scanned PDF size in four steps

For the file you already have, where the scanning decisions are behind you.

  1. ステップ 1: Open the compressor

    Go to Compress PDF. No install, no account, and nothing to configure first.

  2. ステップ 2: Add the scan

    Drag the PDF in. Up to 20 MB a file on the free plan, 250 MB on Pro, which covers most scanned batches.

  3. ステップ 3: Choose a level

    Balanced redraws images at 150 dpi and is right for almost every text scan. Smaller goes to 72 dpi for email. Higher quality holds 300 dpi for anything that will be printed.

  4. ステップ 4: Open the result before you send it

    Find the smallest print on the busiest page. If it is readable, the file is fine. If not, run the original again one level up.

Three numbered panels showing how to shrink an existing scan: divide size by pages, start at the Balanced level, then open the result and check the smallest print
The check at the end is the part people skip, and the only one that tells you whether the level was right.

What each Compress level does to a scan

The three levels are three target resolutions: Smaller redraws every page image at 72 dpi, Balanced at 150 dpi, Higher quality at 300 dpi. Text and vector drawing are never touched, but a scan has none of either, so here the level decides everything.

Run a 600 dpi colour scan through Balanced and each page goes from 34.8 million pixels to 2.2 million: a sixteenth of the pixels, and in practice a tenth to a twentieth of the file. The 40 MB batch comes back at two or three megabytes and still reads perfectly, because a screen was never going to show you 600 dpi anyway.

Smaller, at 72 dpi, is a real loss of legibility for scanned text. Most of a page survives; footnotes, small print and anything handwritten in biro do not. Use it when a hard upload limit is the constraint, and read the guide to compressing without losing quality if you want the levels compared one by one.

Higher quality is for a scan that will be printed: 300 dpi is the point below which printed text starts to look soft. It still helps on a lossless scan, because the re-encode alone takes a large bite out of it.

How to scan properly the first time

Every megabyte you avoid creating is one you never have to compress away. These are the defaults worth setting once on the office scanner.

  • 200 to 300 dpi for anything made of text. 300 dpi is the printing standard, 200 dpi is fine for reading, and almost nothing you own needs 600 dpi. A receipt certainly does not.
  • Greyscale unless colour carries meaning. A signature in blue ink, a colour-coded chart or a stamp that must be seen as red are reasons for colour. A slightly beige page is not.
  • Black and white for clean typed originals only. Smallest mode by a distance, and it destroys pencil, highlighter, thermal receipts and anything photographic.
  • Turn off the 'best quality' or 'archival' preset unless you mean it. It usually means lossless colour at high dpi, and it is where the 40 MB files come from.
  • Check the first page before running eighty through the feeder.

When compressing is not the right answer

Sometimes a file is large because it holds more than anyone needs. A signed agreement with sixty pages of scanned appendices has a page count problem, not a compression problem, and pulling out the pages that matter with Split PDF takes a 40 MB scan to a 3 MB one without touching a pixel.

Scanned batches also arrive in the wrong shape: fed in backwards, every other page blank because the feeder was set to double-sided, a few sheets sideways. Organize PDF drops the blanks and fixes the rotation, and dropping the blank backs alone halves a double-sided scan. Do that first, then compress what is left.

And if a document is mostly typed with a couple of scanned pages stapled on, those pages are the entire weight. The general compression guide covers the mixed case.

What happens to the file you upload

Compression runs on a server, because redrawing every page image of a large scan is more work than a browser tab should do. The upload is streamed to disk, the images are redrawn, and the input is deleted the moment the job finishes. The result is available for thirty minutes and goes as soon as you have it, as the security page sets out.

Scans are often the most sensitive documents anybody owns: passports, bank statements, medical letters, signed contracts. Nothing is kept, no account is attached, no watermark is added.

Frequently asked questions

Why is my scan so big when the document is only ten pages?

Because page count does not decide the size of a scan. Each page is one photograph, and its weight comes from the resolution and colour mode. Ten pages at 600 dpi in colour is around 40 MB; the same ten at 200 dpi in greyscale is around 2 MB.

What scan dpi should I use for documents?

200 to 300 dpi for text. 300 dpi matches what a printer needs; 200 dpi is fine for reading on a screen. 600 dpi is for photographs, and quadruples the data over 300 dpi.

How do I compress a scanned PDF without making the text unreadable?

Use the Balanced level in Compress PDF, which redraws page images at 150 dpi and reads cleanly on any screen. Then open the result, find the smallest print on the busiest page and confirm it is legible. Drop to Smaller only if an upload limit forces you to.

Will compressing a scanned PDF make the text searchable?

No. A scan has no text in it to begin with - the letters are pixels in an image - and compression only changes how many pixels there are, so a compressed scan stays a picture. We do not add a searchable text layer to a PDF. If you need the words, PDF to Word and PDF to Excel can read a scan with OCR on Pro and give you an editable file instead.

Should I scan in greyscale or black and white?

Greyscale for almost everything: a third the size of colour, and it keeps the soft edges that make letters look like letters. Black and white is smaller still and excellent for clean typed pages, but destroys pencil, highlighter and anything photographic.

Why did my scan barely shrink when I compressed it?

Almost always because it was already low resolution: a scan made at 150 dpi has nothing to lose at the Balanced level, which targets 150 dpi. Check the megabytes per page. Under 100 KB a page, the size is coming from the page count instead.

In short

A scanned PDF is too large because it is a stack of photographs rather than a document of text. Its size comes from the resolution and the colour mode, and dpi spends data in squares, so 600 dpi costs four times what 300 dpi does for the same sheet. Scan text at 200 to 300 dpi in greyscale and the problem never appears. For the file you already have, Balanced redraws the page images at 150 dpi and usually leaves a heavy colour scan at a tenth of its size.

One honest limit: a scan compressed here is still a picture, and its text does not become searchable. To get the words out, PDF to Word and PDF to Excel offer OCR on Pro. Compression itself is free and needs no account - see every tool.

役に立ちましたか。共有してください。

コメント

コメントはアカウントをお持ちの方どなたでも書けます。無料で、この欄を迷惑投稿から守っている唯一の仕組みです。

ログインしてコメントする

まだ誰も書いていません。このガイドに足りないところがあれば、黙っているより教えていただけると助かります。

関連するガイド

すべての記事
A PDF document with a dashed measuring frame around a small solid block, standing for a file squeezed to fit a size limit
PDF を圧縮

PDF under 100 KB

How to Compress a PDF to 100 KB (or Any Size a Form Demands)

A visa portal wants 100 KB and your scan is 1.8 MB. Here is how to find out what is making the file heavy, which level to pick, and what to do when Smaller still is not small enough.

10 min readガイドを読む
A PDF document with a red warning triangle over it, marking a file that will not send as an email attachment
PDF を圧縮

Too big to email

Your PDF Is Too Big to Email: The Limits and What Actually Works

The send button failed, or the message bounced an hour later. Here are the real attachment limits at Gmail, Outlook and a typical company server, why your file gets a third bigger in transit, and the four ways to get the document there.

11 min readガイドを読む
A PDF page squeezed between two arrows beside a badge reading 78 percent smaller
PDF を圧縮

How to compress a PDF

How to Compress a PDF: Reduce File Size Without Losing Quality

A 12 MB PDF that will not send is almost always a few oversized images in a wrapper. Here is what compression actually changes, which level to pick, and how to get under a limit without turning your text to mush.

10 min readガイドを読む