Skip to main content
Content

How to Make PDFs Readable for Chatbots and Search

Scanned PDFs look fine to people and are blank to software. How to check for a text layer, fix it with OCR, and publish documents a chatbot and search can use.

A laptop beside printed documents, with a document icon

Many websites keep their most useful information in PDFs: price lists, menus, agendas, handbooks, application forms. To a person they look fine. To a chatbot or a search engine, a surprising number are blank, because they’re pictures of text rather than text.

This guide explains the difference, how to check your documents in a minute, and how to fix the ones that need it.

Text PDFs and scanned PDFs

A PDF can hold text in two ways.

  • Real text. Documents saved from a word processor or exported as PDF contain the actual characters. Software can read, search and quote them.
  • A picture of text. A document that was printed and scanned, or photographed, is an image. It looks the same, but there are no characters inside, only pixels.

A scanned PDF can be made readable with a text layer: an invisible layer of recognized text laid over the image. Many scanners and PDF editors add one automatically using optical character recognition (OCR); older scans often don’t have one.

Why it matters

  • Chatbots can only answer from text they can read. A price list without a text layer is invisible to them.
  • Search engines can index the text in PDFs, which helps people find your documents. No text, nothing to index.
  • Accessibility. Screen readers need real text too. A scan without a text layer is unusable for visitors who rely on them.
  • Your visitors. Real text can be searched with Ctrl+F or Cmd+F and copied. A picture can’t.

How to check a PDF in one minute

  1. Open the PDF in any viewer.
  2. Try to select a sentence with your mouse. If the text highlights word by word, it has text. If you can only draw a box, or nothing happens, it’s a picture.
  3. Use the viewer’s search for a word you can see on the page. If it finds nothing, the text isn’t there.
  4. Copy a paragraph and paste it into a text editor. Garbled characters mean the text layer is poor quality.

If you use GoChatterBot, the plugin does this count for you. Under Settings → Content → Documents it shows how many of your PDFs have text the chatbot can read. See Documents and PDFs.

How to fix a scanned PDF

Best: go back to the original

If the document was written in a word processor, export it again as a PDF from the original file. You get clean text, smaller files and better accessibility than any scan.

Next best: add a text layer with OCR

If only a paper copy or a scan exists, run it through OCR. Most PDF editors have a “recognize text” or “make searchable” option, and many scanner apps can add a text layer when scanning. After OCR:

  • Check a few paragraphs, numbers and dates. OCR can misread characters, especially in tables, small print or poor scans.
  • Rescan badly skewed or low-resolution pages before running OCR again; clean input gives far better results.
  • Upload the new file. If you replace a document, remove the old version so there aren’t two copies with different content.

Also good: publish the key facts as a web page

For information people ask about often, such as opening hours, fees or a short menu, a web page is better than a PDF. It’s faster to read on a phone, easier for everyone to find, and simple to keep up to date. Link to the PDF from the page for anyone who wants the full document.

Make documents easier to find and cite

  • Give each PDF a clear title in the Media Library, such as “Fee schedule 2026” rather than “scan_0034”. Chatbot answers and search results show the title.
  • Use descriptive file names before you upload.
  • Put the date or year in the title of documents that change, so it’s obvious which is current.
  • Remove out-of-date versions. Old documents with old prices lead to wrong answers.
  • Use real headings inside the document. Structured documents are easier for people and software to follow.

Turning on PDF reading in GoChatterBot

In Settings → Content, under Documents, tick:

  1. Read documents in the Media Library
  2. Documents only, not photos and other images
  3. Read the text inside PDFs

New PDFs are read automatically when you upload them. For a large existing library, Read PDFs now reads a few straight away and the rest in the background. Turn on Let searches match the fields below and PDF text so a question can find a PDF by what it says, not only by its title.

The plugin reads PDFs using a common server tool called Ghostscript, and the settings page tells you if your host doesn’t have it. If another plugin already stores each PDF’s text in a custom field, you can point the chatbot at that field under PDF text from instead.

When an answer comes from a PDF, the chatbot links to the document, so visitors can open the full file. More on the Document & PDF Search feature page.

A quick audit you can do this week

PDF audit

  • List the ten PDFs visitors ask about most.
  • Run the one-minute check on each.
  • Re-export or OCR the ones without text.
  • Rename any with unclear titles, and add the year where it matters.
  • Delete or replace out-of-date versions.
  • Move the most-asked facts onto a web page as well.

Pair this with our guide to writing content an AI chatbot can use and you’ll cover both halves of your site: the pages and the documents.