Tips & Tricks

How to Translate Text Inside Embedded Images in a PDF

A PDF translation tool processes the text stream in the document, converting words from one language to another. It reads the paragraphs, the headings, and the captions. But it does not read the text that is inside images. A photograph of a sign, a screenshot of a user interface in a different language, a chart with axis labels rendered as part of the image, all of these contain text that standard translation tools ignore because the text is not in the PDF text stream. It is embedded in the image pixels. Translating text inside embedded images requires extracting the images, running OCR on them, translating the recognized text, and then either re-embedding the translated images or adding the translation as an overlay. The Translate PDF workflow for image-embedded text is more involved than standard text translation, but it covers content that would otherwise remain untranslated.

How to Translate Text Inside Embedded Images in a PDF

Identifying Images That Contain Translatable Text

Not every image in a PDF contains text worth translating. A decorative photograph, a background texture, and a company logo may contain no translatable content. A screenshot of a software application, a diagram with labeled components, a photograph of a document or a sign, and a chart with embedded axis labels all contain text that a reader in the target language needs to understand. Scan through the PDF and identify every image that contains text. Flag these images for the translation workflow.

Images that contain text in the PDF fall into two categories: those where the text is part of the image by design, such as a screenshot or a labeled diagram, and those where the text appears in the image because the document was scanned rather than digitally created. A scanned PDF is one large image per page. All the text on a scanned page is embedded in the page image. A digitally created PDF may have a single screenshot on an otherwise text-based page. The PDF Images that need translation are the ones where text appears in the image pixels and not in the PDF text stream.

WukongPDF

Try Translate PDF

No installation needed. Works directly in your browser.

Get Started โ†’

Extracting Images and Recognizing the Text

Extract the images from the PDF at the highest resolution available. Use a PDF image extraction tool that preserves original resolution. Save each extracted image as a PNG file to avoid introducing compression artifacts during the extraction step. Open each extracted image and confirm that the text is legible. If the image resolution is too low for the text to be read clearly, re-extract at a higher resolution if possible, or note that this image may produce unreliable OCR results.

Run OCR on each extracted image to recognize the text. The OCR engine identifies the text regions within the image and outputs the recognized characters along with their positions. Save the OCR output for each image, including the recognized text and the bounding box coordinates of each text region. This information is needed later to place the translated text in the correct position on the image. WukongPDF handles OCR PDF processing that can recognize text in embedded images as part of the document conversion workflow.

Translating the Recognized Text While Preserving Context

The text recognized from images is often fragmentary. A diagram may contain single-word labels. A screenshot may contain menu items and button labels without sentence context. This fragmented text is challenging for machine translation engines because the surrounding linguistic context that disambiguates word meanings is absent. Provide as much context as possible. If a label reads File and appears on a screenshot of a software menu, note in the translation input that this is a menu item, not a physical object.

For images with multiple text regions, maintain the mapping between each region and its translation. A diagram with ten labels produces ten separate text strings, each needing a translation. Keep the original text, the translated text, and the bounding box coordinates together in a structured format such as a spreadsheet. This mapping is used in the final step to place the translated text onto the image. For the Translate PDF output to be useful, the translated text must appear in the correct location on the image, replacing or accompanying the original text.

Placing Translated Text Back Onto the Images

Two approaches exist for placing translated text onto images. The first approach replaces the original text with the translation. Open the image in an image editor. For each text region, erase the original text by filling the region with the background color. Place the translated text in the same region using a font that matches the style of the original as closely as possible. This approach produces a clean translated image that looks like the original, but in a different language.

The second approach adds the translation alongside the original text. Place the translated text in a contrasting color below or beside the original text. This approach preserves the original for readers who understand both languages and want to compare the translation against the source. It is faster than erasing and replacing, but the resulting image is visually busier. Choose the approach based on the audience needs. A training manual for operators who only read the target language benefits from replacement. A bilingual reference document benefits from side-by-side presentation.

Re-embedding Translated Images Into the PDF

After the images have been processed with translated text, they must be placed back into the PDF. Replace each original image with its translated counterpart. The replacement image must fit in the same space on the page as the original. If the translated image has different pixel dimensions than the original, scale it to match the original bounding box in the PDF. The PDF page layout should not change. The only visible difference should be the language of the text within the images.

Save the PDF with the translated images as a new version, keeping the original untranslated PDF as a reference. Verify the translated PDF by opening it and checking each image. Confirm that the translated text is legible, correctly positioned, and free of OCR or translation errors. A Translate PDF with embedded image translation is complete when every text element in the document, in the text stream and in the images, is accessible to a reader in the target language.

When to Choose Manual Translation Over Automated

For images containing critical text, such as safety warnings, legal notices, or technical specifications, consider manual translation by a human translator instead of machine translation. An OCR error combined with a machine translation error can change the meaning of a safety instruction in ways that have real consequences. The cost of human translation for a small number of critical images is small compared to the cost of an incorrect translation.

For large batches of images where manual translation is not practical, implement a review step. After machine translation and re-embedding, have a human reviewer who reads the target language check a sample of the translated images for accuracy. If the sample reveals systemic errors, adjust the OCR settings or the translation engine parameters and reprocess the batch.

The translation of text in embedded images is often the last step in making a PDF fully accessible to an international audience. The body text, headings, and captions have been translated. The images remain as the final untranslated content. Completing this step produces a document where every piece of text, in every medium, is available in the target language.

For frequently updated documents with embedded images, consider whether the image text can be moved into the PDF text stream during the document creation process. A diagram with labels stored as text overlays rather than as part of the image can be translated using standard text translation tools, avoiding the extract-OCR-translate-reembed workflow entirely.

The quality of the final translated image depends on the quality of each step in the pipeline. A high-resolution image extraction produces clear OCR input. Accurate OCR produces correct text for translation. A good translation engine preserves meaning. Careful re-embedding maintains the visual quality. Each step depends on the previous one.

The Translate PDF workflow for image-embedded text is the most labor-intensive translation scenario because it combines image processing, OCR, translation, and re-embedding into a single pipeline. Automating as many steps as possible reduces the manual effort.

For images that appear identically in multiple PDFs, such as a company logo with a tagline or a standard diagram, translate the image once and reuse it. The translated image can be substituted into every PDF where the original image appears.

This article covers the key techniques and considerations for working with PDFs in this specific scenario. The methods described here apply across different tools and platforms, giving readers the flexibility to choose the approach that best fits their workflow and technical environment.

The practical steps outlined provide a clear path from the initial challenge to a working solution. Each technique has been selected for its reliability and accessibility, ensuring that readers can achieve consistent results regardless of their prior experience with PDF tools.

WukongPDF and similar platforms provide the OCR and document processing tools that support the image text translation workflow described in this article. The combination of these tools enables comprehensive translation coverage.

WukongPDF

Try Translate PDF

No installation needed. Works directly in your browser.

Get Started โ†’