Tips & Tricks

How to Translate Only Figure Captions in a PDF

Scientific papers, technical reports, and illustrated documents contain figure captions that explain what each image, chart, or diagram shows. When the body text of such a document is already in a language the reader understands, but the figure captions are not, translating only the captions avoids the unnecessary work of translating content the reader can already follow. The challenge is isolating the caption text from the surrounding body text so that only the captions pass through the translation engine.

Captions in PDFs are usually set in a slightly different font size or style than the body text. They sit below or beside images, often beginning with Figure or Table followed by a number. These visual and positional cues help a human reader identify captions at a glance. Translate PDF workflows that target only captions use these cues to select the right text, either through manual selection, style-based filtering, or export-and-extract approaches.

WukongPDF's PDF Export and translation tools handle document-wide language conversion. For selective caption-only translation, the following manual and semi-automated workflows isolate caption text from the surrounding content.

How to Translate Only Figure Captions in a PDF

Method 1: Manual Copy-Translate-Paste for Captions

For a document with a manageable number of captions, usually under 20, the most reliable method is manual. Open the PDF in any editor that supports text selection. For each caption, select the text, copy it, paste it into a translation tool such as Google Translate or DeepL, copy the translated output, and paste it back into the PDF at the same location, replacing the original caption text.

The manual method is precise. You control exactly which text gets translated and where the translation is placed. It avoids the risk of accidentally translating body text or missing captions that automated detection might skip. The precision comes at the cost of time. Each caption takes about 30 to 60 seconds to select, translate, and replace. A document with 15 captions takes roughly 10 to 15 minutes. For a one-time translation of a single document, the manual method is adequate. For recurring translations or documents with hundreds of captions, the automated methods below are more practical.

WukongPDF

Try Translate PDF

No installation needed. Works directly in your browser.

Get Started โ†’

Method 2: Style-Based Selection in a PDF Editor

Adobe Acrobat Pro's Edit PDF tool includes a feature that selects all text matching a specific font, size, or color. If the captions in the PDF use a consistent style, such as 9-point italic, you can select all caption text in one operation. Open the PDF in Acrobat Pro, go to Edit PDF, click in a caption text area, then right-click and choose Select All With This Font. Acrobat selects every text block on the page that uses the same font and size. Copy the selected text and paste it into a text document.

Review the pasted text. Body text that accidentally uses the same font as captions will have been selected alongside the actual captions. Remove any non-caption text manually. The text document now contains only caption content. Paste it into the translation tool, translate, and then paste each translated caption back into its original position in the PDF. The style-based selection speeds up the extraction step significantly compared to selecting each caption individually.

Method 3: Export to Word and Filter by Style

Convert the PDF to Word format using Acrobat Pro's Export to Word feature or an online PDF-to-Word converter. Open the resulting DOCX in Microsoft Word. Word preserves the font styling from the PDF, so captions appear in the same font and size as in the original. Use Word's Select All Text With Similar Formatting feature to select all caption text across the document. Right-click a caption paragraph, go to Styles, then Select All. Word selects every paragraph with matching formatting.

Copy the selected text into a table in a new document. Each caption occupies one row. Add a second column for the translated text. Translate each caption using the translation tool, pasting the translation into the adjacent column. After all captions are translated, use the two-column table as a reference to replace each caption in the PDF with its translation. The Word-based approach provides a structured workflow that is easier to track than the PDF-based methods because the translation progress is visible in the table.

MethodHow It WorksBest For
Selective copy-pasteCopy only caption text from PDF, translate, paste backShort documents with 5-10 captions
PDF editor find-replace by styleSelect all text matching caption font, copy to translatorDocuments with styled captions
Export captions to text fileExport PDF to Word, isolate captions by formattingDocuments with many captions, recurring workflow

Handling Captions That Span Multiple Lines

A multi-line caption in the PDF may be split into separate text blocks, one per line. When you select and copy the caption, you get three separate text selections rather than one continuous paragraph. Paste all three into the translation tool as a single block by joining them with a space. The translation will be a single paragraph. When pasting the translation back into the PDF, split it across the same number of text blocks as the original to maintain the layout.

When it comes to document workflows, for captions that contain technical terms, abbreviations, or measurement units, protect those terms from translation by flagging them as do-not-translate in the translation tool if the tool supports it, or by replacing them with placeholder tokens, running the translation, and restoring the originals. A caption that reads Figure 3: SEM image of TiO2 nanoparticles at 50kx magnification should preserve TiO2 and 50kx exactly as written.

Translating only figure captions leaves the body text in the original language while making the visual content understandable to readers of the target language. The selective approach saves time compared to full document translation and produces a document that bilingual readers, the most common audience for partially translated technical content, can navigate efficiently.

Batch-Translating Captions Across Multiple PDFs

For researchers or technical writers managing a library of documents with captions in a foreign language, batch translation across multiple PDFs follows the same workflow scaled up. Export all PDFs to Word format using a batch converter. Use Word's Find and Replace with formatting filters to locate all text matching a specific font and size, which should capture captions across all the exported documents. Copy all found text into a single translation batch.

Translate the entire batch in one operation through DeepL or Google Translate. Both services accept large text blocks and preserve paragraph breaks, which correspond to individual caption boundaries. After translation, split the translated text back into individual captions by paragraph break and paste each one into its corresponding PDF. The batch approach reduces the per-caption translation overhead from 30 seconds to roughly 5 seconds, making it practical for document collections with hundreds of captions.

Preserving Layout When Replacing Translated Captions

Translated captions are often longer or shorter than the originals. German translations run 20 to 30 percent longer than English. Chinese translations are often shorter. Adjust the font size of the translated caption by half a point to absorb the length change without altering text box dimensions. Small font adjustments preserve layout while accommodating translation length differences.

Across document workflows, for significant length changes, split the translated caption across multiple lines manually. Match line breaks approximately to the visual width of the original caption. The goal is for the translated caption to occupy roughly the same visual space as the original, preserving the relationship between the caption and the figure it describes.

Selective translation of figure captions is a precision task that respects the reader existing language skills. Translating only what needs translation saves time and produces a document that bilingual readers navigate comfortably. The manual, style-based, and export-based methods scale from a single document with five captions to a document library with hundreds.

Selective translation of captions is a precision task that produces a document usable by bilingual readers who can follow the body text but need captions in their preferred language. The workflow scales from a single document with five captions to a library with hundreds, and the manual, style-based, and export methods each serve different scales.

Selective caption translation respects the reader existing language skills while making visual content accessible. The precision of translating only what needs translation, rather than the entire document, saves time and produces a result that bilingual readers navigate comfortably. The manual, style-based, and export-based methods scale from single documents to entire libraries.

WukongPDF

Try Translate PDF

No installation needed. Works directly in your browser.

Get Started โ†’