Many outputs are delivered as PDF documents: tables, listings and figures in clinical reporting, Quarto or R Markdown documents rendered to PDF, or reports produced by scheduled jobs. Checking that such outputs did not change is often done by eye, which is slow and easy to get wrong when there are hundreds of pages.
Comparing the underlying data catches many problems, but not all of them: a changed font, a shifted legend, a different default theme after a package upgrade, or a truncated table only show up in the rendered output.
compare_pdfs() renders each page of two PDF files to an
image and compares them page by page with odiff. You get a result per
page, with odiff’s tolerance (threshold) and antialiasing
handling, optional diff images highlighting the changed pixels, and the
same summaries and reports as for image comparisons.
Typical uses:
PDF rendering uses the pdftools package, which needs to be installed:
library(odiffr)
res <- compare_pdfs(
baseline = "outputs-before/f_km_os.pdf",
current = "outputs-after/f_km_os.pdf",
diff_dir = "pdf-diffs"
)
resThe result is an odiffr_batch with one row per page and
a page column. img1 and img2 are
the rendered page images; diff_output holds the diff image
for pages that differ. When diff_dir is given, the rendered
pages are kept in pdf-diffs/pages/; otherwise they are
written to the session’s temporary directory and removed when the R
session ends.
Use the usual tools to inspect the results:
summary(res)
# Which pages differ?
failed_pairs(res)[, c("page", "reason", "diff_percentage", "diff_output")]To compare only some pages, use pages:
If the two files have a different number of pages,
compare_pdfs() emits a message and reports the pages
present in only one file as failures, using the existing reasons so that
all reports work unchanged:
baseline but not in current has
reason = "missing";current but not in baseline has
reason = "error", with the error
"Page N not present in baseline PDF".Pages of different sizes (for example portrait versus landscape) are
compared as images of different dimensions and reported as different.
Pass fail_on_layout = TRUE to report them with
reason = "layout-diff" instead:
compare_pdf_dirs() compares every PDF file in a baseline
directory with the file of the same name in a current directory and
returns one combined result with a file column. This fits
the “re-run all outputs after an upgrade” workflow:
res <- compare_pdf_dirs(
"outputs-r4.3/",
"outputs-r4.4/",
recursive = TRUE,
diff_dir = "pdf-diffs"
)
summary(res)
# Files with at least one differing page
unique(res$file[!res$match])A file missing from the current directory is reported as a single
"missing" row, and a file that cannot be read (for example
a corrupt file) as a single "error" row, so one bad file
does not stop the whole comparison.
The results can be passed to the reporting functions. An HTML report gives reviewers a quick overview of which pages changed:
For continuous integration, write a JUnit XML file so that each page shows up as a test case:
Outputs often contain content that changes on every run, such as a
run-date footer, a program path or a page header with a time stamp.
Exclude such areas with ignore_regions. Coordinates are in
pixels of the rendered page, so they depend on
dpi: a position of x inches from the left edge
is at pixel x * dpi.
For a US Letter page in portrait orientation (8.5 x 11 inches) rendered at 150 dpi, the page is 1275 x 1650 pixels. To ignore the bottom half inch (the footer):
dpi <- 150
footer <- ignore_region(
x1 = 0,
y1 = (11 - 0.5) * dpi,
x2 = 8.5 * dpi,
y2 = 11 * dpi
)
compare_pdfs("before.pdf", "after.pdf", dpi = dpi, ignore_regions = footer)The same regions are applied to every page. If you need to find the
right coordinates, open one of the rendered pages (in img1)
in an image viewer that shows pixel positions.
dpi controls how finely pages are rendered:
Rendering is deterministic for a given file and poppler version, so
identical PDFs give identical images. If you see small differences
caused by antialiasing of text or lines, try
antialiasing = TRUE or a slightly higher
threshold.
Only PDF files are supported. Outputs in other formats such as RTF or DOCX need to be converted to PDF first, for example with LibreOffice:
Use the same converter (and version) for the baseline and the current outputs, so that differences come from the outputs and not from the conversion.
A visual comparison shows whether the rendered pages differ, and where. It does not replace checking the content of the outputs, and it does not by itself make a process compliant with any regulation: it is a tool to make reviews faster and more systematic.