--- title: "Comparing PDF outputs" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Comparing PDF outputs} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", eval = FALSE ) ``` ## Why compare PDFs visually? Many outputs are delivered as PDF documents: tables, listings and figures in clinical reporting, Quarto or R Markdown documents rendered to PDF, or reports produced by scheduled jobs. Checking that such outputs did not change is often done by eye, which is slow and easy to get wrong when there are hundreds of pages. Comparing the underlying data catches many problems, but not all of them: a changed font, a shifted legend, a different default theme after a package upgrade, or a truncated table only show up in the rendered output. `compare_pdfs()` renders each page of two PDF files to an image and compares them page by page with odiff. You get a result per page, with odiff's tolerance (`threshold`) and antialiasing handling, optional diff images highlighting the changed pixels, and the same summaries and reports as for image comparisons. Typical uses: - **Production versus QC outputs**: compare the figure produced by the production program with the one produced by an independent QC program. - **Re-running outputs after an upgrade**: re-run all outputs after upgrading R or packages and compare them with the outputs from before the upgrade. - **Rendered documents**: compare a Quarto or R Markdown document rendered to PDF before and after a change. PDF rendering uses the [pdftools](https://docs.ropensci.org/pdftools/) package, which needs to be installed: ```{r} install.packages("pdftools") ``` ## Comparing two PDF files ```{r} library(odiffr) res <- compare_pdfs( baseline = "outputs-before/f_km_os.pdf", current = "outputs-after/f_km_os.pdf", diff_dir = "pdf-diffs" ) res ``` The result is an `odiffr_batch` with one row per page and a `page` column. `img1` and `img2` are the rendered page images; `diff_output` holds the diff image for pages that differ. When `diff_dir` is given, the rendered pages are kept in `pdf-diffs/pages/`; otherwise they are written to the session's temporary directory and removed when the R session ends. Use the usual tools to inspect the results: ```{r} summary(res) # Which pages differ? failed_pairs(res)[, c("page", "reason", "diff_percentage", "diff_output")] ``` To compare only some pages, use `pages`: ```{r} compare_pdfs("before.pdf", "after.pdf", pages = c(1, 5)) ``` ### Different numbers of pages If the two files have a different number of pages, `compare_pdfs()` emits a message and reports the pages present in only one file as failures, using the existing reasons so that all reports work unchanged: - a page in `baseline` but not in `current` has `reason = "missing"`; - a page in `current` but not in `baseline` has `reason = "error"`, with the error `"Page N not present in baseline PDF"`. ### Different page sizes Pages of different sizes (for example portrait versus landscape) are compared as images of different dimensions and reported as different. Pass `fail_on_layout = TRUE` to report them with `reason = "layout-diff"` instead: ```{r} compare_pdfs("before.pdf", "after.pdf", fail_on_layout = TRUE) ``` ## Comparing whole directories `compare_pdf_dirs()` compares every PDF file in a baseline directory with the file of the same name in a current directory and returns one combined result with a `file` column. This fits the "re-run all outputs after an upgrade" workflow: ```{r} res <- compare_pdf_dirs( "outputs-r4.3/", "outputs-r4.4/", recursive = TRUE, diff_dir = "pdf-diffs" ) summary(res) # Files with at least one differing page unique(res$file[!res$match]) ``` A file missing from the current directory is reported as a single `"missing"` row, and a file that cannot be read (for example a corrupt file) as a single `"error"` row, so one bad file does not stop the whole comparison. ## Reports and CI The results can be passed to the reporting functions. An HTML report gives reviewers a quick overview of which pages changed: ```{r} batch_report( res, output_file = "pdf-diffs/report.html", images = "all", show_all = TRUE ) ``` For continuous integration, write a JUnit XML file so that each page shows up as a test case: ```{r} batch_junit(res, "pdf-diffs/junit.xml") ``` ## Ignoring dynamic content Outputs often contain content that changes on every run, such as a run-date footer, a program path or a page header with a time stamp. Exclude such areas with `ignore_regions`. Coordinates are in **pixels of the rendered page**, so they depend on `dpi`: a position of `x` inches from the left edge is at pixel `x * dpi`. For a US Letter page in portrait orientation (8.5 x 11 inches) rendered at 150 dpi, the page is 1275 x 1650 pixels. To ignore the bottom half inch (the footer): ```{r} dpi <- 150 footer <- ignore_region( x1 = 0, y1 = (11 - 0.5) * dpi, x2 = 8.5 * dpi, y2 = 11 * dpi ) compare_pdfs("before.pdf", "after.pdf", dpi = dpi, ignore_regions = footer) ``` The same regions are applied to every page. If you need to find the right coordinates, open one of the rendered pages (in `img1`) in an image viewer that shows pixel positions. ## Choosing the resolution `dpi` controls how finely pages are rendered: - **72-100 dpi** is fast and catches layout changes, missing elements and changed colours. - **150 dpi** (the default) is a good compromise for tables and figures. - **300 dpi** detects small changes such as a single changed digit in a small font, at the cost of more time and disk space. Rendering is deterministic for a given file and poppler version, so identical PDFs give identical images. If you see small differences caused by antialiasing of text or lines, try `antialiasing = TRUE` or a slightly higher `threshold`. ## RTF and DOCX outputs Only PDF files are supported. Outputs in other formats such as RTF or DOCX need to be converted to PDF first, for example with LibreOffice: ```bash soffice --headless --convert-to pdf --outdir outputs-pdf outputs/*.rtf ``` Use the same converter (and version) for the baseline and the current outputs, so that differences come from the outputs and not from the conversion. ## Limitations A visual comparison shows *whether* the rendered pages differ, and where. It does not replace checking the content of the outputs, and it does not by itself make a process compliant with any regulation: it is a tool to make reviews faster and more systematic.