A PDF report sounds like a small feature. Render some HTML, hand it to a browser, save the result.
I’ve spent years working on reporting in Spartacus Insights, the cybersecurity assessment platform we build at Digital Marmalade. The reports combine assessment results, graphics from the analysis screens, recommendations, tables and content written by users. Some of them run past 100 pages.
Getting a PDF out of the application was the first step. Keeping a long table readable across several pages, keeping a chart legible on A4, and letting someone find a section halfway through the document took much longer, and most of those problems didn’t appear until real content went in.
The content is the difficult part
With a fixed document, I control the length of every paragraph and the size of every table. I can adjust the template until the sample looks right.
An assessment report doesn’t give me that. One user writes a short executive summary, another writes several pages. One assessment has a handful of recommendations, the next has a long table. A section that fits on one page in a demo can spread across several in a real export.
Nobody should have to tidy every export by hand, so the template has to cope on its own. I have to decide what can split, what should stay together, how a heading behaves at the top of a page, and whether a table still makes sense when it’s cut in half.
What the browser gives us
Our stack is Laravel and Blade, JavaScript, Paged.js and Browsershot. A lot of the reporting work has gone into the code around them rather than the tools themselves.
We write the report in HTML and CSS, and the preview and the export use the same templates. Paged.js handles pagination and the visible contents, figures and tables lists. Browsershot produces the PDF.
The browser still has to be told what kind of document it’s dealing with. An A4 page is 210 by 297 mm. With our 20 mm margins, the content area is 170 by 257 mm, and every bit of text and every graphic has to fit inside it. A layout that looks comfortable in a wide application panel tells me very little about how it will sit on paper.
Within that space, a heading should normally stay with some of the content it introduces. A long table needs its headers repeated on the following pages. A chart needs enough room for its labels and caption to stay readable.
Keeping a block together is a sensible default until the block is taller than the printable page, and then the rule can’t be satisfied. In a long document I have to tell apart a small unit that should stay together and a section that has to be allowed to continue.
Where Laravel comes in
Generating a large report is real work for the application. It pulls together assessment data, user content and graphics from the analysis pages. Some of those graphics already exist; others need preparing.
We use Laravel queues and job batching for this. Graphics can be prepared while users are still editing the report, and the generated output is cached and reused, so an export doesn’t have to rebuild every graphic from scratch. The export still has to coordinate graphics preparation with report generation, and job batching lets us run that dependent work without the user holding an HTTP request open the whole time.
Producing the file is one step in the export. The application also stores the report, records the download, updates the assessment’s report references and tells the user whether the export succeeded or failed. The user requests a report and gets a download, but behind that the file and the application’s records have to agree about what was produced.
Graphics have to work on paper
The reports reuse graphics from the assessment’s analysis pages, which keeps the document tied to the results users are working with in the application.
A graphic that works on a large screen can be hard to read on A4, though. I check labels, legends, captions and proportions at the size people will see them in the report. If the labels are too small, or the explanation lands on a different page from the chart, the report is harder to read even though nothing is missing.
Making a long PDF navigable
Sensible page breaks aren’t enough for a hundred-page document. People need a way to get around it.
We’ve written custom table of contents (TOC) linking code for the generated contents, and the report also has lists of figures and tables. The contents entries take readers to the relevant sections, and the page references come from the finished layout rather than a fixed template. With user content, that’s necessary: a longer executive summary or an expanded findings section pushes everything after it along, and the navigation has to describe the document the reader actually received.
We’ve also built custom PDF bookmark construction. The bookmarks show the report’s section hierarchy in the reader’s PDF viewer, so they can move between sections without going back to the contents page or scrolling through dozens of pages.
Details that travel with the file
Once downloaded, the report leaves our application. Users send it to colleagues and clients and keep it alongside other documents.
Our custom PDF work supports meaningful titles, authors, subjects, dates and other PDF metadata, alongside the visible branding and version information in the report itself. When the file is sitting in a folder with several similar reports, the metadata helps identify it.
Preview and export have to agree
When users open a report preview, they’re checking what they’re about to send to someone else. They might look over a chart, make sure a recommendation fits, or check that a section starts in a sensible place. If the exported PDF behaves differently, those checks are worth much less.
Over the years we’ve tried different approaches to splitting pages, including our own logic. A layout that looks right on screen doesn’t necessarily print as the same document, and that took a while to learn.
We’ve worked to keep the preview representative of the exported document. The screen can make the pages easier to inspect, but the dimensions and layout have to stay meaningful. The editor is where people write and arrange content; the paginated preview is where they judge the finished layout.
Checking the document people actually read
A generated PDF isn’t automatically a good report. I still check whether the graphics are readable, whether the captions help, and whether a table makes sense when it continues onto another page.
For a long document, I look at the final page, the page count, the contents links, the bookmark destinations and the points where tables or sections cross a page boundary. The cover and the first few pages don’t tell me much about any of that.
Short documents are fine for quick checks, but the longest user descriptions, the largest tables and the awkward combinations of graphics are where the layout rules stop working.
Automated checks help verify content, navigation and document properties. I still look at the pages myself to see whether they’re comfortable to read, because a correct page count says nothing about the size of a chart’s labels.
Keeping the mechanics shared
Spartacus supports several assessment frameworks. The content varies, but the mechanics of rendering a report are largely the same, so we keep them shared. A fix to table layout or an improvement to navigation reaches every product that uses them. It also means I can’t look only at the report that prompted a change, because different assessment content puts pressure on the same layout rules in different ways.
The reporting system has a lot of accumulated behaviour, and much of it handles difficult content well, so I’m careful about replacing parts of it. When I change it now, I check what the change does to a long report with real content, especially the pages where tables continue or graphics barely fit.