Skip to content
moisis. ENGINEERING & INDEPENDENT WORK
← All writing
Laravel & PHP · 5 min read

Building a Queued PDF Export Pipeline in Laravel

How we coordinate graphics preparation, cached output and queued PDF generation in Spartacus Insights, and why exporting includes more than creating the file.

In this article 7 sections

In Spartacus Insights, generating a report involves more than sending HTML to a PDF tool. The document contains assessment results, content written by users and graphics from the analysis pages. Some reports run past 100 pages.

I’ve written about the layout and navigation work behind those reports. This post is about the export around it, which runs on Laravel queues: a batch of graphics jobs that has to finish before the report job starts.

Start with the dependencies

The report can’t be generated until its graphics exist. Each graphic is its own piece of work, but the report needs all of them.

If I dispatch the graphics jobs and the report job side by side, nothing says the graphics come first. With more than one worker, dispatch order doesn’t guarantee that one job finishes before the next one starts.

So the export groups the graphics jobs into a batch and dispatches the report job from the batch’s success callback.

The request returns while workers prepare the graphics and generate the PDF, and the user gets a notification when the download is ready.

A batch for preparation, a job for the report

This example puts the coordination in an action called by the controller, using simplified job names and an export reference.

<?php

declare(strict_types=1);

namespace App\Actions;

use App\Jobs\GenerateReportFile;
use App\Jobs\PrepareReportChart;
use Illuminate\Bus\Batch;
use Illuminate\Support\Facades\Bus;

class QueueReportExport
{
    public function handle(int $exportId): void
    {
        Bus::batch([
            new PrepareReportChart($exportId, 'summary'),
            new PrepareReportChart($exportId, 'recommendations'),
        ])
            ->then(function (Batch $batch) use ($exportId): void {
                GenerateReportFile::dispatch($exportId)
                    ->onQueue('reports');
            })
            ->onQueue('reports')
            ->dispatch();
    }
}

The chart jobs implement ShouldQueue and use Laravel’s Batchable trait. The report job is queued as well, but it only runs once the preparation batch is done. This assumes an asynchronous queue connection and a worker listening on the reports queue.

I use then() because the report depends on every graphic succeeding. finally() would also run after failed jobs.

When then() fires, the graphics are ready, but the export isn’t finished. The report job still has to generate the PDF, store it and make it available for download.

Prepare graphics before the user asks for an export

We dispatch graphics preparation while users work on a report, including when they open the editing and preview screens. By the time someone requests a PDF, some of the output may already exist.

The export runs its preparation stage anyway. Each graphics job checks for cached output and only renders the graphic if it isn’t there, so a job can complete just by finding what it needs.

That way editing and exporting share the same preparation work, and the export keeps an explicit dependency on its graphics. It doesn’t have to assume that an earlier screen visit prepared everything.

The cache check also has to account for changes to the data or graphic settings. Output that exists isn’t necessarily output that’s still usable.

Give exports their own workers

Exports run on a separate queue with their own Horizon worker configuration, apart from notifications. That’s where we set the resources and time allowed for longer-running work.

A graphics job and a notification need different things. Sharing one worker configuration would make those trade-offs harder to manage. Separate queues control which workers take which jobs, although the processes still share the machines they run on.

The worker timeout is shorter than the connection’s retry_after, so a stalled worker is killed before its job becomes available for another attempt.

Generating the PDF is only part of the report job

Our report job uses a shared export lifecycle. It generates the document, stores the output, creates the download record and updates the assessment’s report reference. It also sends the success notification and cleans up temporary output.

A PDF sitting in temporary storage is no use to the user yet. The application needs a download record that points at the stored file.

Several assessment products share this lifecycle. The report content differs, but storing a file and making it available shouldn’t need its own implementation for every framework.

The ready notification goes out only after the file is stored and the records are updated.

Failures happen at more than one stage

The shared report job logs failures and can notify the user when the export fails. But that job only starts after the graphics batch succeeds.

If preparation fails, the report job may never be dispatched, so its failure handler never gets a chance to report anything. Preparation failures need their own path to the user.

Retries raise a separate concern, which is what had already happened before the job failed.

For a graphics job, usable output that already exists means a retry can skip the rendering. The report job has more side effects. It may have stored a file or created a download record before a later step failed, so retrying it means looking at those effects as well as whether the PDF can be generated again.

ShouldBeUnique can limit duplicate dispatches, but it doesn’t make those side effects safe to repeat. Its uniqueness constraints also don’t apply to jobs inside batches.

What I look for when reviewing an export

When I review an export, or debug one that went wrong, I follow a single request through the jobs. Which graphics might already be cached? Which stage dispatches the report? When does the download become available, and what happens if a worker fails before then?

Graphics preparation, report generation, storage and the download record each fail in their own way, so I check them one at a time instead of treating the export as a single step.