PDFPipe

Python / contracts

Contract PDFs in Python

Post your contract HTML to a render API from an async endpoint that returns a streaming response, then stream the response with httpx.stream and yield chunks. No browser binary in your Python deployment.

Why contracts become PDFs

A contract has to be fixed at signature. Both parties need to be certain they are looking at identical words, which a rendered page cannot guarantee. That is the difference between a document and a view of a database, and it is why legal and sales operations keep asking for this.

What a contract has to carry

Before any of the code below matters, the template has to produce a document that is actually a contract. These are the fields that make it one, and the ones a reader or an auditor will look for first.

  • the parties, by legal entity name and registered number
  • the effective date and the term
  • numbered clauses with stable numbering, since amendments reference them
  • the commercial schedule, usually as an annexe rather than in the body
  • signature blocks with name, title and date for each party

The Python implementation

httpx is the client most projects here already have. The render happens in an async endpoint that returns a streaming response, and you stream the response with httpx.stream and yield chunks.

python
import os, httpx
from fastapi import FastAPI
from fastapi.responses import StreamingResponse

app = FastAPI()

@app.get("/contract/{contract_id}.pdf")
async def contract_pdf(contract_id: str):
    html = render_contract(contract_id)

    client = httpx.AsyncClient(timeout=60.0)
    request = client.build_request(
        "POST",
        "https://api.pdfpipe.xyz/v1/pdf",
        headers={"Authorization": f"Bearer {os.environ['PDFPIPE_KEY']}"},
        json={"html": html, "options": {"format": "A4", "printBackground": True}},
    )
    response = await client.send(request, stream=True)

    # stream=True keeps a long contract off the worker's heap.
    return StreamingResponse(
        response.aiter_bytes(),
        media_type="application/pdf",
        background=response.aclose,
    )

The CSS that makes a contract page correctly

The layout problem specific to this document is that numbered clauses with cross-references, where a page break in the wrong place separates a clause from the term it defines. These rules handle it.

css
/* A clause must not be separated from the term it defines, and a
   signature block alone on a final page invites a dispute. */
.clause     { break-inside: avoid; }
.signatures { break-inside: avoid; break-before: avoid; }

p { orphans: 3; widows: 3; }

@page { size: A4; margin: 25mm 20mm; }

What goes wrong in Python

requests holds the entire body in memory before you touch it. httpx.stream does not, and on documents past a few megabytes that difference decides whether your worker survives a burst.

What people try first

Most Python projects reach for WeasyPrint, which is excellent at CSS Paged Media but needs Cairo, Pango and GDK-PixBuf present in the image, and diverges from a browser on modern layout. That works until it is running on more than one machine, at which point the browser becomes the thing you operate rather than the thing you use.

Where the contract lives afterwards

Rendering is the short part. A contract is retained for the term plus the limitation period, which typically means well over a decade. Immutability matters more here than anywhere else on this list: both parties must be able to prove the words have not moved.

Getting the document right

  • Check the parties and the effective date, which is what every later dispute turns on.
  • Generation is triggered by agreement being reached, often with a signature workflow following, so size the timeout for that path rather than for a health check.
  • These are produced one at a time by a person who is waiting, so latency is felt directly.
  • Backgrounds are painted by default here, so a design that uses colour needs nothing set. Only an explicit print_background of false turns them off.

Frequently asked

Do I need Chromium installed to generate contracts from Python?

No. The render happens over HTTP, so your Python deployment stays the size it is now. That is the main reason to use an API rather than WeasyPrint, which is excellent at CSS Paged Media but needs Cairo, Pango and GDK-PixBuf present in the image, and diverges from a browser on modern layout.

How do I stop a long contract using all the memory?

Stream the response with httpx.stream and yield chunks. requests holds the entire body in memory before you touch it. httpx.stream does not, and on documents past a few megabytes that difference decides whether your worker survives a burst.

What has to be on a contract?

At minimum: the parties, by legal entity name and registered number; the effective date and the term; numbered clauses with stable numbering, since amendments reference them. The one to get right before anything else is the parties and the effective date, which is what every later dispute turns on.

Do I need to store the generated contracts?

Depends on the document, and this one has a clear answer: a contract is retained for the term plus the limitation period, which typically means well over a decade. Immutability matters more here than anywhere else on this list: both parties must be able to prove the words have not moved.

Can I keep my existing contract template?

Yes, if it produces HTML. Whatever renders your contract view today can render the same markup for the PDF, which is why the CSS above is the only new thing you write.

Related

Other Python documents, and the same contract in other stacks.

100 free documents a month, no card.