Manifesto: Freight Contract Data Extraction Pipeline · Matt Alexius

Manifesto pulls rates and surcharges out of freight contract PDFs and links each value back to its source line, cutting manual data entry from hours to minutes. Next.js 16, FastAPI, PostgreSQL.

Manifesto reads freight contracts and turns them into structured rate data, cutting manual data entry from hours to minutes. Every rate, arbitrary, and surcharge it extracts links back to its line in the source PDF, so pricing teams can trust the output without re-checking every row by hand.

The problem

Carrier contracts arrive as long, dense PDFs, with hundreds of rates, arbitraries, and surcharges spread across tables and footnotes. Someone on the pricing team has to type all of it into a rate sheet. That takes hours per contract, typos creep in, and when a quote gets disputed there's no easy way to see where a number came from.

The approach

The extraction is automated, and every value keeps a pointer to where it came from. A FastAPI pipeline parses each contract PDF and stores structured rows in PostgreSQL, each linked to its line in the source document. On top of that is a Next.js 16 app where reviewers search contracts, compare extracted rows against the PDF side by side, and approve them. Several companies use the same platform, so it was multi-tenant from the start: data is isolated per organization, access is role-based, and new teams go through a guided onboarding flow.

What I did

  • Built the PDF extraction pipeline with FastAPI and PostgreSQL.
  • Wrote the whole web UI in Next.js 16: dashboard, contract viewer, search, and reviews.
  • Implemented a 6-phase onboarding flow using Supabase Auth and org memberships.
  • Designed the multi-tenant setup, with org-scoped data isolation and RBAC.
  • Added a platform admin panel with impersonation, feature flags, and lifecycle tracking.

What I learned

People only adopt automation they can check. A fast extractor saves nobody time if the pricing team still re-checks every row, so showing exactly where each number came from mattered as much as accuracy. Building multi-tenancy early paid off too. Org isolation and RBAC are painful to retrofit, and because they were there from day one, onboarding, impersonation, and feature flags were easy to add later.

Champion of the "From Paper to Data" Logistics Challenge