
Closed
Posted
Paid on delivery
I have roughly 5 000 PDFs that are a mix of Bills of Lading and Commercial Invoices. I need a reliable script—Python preferred, but any language is fine—that can open each file, read the key parties on the document, and aggregate everything into a single Excel workbook. The script must capture: • Shipper details • Receiver details • Broker details Accuracy matters more than speed; some files are machine-readable, others are scanned, so you may have to blend text parsing with OCR (think [login to view URL], PyPDF2, Camelot, Tesseract, or any stack you trust). The output should be a clean .xlsx file with one row per shipment and clearly labeled columns for each data point. Please send a brief but detailed proposal that explains: – The libraries or tools you will use – How you will handle both text and scanned PDFs – A plan for testing accuracy on a sample set before running the full batch – An estimated timeline for completion Deliver a runnable script, clear setup instructions, and the final Excel file. Let me know any questions you have so we can get started quickly.
Project ID: 40582233
10 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
10 freelancers are bidding on average ₹2,380 INR for this job

Hey, this is a great fit - I build this kind of PDF-to-Excel extraction pipeline regularly, including mixed batches of machine-readable and scanned docs like BOLs and commercial invoices. Approach: pdfplumber/PyPDF2 for machine-readable PDFs, Tesseract OCR (via pdf2image) for scanned ones. Layouts vary by shipper/broker, so I'll use keyword/positional anchors ("Shipper", "Consignee", "Broker") instead of rigid templates so it holds up across formats. Camelot if any fields need table extraction. Testing plan: before the full 5,000-file run, I'll process a sample of 30-50 files across your formats, verify shipper/receiver/broker fields against source PDFs, and refine logic based on mismatches. Once sample accuracy is solid, I run the full batch, with a status column flagging low-confidence files for review. Output: single .xlsx, one row per shipment, columns for Shipper, Receiver, and Broker (name/address/contact), plus a confidence/flag column. Deliverables: runnable Python script, short setup guide (dependencies, pointing at your PDF folder, rerunning on new batches), and the final Excel file. Timeline: sample + validation in the first couple days, then full run once logic's confirmed - roughly 5-7 days depending on how many scanned/low-quality files show up. Question: are shipper/receiver/broker fields labeled consistently across documents, or does it vary a lot by carrier/broker? Affects how much I lean on OCR + fuzzy matching vs straightforward parsing.
₹1,500 INR in 2 days
5.6
5.6

Hi, Could you provide a sample of the PDFs so I can assess the complexity involved? To tackle your project, I’ll use Python with libraries like PyPDF2 for machine-readable PDFs and Tesseract for OCR on scanned files. My approach involves initially parsing documents, extracting shipper, receiver, and broker details, then aggregating everything into an organized Excel workbook. I will implement a testing phase on a small sample of PDFs to ensure accuracy before processing the entire batch. The complete project should take around 5-7 days, including testing and revisions. You’ll receive a runnable script with thorough setup instructions and the final .xlsx file. Looking forward to your response! Best Regards, Ahmad
₹1,050 INR in 7 days
0.0
0.0

Hi, this is squarely what I do. I build Python pipelines that read messy PDF batches and turn them into one clean Excel sheet, so 5000 mixed Bills of Lading and Commercial Invoices into a single workbook is a job I can run reliably. My approach is text first, OCR fallback. I parse the machine readable files directly with pdfminer and Camelot for tables, and route the scanned ones through Tesseract, so both kinds get read accurately. For each file I pull shipper, receiver and broker details and write one row per shipment with clearly labelled columns. I add validation flags so any low confidence read is easy to spot and fix. Since you said accuracy beats speed, I will run a small sample first so you can confirm the columns before the full batch. Send me a few sample PDFs and I will start.
₹1,050 INR in 7 days
0.0
0.0

Moga, India
Member since Mar 17, 2025
₹600-700 INR
$30-250 USD
₹1500-12500 INR
$30-250 USD
₹1500-12500 INR
$2-8 USD / hour
₹400-750 INR / hour
$250-750 USD
€30-120 EUR
$30-250 NZD
£20-250 GBP
$30-250 USD
₹600-1500 INR
£20-250 GBP
₹600-1500 INR
$250-750 USD
$15-25 USD / hour
$250-750 USD
₹1500-12500 INR
£250-750 GBP
₹1500-12500 INR