
Completed
Posted
We are building Pharmelis Registry — a canonical database for pharmaceuticals. To make any pharmaceutical product understandable, anywhere, in any language. Pharmaceutical data today is fragmented, inconsistent, and multilingual. There is no reliable way to identify and reconcile products globally across sources and markets. Pharmelis Registry aims to capture real-world data and resolve it into clear, unambiguous product identities. Objective Design a system that: - ingests heterogeneous pharmaceutical data (CSV, APIs, websites, PDFs, images) - works across countries and languages - handles messy, inconsistent data - resolves product identity across sources - produces a consistent global reference - includes an internal dashboard to operate the system The dashboard must allow: - inspecting ingested data - reviewing identity decisions - monitoring system coverage - adding and managing data sources Deliverables Provide a concise architecture document covering: - identity resolution strategy (core part) - ingestion approach for different source types - core data model (what is stored vs computed) - system architecture and data flow - dashboard design and capabilities - recommended tech stack and trade-offs Required Profile Strong in: - Python - data engineering / pipelines - web scraping / ingestion - API/backend development - dashboard or web app development Experience with messy data, search systems, or entity resolution is expected.
Project ID: 40411067
26 proposals
Remote project
Active 2 mos ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
26 freelancers are bidding on average $23 USD/hour for this job

Systems like this often look like data pipelines, but the real challenge is accurate identity resolution across messy, multilingual pharmaceutical data—that’s where most solutions fail. I’ve worked on similar data/ER systems where the focus was building reliable matching logic (rules + ML) and clean ingestion pipelines, not just collecting data. I can help with: Designing a robust identity resolution strategy (core matching + scoring + human-in-loop) Building ingestion pipelines (CSV, APIs, scraping, PDFs/OCR) Defining a clean canonical data model vs raw data layer System architecture (Python-based pipelines, API, search, storage) Dashboard design for review, monitoring, and source management Tech stack recommendations with clear trade-offs The goal is a scalable, trustworthy registry, not just a data aggregator. For similar work and case studies: https://www.freelancer.com/u/Microlent Happy to draft a concise architecture doc tailored to your use case quickly. — Rajesh Rolen
$20 USD in 40 days
9.3
9.3

⭐⭐⭐⭐⭐ Design the Architecture of Pharmelis Registry (Global Pharmaceutical Data System) -- 2 ❇️ Hello! I have thoroughly reviewed the requirements for the Pharmelis Registry project and am excited about the opportunity to architect a robust, global pharmaceutical data system. With my extensive experience in data engineering, particularly in handling diverse and inconsistent datasets, I am confident in delivering a system that not only meets but exceeds your expectations. ➡️ Why Me? I hold a PhD in Computer Science with a specialization in data systems and have over 10 years of experience in data engineering, API development, and complex system architecture. My expertise in Python, coupled with a profound understanding of web scraping, data ingestion, and backend development, positions me perfectly to tackle the challenges presented by Pharmelis Registry. ➡️ Lets discuss this project in more detail at your earliest convenience. I am eager to share insights and explore the specifics of the architecture and technologies that will best serve this project. ➡️ Some of my relevant work: ✅ Development of a Multilingual Content Aggregator for a major news outlet ✅ Architecting a Data Reconciliation System for Financial Institutions ✅ Creation of a Customized Dashboard for Real-Time Data Monitoring in Healthcare ✅ Implementation of Entity Resolution Systems for E-commerce Giants ✅ Comprehensive Data Pipeline Solutions for Multinational Corporations ✅ Design and Deployment of API Services for Cross-Platform Data Integration Each of these projects involved complex data issues similar to those expected with the Pharmelis Registry, particularly in terms of data ingestion from heterogeneous sources and the resolution of entity identities across various systems and languages. I look forward to bringing my background in these areas to the Pharmelis Registry project, ensuring a successful design and implementation of a global reference system for pharmaceuticals. Waiting for your response! Best Regards, Dr. Muhammad Asad
$18 USD in 30 days
6.9
6.9

Greetings! I specialise in data engineering and entity resolution, with over 9 years of experience building pipelines that ingest messy heterogeneous data and resolve it into clean, canonical identities across languages and sources. Here's how I can help: - Design a robust identity resolution strategy using probabilistic and deterministic matching across pharmaceutical attributes — name variants, active ingredients, dosage forms, and market codes — reconciling records into unambiguous global identities - Build ingestion pipelines for every source type — CSV, REST APIs, web scrapers, PDF extractors, and OCR for images — normalising everything into a consistent staging schema before resolution runs - Define a clear data model separating raw ingested records from computed canonical entities, with full lineage tracking so every identity decision is auditable from the dashboard - Deliver an internal dashboard covering ingested data inspection, identity decision review, source coverage monitoring, and data source management connected to live pipeline state My recommended stack is Python, PostgreSQL with pgvector, dbt for transformations, and Airflow for orchestration — with Streamlit or FastAPI plus React for the dashboard. How many data sources and approximate product records are you targeting for the initial ingestion phase?
$20 USD in 40 days
6.7
6.7

Hello, I see you’re building Pharmelis Registry with a focus on global identity resolution across heterogeneous pharmaceutical data, and that signals the need for a robust architecture that can systematically normalize and unify messy sources. I’ve designed similar cross‑market data systems where I delivered a multilayer ingestion pipeline and an entity resolution engine that reduced duplicate product identities by over 40%. The real challenge here is enforcing a consistent identity model when sources differ not only structurally but semantically across regions, languages, and regulatory models. A junior engineer might focus on parsing, but the real risk is weak feature extraction, without a reliable canonicalization layer, identity resolution becomes brittle. I’ll map out an ingestion framework covering structured files, scrapers, OCR pipelines, and API connectors. I’ll define the core entities, derived attributes, and match rules, then outline batch and incremental workflows. For the dashboard, I’ll specify review interfaces, coverage metrics, and source management tools aligned with your operational needs. Before drafting the architecture, I need to confirm expected data volume ranges, preferred deployment environment, and whether the resolution engine must support active learning. Thanks, John allen.
$15 USD in 27 days
5.8
5.8

I’ll design a scalable pharma registry with robust entity resolution (fuzzy matching + rules + embeddings), multi-source ingestion (APIs/CSV/PDF/OCR), and a clean dashboard to review identities and coverage—optimized for messy global data.
$15 USD in 40 days
5.4
5.4

With over 7 years in the field of Full Stack Development - specialized in PHP, Laravel, CodeIgniter and API development, I offer a robust suite of skills that merge perfectly with your project requirements for Pharmelis Registry. Having mastered the exact core data handling you are seeking - from ingestion to search systems, and ultimately to identity resolution- I am confident that my expertise will provide you with a well-rounded and efficient system. Moreover, I fully understand the gravity of clean data in such a project and have had comprehensive experience managing messy data in different sources- CSVs to APIs, websites to PDFs. This includes proficient understanding of multilingual challenges, ensuring that Pharmelis will be able to operate seamlessly across languages and borders. Additionally, my web scraping proficiencies guarantee accurate collection of data from various sources augmenting the reliability of your database. My well-honed dashboard design bestows you visibility into what goes into the system throughout its composition process. Ultimately, I aim to employ my 1000 character space to assert that when it comes to designing an architecture for Pharmelis Registry, your best fit is Sachin.
$15 USD in 40 days
5.0
5.0

Hi there, ❤️❤️❤️ I’ve reviewed your Pharmelis Registry project and it aligns well with my experience in Python data engineering, entity resolution, ingestion pipelines, and backend/dashboard architecture. I can help you design a clear architecture for a global pharmaceutical reference system that handles messy multilingual data and resolves product identities across sources. How I can help: • Define a practical identity resolution strategy using canonical product entities, source records, confidence scoring, human review, and audit trails. • Design ingestion flows for CSV, APIs, websites, PDFs, and images, including normalization, validation, deduplication, and enrichment. • Propose the core data model, system data flow, dashboard capabilities, and a realistic tech stack with trade-offs. Relevant experience: I’ve worked on similar systems involving Python pipelines, web scraping, API backends, search/indexing, messy data reconciliation, and operational dashboards, and I can start working immediately. Approach: I will keep the architecture concise, implementation-oriented, and focused on scalable identity decisions, maintainable pipelines, and operational visibility. Best regards,
$35 USD in 20 days
4.2
4.2

Hi, this is Kris from McKinney, Texas, I've reviewed your project requirements and understand that the key challenge lies in designing a system that can effectively ingest heterogeneous pharmaceutical data, resolve product identities globally, and provide a consistent reference across sources and markets. My approach would involve developing a robust data engineering pipeline using Python, implementing web scraping techniques for data ingestion, and creating a user-friendly dashboard for monitoring and managing the system operations. A few additional questions: Q1: Have you already identified specific data sources that need to be integrated into Pharmelis Registry? Q2: What level of scalability and performance are you expecting from the system? Q3: Do you have any preferences for the tech stack to be used in this project? Best regards, Kris Kramer
$20 USD in 40 days
4.8
4.8

⭐⭐⭐⭐⭐ ✅Hi there, hope you are doing well! I have experience designing data ingestion and identity resolution systems that handle heterogeneous, multilingual, and inconsistent data inputs, streamlining global product datasets seamlessly. The crucial part to successfully complete this project is creating a robust identity resolution strategy that accurately reconciles fragmented pharmaceutical data across diverse sources and languages. Approach: ⭕ I will architect a scalable ingestion pipeline supporting CSV, APIs, web scraping, and OCR on PDFs/images. ⭕ Design the core data model distinguishing stored raw data vs computed identities. ⭕ Develop an internal dashboard for data inspection, identity decision review, system monitoring, and source management. ⭕ Recommend and justify a modern tech stack with Python, Django, and relevant data engineering frameworks. ❓ Could you clarify the expected volume of data ingestion and preferred tech preferences if any? I am confident my expertise in Python, data engineering, and entity resolution will deliver a comprehensive architecture document that meets all your objectives precisely. Looking forward to collaborating! Best regards, Nam
$25 USD in 27 days
3.9
3.9

Designing a canonical registry like Pharmelis requires a precise architectural balance between global data integrity and the flexibility to accommodate diverse regulatory standards. I have architected high-compliance data systems where master data management and a "single source of truth" were the primary drivers. My focus is on creating a schema that remains performant as it aggregates complex, multi-national pharmaceutical datasets, ensuring no loss of granularity during normalization. I will architect Pharmelis using a Master Data Management (MDM) framework centered on ISO IDMP standards to ensure interoperability between FDA and EMA datasets. The core will leverage a polyglot persistence strategy, utilizing a relational engine for ACID transactions alongside a Graph database to map drug-to-ingredient relationships. I’ll design automated ETL pipelines with strict validation layers to ingest data from heterogeneous sources while maintaining canonical record quality. This infrastructure will be containerized and optimized for high-throughput API access, providing stakeholders with reliable, low-latency data via a secure GraphQL or REST interface. Are you prioritizing specific regional registries for the initial build, and what are your expectations for data refresh frequency? Understanding the scale of anticipated third-party integrations will also help me refine the caching and rate-limiting strategy. I’d love to discuss your long-term scalability goals; I’m available for a brief chat or call to align on the architectural roadmap for the Pharmelis ecosystem.
$25 USD in 7 days
3.3
3.3

Hello, I will design a robust architecture for Pharmelis Registry, ensuring seamless ingestion of pharmaceutical data, multilingual support, and identity resolution across sources. The system will feature an intuitive dashboard for data management and monitoring. Let's discuss the project further. Thanks
$15 USD in 40 days
3.3
3.3

With a strong background in several essential skills needed for this project -- backend development, data management, and data processing -- I am confident that I can offer a unique perspective and expertise to redefine how global pharmaceutical data is managed. My proficiency in Python, specifically data engineering and pipeline development, will be valuable as we tackle the challenge of ingesting and managing heterogeneous pharmaceutical data across multiple sources. My solid experience in web scraping will also come in handy in ensuring the system collects accurate and up-to-date information. Regarding the dashboard, I can design an interface that is intuitive, user-friendly, and informative. Being proficient in Django gives me an edge in web app development so designing an efficient dashboard for monitoring system coverage, reviewing identity decisions, and managing data sources will be within my skillset. I always ensure that the tech stack I recommend prioritizes both functionality and trade-offs that meet clients' needs efficiently. Choose me to envision a whole new standard for global pharmaceutical data management with Pharmelis Registry!
$20 USD in 40 days
2.8
2.8

Hi there, I’m Everett, and I understand the importance of creating a unified and comprehensible pharmaceutical database through the Pharmelis Registry. With extensive experience in data engineering, particularly in handling complex and messy datasets, I am confident in my ability to craft an architecture that efficiently ingests various data types while resolving product identities. I aim to develop a robust ingestion framework that supports CSV, APIs, and more, ensuring the system can navigate inconsistencies and language differences. The architecture will focus on a solid core data model, allowing for clear data flow and identity resolution across platforms. The internal dashboard will be user-friendly, enabling easy inspection, identity review, and source management. I am available for real-time communication aligned with your time zone and can provide a simple demo within 12 hours of project commencement. Q1: What specific challenges have you faced with existing pharmaceutical data? Q2: Are there particular data sources or formats you prioritize for initial ingestion? Q3: What are your key metrics for evaluating the success of the system? Looking forward to collaborating on this impactful project! Best regards, Everett
$50 USD in 9 days
2.9
2.9

Hello, I am Vishal Maharaj, with 20 years of experience in Python, Django, Software Architecture, API Development, Data Management, Backend Development, Data Integration, Web Scraping, and Data Modeling. I have carefully reviewed the requirements for designing the architecture of Pharmelis Registry. For this project, I propose to create a robust system that can efficiently ingest heterogeneous pharmaceutical data from various sources, handle multilingual challenges, resolve product identities, and provide a consistent global reference. The architecture will include a comprehensive identity resolution strategy, adaptable ingestion methods, a detailed data model, a structured system architecture, and a user-friendly dashboard for data management. I am confident in my ability to deliver a solution that meets the project objectives and requirements. Please initiate a chat to discuss further details. Cheers, Vishal Maharaj
$20 USD in 40 days
2.6
2.6

What are the key technical challenges in ensuring the success of the Pharmelis Registry project, especially when dealing with fragmented pharmaceutical data sources globally? Hey, I understand the critical need to design a system that can effectively ingest and resolve heterogeneous pharmaceutical data while maintaining consistency across languages and regions. My strengths lie in: - Python expertise for robust data engineering and backend development - Experience in web scraping and API integration for data ingestion - Proficiency in designing intuitive dashboards for efficient system monitoring and management My approach involves creating a streamlined architecture that prioritizes identity resolution strategies, adaptable ingestion methods, and a user-friendly dashboard interface. Let's discuss how we can tackle the complexities of messy data and ensure seamless product identity reconciliation on a global scale. Best, Daniela
$25 USD in 40 days
0.0
0.0

I can help you design a robust, scalable architecture for Pharmelis Registry as a global canonical pharmaceutical database. My focus would be on data integrity, interoperability, and performance, ensuring your system can reliably support complex queries and integrations across markets and regulatory environments. I’ve designed data platforms and registries for regulated domains, including healthcare and life sciences, involving strict validation rules, auditability, versioning, and secure access layers. This background is directly relevant to building a trustworthy, authoritative pharmaceutical registry. My approach would start with requirements clarification, then a clear domain model, followed by a modular system architecture: data ingestion and normalization, master data management, APIs, security, and observability. I’d also define technology choices and scaling patterns to support global usage. I would love to chat more about your project! Regards
$20 USD in 7 days
0.0
0.0

Paris, France
Payment method verified
Member since Jul 29, 2017
$5000-10000 USD
$30-250 USD
$250-750 USD
min €36 EUR / hour
$30-250 USD
₹250000-500000 INR
$30-250 USD
₹12500-37500 INR
$8-15 USD / hour
₹12500-37500 INR
$100-225 USD / hour
$5000-10000 USD
$10-30 USD
$10-100 USD
$10-30 USD
$1500-3000 SGD
₹12500-37500 INR
₹1500-2000 INR
₹5000-20000 INR
₹600-1500 INR
$250-750 USD
$2-8 USD / hour
£20-250 GBP
$250-750 USD
$250-750 USD