Python Data Extraction Engineer
VSERVE EBUSINESS SOLUTIONS INDIA PRIVATE LIMITED Coimbatore North, Tamil Nadu, India
Outsourcing and Offshoring Consulting · 51-200 employees
About the role
You will build and maintain scalable data-acquisition pipelines to extract, clean, and normalize fragmented information from government portals and public websites. This involves developing automated processes for entity matching, record linkage, and quality assurance to integrate data into a centralized intelligence platform.
What they look for
Requirements
The role requires 3-7 years of experience with strong proficiency in Python, web scraping libraries, and data manipulation tools. Candidates must demonstrate problem-solving skills in handling complex, semi-structured data sources like PDFs, JavaScript-rendered pages, and inconsistent government schemas.
Full description
Python Data Extraction Engineer – Web Scraping & Government Data Location: Coimbatore Experience: 3–7 years Role Type: Full-time
About the Role We are building a data intelligence platform that looks to analyze fragmented public information into structured, actionable business data. We are looking for a strong Python Data Extraction Engineer who can discover, extract, clean, normalize and integrate data from government portals, public websites, PDFs, APIs and other open data sources. This is not a conventional application-development role. The ideal candidate enjoys solving difficult data-acquisition problems involving poorly structured websites, inconsistent government portals, changing schemas, PDFs, JavaScript-rendered pages and large volumes of semi-structured information. What You Will Own You will build and maintain the data-acquisition layer of the platform. Key responsibilities include:
Identify and evaluate government and public data sources relevant to property, businesses and commercial activity.
Build Python-based crawlers and extraction pipelines for government portals and public websites.
Extract structured information from HTML pages, tables, PDFs, downloadable files and publicly accessible APIs.
Work with JavaScript-rendered websites and multi-step public search interfaces.
Automate recurring extraction from multiple sources while respecting applicable access rules, rate limits, and terms.
Clean, standardize and normalize inconsistent data from different government authorities.
Perform entity matching and record linkage across datasets using fields such as owner name, company name, address, survey number, coordinates and project information.
Develop mechanisms to detect website/schema changes and extraction failures.
Build validation and QA processes to measure completeness and accuracy.
Store extracted information in structured databases and expose clean datasets to downstream applications.
Work closely with GIS, product and engineering teams to combine location-based signals with government/public records.
Research new public and open-data sources that can improve the accuracy and completeness of our intelligence.
Examples of Data Sources The work may involve sources such as:
Municipal corporation portals
State planning and development authorities
DTCP and similar planning authorities
RERA databases
Building and planning permission records
Land and property records
Tender and procurement portals
Company/business registries
Government open-data portals
Environmental and regulatory approvals
Public notices and downloadable government documents
Maps and geospatial datasets
Other legally accessible public and open-source information
Required Technical Skills Strong hands-on experience with:
Python
Web scraping and crawling
Requests / HTTP clients
BeautifulSoup / lxml
Selenium and/or Playwright
REST APIs and JSON
HTML/XML parsing
Pandas
SQL
Data cleaning and transformation
Regex and text processing
ETL/data pipelines
Git
What We Are Looking For We particularly want someone who is a problem solver rather than simply a Python programmer. The candidate should be able to investigate the available sources, understand how the underlying website works, determine the best extraction approach, build the pipeline and validate the resulting data. Ideal Background Candidates may come from backgrounds such as:
Web scraping / data extraction companies
Alternative-data companies
PropTech / real-estate data companies
Market-intelligence companies
OSINT/data intelligence companies
Government-data projects
Data aggregation platforms
Lead/data enrichment companies
GIS/location-intelligence companies
Success in This Role Within the first few months, the successful candidate should be able to:
Map relevant government/public data sources.
Build reliable extraction pipelines across multiple portals.
Convert fragmented information into standardized records.
Cross-reference records from multiple sources.
Establish automated QA and monitoring.
Continuously discover additional datasets that improve our product's coverage and accuracy.
The objective is not simply to scrape websites. It is to build a scalable public-data acquisition and enrichment engine that becomes a core component of our intelligence platform.
Requirements
Python, pandas, webscraping, web crawling,selenium postman
Similar roles
-
Software Engineer Intern (Python)
Thales Singapore, Singapore
-
Data Python Scientist
Ford Dearborn, Michigan, United States · $100K–$193K/yr
-
GHR AI Engineer – Agentic Integrations (MCP / A2A / Python / APIs) AVP
State Street Princeton, New Jersey, United States · $90K–$158K/yr
-
Praktikum Data driven Machine Diagnosis in Python (SS27)
TRUMPF Ditzingen, Baden-Württemberg, Germany
-
Senior Python Full Stack Developer – AWS Cloud & AWS Deployments
Derex Technologies Inc Florida City, Florida, United States
-
Technology Architect Python Big Data
Timebox Pro Charlotte, North Carolina, United States