
Job description
Senior Data Engineer (Scraping) Job Title: Senior Data Engineer (Scraping) Employment Type: Full-Time Job Summary We are looking for an experienced Senior Data Engineer (Scraping) to design, develop, and maintain scalable web scraping and data harvesting pipelines. The ideal candidate should have strong expertise in Python, web scraping frameworks, ETL processes, distributed data processing, and workflow orchestration.
This role involves working closely with the Data Harvest team to collect, process, and deliver high-quality data from multiple web sources and APIs.
Requirements
Key Responsibilities Design, develop, and maintain scalable web scraping and data harvesting pipelines. Build and maintain web scrapers using Python frameworks such as Scrapy, BeautifulSoup, Selenium, and Playwright. Handle dynamic JavaScript-rendered websites and overcome anti-bot mechanisms including proxy rotation, IP rotation, user-agent rotation, rate limiting, and CAPTCHA handling.
Develop ETL workflows for data extraction, parsing, transformation, and cleaning. Process and transform large-scale datasets using PySpark and distributed computing. Design and schedule workflows using Apache Airflow (Dagster, Prefect, or Luigi experience is an added advantage). Store and manage data in SQL and NoSQL databases.
Work with data formats such as CSV, JSON, XML, and Parquet. Implement monitoring, logging, retry mechanisms, and error handling to ensure pipeline reliability. Ensure compliance with website policies, robots.txt, and data privacy regulations such as GDPR. Collaborate with technical teams and data consumers to maintain data quality and timely delivery.
Benefits
Required Skills Strong programming experience in Python. Knowledge of Node.js or JavaScript is an added advantage. Hands-on experience with: Scrapy BeautifulSoup Selenium Playwright lxml requests/httpx Puppeteer (preferred) Strong understanding of: HTML CSS DOM XPath CSS Selectors HTTP Protocol Experience with REST APIs and GraphQL APIs.
Experience parsing JSON, XML, and HTML data. Strong ETL development experience. Experience with PySpark for distributed data processing. Hands-on experience with Apache Airflow. Good knowledge of SQL and NoSQL databases including PostgreSQL, MySQL, and MongoDB. Experience with asynchronous programming, concurrency, and distributed scraping.
Knowledge of Git, Docker, and cloud platforms such as AWS, Azure, or GCP. Experience implementing monitoring and alerting for production pipelines. Understanding of legal and compliance requirements related to web scraping.
Preferred Skills
Experience with Dagster, Prefect, or Luigi. Knowledge of serverless and cloud-native architectures. Experience with large-scale data engineering projects. Strong debugging, troubleshooting, and analytical skills. Educational Qualification Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
Equivalent practical experience will also be considered. For more information call @ 9952978818 Mail your resume @ anbumozhi.p@verifitech.com