Senior Data Infrastructure Engineer (Scraping & Scale)
AWISEE · United States · posted today ago
Listing supplied by Himalayas. 101 Careers did not originate this post and applications are handled by the employer.
About this role
This is a remote position.
We are looking for a Senior Data Infrastructure Engineer specializing in web scraping, anti-bot evasion, and large-scale data ingestion. In this role, you will build and maintain the core ingestion engine powering our social media data lab. You will overcome complex platform defenses to deliver millions of profile, post, and video records daily with near-zero downtime.Responsibilities
- Anti-Bot Evasion Architecture: Build and manage stealth scraping clusters using residential proxy networks, TLS fingerprinting, headful/headless browser farms (Playwright, Puppeteer), and session rotation.
- High-Throughput Pipelines: Build fault-tolerant, scalable web-scraping pipelines that extract data from Instagram, TikTok, YouTube, X, and web sources.
- Pipeline Orchestration: Design distributed queues and workflow engines (Temporal, Ray, Apache Kafka, Celery) to manage millions of asynchronous scraping tasks daily.
- Storage & Data Lake Management: Architect structured and unstructured storage environments (Parquet, Apache Iceberg, Snowflake, S3) for downstream AI modeling.
- Monitoring & Evasion Recovery: Implement automated alerts for platform UI/API changes, blocking patterns, and proxy failures.
Requirements
- 4+ years of experience in high-volume web scraping, data engineering, or reverse engineering.
- Deep experience defeating advanced anti-bot providers (Cloudflare, Akamai, PerimeterX) via TLS impersonation, browser automation, and proxy management.
- Mastery of Python or Go, with deep knowledge of Playwright, Puppeteer, Scrapy, or Selenium.
- Experience with distributed systems and task queues (Temporal, Ray, Kafka, Redis).
- Strong SQL skills and experience with modern analytical data lakes / data warehouses.
- Direct experience extracting short-form video content and user metadata from TikTok, Instagram, and YouTube.
- Experience integrating scraping outputs directly into vector databases and real-time AI processing queues.
Originally posted on Himalayas