Proof of Skill Logo

Proof of Skill

Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)

Posted One Month Ago
Be an Early Applicant
In-Office
Hyderabad, Telangana, IND
Expert/Leader
In-Office
Hyderabad, Telangana, IND
Expert/Leader
Own and productionize a web data acquisition platform combining resilient crawling with AWS Bedrock-based verification. Responsibilities include modernizing HTTP and browser crawling, implementing proxy and egress management, reverse-engineering APIs, respecting robots.txt and legal requirements, containerizing and scheduling pipelines, and building testing, observability, alerting, and canary-crawl systems. The role requires deep expertise across legacy and modern scraping stacks, anti-bot reliability, data extraction, distributed crawling, compliance, and large-scale platform ownership.
The summary above was generated by AI
About Chryselys

Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.

Role Summary

Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities
  • Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
  • Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
  • Introduce proxy rotation and egress management; retire the single-IP failure mode.
  • Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
  • Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
  • Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
  • Containerize and schedule the pipeline; add CI running the offline tests on every change.
  • Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Skills

Web & protocol fundamentals — HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]

Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]

Modern stack — Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]

Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]

Methodology breadth — API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]

Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]

Anti-bot & reliability — Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]

Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]

Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]

Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]

Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]

Experience — Required
  • 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
  • Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
  • Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
  • Mentored engineers; set crawl standards, review practice, and on-call runbooks.
  • Degree optional — equivalent practical experience is fully accepted.
Nice-to-Have
  • US payer policy, formulary, or prior-authorization document domain knowledge.
  • Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
  • LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/.
  • Compliance or legal-review exposure on data acquisition programmes.
Equal Employment Opportunity

Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Similar Jobs

Yesterday
In-Office
Hyderabad, Telangana, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads the strategy, architecture, delivery, and operationalization of agentic AI solutions in enterprise healthcare environments. Manages engineering teams building LLM workflows, copilots, intelligent automation, and integrated AI platforms. Responsibilities include cloud architecture, secure deployment, AI evaluation, stakeholder alignment, delivery planning, technical governance, executive communication, and coaching engineering leaders. The role also drives scalable engineering practices and production-ready solutions across marketing technology deployments.
Top Skills: Adobe Experience PlatformAPIsAzureData PlatformsInfrastructure As CodeLlmsMachine LearningNextjsReact
Yesterday
In-Office
Hyderabad, Telangana, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads a team of AI engineers while designing, deploying, monitoring, and optimizing agentic AI systems and AI-driven developer lifecycle tools. Oversees production delivery, team development, PI planning, resource forecasting, roadmaps, stakeholder collaboration, and healthcare provider platform enhancements. The role requires hands-on expertise with agent orchestration, LLM platforms, observability, performance tuning, and AI-enabled SDLC automation, along with strong leadership and delivery management.
Top Skills: Adk ToolkitsAgentic ArchitecturesAIAi Developer Lifecycle FrameworksAi System ObservabilityAnthropicClaudeCopilotGeminiLlmsMcp ServersOpenai
Yesterday
In-Office
Hyderabad, Telangana, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads enterprise AI/ML and Generative AI engineering strategy, architecture, delivery, MLOps, governance, and production deployment. Builds and mentors engineering teams, drives LLM, AI agent, predictive analytics, and automation initiatives, evaluates emerging technologies, and partners with business, product, architecture, and technology stakeholders to deliver secure, responsible AI solutions in a healthcare technology environment.
Top Skills: Ai AgentsAWSAzureAzure OpenaiDatabricksGCPGenerative AiLangchainLarge Language Models (Llms)MlflowMlopsPythonRag ArchitecturesSemantic Kernel

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account