Techdome Logo

Techdome

Data Engineer

Posted 5 Days Ago
In-Office
Hyderabad, Telangana, IND
Mid level
In-Office
Hyderabad, Telangana, IND
Mid level
Build and maintain data foundations for AI agents, analytics, scorecards, and dashboards. Responsibilities include developing ingestion pipelines and enterprise connectors, modeling raw-to-mart layers, managing metadata and vector indexes, implementing data quality and lineage controls, and supporting governed data stores. The role also covers API integrations, incremental ingestion, security controls, testing, release activities, and data validation. Preferred experience includes RAG preparation, unstructured content processing, knowledge graphs, LLM applications, and regulated life sciences data.
The summary above was generated by AI
Role summary: The Data Engineer builds the data foundation that AI agents, scorecards, and dashboards run on: ingestion pipelines, connectors, metadata and search indexes, vector stores, data marts, and quality instrumentation.

Experience: 3 to 6 years in data engineering. Senior Data Engineer: 2+ years, including ownership of data foundation architecture and connector frameworks.
Key responsibilities
  • Design and build ingestion pipelines for structured data, unstructured content (documents, PDFs), and metadata.
  • Build reusable connectors to enterprise systems, catalogs, content repositories, and third-party or licensed sources via APIs.
  • Model and build raw-to-mart data layers that serve analytics and AI use cases.
  • Implement metadata extraction, enrichment, and search indexing, including semantic and vector indexes for RAG.
  • Set up and manage vector databases and knowledge repositories used by LLM agents.
  • Implement data quality rules, profiling, scoring outputs, and exception handling.
  • Register lineage and maintain source registries, version tracking, and refresh controls.
  • Design data stores for signals, findings, audit trails, and user feedback loops.
  • Apply security and governance controls: RBAC, PII/sensitivity flagging, and approved data handling.
  • Support SIT/UAT data validation, defect fixes, and production release activities.
Required skills
  • Strong Python and SQL; solid data modeling (dimensional and normalized).
  • Hands-on experience with a modern data platform such as Databricks, Snowflake, or Azure/AWS data services.
  • Pipeline orchestration and transformation (Spark, Airflow, ADF, dbt, or similar).
  • API-based integration (REST), JSON handling, and incremental/CDC ingestion patterns.
  • Data quality frameworks and testing practices for pipelines.
  • Version control (Git) and CI/CD for data workloads.
Preferred skills
  • RAG data preparation: chunking, embeddings, vector databases (Azure AI Search, pgvector, Pinecone, or similar).
  • Unstructured content processing: text extraction, OCR, document parsing.
  • Metadata management, data catalogs, ontologies, or knowledge graphs (for example Neptune or other graph databases).
  • Experience supporting LLM or agentic applications with grounded, traceable data.
  • Life sciences data exposure (commercial, medical, regulatory, or launch data) and regulated-data handling.
  • Cloud certification (Azure Data Engineer, Databricks, AWS, or Snowflake).

Similar Jobs

2 Days Ago
In-Office
Hyderabad, Telangana, IND
Mid level
Mid level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and maintains scalable ETL/ELT pipelines using Databricks, Snowflake, Python, Teradata, and Airflow. Builds data solutions for reporting, analytics, and AI/ML workloads; performs data quality validation, analysis, tuning, and operational monitoring. Collaborates with business and technical teams, supports deployments and issue resolution, applies CI/CD practices, and ensures data governance, security, and healthcare regulatory compliance.
Top Skills: Apache AirflowCi/CdDatabricksEltETLGitGithub CopilotPythonSnowflakeSQLTeradataUnix
17 Days Ago
Remote or Hybrid
India
Senior level
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Design and build data services and pipelines supporting machine learning products. Develop automation tools for model deployment, maintain production data systems, improve infrastructure, participate in code reviews, and collaborate across engineering, data science, and product teams. The role requires expertise in distributed systems, large-scale data processing, CI/CD, container orchestration, and AI-enabled workflow improvements.
Top Skills: AirflowAWSAws BatchCi/CdDockerEmrGlueGoKafkaKubernetesKv StoresLinuxPythonRelational DatabasesSagemakerSparkSpinnaker
19 Days Ago
Hybrid
Hyderabad, Telangana, IND
Senior level
Senior level
Financial Services
Lead data engineering for JPMorganChase’s consumer banking technology team. Design scalable batch and real-time data pipelines, data lake architectures, and end-to-end ingestion-to-analytics solutions using AWS, Spark, Glue, Iceberg, Kafka, and Flink. Build self-healing data operations with agentic AI workflows, establish data architecture standards, and promote secure AI-assisted engineering practices, governance, observability, resiliency, and automated deployment.
Top Skills: Agentic Ai FrameworksAmazon EksAmazon S3Apache FlinkApache IcebergApache KafkaSparkAWSAws GlueAws LambdaAws Step FunctionsData GovernanceData LakesData ObservabilityDataopsInfrastructure As CodeJavaPythonScalaSQL

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account