Padmi

Data Engineer

Dallas–Fort WorthPosted 1 month ago
Software engineeringUnspecified
Apply at Thoughtwave Software and Solutions

Opens the source posting on atsapp.swarmhr.com

Source description

About the role

View original

Role : Data Engineer

Location : Dallas (1st choice ), Chicago , NYC – Onsite Term: 4-6 months.

We are seeking an experienced Data Engineer to design, build, and maintain high-quality, scalable data pipelines. This role requires not just technical expertise but also a strong analytical mindset and a deep sense of ownership. The ideal candidate thrives on solving complex data problems, understands both structured and unstructured data sources, and can take pipelines from concept through production with accountability for performance, quality, and reliability.

Responsibilities

Design, build, and optimize robust, scalable pipelines in Databricks (PySpark, SQL, Delta Lake) for structured, semi-structured, and unstructured data. Ingest and process data from diverse sources: relational databases, APIs, PDFs, Excel, flat files, and web scraping. Implement data quality frameworks, validation checks, and automated tests to ensure reliability across the pipeline lifecycle. Conduct UAT/UAD processes to align outputs with business requirements and ensure trust in data. Own pipelines end-to-end—from development through deployment, monitoring, scaling, and continuous improvement. Apply best practices for dataset versioning, reproducibility, and lineage (Delta Lake, MLflow, or equivalent). Optimize pipelines for large-scale performance and cost-efficiency in Databricks and AWS environments. Collaborate with data science and analytics teams to deliver model-ready datasets and ensure schema stability. Produce clear technical documentation (design docs, lineage diagrams, data dictionaries) and communicate effectively with technical and business stakeholders.

Required Qualifications

5+ years of professional experience in Data Engineering or related roles. Hands-on expertise in Databricks (PySpark, SQL, Delta Lake). Proven track record of building and deploying end-to-end production pipelines at scale. Experience working with structured and unstructured data, including extracting information from PDFs, Excel, web scraping, and APIs. Strong foundation in data quality, testing, and observability (e.g., Great Expectations, Deequ, dbt tests, or similar). Ability to troublshoot, debug, and optimize Spark/Databricks workloads for scale and performance.

Familiarity with dataset versioning and reproducibility (Delta Lake time travel, MLflow, or similar). Excellent problem-solving skills with the ability to analyze complex data issues and implement long-term fixes. Strong written and verbal communication skills with a focus on clear documentation and stakeholder alignment.

Preferred Qualifications

Experience with AWS services (S3, Glue, Redshift, Lambda, Step Functions). Exposure to MLOps and ML pipeline integration, including data preparation for feature stores and model training. Familiarity with CI/CD practices for data pipelines (GitHub Actions, Jenkins, Azure DevOps, etc.). Understanding of data governance, compliance, and security best practices. Exposure to AI/LLM-based data processing (text extraction, embeddings, entity recognition, summarization). Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate, pgvector) and retrieval pipelines. Awareness of AI-driven automation tools (LangChain, LangGraph, MCP) for validation, metadata extraction, or observability. Interest in responsible AI/data governance practices such as bias checks, explainability, and auditability.

Soft Skills

Strong analytical and critical-thinking skills, able to go beyond surface-level issues. Ownership mentality—responsible for pipeline correctness, performance, and usability. Ability to balance independent execution with collaborative teamwork across data, engineering, and business stakeholders.

One address, no account. We’ll tell you when matching roles go live.

More at Thoughtwave Software and Solutions

Related open roles

View all roles