Hiring Resources

Data Engineers Job Description Template (Copy-Paste Ready)

Brian Hunt
Brian Hunt
CEO & Founder, Kore BPO
August 31, 2026 9 min read Reviewed 2026
Hiring manager reviewing a data engineer job description document on a laptop at a modern desk with an orange coffee mug
Quick Answer
What should a data engineer job description include?

A strong data engineer job description must specify your pipeline stack (Python, dbt, Airflow, Spark, or similar), your cloud data warehouse (Snowflake, BigQuery, or Redshift), required experience level, data quality and orchestration ownership, and whether streaming or batch pipelines are in scope. Generic JDs that list every data tool attract unfocused candidates and slow screening. Stack-specific JDs attract engineers who match your environment and can contribute within their first sprint.

JDs that name specific tools (dbt, Airflow, Snowflake) receive 40% more qualified applications than generic ones
Data quality ownership requirements filter for senior engineers who have operated production pipelines
Compensation transparency in JDs reduces time-to-offer by removing mismatched salary expectations late in the process
See hiring steps at our full hiring guide

Data engineer job descriptions fail before a single candidate reads them when they list 20 tools, confuse data engineering with data analysis, and give no clarity on whether the role is batch or streaming, Snowflake or Databricks, greenfield or maintenance-heavy. The result is an unfocused applicant pool that takes weeks to screen down to viable candidates.

A well-written data engineer JD does three things: it specifies your exact stack so candidates self-select accurately, it communicates what ownership actually looks like in this role, and it signals the engineering culture in enough detail that qualified engineers want to apply. This guide breaks down what to include and gives you a copy-paste template ready for your next search.

What Makes a Strong Data Engineer JD

The best data engineering JDs share four characteristics. First, they name specific tools rather than generic categories. "Experience with data orchestration tools" is useless signal. "3+ years of production Apache Airflow experience including DAG design, task dependency management, and MWAA operations" is precise enough to be useful. Second, they distinguish between the transformation layer and the infrastructure layer. Are you hiring someone to own dbt models and data quality, or someone to build and maintain the Spark cluster and data lake architecture? These are different roles. Third, they specify the cloud platform and version where relevant. Snowflake and BigQuery have meaningfully different SQL dialects, cost models, and optimization levers. Fourth, they communicate whether this role is greenfield or legacy-heavy. Engineers optimize for different things depending on whether they are building from scratch or maintaining existing production systems.

Length and Format

Effective data engineer JDs run 400 to 600 words. Long enough to be specific, short enough to be read. Structure them with three sections: About the Role (2-3 sentences on the team and context), Responsibilities (8-12 bullet points), and Requirements (split into must-have and nice-to-have). Avoid "About Us" sections longer than 3 sentences. Engineers read those last and they rarely influence apply decisions.

Team reviewing data engineering job requirements on a whiteboard with dbt and Snowflake diagrams

Full JD Template (Copy-Paste)

Customize the bracketed sections for your stack and context. Remove requirements that do not apply and add any platform-specific tools your team uses.

Data Engineer

Location: [Remote / Nearshore Latin America preferred / Hybrid]
Type: Full-Time
Team: [Data Platform / Analytics Engineering / Data Infrastructure]

About the Role
We are looking for a Data Engineer to join our [Data Platform / Analytics] team and own the pipelines that move data from [source systems] through our [Snowflake / BigQuery / Redshift] data warehouse to [BI tools / ML feature store / downstream consumers]. You will design and maintain production-grade data pipelines, ensure data quality at every layer, and partner with analytics engineers and data scientists to support reliable, well-modeled data across the organization.

Responsibilities

  • Design, build, and maintain batch and/or streaming data pipelines using Python and [Airflow / Prefect / Dagster]
  • Own and extend our dbt project, including model design, testing, documentation, and performance optimization
  • Develop and maintain ELT/ETL processes ingesting data from [APIs / SaaS tools / operational databases] into [Snowflake / BigQuery / Redshift]
  • Define and enforce data quality standards: write dbt tests, implement freshness checks, and build monitoring and alerting for pipeline failures and data anomalies
  • Optimize query performance and manage compute costs in [Snowflake / BigQuery / Redshift] through partitioning, clustering, and query tuning
  • Collaborate with analytics engineers and data scientists to understand data requirements and translate them into reliable, well-documented data models
  • Maintain and improve data lineage documentation and column-level data dictionaries
  • Participate in on-call rotation for data pipeline incidents [if applicable]

Requirements

  • 4+ years of experience in a data engineering role with production pipeline ownership
  • Strong Python proficiency for data pipeline development (pandas, PySpark, or similar)
  • Advanced SQL skills including window functions, CTEs, and query optimization
  • Hands-on experience with dbt Core or dbt Cloud (model design, testing, macros)
  • Experience with [Apache Airflow / Prefect / Dagster] for pipeline orchestration
  • Proficiency with [Snowflake / BigQuery / Redshift / Databricks] as the primary data warehouse
  • Experience with cloud infrastructure on [AWS / GCP / Azure]
  • Strong communication skills in English for cross-functional collaboration

Nice-to-Have

  • Experience with streaming pipelines (Kafka, Flink, or Spark Streaming)
  • Familiarity with data observability tools (Monte Carlo, Bigeye, or Great Expectations)
  • Exposure to MLOps or feature store pipelines (Feast, Tecton)
  • Experience with data contracts or schema registry tooling
  • Contributions to open source data tooling

Compensation
[$XX,000 to $XX,000 annually] depending on experience. Full benefits, 90-day replacement guarantee through our staffing partner.

Need Candidates Who Match This JD?

Share your JD with Kore BPO and we will have pre-screened nearshore data engineers matching your stack within 72 hours.

GET STARTED

Responsibilities Breakdown

Not every responsibility applies to every data engineering role. Use this breakdown to select what is actually in scope for your hire rather than listing everything and inflating the JD.

Pipeline Development and Ownership

This is the core of most data engineering roles: designing, building, and maintaining the pipelines that move and transform data. Be specific about whether your pipelines are batch (daily or hourly Airflow DAGs), micro-batch (10-minute intervals), or streaming (sub-second Kafka consumers). The operational model is completely different across these three paradigms and candidates who have only done batch work will struggle immediately in a streaming environment.

Transformation Layer (dbt)

If dbt is central to your stack, dedicate at least two responsibility bullets to it. Most data teams want their data engineer to both write and maintain dbt models and to enforce modeling standards (grain definitions, naming conventions, test coverage) across the project. If your dbt project has technical debt, say so. "Refactor and document existing dbt models" is a realistic and honest responsibility that senior engineers will actually appreciate because it signals you have room for impactful work.

Data Quality and Observability

Data quality ownership is what separates senior data engineers from pipeline builders. If you want an engineer who will own the reliability of your data, not just the uptime of your pipelines, include explicit responsibility bullets for data quality: writing tests, defining freshness SLAs, building alerting for anomalies, and communicating data incidents to business stakeholders. Candidates who have owned this in production will recognize it immediately and self-select in.

Data engineer reviewing pipeline test results and data quality dashboards on a large monitor

Requirements: Must-Have vs. Nice-to-Have

Splitting requirements into must-have and nice-to-have is one of the highest-leverage things you can do in a data engineering JD. It immediately narrows the screening pool, reduces time spent on candidates who are close but not quite right, and communicates to senior candidates that you have thought carefully about what this role actually needs.

Must-Have: Production Experience, Not Just Familiarity

Every requirement in the must-have list should have "in production" implied. "Familiarity with dbt" is a nice-to-have. "Hands-on experience with dbt Core or dbt Cloud including model design, macros, and production deployment" is a must-have for a role where dbt ownership is central. The distinction matters because candidates with only tutorial or side-project experience with a tool will take significantly longer to be productive in a role where it is a daily responsibility.

Nice-to-Have: Adjacent Skills That Accelerate Ramp

Nice-to-have requirements should represent skills that make a candidate more immediately productive or valuable in your context without being blocking. Streaming experience is a legitimate nice-to-have for a predominantly batch-focused team that has plans to add streaming later. MLOps familiarity is a nice-to-have for a team that collaborates closely with data scientists but does not own the ML infrastructure. Keep the nice-to-have list short: more than five items signals you are not clear on what you actually need.

Adapting for Nearshore Candidates

A nearshore-adapted data engineer JD makes two additions. First, a clear statement on working hours: "This role works US business hours (Eastern or Central Time). Candidates in Latin America time zones are strongly preferred." This removes ambiguity for candidates and ensures the staffing partner filters for geographic fit. Second, an explicit English fluency requirement with the expected collaboration contexts: "Strong written and spoken English required for daily collaboration with US-based data analysts, analytics engineers, and business stakeholders." Data engineers in nearshore roles spend significant time in stakeholder-facing communication, not just heads-down pipeline work, and English fluency at the collaboration level needs to be stated explicitly.

Avoid adding requirements that are not actually relevant to nearshore data engineering work. Certifications, US degree requirements, and in-person availability requirements are common additions that reduce the qualified pool without improving hire quality for remote roles. If certifications are genuinely required for compliance or client-facing contexts, include them in the must-have section with an explanation.

Common JD Mistakes

Listing every data tool your company has ever used. A JD that requires Spark, dbt, Airflow, Kafka, Flink, Databricks, Snowflake, BigQuery, Redshift, and Python proficiency in one role is not describing a real position. It signals that you have not thought carefully about the role and will attract candidates who exaggerate their experience across all of them. Pick the tools that are actually central to day-one responsibilities.

Confusing data engineering with data science or analytics. Data engineering is about pipelines, orchestration, data quality, and modeling infrastructure. If your JD asks for machine learning experience, statistical analysis, or data visualization skills as requirements for a data engineering role, you are either hiring two roles at once or you are not clear on which role you need. This confusion leads to mismatched expectations and early attrition.

Omitting compensation range. Senior data engineers in Latin America know what market rates look like. A JD without a compensation range will generate applications from candidates across a wide salary spectrum, and mismatches discovered late in the process waste both parties' time. Include a range. It improves application quality and reduces time-to-offer.

Writing responsibilities as requirements. "Must have experience building dbt models" belongs in requirements. "Design and maintain our dbt models" belongs in responsibilities. Mixing them creates a confusing read and makes it harder for candidates to assess fit quickly.

Frequently Asked Questions

Should I require a specific number of years of experience?

Years of experience requirements are useful as a rough filter but should not be the primary screen. A candidate with 3 years of deep, production-accountable data engineering experience will typically outperform one with 7 years of peripheral data work. Use years as a minimum threshold to reduce the pool, but weight the quality and scope of production experience more heavily in your actual screening. For senior data engineering roles, we recommend 4+ years as the minimum, with an emphasis on production pipeline ownership rather than total years in adjacent roles.

How specific should I be about the data warehouse platform?

As specific as possible. Snowflake, BigQuery, and Redshift have different SQL dialects, different cost models, and different optimization approaches. An engineer who has done deep Snowflake work (resource monitors, micro-partitions, zero-copy cloning, dynamic data masking) will have meaningful ramp-up time in a BigQuery environment. If you are committed to one platform, name it in the requirements. If you are platform-agnostic, list your primary platform as required and others as nice-to-have.

Do I need to list dbt and orchestration separately?

Yes, and they are distinct enough to warrant separate bullet points. dbt expertise is about SQL modeling, testing, and documentation. Orchestration expertise (Airflow, Prefect, Dagster) is about dependency management, scheduling, retries, alerting, and pipeline reliability. Some data engineers are deep in one and weak in the other. Separate requirements allow candidates to self-assess accurately and help your screener ask targeted questions.

Should I include streaming in the JD if it is only occasional work?

Put it in nice-to-have if it is occasional or aspirational. Streaming pipelines (Kafka, Flink, Spark Streaming) require meaningfully different skills from batch pipelines, and listing streaming as a must-have when it is only 10% of the role will exclude well-qualified batch-focused engineers and attract streaming specialists who may be bored by predominantly batch work. If streaming is a significant part of the role today, list it in must-have. If it is a future roadmap item, put it in nice-to-have with a brief note.

Can a staffing partner use this JD to source nearshore candidates?

Yes. Kore BPO takes your JD as the primary sourcing brief. We filter our bench and active candidate pipeline against your must-have requirements and present 2-3 profiles that match within 72 hours. The more specific your JD, the faster and more accurate our matching. Bring the JD template from this page with your stack details filled in to your discovery call and we can begin sourcing immediately.

Brian Hunt
Brian Hunt
CEO & Founder, Kore BPO

Brian Hunt is the CEO and Founder of Kore BPO, a US-owned nearshore and offshore staffing firm headquartered in Dallas. He has spent over two decades building and scaling distributed engineering teams for US companies across Latin America and Southeast Asia.

HIRE YOUR NEARSHORE DATA ENGINEER

Get pre-screened candidates from Latin America on your desk within 72 hours. 90-day replacement guarantee on every placement.

GET STARTED TODAY

No upfront fees  |  90-day replacement guarantee