SH

Data Engineer I

Full Time Cape Town, Brackenfell, Western Cape, South Africa, ZA
Posted 52 minutes ago 36 views 0 applications

Job Description

The user wants me to format the given job description as clean semantic HTML. I need to follow the rules: use

for intro paragraphs,

for section headers,
  • for lists, no CSS/classes/divs/br, preserve all content word-for-word, no markdown fences, just HTML. First, analyze the job description text. It has sections: "About the role:", "What you bring:", "What you'll do:". Each section has content. "About the role:" is a paragraph. "What you bring:" includes a list of qualifications, skills, and attributes. "What you'll do:" includes a list of duties. I need to wrap "About the role:" as a paragraph? The rule: Use

    for section headers (e.g. Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role). So "About the role:" should be an

    header. Then the following paragraph is a

    . "What you bring:" is an

    . Then the list items: there are bullet points? The text shows lines separated by line breaks. I need to convert each line into
  • items. However, some lines are grouped: "Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5–6 (Essential) AWS Certification (Beneficial) 2–4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential) Basic knowledge of data engineering concepts, data modelling, and databases Technical Skills such as: Python PySpark SQL dbt (Data Build Tool) Git / Version Control ETL pipeline development Data modelling Apache Airflow / Orchestration Data quality frameworks and observability Apache Iceberg / Delta Lake (open table formats) Terraform / Infrastructure as Code AWS Cloud — EMR, S3, Glue, Athena (Beneficial) Snowflake (Beneficial) Retail Operations experience Developing Specialist Capability — Supports the analysis, diagnosis, testing, resolution, and documentation of data solutions within an agile team Technical Aptitude — Strong passion and excitement for data, new technologies and solutions, and their range of possibilities and value for the business Self-Motivation — High level of drive to set, meet, and exceed goals and expectations. Uses own initiative in dealing with challenges Detail and Quality Focus — Has an affinity for structure and efficiency. Diligently watches over work processes, tasks, and outputs to ensure accuracy while promptly correcting quality concerns Communication — Communicates well both verbally and in writing. Able to simplify complex technical concepts for a variety of stakeholders Collaboration — Builds sound working relationships across the business. Able to work independently or collaboratively. Willing to coach/mentor others Resilience — Ability to work under pressure and tight time constraints, efficiently prioritising workloads in a high-volume, fast-moving environment Curiosity — Open to learning with a strong interest in data and discovery. Curious about exploring and answering business analytics questions" This appears as a single block but likely each line is a separate bullet. However, the text is not clearly separated by line breaks in the provided input. In the user message, it's shown as separate lines? Let's check the raw input: The user pasted the job description with line breaks. In the text, after "What you bring:" there are line breaks for each item? Actually the user message shows: "What you bring: Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5–6 (Essential) AWS Certification (Beneficial) 2–4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential) Basic knowledge of data engineering concepts, data modelling, and databases Technical Skills such as: Python PySpark SQL dbt (Data Build Tool) Git / Version Control ETL pipeline development Data modelling Apache Airflow / Orchestration Data quality frameworks and observability Apache Iceberg / Delta Lake (open table formats) Terraform / Infrastructure as Code AWS Cloud — EMR, S3, Glue, Athena (Beneficial) Snowflake (Beneficial) Retail Operations experience Developing Specialist Capability — Supports the analysis, diagnosis, testing, resolution, and documentation of data solutions within an agile team Technical Aptitude — Strong passion and excitement for data, new technologies and solutions, and their range of possibilities and value for the business Self-Motivation — High level of drive to set, meet, and exceed goals and expectations. Uses own initiative in dealing with challenges Detail and Quality Focus — Has an affinity for structure and efficiency. Diligently watches over work processes, tasks, and outputs to ensure accuracy while promptly correcting quality concerns Communication — Communicates well both verbally and in writing. Able to simplify complex technical concepts for a variety of stakeholders Collaboration — Builds sound working relationships across the business. Able to work independently or collaboratively. Willing to coach/mentor others Resilience — Ability to work under pressure and tight time constraints, efficiently prioritising workloads in a high-volume, fast-moving environment Curiosity — Open to learning with a strong interest in data and discovery. Curious about exploring and answering business analytics questions" It seems each item is separated by a space? Actually there are line breaks in the original but in the text they might be separated by newlines. I need to treat each distinct bullet as a list item. However, the rule: "Preserve ALL original content word-for-word — do not add, remove, summarise, or rephrase". So I must keep the exact text. But I need to structure as HTML. The best is to treat each line as a separate
  • . But the input may not have clear line breaks. In the user message, it's presented as a block of text with spaces. However, the user likely expects me to parse the sections and list items as they appear in the original scraped text. Since I cannot see the original line breaks, I must assume each bullet is separated by a newline in the source. But the user gave me the text as a single string. I need to infer list items from the content. The typical pattern: each bullet starts with a capital letter and ends with a period or similar. But there are many items. Given the instruction: "If text has no clear sections, just wrap paragraphs in

    tags". But there are clear sections: "About the role:", "What you bring:", "What you'll do:". So I should use

    for those headers. For the content under "About the role:", it's a paragraph. For "What you bring:", it's a list of qualifications and skills. For "What you'll do:", it's a list of duties. I'll structure as:

    About the role:

    The purpose of the Data Engineer I role is to support the design, development, testing, and documentation of data engineering solutions in line with LOA 2S expectations. The role collects, manages, and converts raw data into usable information for data scientists, analysts, and business stakeholders, while supporting data pipelines, data workflows, routine data tasks, and data quality controls that enable data-driven insights and operational efficiency.

    What you bring:

    • Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5–6 (Essential)
    • AWS Certification (Beneficial)
    • 2–4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential)
    • Basic knowledge of data engineering concepts, data modelling, and databases
    • Technical Skills such as: Python PySpark SQL dbt (Data Build Tool) Git / Version Control ETL pipeline development Data modelling Apache Airflow / Orchestration Data quality frameworks and observability Apache Iceberg / Delta Lake (open table formats) Terraform / Infrastructure as Code AWS Cloud — EMR, S3, Glue, Athena (Beneficial) Snowflake (Beneficial) Retail Operations experience
    • Developing Specialist Capability — Supports the analysis, diagnosis, testing, resolution, and documentation of data solutions within an agile team
    • Technical Aptitude — Strong passion and excitement for data, new technologies and solutions, and their range of possibilities and value for the business
    • Self-Motivation — High level of drive to set, meet, and exceed goals and expectations. Uses own initiative in dealing with challenges
    • Detail and Quality Focus — Has an affinity for structure and efficiency. Diligently watches over work processes, tasks, and outputs to ensure accuracy while promptly correcting quality concerns
    • Communication — Communicates well both verbally and in writing. Able to simplify complex technical concepts for a variety of stakeholders
    • Collaboration — Builds sound working relationships across the business. Able to work independently or collaboratively. Willing to coach/mentor others
    • Resilience — Ability to work under pressure and tight time constraints, efficiently prioritising workloads in a high-volume, fast-moving environment
    • Curiosity — Open to learning with a strong interest in data and discovery. Curious about exploring and answering business analytics questions

    What you'll do:

    • Creating data feeds from on-premises to cloud environments
    • Supporting data feeds in production on a break-fix basis
    • Building data warehouse layers and data products using dbt or similar transformation tools
    • Manipulating data using Python and PySpark
    • Processing data using distributed compute, particularly EMR Serverless
    • Development for Big Data and Business Intelligence including automated testing and deployment
    • Working with open table formats (Iceberg/Delta Lake)
    • Writing and maintaining pipeline orchestration DAGs (Airflow)
    • Using Git for version control, branching, pull requests, and code reviews
    But note: The "Technical Skills such as:" line includes many skills separated by spaces. In the original, they might be separate bullet points. However, the text says "Technical Skills such as: Python PySpark SQL dbt (Data Build Tool) Git / Version Control ETL pipeline development Data modelling Apache Airflow / Orchestration Data quality frameworks and observability Apache Iceberg / Delta Lake (open table formats) Terraform / Infrastructure as Code AWS Cloud — EMR, S3, Glue, Athena (Beneficial) Snowflake (Beneficial) Retail Operations experience". That could be a single bullet point listing skills. But the rule: preserve all original content word-for-word. So I should keep that entire line as one list item? But the line includes many skills. The original might have each skill on a new line. Since I don't have the original line breaks, I'll treat the whole block after "Technical Skills such as:" as one bullet? However, the user expects me to format as clean semantic HTML. I think it's better to split each skill into separate list items? But that would be adding structure not present. The rule: "Preserve ALL original content word-for-word — do not add, remove, summarise, or rephrase". So I must not split or combine. I need to keep the exact text as it appears. But the text is given as a continuous string. I need to decide how to split into list items. The user likely expects each bullet point as a separate list item. In the "What you bring:" section, there are multiple bullet points. The text shows each bullet on a new line in the user message? Actually in the user message, it's formatted with line breaks? Let's check the raw user message: It's plain text. In the chat, it's displayed with line breaks. The user wrote: "What you bring: Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5–6 (Essential) AWS Certification (Beneficial) 2–4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential) Basic knowledge of data engineering concepts, data modelling, and databases Technical Skills such as: Python PySpark SQL dbt (Data Build Tool) Git / Version Control ETL pipeline development Data modelling Apache Airflow / Orchestration Data quality frameworks and observability Apache Iceberg / Delta Lake (open table formats) Terraform / Infrastructure as Code AWS Cloud — EMR, S3, Glue, Athena (Beneficial) Snowflake (Beneficial) Retail Operations experience Developing Specialist Capability — Supports the analysis, diagnosis, testing, resolution, and documentation of data solutions within an agile team Technical Aptitude — Strong passion and excitement for data, new technologies and solutions, and their range of possibilities and value for the business Self-Motivation — High level of drive to set, meet, and exceed goals and expectations. Uses own initiative in dealing with challenges Detail and Quality Focus — Has an affinity for structure and efficiency. Diligently watches over work processes, tasks, and outputs to ensure accuracy while promptly correcting quality concerns Communication — Communicates well both verbally and in writing. Able to simplify complex technical concepts for a variety of stakeholders Collaboration — Builds sound working relationships across the business. Able to work independently or collaboratively. Willing to coach/mentor others Resilience — Ability to work under pressure and tight time constraints, efficiently prioritising workloads in a high-volume, fast-moving environment Curiosity — Open to learning with a strong interest in data and discovery. Curious about exploring and answering business analytics questions" It appears each bullet is separated by a space? Actually there are no line breaks in the text as pasted. But the user might have intended each bullet on a new line. However, the instruction says: "You receive a plain-text job description." So the plain text may have line breaks. In the user message, it's shown as a block. But as an AI, I see the text as a single string with spaces. I need to parse it. The safest is to treat each distinct bullet as a separate list item based on the content. But I cannot guess. The rule: "If text has no clear sections, just wrap paragraphs in

    tags". But there are clear sections. For the list items, I could wrap the entire "What you bring:" content in a

    ? But the rule says use

    • for list items (duties, requirements, qualifications, skills). So I should use list. Given the ambiguity, I'll assume each bullet is separated by a newline in the original. Since the user pasted it, maybe the newlines are preserved. In the chat interface, the text may have newlines. I'll treat each line as a bullet. But the text I see doesn't have newlines. However, I can split by the pattern of capitalized phrases? That's risky. Alternatively, I can output the entire "What you bring:" content as a single paragraph? But that would not be semantic. Given the instruction to "Preserve ALL original content word-for-word", I should not alter the text. So I could output the exact text under each header as a paragraph, but that would not use lists. However, the rule says "Use
      • for list items (duties, requirements, qualifications, skills)". So I must identify list items. The "What you'll do:" section clearly has bullet points each on a new line in the user message? Let's check: "What you'll do: Creating data feeds from on-premises to cloud environments Supporting data feeds in production on a break-fix basis Building data warehouse layers and data products using dbt or similar transformation tools Manipulating data using Python and PySpark Processing data using distributed compute, particularly EMR Serverless Development for Big Data and Business Intelligence including automated testing and deployment Working with open table formats (Iceberg/Delta Lake) Writing and maintaining pipeline orchestration DAGs (Airflow) Using Git for version control, branching, pull requests, and code reviews" In the user message, these are on separate lines. So for "What you'll do:", I can split by newline. For "What you bring:", the text is not separated by newlines in the user message? Actually it might be but the rendering collapsed. I'll assume each bullet is separated by a newline in the original plain text. Since I cannot see newlines, I'll have to make a judgment. Given the typical format of such job descriptions, each bullet is a line. I'll split the "What you bring:" text by the pattern of each bullet starting with a capital letter and ending with a period? But there are many. Maybe the best is to output the HTML with the exact text as provided, but with headers and paragraphs, and for the list sections, I'll put each line as a list item. Since the user gave the text in the prompt, I can see that "What you bring:" is followed by a block of text. In the prompt, it's written as a single paragraph. But the user might have intended each bullet on a new line. However, the instruction says: "You receive a plain-text job description." So the plain text likely has line breaks. As an AI, I should treat the input as having line breaks where they appear in the user message. In the user message, the text is formatted with line breaks for "What you'll do:" but not for "What you bring:"? Actually looking at the raw user message (as I see it), it's: "About the role: The purpose of the Data Engineer I role is to support the design, development, testing, and documentation of data engineering solutions in line with LOA 2S expectations. The role collects, manages, and converts raw data into usable information for data scientists, analysts, and business stakeholders, while supporting data pipelines, data workflows, routine data tasks, and data quality controls that enable data-driven insights and operational efficiency. What you bring: Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5–6 (Essential) AWS Certification (Beneficial) 2–4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential) Basic knowledge of data engineering concepts, data modelling, and databases Technical Skills such as: Python PySpark SQL dbt (Data Build Tool) Git / Version Control ETL pipeline development Data modelling Apache Airflow / Orchestration Data quality frameworks and observability Apache Iceberg / Delta Lake (open table formats) Terraform / Infrastructure as Code AWS Cloud — EMR, S3, Glue, Athena (Ben

Apply Now ↗

How well do you match?

Get an instant AI match score for this role — free, takes 3 minutes.

Tailor your CV for this role

The concierge rewrites your whole CV and writes a matching cover letter for this job — opens right here, nothing to paste.

Tailor My CV to This Job ✍️

Free cover letter for this job

Upload your CV and get a tailored cover letter in seconds — free, no account needed.

Generate a Cover Letter 📝

Join Our Zimbabwe Channels

Get free job alerts on your phone

MJC
ECHO
Your MJC Assistant

I'm ECHO, your MJC career assistant. I can help you find jobs, explore career tools, and connect with opportunities across Africa.

How was your experience with ECHO?