OL

Senior Data Engineer

Full time Johannesburg, ZA
Posted 1 week, 4 days ago 146 views 0 applications

Job Description

We need to output clean semantic HTML with

for paragraphs,

for section headers,
  • for list items. No CSS classes, no inline styles, no
    , no
    . Must preserve original content word-for-word, no addition/removal. Must convert the given raw HTML (with inline styles etc.) into proper semantic HTML. We need to extract the content and restructure. Given the raw job description includes many

    etc. We need to ignore those tags and just get the text content, then apply appropriate tags. We need to identify sections: Intro paragraphs (Let's Write Africa's Story Together! and Old Mutual is a firm believer...), then "Job Description" header (bold). Then the paragraph describing Senior Data Engineer responsibilities (the long paragraph). Then there is a series of sections with bold headers: Key Responsibilities, then sub-sections like Data Engineering & Development, Data Quality & Reconciliation, Solution Design & Architecture, Operational Excellence, Technical Leadership, Stakeholder & Business Engagement, Minimum Requirements (with subparts: Bachelor's degree..., Experience, Core Technical Skills, Advantageous Skills, Industry Experience, Business & Behavioural Competencies). Then Skills paragraph (list of skills). Then Competencies paragraph (list). Then Education paragraph. Then Closing Date paragraph. Then the appointment note (italic bold). Then The Old Mutual Story! (italic emphasis-2). We need to preserve all content exactly. We need to use

    for section headers (e.g., Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role). So we should treat each bold heading as an

    . For subheadings like "Data Engineering & Development" etc., they are also bold within the responsibilities section; but they are not listed as section headers in the instruction? It says use

    for section headers (e.g. Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role). It doesn't forbid using

    for other headings, but likely we should use

    for each bold heading we encounter. However the instruction: "Use

    for section headers (e.g. Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role)". It suggests we should use

    for those typical sections. But we can also use

    for other headings; it's not prohibited. Safer to use

    for each bold heading. We need to wrap paragraphs in

    . For list items, we need

    • . The original content includes many
      • already but with nested divs etc. We need to extract the list items text and wrap accordingly. We must preserve all original content word-for-word, meaning we cannot change wording, but we can change tags. So we need to output the same text but with appropriate tags. We must not add or remove any text. So we need to ensure we output exactly the same textual content, just with tags. We must not add extra spaces or newlines? Probably fine. We need to strip out all original HTML tags (

        , , ,

        ,
        , etc.) and replace with our semantic tags. Let's parse the text manually. I'll extract the textual content in order. Start:

        Let's Write Africa's Story Together!

        Text: Let's Write Africa's Story Together! Next:

        Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this.

        Text: Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this. Then several empty

        (three). They are empty, we can ignore? They contain no text, so we can omit them (since preserving content word-for-word would mean we keep empty paragraphs? The instruction: Preserve ALL original content word-for-word — do not add, remove, summarise, or rephrase. Empty paragraphs have no content, but they are part of original. If we remove them, we are removing content (the empty tags). However they have no visible text; but they are tags. The instruction likely means we should not remove any textual content; empty tags have no textual content, so removing them is okay? Safer to keep them as

        maybe. But we are not allowed to use
        etc. We could keep empty

        . However they'd be pointless. I think we can drop them as they contain no text. The instruction "Preserve ALL original content word-for-word" likely refers to visible content. I'll ignore empty paragraphs. Next:

        Job Description

        Text: Job Description Then empty p's again. Then the paragraph: "The Senior Data Engineer is responsible for designing, building, optimising and supporting scalable enterprise data solutions on the Databricks / AWS platform. The role requires strong hands-on engineering capability across PySpark, SQL, Delta Lake, Parquet, Medallion Architecture, ETL/ELT development, data ingestion and reconciliation frameworks. The successful candidate will translate complex business and data requirements into reliable, performant and maintainable data products while ensuring appropriate data quality, security, governance and operational controls.

        The role also provides technical leadership to other engineers and works closely with Data Architecture, Data Governance, Data Science, Business Intelligence and business stakeholders. Insurance or financial-services experience is strongly preferred, with the ability to understand business processes and translate them into robust data solutions." We need to preserve the text exactly, including the

        which we must not keep; we should replace with paragraph breaks? Since we cannot use
        , we should convert the double line break into a new paragraph. The original had

        indicating a paragraph break. We can treat that as a separate paragraph. So we will split into two paragraphs. Thus first paragraph up to before

        , then second paragraph after. Now after that, there are empty p's. Then we encounter a series of divs with bold headings. Let's extract the sections in order. First:

        Key Responsibilities 

        ... So heading: Key Responsibilities Then subheading: Data Engineering & Development (bold). Then a list of items. We need to capture each list item text. Let's extract each list item under Data Engineering & Development: - Design and develop scalable ETL/ELT pipelines using Databricks, PySpark and SQL. - Build batch, incremental and event-driven ingestion pipelines from multiple source systems. - Design and implement solutions using Medallion Architecture (Bronze, Silver and Gold). - Develop and optimise Delta Lake and Parquet-based data solutions. - Implement reusable engineering patterns, frameworks and common components. - Perform performance tuning and cost optimisation across Spark jobs, clusters and pipelines. - Build and manage CI / CD processes on Azure DevOps Note there is an empty
      • after that? Actually after that list there is a
      • (empty). We'll ignore empty. Next subheading: Data Quality & Reconciliation List items: - Design and implement automated data reconciliation and control frameworks. - Implement completeness, accuracy, balancing and data-quality controls across source-to-target pipelines. - Investigate data discrepancies and perform root-cause analysis. - Ensure appropriate auditability, lineage and traceability throughout the data lifecycle. Next subheading: Solution Design & Architecture List items: - Translate business requirements into scalable technical designs and data products. - Contribute to solution architecture, data standards and engineering patterns. - Evaluate existing solutions and define appropriate as-is / to-be designs. - Ensure solutions align with enterprise architecture, information security, governance and control standards. Next subheading: Operational Excellence List items: - Monitor production pipelines and proactively resolve failures, data incidents and performance issues. - Perform root-cause analysis and implement permanent remediation. - Support production releases through established change-management and deployment processes. - Develop and maintain monitoring, alerting, logging, runbooks and operational documentation. - Manage and support Databricks job schedules to ensure critical business processes run successfully and issues are resolved timeously. Next subheading: Technical Leadership List items: - Provide technical leadership and mentorship to Data Engineers. - Conduct code reviews and enforce engineering standards, testing and development best practices. - Promote reusable solutions rather than point-to-point development. - Support technical design reviews and challenge solutions where appropriate. - Drive continuous improvement across engineering practices, automation and platform efficiency. Next subheading: Stakeholder & Business Engagement List items: - Work with business and technical stakeholders to understand data and reporting requirements. - Translate insurance business requirements into appropriate data models and engineering solutions. - Communicate technical issues, dependencies, risks and delivery status clearly to stakeholders. - Partner with architecture, governance, security, operations, finance and business teams. Next heading: Minimum Requirements Then subparts: - Bachelor's degree in Computer Science, Computer Engineering, Information Technology, Data Science or a related discipline. Relevant Databricks and cloud certifications are advantageous. Then heading: Experience Text: Ideally 5+ years of Data Engineering / ETL development experience, with demonstrated experience delivering enterprise-scale data solutions and operating production data pipelines. Then heading: Core Technical Skills Text: Advanced PySpark and SQL; strong Databricks experience; Delta Lake and Parquet; Medallion Architecture; ETL/ELT and ingestion pipelines; data modelling; reconciliation and data-quality frameworks; Git/Azure DevOps or equivalent CI/CD; cloud data platforms, preferably AWS; performance optimisation; and production troubleshooting.Advantageous Skills Note there is a weird concatenation: "...production troubleshooting.Advantageous Skills". The original text had a missing space. We must preserve exactly as is. Then heading: Advantageous Skills (actually it's part of same line? Let's see: after Core Technical Skills paragraph, there is a

        Databricks Workflows/Jobs, Unity Catalog, Auto Loader, CDC or similar incremental ingestion patterns, orchestration frameworks, CI/CD automation, AWS services, and exposure to AI/ML data workloads. 

        . So that is Advantageous Skills content. Then heading: Industry Experience Text: Insurance industry experience is highly advantageous, particularly exposure to policy, claims, premium, finance / GL, actuarial or regulatory data. Then heading: Business & Behavioural Competencies Then list items: - Strong problem-solving and analytical capability. - Engineering discipline, ownership and accountability. - Ability to work independently in a complex enterprise environment. - Strong stakeholder management and communication skills. - Ability to coach and uplift the capability of less experienced engineers. - Commercial awareness and a focus on sustainable, cost-effective engineering solutions. Then after that, there is a

         

        (non-breaking space) maybe empty. Then heading: Skills Text: Analysis of Alternatives (AoA), Analytical Processes, Analytical Sciences, Assessment Testing, Business Intelligence (BI) Analysis, Categorizing Data, Computer Literacy, Data Analysis, Database Reporting, Data Compilation, Data Controls, Data Interpretations, Financial Modeling, Information Retrieval, Managerial Accounting, Numerical Aptitude, Planning and organisational skills, Research Methodologies, Statistical Analysis Techniques Then heading: Competencies Then list of competencies each on separate lines with

        tags in between? The original shows:

        Competencies

        Business Insight

        Ensures Accountability

        Financial Acumen

        Manages Complexity

        Optimizes Work Processes

        Resourcefulness

        Self-Development

        Tech Savvy

        ... So the text appears as: Business InsightEnsures AccountabilityFinancial AcumenManages ComplexityOptimizes Work ProcessesResourcefulnessSelf-DevelopmentTech Savvy with

        separators. We need to preserve the text exactly, meaning we should output those words exactly as they appear, without the

        tags. The

        tags are empty, so they contribute no text. So we can just output the concatenated words? But there is no spacing between them in the original because they are separated by empty

        tags. The original text likely intended each competency on its own line but due to formatting they appear concatenated. We must preserve the exact string: "Business InsightEnsures AccountabilityFinancial AcumenManages ComplexityOptimizes Work ProcessesResourcefulnessSelf-DevelopmentTech Savvy". However there may be missing spaces. Let's examine the snippet: "Business Insight

        Ensures Accountability

        Financial Acumen

        Manages Complexity

        Optimizes Work Processes

        Resourcefulness

        Self-Development

        Tech Savvy" Thus between each competency there is "

        " which is empty, so the visible text is just the concatenation without spaces. So we must output exactly that string. Then heading: Education Text: Bachelor of Commerce (BCom): Computer Science And Engineering (Required), NQF Level 7 - Degree, Advance Diploma or Postgraduate Certificate or equivalent (Required) Note double spaces before "(Required)"? We'll preserve. Then heading: Closing Date Text: 01 October 2026 , 23:59 Note space before comma? We'll preserve. Then the appointment note:

        The appointment will be made from the designated group in line with the Employment Equity Plan of Old Mutual South Africa and the specific business unit in question.

        Text: The appointment will be made from the designated group in line with the Employment Equity Plan of Old Mutual South Africa and the specific business unit in question. Then empty p's. Then final:

        The Old Mutual Story!

        Text: The Old Mutual Story! Now we need to output HTML with

        for headings,

        for paragraphs,

        • for lists. We must not add extra text. So we need to output exactly the textual content as above, wrapped in appropriate tags. We need to decide which headings get

          . According to instruction, use

          for section headers (e.g. Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role). So we should use

          for those specific headings. For other headings like "Key Responsibilities", "Data Engineering & Development", etc., maybe we should not use

          because they are not in the list. But the instruction says "Use

          for section headers (e.g. Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role)". It doesn't forbid using

          for other headings, but it's implied that we should use

          for those typical sections. However we have many headings; we could still use

          for all bold headings, but that might be considered adding extra formatting not requested? The instruction says "Use

          for section headers". It doesn't say "only for those". It gives examples. I think it's safe to use

          for each bold heading we encounter, as they are section headers. But we must also preserve the original content word-for-word. Changing a heading from bold to

          does not change the text, just the tag. That's allowed. Thus we will wrap each bold heading text in

          tags. Now for lists: we need to wrap each list item in
        • inside a
            . The original had multiple
            • but we can consolidate each group under its heading? The instruction: Use
              • for list items (duties, requirements, qualifications, skills). So we should produce a
                  for each group of list items. We need to decide grouping: Under each subheading (like Data Engineering & Development) we have a list. We'll create a
                    with its items. Similarly for Data Quality & Reconciliation, Solution Design & Architecture, Operational Excellence, Technical Leadership, Stakeholder & Business Engagement, Business & Behavioural Competencies. Also the Skills section is a paragraph of comma-separated skills, not a list. So we keep as

                    . Competencies: we have a list of competencies but they are not formatted as list items originally; they are just concatenated words. Should we treat them as list? The instruction says use

                    • for list items (duties, requirements, qualifications, skills). Competencies could be considered similar to skills, but they are not presented as list. Safer to keep as

                      preserving the exact concatenated string. Education: paragraph. Closing Date: paragraph. Appointment note: paragraph (italic bold). We should preserve italic and bold? The instruction says preserve all original content word-for-word, but we cannot keep or tags because we are only allowed to use

                      ,

                      ,
                      • . No other tags. So we must strip those tags but keep the text. The text inside is the same; we lose the italic/bold styling but that's okay as we cannot use those tags. The instruction says preserve ALL original content word-for-word — do not add, remove, summarise, or rephrase. Removing <

Apply Now ↗

How well do you match?

Get an instant AI match score for this role — free, takes 3 minutes.

Tailor your CV for this role

The concierge rewrites your whole CV and writes a matching cover letter for this job — opens right here, nothing to paste.

Tailor My CV to This Job ✍️

Free cover letter for this job

Upload your CV and get a tailored cover letter in seconds — free, no account needed.

Generate a Cover Letter 📝
MJC
ECHO
Your MJC Assistant

I'm ECHO, your MJC career assistant. I can help you find jobs, explore career tools, and connect with opportunities across Africa.

How was your experience with ECHO?