Data Engineer I at Shoprite Group of Companies
Not Specified NewBookmark Details
About the role
The purpose of the Data Engineer I role is to support the design, development, testing, and documentation of data engineering solutions in line with LOA 2S expectations.
The role collects, manages, and converts raw data into usable information for data scientists, analysts, and business stakeholders, while supporting data pipelines, data workflows, routine data tasks, and data quality controls that enable data-driven insights and operational efficiency.
What you bring
Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5 – 6 (Essential)
AWS Certification (Beneficial)
2 – 4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential)
Basic knowledge of data engineering concepts, data modelling, and databases
Technical Skills such as:
Python
PySpark
SQL
dbt (Data Build Tool)
Git / Version Control
ETL pipeline development
Data modelling
Apache Airflow / Orchestration
Data quality frameworks and observability
Apache Iceberg / Delta Lake (open table formats)
Terraform / Infrastructure as Code
AWS Cloud — EMR, S3, Glue, Athena (Beneficial)
Snowflake (Beneficial)
Retail Operations experience
Developing Specialist Capability — Supports the analysis, diagnosis, testing, resolution, and documentation of data solutions within an agile team
Technical Aptitude — Strong passion and excitement for data, new technologies and solutions, and their range of possibilities and value for the business
Self-Motivation — High level of drive to set, meet, and exceed goals and expectations. Uses own initiative in dealing with challenges
Detail and Quality Focus — Has an affinity for structure and efficiency. Diligently watches over work processes, tasks, and outputs to ensure accuracy while promptly correcting quality concerns
Communication — Communicates well both verbally and in writing. Able to simplify complex technical concepts for a variety of stakeholders
Collaboration — Builds sound working relationships across the business. Able to work independently or collaboratively. Willing to coach/mentor others
Resilience — Ability to work under pressure and tight time constraints, efficiently prioritising workloads in a high-volume, fast-moving environment
Curiosity — Open to learning with a strong interest in data and discovery. Curious about exploring and answering business analytics questions
What you’ll do
Creating data feeds from on-premises to cloud environments
Supporting data feeds in production on a break-fix basis
Building data warehouse layers and data products using dbt or similar transformation tools
Manipulating data using Python and PySpark
Processing data using distributed compute, particularly EMR Serverless
Development for Big Data and Business Intelligence including automated testing and deployment
Working with open table formats (Iceberg/Delta Lake)
Writing and maintaining pipeline orchestration DAGs (Airflow)
Using Git for version control, branching, pull requests, and code reviews
Closing Date
2026/08/10
Share
Facebook
X
LinkedIn
Telegram
Tumblr
Whatsapp
VK
Bluesky
Threads
Mail