Scrumconnect Limited Logo

Scrumconnect Limited

Data Engineer

Posted 29 Days Ago
In-Office or Remote
Hiring Remotely in Walker-on-Tyne, Newcastle upon Tyne, England, GBR
Entry level
In-Office or Remote
Hiring Remotely in Walker-on-Tyne, Newcastle upon Tyne, England, GBR
Entry level
Build, maintain, and troubleshoot scalable AWS data pipelines using Python, PySpark, Spark, and Airflow. Perform root-cause analysis of pipeline and data-quality issues, apply dimensional modeling and slowly changing dimensions, and deliver governed data assets. Provision infrastructure with Terraform, containerize solutions with Docker, and manage GitLab CI/CD deployments. Work with AWS security, encryption, IAM, monitoring, and analytics services in a regulated government environment.
The summary above was generated by AI

Data Engineer
Up to £65k per annum


Apache Spark Python AWS Cloud Data Pipelines

A hands-on data engineering role within a large-scale cloud data programme, responsible for building, maintaining, and troubleshooting data pipelines using Apache Spark, PySpark, Apache Airflow, and a broad suite of AWS services. You will apply strong analytical and engineering skills to deliver trusted, well-governed data assets in a modern, cloud-native environment.

About Scrumconnect

Scrumconnect is a leading UK technology consultancy delivering digital transformation across public and private sectors, contributing to over 20% of the UK's major citizen-facing public services. We specialise in cloud engineering, data platforms, and agile delivery, helping clients build scalable, secure, and user-centred digital solutions that create real impact.


Working arrangement:
This role is hybrid. Candidates must be willing and able to travel to the Newcastle office once per week. Remaining days may be worked remotely from anywhere in the UK.

About the role

You will work as a Data Engineer on a complex, cloud-based data programme - designing, building, and maintaining data pipelines that process large volumes of data across a modern AWS-native stack. Using Apache Spark and PySpark for distributed data processing, Apache Airflow for orchestration, and a range of AWS services for storage, compute, and analytics, you will help deliver reliable, well-governed data assets to downstream users.

You will apply strong data analysis skills to identify root causes of data issues, work with dimensional data models and slowly changing dimensions, and implement infrastructure as code using Terraform. Familiarity with engineering best practices and the ability to translate customer expectations into applied technical functionality are key to success in this role.

Key responsibilitiesData pipeline development

Build and maintain scalable data pipelines using Apache Spark and PySpark, processing and transforming large datasets across distributed cloud infrastructure.

Workflow orchestration

Configure and manage Apache Airflow DAGs for task orchestration, ensuring reliable scheduling, monitoring, and execution of data processing workflows.

Root cause analysis

Perform data analysis to identify and resolve root causes of pipeline failures and data quality issues - including reviewing EMR output logs and CloudWatch metrics.

Data modelling

Apply understanding of dimensional data models and slowly changing dimensions (SCD) to design and maintain well-structured, analytically trusted data assets.

Infrastructure as code

Provision and manage cloud infrastructure using Terraform. Containerise solutions using Docker and manage deployments through GitLab CI/CD pipelines and release tagging.

Security & encryption

Apply understanding of both Server Side and client-side encryption patterns within AWS. Work within IAM policies and data governance standards appropriate to a regulated government environment.

Technical skills requiredLanguages & analytics

  • Python - primary language for pipeline development and data processing
  • SQL - used for querying, transformation, and validation across data stores
  • PySpark - Power BI for distributed data processing using Apache Spark on AWS EMR
  • Familiarity with basic data structures for constructing robust, scalable solutions

Data processing & orchestration

  • Apache Spark - understanding of distributed data processing architecture and execution
  • Apache Airflow - configuring DAGs and managing task orchestration at scale
  • Jupyter Notebooks - for exploratory data analysis and pipeline prototyping
  • Understanding of dimensional data models and slowly changing dimensions (SCD Types 1, 2, 3)
  • Data analysis skills to identify root cause of issues within pipelines and data assets

AWS services

  • Amazon EMR - running Spark workloads and reviewing output logs
  • Amazon Athena - ad hoc querying of data in S3
  • Amazon Textract and Comprehend - familiarity with AI/ML document extraction and NLP services
  • AWS S3, IAM, CloudWatch, EC2, ECR - core platform services used day-to-day
  • AWS console proficiency - navigating, configuring, and monitoring services
  • Understanding of Server Side and client-side encryption within AWS

Infrastructure, DevOps & delivery

  • Terraform - Infrastructure as Code for provisioning and managing AWS environments
  • Docker - containerisation of data engineering solutions
  • GitLab - source code management, CI/CD pipeline configuration, release tagging, and component versioning
  • Familiarity with engineering best practices
  • Ability to translate customer expectations into applied, functional technical solutions

Technology stack at a glance

PythonPySparkSQLApache, Power BI SparkApache AirflowJupyter NotebooksDimensional modelling/SCDAWS EMRAmazon AthenaAWS S3AWS IAMAWS CloudWatchAWS EC2/ECRAmazon TextractAmazon ComprehendTerraformDockerGitLab CI/CDGitLab Tags



Similar Jobs

Yesterday
In-Office or Remote
Bristol, England, GBR
Mid level
Mid level
Consulting
Designs, delivers, and optimizes scalable cloud-based data platforms and solutions. Responsibilities include building ingestion, transformation, modeling, and integration pipelines; ensuring performance, security, governance, reliability, and data quality; collaborating with analysts, scientists, architects, and project managers; supporting feature engineering and analytical projects; implementing monitoring, testing, CI/CD, documentation, and data lineage; and troubleshooting complex data and infrastructure issues.
Top Skills: Apache AirflowAzure Data FactoryAzure FunctionsC#Ci/CdData WarehousesEtl/EltJavaNon-Relational DatabasesPythonRelational DatabasesScalaSQL
12 Days Ago
Remote or Hybrid
2 Locations
Senior level
Senior level
Enterprise Web • HR Tech • Information Technology • Software • Cybersecurity
Develop and maintain customer-facing analytical reporting, data pipelines, and data models. Build Python applications, ensure data quality and consistency, collaborate with analytics engineers, and enable business teams to access actionable data. The role also involves cloud data warehouse optimization, infrastructure as code, software engineering best practices, and cross-functional product collaboration within an agile SaaS environment.
Top Skills: AWSAzureBigQueryCloudFormationContinuous IntegrationDbtFlaskGCPGitLookerPlotlyPower BIPythonRedshiftSnowflakeSQLSqlalchemyTerraform
11 Days Ago
Remote
GBR
Junior
Junior
Information Technology • Analytics • Business Intelligence • Big Data Analytics
Build and maintain reliable data pipelines using Python and SQL. Integrate REST APIs, databases, files, cloud storage, and third-party platforms; process and load data into warehouses for analytics and BI. Implement incremental and fault-tolerant loading, logging, monitoring, retries, error handling, and data quality checks. Work with ETL/ELT, Git, and potentially dbt, Spark, orchestration tools, cloud platforms, Docker, dimensional modeling, and BI datasets.
Top Skills: AirflowAWSAws S3AzureAzure Blob StorageBigQueryCi/CdDagsterDbtDockerDomoGCPGitGoogle Cloud StorageLookerMs SqlPostgresPower BIPrefectPysparkPythonQuicksightRedshiftRest ApisSnowflakeSparkSQLTableau

What you need to know about the Bristol Tech Scene

Along with Gloucester, Swindon and Bath, Bristol is part of the "Silicon Gorge" tech hub, a region in the U.K. renowned for its high-tech and research-driven industries, with a particular emphasis on sustainability and reducing environmental impact. As the European Green Capital, Bristol is home to 25,000 cleantech companies, including Baker Hughes and unicorn Ovo Energy. The city has committed to achieving net-zero emissions within the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account