A
Data Engineer (Spark)
Addepto
Contract type
Ongoing
Work mode
100% remote
Experience
Mid-level · 4+ years
Job description
Key details
- Develop and maintain a high-performance data processing platform for automotive data, ensuring scalability and reliability
- Design and implement data pipelines that process large volumes of data in both streaming and batch modes
- Optimize data workflows to ensure efficient data ingestion, processing, and storage using technologies such as Spark, Cloudera, and Airflow
- Work with data lake technologies (e.g., Iceberg) to manage structured and unstructured data efficiently
- Collaborate with cross-functional teams to understand data requirements and ensure seamless integration of data sources
- Monitor and troubleshoot the platform, ensuring high availability, performance, and accuracy of data processing
- Leverage cloud services (AWS) for infrastructure management and scaling of processing workloads
- Write and maintain high-quality Python (or Java/Scala) code for data processing tasks and automation
- Company mission
- Information not specified
Primary stack
Core technologies
AWSDockerPython (Programming Language)
Benefits
- Work in a supportive team of passionate enthusiasts of AI & Big Data
- Engage with top-tier global enterprises and cutting-edge startups on international projects
- Enjoy flexible work arrangements, allowing you to work remotely or from modern offices and coworking spaces
- Accelerate your professional growth through career development paths, knowledge-sharing initiatives, language classes, and sponsored training and conferences
- Benefit from partnerships with Databricks and Anthropic, which provide access to industry-leading training materials
Requirements & details
- At least 4 years of commercial experience implementing, developing, or maintaining Big Data systems, data governance and data management processes
- Strong programming skills in Python (or Java/Scala): writing clean code, OOP design
- Hands-on with Big Data technologies like Spark, Cloudera, Kafka, Data Platform, Airflow, NiFi, Docker, and Iceberg
- Excellent understanding of dimensional data and data modeling techniques
- Experience implementing and deploying solutions in cloud environments
- Consulting experience with excellent communication and client management skills, including prior experience directly interacting with clients as a consultant
- Ability to work independently and take ownership of project deliverables
- Fluent English (at least C1 level)
- Bachelor's degree in technical or mathematical studies
- Experience with an MLOps framework such as Kubeflow or MLFlow
- Familiarity with Databricks and/or dbt
- Spark, Cloudera, Airflow, Iceberg, Python, AWS, Kafka, NiFi, Docker, Databricks, dbt, Kubeflow, MLFlow
- Python (Programming Language)
- AWS
- Docker
