Back to jobs

DATA ENGINEER (DATABRICKS) | SPECIALIST

Jobgether
Full-timemid
Sign in to applyFree account, takes a minute.

Job description

Accountabilities: • Build and maintain scalable, reliable data pipelines using modern distributed processing technologies to support high-quality data ingestion, transformation, and delivery. • Organize and manage data tables using Delta Lake and Unity Catalog , ensuring effective data management and governance. • Design and implement an automated, native MLOps pipeline on Databricks and Google Cloud Platform (GCP) covering the complete machine learning model lifecycle. • Develop processes for data preparation, feature engineering, model training, validation, registration, deployment, serving, monitoring, and automated retraining. • Participate in technical discovery activities, including inventorying existing machine learning models and assessing their migration requirements. • Develop and validate a standardized MLOps pipeline template through a pilot implementation, followed by progressive migration of models in waves based on business and technical criticality. • Collaborate with engineering and data teams to ensure solutions are scalable, maintainable, reliable, and aligned with technical standards. • Work within an agile delivery model, actively participating in sprints, refinement sessions, reviews, retrospectives, and other team rituals. • Continuously identify opportunities to improve data pipeline performance, automation, reliability, and operational efficiency. Requirements • Demonstrated professional experience working with Databricks and modern data engineering environments. • Strong hands-on experience with PySpark and Apache Spark for distributed data processing. • Experience developing and orchestrating workflows using Apache Airflow . • Practical experience with Google BigQuery and cloud-based data platforms. • Experience integrating and using MLflow for machine learning lifecycle management. • Knowledge of AWS Glue and its application within data integration and processing workflows. • Experience working with both SQL and NoSQL databases , including technologies such as PostgreSQL, MongoDB, and Cassandra. • Ability to design and maintain scalable data pipelines with a strong focus on data quality, performance, automation, and reliability. • Experience working in Agile/Scrum environments , including sprint planning, refinement, reviews, and retrospectives. • Strong analytical and problem-solving abilities, with the capacity to work independently and collaboratively on complex technical challenges. • Nice to have: Knowledge of Apache Kafka for event streaming and real-time data architectures. • Nice to have: Experience with dbt for data transformation and analytics engineering. Benefits • Opportunity to work with modern data engineering, AI, cloud, and MLOps technologies . • Exposure to large-scale Databricks and GCP-based data environments. • Opportunity to contribute to end-to-end machine learning lifecycle automation and reusable technical solutions. • Collaborative and agile working environment. • Continuous learning and professional development opportunities. • Exposure to emerging trends in Artificial Intelligence, Generative AI, and advanced technology. • Opportunity to work on technically challenging projects with meaningful business impact. • Career growth within a technology-focused and innovation-driven environment. • Compensation and benefits package aligned with the role and local market How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether?    Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.     #LI-CL1

Skills

DatabricksPySparkApache SparkApache AirflowGoogle BigQueryMLflowAWS GluePostgreSQLMongoDBCassandraSQLNoSQLDelta LakeUnity CatalogGCPdbtApache Kafka