|
Detailed JD
|
We’re looking for a Big Data Lead Engineer to:
•
engineer reliable data pipelines for sourcing, processing, distributing, and
storing data in different ways, using cloud (Azure) data platform
infrastructure effectively.
•
transform data into valuable insights that inform business decisions, making
use of our internal data platforms and applying appropriate analytical
techniques.
•
develop, train, and apply data engineering techniques to automate manual
processes, and solve challenging business problems.
•
ensure the quality, security, reliability, and compliance of our solutions by
applying our digital principles and implementing both functional and
non-functional requirements.
•
build observability into our solutions, monitor production health, help to
resolve incidents, and remediate the root cause of risks and issues.
•
understand, represent, and advocate for client needs.
• Codify
best practices, methodology and share knowledge with other engineers in UBS
•
have a continuous improvement mindset, who is always on the look out for ways
to automate and reduce time to market for deliveries.
Your expertise
• Extensive experience in building Data Processing pipelines using Apache
Spark/Databricks with Python
and PySpark
• Good Knowledge of inner working on Apache Spark. Structured streaming is a
plus.
• Deep understanding of Python and its ecosystem, principles and tooling that
helps to write production
grade applications e.g. PEP8, MyPy, PyLint, Pytest.
• Good knowledge of data design patterns and methodologies to build a data
lake, based Azure cloud
stack e.g. ADLSv2.
• Experience in creating data structures optimized for storage and various
query patterns for DeltaLake,
Parquet, Avro
• Deep understanding of the SDLC using Gitlab, Github and knowledge of CI/CD
is a plus.
• Good to have working experience in cloud (Azure is preferrable). Knowledge
of (Kafka or Event Hub) is
plus
• Datalakehouses using medallion architecture. Knowledge of DataMesh
principles is a plus.
• Ability to debug using tools Spark UI, Ganglia UI, expertise in Optimizing
Spark Jobs
• The ability to work across structured, semi-structured, and unstructured
data, extracting information and
identifying linkages across disparate datasets.
• Experience of building applications using Polars, Pandas, Numpy is plus.
• Experience of building microservices on Kubernetes is plus
• Experience in traditional data warehousing concepts (Kimball Methodology,
Star Schema, SCD2).
• Experience in orchestration tools like Azure Databricks Workflow, Apache
Airflow is plus.
• Ability to clearly communicate complex solutions.
• Strong problem solving and analytical skills.
• Working experience in Agile methodologies (SCRUM)
• A proven team player with strong leadership skills, who can work in a
collaborative way across business
units, teams and regions
|