Lead Data Engineer (Databricks)
ACAISOFT POLAND Sp. z o.o.
Zobacz wszystkie oferty
170 - 210 PLN
B2B
Lead
Menedżer
Warszawa
Wygasa 7 paź 2026
12 dni temu
W skrócie
Lead Data Engineer needed in Warsaw for data ingestion/processing on Databricks. Requires Python, SQL, Spark, Databricks experience, data modeling, and heterogeneous source integration. B2B contract, PLN 170-210/hr.
Słowa kluczowe
Skrót przygotowany przez AI na podstawie treści ogłoszenia.
Technologies we use
About the project
This is how we organize our work
This is how we work
Your responsibilities
- Build, evolve, and operate data ingestion and processing capabilities for structured, semi-structured, and unstructured data, supporting the transition from early prototypes through early adoption and general release.
- Implement and maintain rich metadata and data quality practices that enable cross- record querying, traceability, and AI-ready data access across experiments, files, inventory, and workflows.
- Work with architects, AI engineers, workflow engineers, and domain experts to ensure data is usable, performant, and trustworthy for downstream GenAI and analytics use cases.
- Support engineering quality through code reviews and mentoring of junior/mid-level engineers.
- Lead architectural decisions and work with senior business and technology stakeholders.Guide and mentor data and engineering teams
Our requirements
- Experience as a Lead / Principal Data Architect in multinational environments.
- Deep knowledge of data modelling, integration, warehousing, and modern data platforms.
- Strong hands-on Python and complex SQL, with production Spark experience on Databricks (Structured Streaming, Auto Loader, Delta Lake, Unity Catalog) to support analytics, AI/ML, and enterprise application use cases.
- Practice working with structured and unstructured data across the full lifecycle: ingestion, transformation, enrichment, indexing, and retention, delivered through disciplined engineering practice including infrastructure-as-code, CI/CD, and automated testing of pipelines.
- Proven experience ingesting from heterogeneous operational sources including relational (Oracle, PostgreSQL) and document stores (MongoDB), using Full load + CDC or equivalent replication patterns, with practical handling of schema drift, deletes, and late or out-of-order data.
- Skills in data modelling across a layered architecture, with clear contracts between raw, curated and serving layers, supporting cross-entity queries and contextual linking, and performant data access patterns.
Optional
- Experience building unstructured document pipelines: parsing and extraction from PDF, Office and scanned formats, chunking, embedding generation, and maintaining searchable indexes at scale.
What we offer
- Private medical care
- Multisport card
- Friendly, informal atmosphere, and direct contact with everyone in the company
- Modern technologies
This is how we work on a project
Benefits
| Opublikowana | 2026-09-07 |
| Źródło |
|
Hexjobs App
Narzędzia dopasowane do tej oferty.
Pozostało 18 dni
07.10.2026
Hexjobs App
Narzędzia dopasowane do tej oferty.
Podobne oferty
Senior AI Engineer (Computer Vision)
AI CLEARING sp. z o.o.
Warszawa, MasovianAI Engineer
WEALTHARC sp. z o.o.
Warszawa, MasovianSolution Architect (Java)
QualityMinds Sp. z o.o.
Warszawa, MasovianExpert IT Network Engineer
CD PROJEKT RED S.A.
Warszawa, MasovianSignal Processing & Data Analysis Specialist (f/m/x)
Sii Sp. z o.o.
Warszawa, Masovian