Skip to content
Apply now

This job is listed on several sites

ACAISOFT POLAND Sp. z o.o.

Lead Data Engineer (Databricks)

ACAISOFT POLAND Sp. z o.o. Show all offers
170 - 210 PLN
B2B
Lead
Manager
Warszawa
Expires Oct 7, 2026
12 days ago

In short

Lead Data Engineer needed in Warsaw for data ingestion/processing on Databricks. Requires Python, SQL, Spark, Databricks experience, data modeling, and heterogeneous source integration. B2B contract, PLN 170-210/hr.

AI-written summary based on the listing content.

Technologies we use

About the project

This is how we organize our work

This is how we work

Your responsibilities

  • Build, evolve, and operate data ingestion and processing capabilities for structured, semi-structured, and unstructured data, supporting the transition from early prototypes through early adoption and general release.
  • Implement and maintain rich metadata and data quality practices that enable cross- record querying, traceability, and AI-ready data access across experiments, files, inventory, and workflows.
  • Work with architects, AI engineers, workflow engineers, and domain experts to ensure data is usable, performant, and trustworthy for downstream GenAI and analytics use cases.
  • Support engineering quality through code reviews and mentoring of junior/mid-level engineers.
  • Lead architectural decisions and work with senior business and technology stakeholders.Guide and mentor data and engineering teams

Our requirements

  • Experience as a Lead / Principal Data Architect in multinational environments.
  • Deep knowledge of data modelling, integration, warehousing, and modern data platforms.
  • Strong hands-on Python and complex SQL, with production Spark experience on Databricks (Structured Streaming, Auto Loader, Delta Lake, Unity Catalog) to support analytics, AI/ML, and enterprise application use cases.
  • Practice working with structured and unstructured data across the full lifecycle: ingestion, transformation, enrichment, indexing, and retention, delivered through disciplined engineering practice including infrastructure-as-code, CI/CD, and automated testing of pipelines.
  • Proven experience ingesting from heterogeneous operational sources including relational (Oracle, PostgreSQL) and document stores (MongoDB), using Full load + CDC or equivalent replication patterns, with practical handling of schema drift, deletes, and late or out-of-order data.
  • Skills in data modelling across a layered architecture, with clear contracts between raw, curated and serving layers, supporting cross-entity queries and contextual linking, and performant data access patterns.

Optional

  • Experience building unstructured document pipelines: parsing and extraction from PDF, Office and scanned formats, chunking, embedding generation, and maintaining searchable indexes at scale.

What we offer

  • Private medical care
  • Multisport card
  • Friendly, informal atmosphere, and direct contact with everyone in the company
  • Modern technologies

This is how we work on a project

Benefits

Published 2026-09-07
Source