Zum Inhalt springen
CodiLime

Data Engineer

Praca zdalna
vor 11 Tagen

Kurzfassung

Data Engineer role focused on building and operating production Python systems for a large-scale data platform. Requires Python, SQL, Snowflake, dbt, Airflow, Spark, Azure, FastAPI, and CI/CD experience. Offers remote/hybrid work, flexible hours, and professional growth. Location: Poland.

Von KI erstellte Kurzfassung des Anzeigentextes.

The project and the team

You will join the team behind a large-scale, centralized data platform built for a global consulting organization. The platform is the shared source of company data behind several of the firm's internal products, and is used daily by consultants for company research, including support for Mergers & Acquisitions (M& A) engagements.

This is fundamentally a data engineering role combined with software engineering: you'll be designing, coding, testing, and operating production Python systems - pipelines, libraries, and services - that move, transform, and serve data at scale. This is not a role focused on configuring tools or writing one-off queries. You'll be building reliable, maintainable software that powers our data platform. The goal is a unified, enterprise-grade dataset of 300M+ company records, integrated from 10+ external and internal sources.

The platform delivers firm-level and site-level data - firmographics, technographics, and hierarchical relationships (parent company, subsidiary, site) - alongside key business metrics such as revenue, CAGR, EBITDA, headcount, M& A activity, competitors, industry classification, and web traffic. Data needs to stay accurate, well-structured, and fast to query as both the dataset and the number of consumers keep growing.

Technology stack:

  • Languages: Python, SQL
  • Data platform: Snowflake, dbt
  • Workflow orchestration: Apache Airflow (complex DAGs), running on Kubernetes
  • Data processing: Apache Spark on Azure Databricks
  • Data tooling and DBs: pandas, Polars, PyArrow, DuckDB, PySpark, PostgreSQL, Redis
  • Cloud: Azure (AKS, Blob Storage, ACR, Databricks, Open

Search, Azure AI Search)

  • API & services: FastAPI (REST, async), API Gateway
  • Testing & code quality: pytest, mypy/pyright, ruff/black, sqlfluff, SonarQube
  • Schema validation: Pydantic
  • Dependency & environment management: uv, Poetry
  • CI/CD & infrastructure: GitHub Actions, Docker, Kubernetes
  • AI-Assisted Development: Cursor, Claude Code, ChatGPT Enterprise
  • Future direction: agentic AI systems, LangChain, Azure OpenAI integration

What else you should know:

  • Team: Data Architecture Lead, Data Engineers, DataOps Engineers, Backend Engineer, Product Owner, collaboration with Frontend Engineers and Data Science and AI Engineers
  • Distributed team across Europe and India
  • Agile, collaborative environment; given the platform's organization-wide impact, we're looking for a mature, proactive, results-driven approach
  • Code quality is enforced through testing, typing, and tooling - not just code review

We work on multiple interesting projects at a time, so it may happen that we’ll invite you to an interview for another project if we see that your competencies and profile are well suited for it.

More reasons to join us

  • Flexible working hours and approach to work: fully remotely, in the office or hybrid
  • Professional growth supported by internal training sessions and a training budget
  • Solid onboarding with a hands-on approach to give you an easy start
  • A great atmosphere among professionals who are passionate about their work
  • The ability to change the project you work on
Veröffentlicht 2026-09-07
Quelle