Senior MLOps Engineer
W skrócie
Senior MLOps Engineer to manage ML infrastructure, LLM serving, fine-tuning pipelines, and multi-provider routing. Requires 5+ years in MLOps/ML Infra, Python, Docker, Kubernetes. Remote, Full-Time. Max 145 PLN/h net B2B.
Słowa kluczowe
Skrót przygotowany przez AI na podstawie treści ogłoszenia.
We are seeking a Senior MLOps Engineer to own end-to-end Machine Learning infrastructure, with a strong focus on high-scale LLM inference serving, distributed fine-tuning pipelines, and multi-provider gateway routing. In this role, you will bridge the gap between ML models and production reliability—deploying self-hosted open-weight models (ranging from ~7B to ~376B parameters), optimizing multi-GPU/distributed topologies, managing cloud cost governance, and treating the internal engineering platform as a product.
Details
Role: Senior MLOps Engineer / ML Infrastructure Engineer
Seniority: Senior / Lead (5+ years in production MLOps/ML Infrastructure)
Allocation: Full-Time (1 FTE)
Rate Cap: Max 145 PLN/h net B2BWork Model: Remote / Cloud-native across major hyperscalers (AWS, GCP, Azure)
Responsibilities
High-Scale Inference Serving
Deploy, scale, and optimize self-hosted open-weight models (~7B to ~376B parameters) using engines such as vLLM, Triton Inference Server, or TGI.
Apply continuous batching, tensor/pipeline parallelism, and quantization strategies (AWQ, GPTQ, FP8) tailored to model size and strict SLA constraints.
Multi-Provider Gateway & Cost Governance
Operate and enhance the intelligent routing layer spanning self-hosted models and external APIs (OpenAI, Anthropic, OpenRouter).
Build token accounting, rate limiting, budget controls, and cost/latency/quality-aware routing logic.
Distributed Training & Fine-Tuning Infrastructure
Build and maintain automated pipelines for fine-tuning, evaluation, versioning, and continuous delivery (MLflow, SageMaker Pipelines, or Kubeflow).
Manage distributed training workloads using DeepSpeed, FSDP, or Accelerate.
Reliability, Observability & Production Ownership
Own production reliability, monitoring, logging, and incident response for ML services (handling GPU OOMs, degraded inference, and latency spikes).
Participate in on-call rotation for the inference serving platform.
Evaluation Harnesses & Platform Engineering
Stand up automated evaluation and verification harnesses to catch quality/performance regressions before deployment.
Deliver infrastructure-as-code (IaC) and CI/CD pipelines to provide self-serve tooling for internal engineering teams.
Requirements
Production Experience: 5+ years of hands-on experience in MLOps, ML Infrastructure, or ML Engineering owning end-to-end model lifecycles in production.
On-Call & Incident Handling: Proven experience being on-call for live ML services and resolving real-world incidents (e.g., GPU OOMs, memory leaks, routing failures, cost blowups).
Inference & Quantization Depth: Deep understanding of quantization and parallelism trade-offs under latency, throughput, and hardware cost constraints.
Core Tech Stack:
Languages: Advanced Python (C/C++ for performance-sensitive paths is a plus).
Containerization & Orchestration: Docker, Kubernetes (EKS/GKE), Helm, and IaC tools (Terraform/Cloud
Formation).
Cloud Ecosystems: Deep experience with hyperscaler ML services (AWS SageMaker, EC2 GPU instances, Lambda).
MLOps & Training Frameworks: MLflow, Kubeflow, PyTorch, DeepSpeed, FSDP, or Hugging Face Accelerate.
Nice to Have / BonusPrior experience building multi-provider LLM gateways with token billing and cost controls.
Track record of building programmatic LLM evaluation/verification frameworks.
| Opublikowana | 2026-08-11 |
| Wygasa | 2026-11-09 |
| Źródło |
|
Hexjobs App
Narzędzia dopasowane do tej oferty.
Hexjobs App
Narzędzia dopasowane do tej oferty.
Podobne oferty
Cloud Data Architect (GCP)
Future Processing
GliwiceSenior AWS DevOps Engineer
co.brick Talents
GliwiceSenior Salesforce Engineer
co.brick Talents
GliwicePySpark Engineer
co.brick Talents
GliwiceSenior Backend Developer
co.brick Talents
Gliwice