Skip to content
Apply now

This job is listed on several sites

T-Mobile

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration

T-Mobile Show all offers
Mid
Warszawa
Expires Oct 1, 2026
18 days ago

In short

AIOps Engineer for AI Infrastructure & Orchestration at T-Mobile in Warsaw. Responsibilities include vLLM service deployment, GPU resource management, and monitoring. Requires 5+ years in DevOps/SRE and 2+ years in MLOps/LLM platforms, with Kubernetes/OpenShift and vLLM experience.

AI-written summary based on the listing content.

Technologies we use

Your responsibilities

  • Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
  • Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
  • Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
  • Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
  • Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
  • Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
  • Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
  • Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
  • Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
  • Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
  • Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
  • Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
  • Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
  • Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.

Our requirements

  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
  • At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
  • Strong experience with Kubernetes and OpenShift administration in production environments.
  • Proven experience deploying and operating vLLM-based inference platforms in production.
  • Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
  • Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
  • Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
  • Strong Python programming skills with experience developing automation and operational tooling.
  • Experience with Bash scripting and Linux systems administration.
  • Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
  • Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.

What we offer

  • Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
  • You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.

Benefits

Published 2026-09-01
Source