T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
T-Mobile
Show all offers
Mid
Warszawa
Expires Oct 1, 2026
18 days ago
In short
AIOps Engineer for AI Infrastructure & Orchestration at T-Mobile in Warsaw. Responsibilities include vLLM service deployment, GPU resource management, and monitoring. Requires 5+ years in DevOps/SRE and 2+ years in MLOps/LLM platforms, with Kubernetes/OpenShift and vLLM experience.
AI-written summary based on the listing content.
Technologies we use
Your responsibilities
- Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
- Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
- Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
- Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
- Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
- Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
- Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
- Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
- Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
- Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
- Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
- Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
- Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
- Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.
Our requirements
- 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
- At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
- Strong experience with Kubernetes and OpenShift administration in production environments.
- Proven experience deploying and operating vLLM-based inference platforms in production.
- Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
- Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
- Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
- Strong Python programming skills with experience developing automation and operational tooling.
- Experience with Bash scripting and Linux systems administration.
- Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
- Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.
What we offer
- Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
- You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.
Benefits
| Published | 2026-09-01 |
| Source |
|
Hexjobs App
Tools tailored to this listing.
12 days left
10/1/2026
Hexjobs App
Tools tailored to this listing.
Similar offers
Senior AI Engineer (Computer Vision)
AI CLEARING sp. z o.o.
Warszawa, MasovianAI Engineer
WEALTHARC sp. z o.o.
Warszawa, MasovianSolution Architect (Java)
QualityMinds Sp. z o.o.
Warszawa, MasovianExpert IT Network Engineer
CD PROJEKT RED S.A.
Warszawa, MasovianSignal Processing & Data Analysis Specialist (f/m/x)
Sii Sp. z o.o.
Warszawa, Masovian