- Location
- София, ул. Кукуш 1, сграда 7, етаж 3, Bulgaria
- Workplace
- Hybrid
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Closing date
- Today
- Source
- Workday
Description
Strength. Care. Growth
A1 Competence Delivery Center is a vital component of A1’s telecommunications business. Acting as an expertise hub, CDC is dedicated to delivering a full range of high-quality IT, network, financial and other services to support A1’s operations across all OpCos, independent of location.
Using the power of being OneGroup and leveraging synergies, CDC enables transparency of resources, key skills and knowledge expansion and personal career growth opportunities’ enhancement, paired with job stability.
You will know we are the right place for you, if you are driven by:
- Opportunities to learn and build your career.
- Meaningful work in a stable and fast-paced company.
- Diversity of people, projects, and platforms.
- A supportive, fun, and inspiring place to work.
Would you like to join us?
Job Purpose:
We are building an enterprise AI-as-a-Service platform that brings together reusable AI services, DataOps, MLOps, LLMOps, Kubernetes-based workloads, APIs, and secure enterprise integrations.
As a Senior Platform Engineer in Data DC, you will focus on the service and integration layer above the cloud foundation. You will help deploy, integrate, automate, and operationalise AI and data-platform services on Kubernetes.
Role insights:
- Deploy and maintain containerised AIaaS, DataOps, MLOps, and LLMOps services on Kubernetes.
- Integrate services such as workflow orchestration, data integration, metadata management, model tracking, feature stores, notebooks, model serving, RAG, and LLM observability.
- Develop and maintain Helm charts, Kubernetes manifests, configuration templates, and deployment pipelines.
- Establish repeatable deployment patterns across DEV, test, and production environments.
- Configure namespaces, service accounts, RBAC, resource quotas, network policies, secrets, certificates, and application ingress.
- Integrate platform services with enterprise APIs, databases, object storage, identity providers, monitoring systems, and internal data sources.
- Implement application-level observability using metrics, logs, traces, health probes, dashboards, and actionable alerts.
- Support vulnerability remediation, container-image governance, certificate renewal, secrets rotation, backup integration, and platform hardening.
- Diagnose failures spanning Kubernetes workloads, service configuration, APIs, identity, networking, storage, and external dependencies.
- Produce deployment documentation, technical runbooks, troubleshooting guides, and handover material.
- Review vendor deliverables and identify hidden infrastructure assumptions, privileged dependencies, portability gaps, and operational risks.
- Help translate pilot implementations into repeatable, supportable production services.
What makes you unique:
- Strong practical experience operating applications on Kubernetes.
- Sound knowledge of Helm, Kubernetes manifests, services, ingress, RBAC, secrets, ConfigMaps, persistent volumes, and network policies.
- Experience with CI/CD and infrastructure or configuration automation.
- Proficiency in at least one scripting or programming language, preferably Python, Go, or Bash.
- Experience integrating APIs, databases, identity services, object storage, and enterprise applications.
- Working knowledge of monitoring and observability concepts across metrics, logs, traces, dashboards, and alerting.
- Ability to troubleshoot across application, Kubernetes, network, identity, and storage boundaries.
- Clear technical communication and the ability to collaborate across engineering, security, architecture, and operations teams.
Nice to have:
- Experience with Cilium, CRI-O, Traefik, cert-manager, Kyverno, Falco, Harbor, Velero, Prometheus, Grafana, Loki, or OpenTelemetry.
- Exposure to Airflow, Airbyte, Meltano, OpenMetadata, MLflow, Feast, JupyterLab, LangFlow, vLLM, or RAG platforms.
- Experience with GPU scheduling, NVIDIA GPU Operator, model serving, or AI inference workloads.
- Familiarity with Exoscale or another European or sovereign-cloud provider.
- Experience assessing workload portability between Kubernetes distributions.
Our gratitude for the job done will be eternal, but we’ll also offer you:
- Innovative technologies and platforms to work with.
- Modern working environment for your comfort.
- Friendly, ambitious, and motivated teammates to support each other.
- Thousands of online and in-person learning opportunities to grow.
- Challenging assignments and career development opportunities in multinational environment.
- Attractive remuneration package.
- Flexible working schedule and opportunity for home office.
- Numerous additional goodies, including, but not limited to free A1 services, discounts, health insurance and services, sports center, childcare, team and family events, etc.
If you have any questions, please do not hesitate to contact Mariya Ivanova.