Senior DevOps Engineer – Large-scale cloud-native platform
Indexed description
Want to secure a large-scale cloud-native platform with 100+ production services?
Want to turn DevOps strategy into practical controls across Kubernetes, CI/CD, cloud infrastructure, secrets management, observability, and public-facing systems?
Come to us to build and improve security for a modern microservices ecosystem using GitLab CI, Flux CD, Kubernetes, Vault, Cloudflare, OpenTelemetry, SigNoz, ELK, VictoriaMetrics, and other production-grade technologies. This is a hands-on engineering role, not a pure compliance or advisory role. You will work directly with a small DevOps team that owns both platform operations and security improvements, with autonomy to build scripts, CI templates, dashboards, runbooks, policies-as-code, and practical controls that reduce real production risks.
Working model: remote-first or hybrid, with planned offline meetings twice each week.
Reporting line: DevOps lead
Environment
- 3 cloud providers: Google Cloud, Viettel Cloud, and CMC Cloud.
- Around 50 million requests per day.
- 100+ production services on Kubernetes.
- Multiple WordPress deployments on Kubernetes.
- Production databases across MSSQL, MongoDB, PostgreSQL, ClickHouse, and MySQL.
- GitLab CI, Flux CD, Terraform/Terragrunt, Ansible, Vault, OpenTelemetry, SigNoz, ELK, VictoriaMetrics, and Grafana.
- Self-hosted stateful infrastructure and high-impact production workloads.
- Operate Kubernetes clusters, ingress, DNS, networking, storage, workload scheduling, and shared runtime components.
- Investigate incidents using metrics, logs, traces, Kubernetes events, infrastructure state, and application signals.
- Improve alerts, dashboards, escalation paths, and runbooks for critical systems.
- Participate in production escalation, including after-hours incidents when required.
- Convert repeated incidents into automation, standards, capacity changes, or documented prevention work.
- Build and maintain infrastructure using Terraform/Terragrunt, GitLab CI, Flux CD, and Ansible.
- Standardize cloud patterns across Google Cloud, Viettel Cloud, and CMC Cloud where practical.
- Manage network connectivity, firewall rules, VPN, DNS, service exposure, certificates, and cloud resource lifecycle.
- Improve rollback, break-glass access, production change review, and GitOps recovery procedures.
- Review infrastructure changes for reliability, security, cost, and operational impact.
- Maintain reusable GitLab CI templates and deployment patterns.
- Help application teams adopt standard patterns for configuration, secrets, rollouts, observability, and production readiness.
- Reduce one-off support through self-service workflows, checks, and concise documentation.
- Support operational work around MSSQL, PostgreSQL, MongoDB, MySQL, ClickHouse, Redis, Kafka, and Elasticsearch.
- Improve replication, capacity planning, performance investigation, and recurring data-pipeline operations.
- Coordinate high-risk data changes with service owners and database stakeholders.
- Maintain runbooks for lag investigation, restore testing, failover, and data incident response.
- Improve availability patterns for workloads, ingress, databases, caches, queues, storage, and cloud infrastructure.
- Validate backup, restore, failover, rollback, and recovery procedures for critical systems.
- Run disaster recovery drills and document RTO, RPO, and remediation gaps.
- Review architecture changes for single points of failure, capacity risk, zone dependency, and recovery impact.
- Build scripts or tools for inventory, drift detection, alert review, cost review, backup checks, and recurring operations.
- Use AI-assisted workflows for log analysis, incident summaries, runbook drafts, infrastructure review, and operational reporting.
- Keep human approval for production changes, database recovery, security exceptions, and disaster recovery execution.
- Production Kubernetes experience covering deployments, services, ingress, autoscaling, resource limits, troubleshooting, and recovery.
- Hands-on Linux, networking, DNS, TLS, load balancing, firewall, VPN, and production debugging experience.
- Cloud infrastructure experience. Google Cloud is preferred; multi-cloud experience is a strong plus.
- Infrastructure as code experience with Terraform or Terragrunt.
- CI/CD experience. GitLab CI is preferred.
- GitOps or Kubernetes deployment automation experience. Flux CD or Argo CD is preferred.
- Observability experience with logs, metrics, traces, alerts, dashboards, and incident analysis.
- Scripting or automation ability using Bash, Python, Go, or similar tools.
- Production experience with at least one database, cache, queue, or search system.
- Ability to write clear runbooks, incident notes, technical decisions, and async updates.
- Senior enough to own problems end to end without waiting for detailed task breakdown.
- Comfortable working in a small DevOps team where platform, cloud, security, database, observability, and support work compete for capacity.
- Prioritizes production safety, automation, documentation, and repeatable standards.
- Communicates clearly in writing and works effectively with distributed engineering teams.
- Can explain tradeoffs to developers, product teams, and leadership without hiding behind tool names.
- Accepts operational responsibility during incidents.
- Has operated Kubernetes for a production microservices platform with meaningful traffic.
- Has reduced alert noise, recurring incidents, manual changes, or deployment failures.
- Has built reusable CI/CD templates, GitOps workflows, Terraform modules, runbooks, or internal tools.
- Has handled incidents involving Kubernetes, cloud networking, databases, Kafka, Redis, Elasticsearch, or observability systems.
- Has performed backup and restore testing, disaster recovery planning, failover drills, or postmortems.
- Has worked at a startup, marketplace, fintech, ecommerce, travel, logistics, SaaS, or other high-traffic online platform.
- Has practical security experience around IAM, Vault, Kubernetes hardening, CI/CD controls, secrets rotation, or public endpoint protection.
- Has used AI tools responsibly for investigation, documentation, code review, automation, or operational analysis.
- Attractive salary package with performance-based monthly bonuses.
- 100% salary during the probation period.
- 13th-month salary to ensure financial stability.
- Performance reviews twice a year with opportunities for salary adjustments.
- Hybrid working model with remote work and offline in-person meetings twice each week.
- Work directly with the DevOps Lead and collaborate with experienced CTOs, architects, and tech leaders at Vexere.
- Codex AI access to support research, analysis, automation, and documentation work.
- Opportunity to secure a cloud-native microservices platform with meaningful production scale.
- Startup environment with low bureaucracy, no micromanagement, and direct ownership.
- Freedom to research, test, and apply new technologies when they provide practical value.
- Build your public profile through publishing technical articles, contributing to open-source projects managed by Vexere, and joining tech talks or industry events.
- Attractive stock options for dedicated developers and team leads.
- Up to 30% discount on bus tickets for Vexere employees and their families.
- 12 days of annual leave, with 1 extra day off every 3 years, convertible into salary.
- Vibrant and dynamic work environment with a friendly, supportive team.
- Training sessions on negotiation, communication, work management, interpersonal skills, and software technology.
- Exciting company activities: annual trips, team-building events, year-end parties, and more.
- Send your Resume to email: [email protected], with title: Fullname – Applied Position
- Phone/ Zalo: 0966 197 741 ( Mr. Anh)
- Office Location: Vexere Trading Service Co., Ltd – 2nd floor – Building H3, 384 Hoang Dieu, Ward 6, District 4, HCMC
- Working Hours: Hybrid Working 8:30 am – 6:00 pm from Monday to Friday, and Saturday morning.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search