Back to search
Reclub Linkedin · Posted 2d ago

Senior DevOps & Infrastructure Engineer

Ho Chi Minh City

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About the Job:


This role owns infrastructure, reliability, deployments, monitoring and on-call. It is our first dedicated infrastructure hire. Today, the platform is run part-time by the engineers who built it. It works, but parts of the estate are still manual, backups are not drill-tested enough, and access controls need tightening.

The focus of this role is to close that gap: take ownership of production infrastructure, turn repeated manual work into code, improve operational safety, and build the runbooks and documentation the team can rely on.

This role includes three onsite days shared with the engineering team. Releases and on-call are remote. Releases are currently attended and often happen outside business hours, with time shifted rather than added on top of a full day. Part of the role is making them safe enough to run in daylight.


About Reclub

Reclub is a social sports platform that helps people organize activities, discover communities, and play together more easily. Our mission is to make recreational sports more accessible, social, and enjoyable.


Scope of Job

● Own production infrastructure across cloud environments, containers, networking and self-hosted servers.

● Run deployments and improve CI/CD pipelines, staged releases, migrations and rollbacks.

● Operate datastores, including backups, restores, replication, upgrades and capacity planning.

● Own observability across metrics, traces, logs, dashboards and alerting.

● Carry primary on-call and build the runbooks, restore drills, escalation paths and postmortem habits around it.

● Own security and access, including secrets management, patching, cloud and CI permissions, and least-privilege reviews.

● Replace repeated manual work with infrastructure as code.

● Manage infrastructure cost and capacity without compromising reliability.

Key Success Indicators for this Role :

● More infrastructure and access setup is documented, current and managed as code.

● Backup, restore and release processes are more reliable, better tested and safer to run.

● On-call and incident response are more sustainable, with clearer alerts, stronger runbooks and better follow-through on root causes.

RequirementsMust Have1. Professional Expertise (Hard Skills & Experience):

● 5+ years in DevOps, SRE, platform or infrastructure engineering with real production ownership.

● Strong Linux fundamentals and confident shell scripting.

● Strong hands-on cloud experience, preferably AWS, across compute, networking, storage and IAM.

● Strong hands-on infrastructure as code experience is a hard requirement, along with CI/CD ownership experience.

● Production experience with containers and orchestration such as ECS, Kubernetes or equivalent.

● PostgreSQL operations experience, including replication, backups, restores, upgrades and production migrations.

● Practical experience with secure remote access, VPN or zero-trust mesh networking, SSH access models and firewall policy.

● Experience being primary on-call for a production system and handling real incidents calmly.

● Strong observability and alerting judgment, with the ability to design alerts people trust.

2. Communication & Documentation:

● Documentation-first mindset.

● Able to keep runbooks, architecture notes, access lists and SOPs current.

● Able to turn operational knowledge into clear written standards the team can follow.

3. Personality & Behavioral Traits (The "Reclub" Fit):

● Comfortable operating as a single owner with autonomy and accountability.

● Calm and pragmatic under pressure.

● Focused on fixing root causes, not just absorbing repeated operational pain.

● Cost-conscious and deliberate in infrastructure decisions.

4. Culture & Career :

● Comfortable working in a hybrid setup with three onsite days shared with the engineering team.

● Willing to take ownership in an environment that works today but needs stronger infrastructure foundations.

● Motivated by building long-term systems and standards, not just clearing a short-term issue backlog.

Nice to Have

● Pulumi with TypeScript experience.

● Experience with self-hosted datastores or brokers such as ClickHouse, Scylla, Cassandra, MQTT or Kafka.

● Mobile release infrastructure experience across App Store and Play pipelines, build services and code signing.

● Security or compliance experience such as SOC 2, ISO 27001, endpoint hardening or MDM policy.

● Experience as a first dedicated infrastructure hire.

● Experience with identity-based mesh access such as Tailscale or similar.

Compensation & Benefits

● 100% salary during probation

● Competitive salary + 13th-month bonus

● Annual salary review

● Premium health insurance from Bao Viet after probation

● Transportation support

● Lunch, coffee, and snacks

● All statutory benefits per Vietnamese Labor Law


Work Environment

● In-office, collaborative working environment

● Professional English-speaking workplace

● Supportive, fast-moving team

● Opportunity to build new product areas from the ground up

● Working time : 9am-6pm, Mon - Fri

● Working Location: The Sun Avenue, 28 Mai Chi Tho, Binh Trung ward, HCM City


Apply
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search