ProdOps Engineer III
Indexed description
Credit One Bank is seeking a highly motivated Production Operations (ProdOps) Engineer to manage and support mission-critical production platforms. The ideal candidate will have strong expertise in GitLab CI/CD, Kubernetes, PostgreSQL, Database Deployments, Application Troubleshooting, Dynatrace Monitoring, and On-Call Production Support. This role requires a proactive individual who can ensure platform stability, automate operational processes, improve deployment efficiency, and rapidly resolve production incidents.
Summary Of Essential Job Functions
Production Operations & Support
- Provide operational support for production and non-production environments.
- Participate in a 24x7 on-call rotation and respond to critical production incidents.
- Monitor system health, application performance, and infrastructure availability.
- Perform root cause analysis (RCA) and implement preventive measures.
- Troubleshoot complex applications, databases, infrastructure, and deployment issues.
- Design, maintain, and optimize GitLab CI/CD pipelines.
- Automate application deployment, testing, and release processes.
- Implement deployment best practices to improve reliability and reduce downtime.
- Collaborate with development teams to streamline DevOps processes.
- Deploy, configure, and manage containerized applications on Kubernetes.
- Troubleshoot pod, service, ingress, and cluster-level issues.
- Manage scaling, availability, and performance optimization of Kubernetes workloads.
- Work with Helm charts and Kubernetes manifests for application deployments.
- Plan and execute database deployment activities across environments.
- Support database schema changes, migrations, and rollback procedures.
- Monitor and optimize database performance.
- Collaborate with development teams on database release strategies.
- Support and maintain PostgreSQL databases.
- Troubleshoot database performance bottlenecks and connectivity issues.
- Manage backup, recovery, replication, and high-availability configurations.
- Ensure database security and compliance standards are maintained.
- Configure and maintain monitoring dashboards using Dynatrace.
- Analyze application and infrastructure performance metrics.
- Create alerts and proactive monitoring strategies to identify issues before customer impact.
- Use observability tools to drive system reliability and operational excellence.
- Develop automation scripts and operational tooling.
- Create and maintain operational runbooks and standard operating procedures.
- Identify opportunities for process improvement and operational efficiency.
- Partner with engineering teams to improve platform reliability and scalability.
Technical Skills
- GitLab CI/CD
- Pipeline creation and optimization
- GitLab Runners
- Deployment automation
- Kubernetes
- Cluster operations and administration
- Helm Charts
- Pod, Service, Ingress troubleshooting
- Database Deployments
- Schema migrations
- Release management
- Rollback strategies
- Monitoring & Observability
- Dynatrace
- Application Performance Monitoring (APM)
- Log analysis and alerting
- Troubleshooting
- Application debugging
- Infrastructure issue diagnosis
- Performance analysis
- Production incident management
- 24x7 On-Call Support
- Incident Management
- Root Cause Analysis (RCA)
- Problem Management
- Change Management
- Release Coordination
- Bachelor’s degree in computer science, Information Technology, or related field.
- 5+ years of experience in Production Support, DevOps, SRE, or Platform Engineering roles.
- Experience with Linux administration and shell scripting.
- Familiarity with cloud platforms such as Azure, AWS, or GCP.
- Experience with Infrastructure as Code (Terraform preferred).
- Understanding of microservices architecture and containerization technologies.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search