Back to search
Cognizant Linkedin · Posted 20d ago

Infrastructure Engineer

Federal Territory of Kuala Lumpur

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Role purpose

This role is intended for senior engineers who can independently troubleshoot complex production issues, improve stability through automation and guide junior support resources.


Job Responsibilities

  • Own deeper technical triage across OS, database, batch, middleware/application logs, observability and platform dependencies.
  • Design and enhance automation scripts for recurring support tasks, health checks, log extraction, alert enrichment and manual process reduction.
  • Lead problem analysis, trend identification and preventative actions for recurring incidents.
  • Support containerized or modernized environments where Docker/Kubernetes are part of the production or transformation landscape.
  • Support the communication surveillance and related upstream / downstream applications through incident triage, root-cause analysis and service restoration.
  • Maintain production stability through monitoring, alert analysis, capacity awareness, runbook execution and risk escalation.
  • Work with CTB, RTB, vendor and cross-functional technology teams to support production fixes, enhancements, transition readiness and release/change activities.
  • Use Jira and Confluence to maintain traceability of incidents, problems, changes, risks, user stories, knowledge articles and operational procedures.


Must-have skills

  • Minimum 5 to 8 years of infrastructure/application production support experience, preferably in banking, capital markets or regulated technology.
  • Strong scripting depth in at least one language with ability to write, maintain and troubleshoot automation scripts independently.
  • Ability to write and optimize SQL queries, investigate database-related production issues and work with DBAs when deeper administration is required.
  • Hands-on experience in ITIL-based Incident, Problem and Change Management in production environments.
  • Strong operating system exposure across Linux / Unix / RHEL and Windows, with ability to troubleshoot application and infrastructure issues.
  • Scripting capability in at least one relevant language, preferably Bash, Python, PowerShell, Perl or batch scripting, with evidence of automation or support tooling.
  • Working knowledge of relational databases such as MS SQL and/or Sybase, including query execution, query writing and production issue investigation.
  • Experience using enterprise monitoring and observability tools such as Splunk, ITRS Geneos, AppDynamics, Dynatrace, Nagios, ELK or Grafana.
  • Ability to document support procedures, incident notes, change records and runbooks using Jira and Confluence or equivalent tools.
  • Banking, financial services, capital markets or regulated production support exposure.


Preferred / nice-to-have skills

  • Exposure to communication surveillance, conduct surveillance or regulatory surveillance platforms such as Smarsh Enterprise Conduct, Csurv or Theta Lake.
  • Experience with batch scheduling tools such as BMC Control-M, Autosys or Tidal.
  • Exposure to Docker and Kubernetes, especially troubleshooting deployed containers or supporting application platforms in production.
  • Understanding of sales tooling, CRM, trading or capital markets technology workflows.
  • Experience supporting disaster recovery, resilience testing, configuration management and production readiness reviews.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search