Site Reliability Engineer
Indexed description
Hi,
Hope you are doing good!
We have a Full-time opportunity for you as Site Reliability Engineer @ Chandler, AZ
Role: Site Reliability Engineer
Locations: Chandler, AZ
Type of Hiring: FTE
Required Skill and Experience
• Strong experience in Site Reliability Engineering / Production Engineering.
• Hands-on expertise with: IBM MQ (queue managers, clustering, channels, DLQ management).
• Kafka / Confluent platform (topics, brokers, partitions, consumer groups).
• Large-scale distributed messaging systems and runtime management.
• Deep understanding of: System reliability, scalability, and high availability design.
• Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering).
• Incident management, root cause analysis, and problem management.
• Experience with: Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms.
• Event and anomaly detection in high-volume systems.
• Strong scripting/automation skills: Shell, Python, PowerShell.
• Experience managing Linux/Unix and Windows production environments.
• Knowledge of: Event-driven architecture and messaging-based integration patterns.
• Understanding of: Messaging platform security (TLS, certificates, channel auth, encryption).
• Vulnerability remediation and risk mitigation in production systems.
• Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures.
Preferred Skill and Experience
• Clear, concise communication with technical and non-technical stakeholders.
• Ability to work effectively across engineering, infrastructure, security, and application teams
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search