Operations Engineer –(24/7 Platform Support)
Indexed description
Reports To: Operational Support Lead / NOC Manager
Shift Pattern: Rotating 24/7 Shift Roster (Includes Day, Night, and Weekend coverage with Shift Differential Pay)
Employment Type: Full-Time
🔹 Role Overview
We are seeking a proactive and detail-oriented Tier 1 Operations Engineer to join our 24/7 Platform Support team. In this role, you will serve as the first line of defense for our high-throughput, real-time trading and financial data ingestion platform.
You will monitor platform health, triage incoming telemetry alerts, execute initial troubleshooting, and uphold continuous operational uptime. This is an exceptional opportunity for an ambitious operations professional eager to accelerate their career in financial technology, cloud operations, and enterprise observability systems.
🔹 Key Responsibilities
- Maintain active, real-time monitoring of live trading infrastructure, low-latency APIs, and high-volume data pipelines using Grafana and observability dashboards. Identify and triage alerts strictly following standard operational runbooks.
- Perform rapid initial diagnostic checks across application, infrastructure, and network layers during operational events to isolate failing components and root causes.
- Accurately log, classify, and track operational issues through their full lifecycle within our ITSM tooling. Maintain clear, timely communication with internal business stakeholders throughout incident resolution.
- Timely escalate unresolved, complex, or high-severity technical incidents to Tier 2/3 engineering leads, providing structured diagnostic summaries and context.
- Maintain detailed real-time shift logs and lead thorough operational handovers to ensure seamless 24/7 continuous platform coverage. Proactively contribute to updating and refining runbooks for our evolving platform architecture.
- Assist senior engineering teams with scheduled maintenance windows, system health checks, pre/post-release validation, and sanity testing.
🔹 Required Skills & Experience
- 1–3 years of hands-on technical experience in a Technical Support, Network Operations Center (NOC), or Tier 1 IT Operations environment.
- Solid foundational understanding of Linux environments and command-line tools (CLI) for system inspection, process management, log navigation, and file analysis.
- Direct experience reading, validating, and working with JSON and YAML configuration files.
- Practical exposure to enterprise observability and monitoring platforms (e.g., Grafana, Prometheus, Datadog).
- Familiarity with structured incident management processes and platforms (e.g., Jira Service Management, ServiceNow, Linear.app).
- Excellent written and verbal English communication skills for incident logging, escalation drafting, and shift handovers. Full flexibility and willingness to work a rotating 24/7 shift pattern (including day, night, and weekend rotations).
🔹 Preferred Qualifications
- Fintech Domain Knowledge: Familiarity with financial trading platforms, market data feeds (FIX protocol, market feeds), or real-time data streaming architectures.
- Basic hands-on exposure to public cloud platforms (AWS, GCP, or Azure) and containerized environments (Docker, Kubernetes).
- Basic Shell scripting (Bash) or Python skills for automating minor routine operational tasks.
- Basic ability to execute standard SQL or NoSQL queries to verify data health and pipeline updates.
- ITIL v4 Foundation certification or relevant entry-level Cloud/Linux certifications (e.g., AWS Certified Cloud Practitioner, Linux Essentials).
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search