CloudOps Engineer (L1)
Indexed description
Roles & Responsibilities
Infrastructure Monitoring
- Continuously monitor the health, availability, and performance of servers, virtualization, Kubernetes, storage, backup, network, and GPU infrastructure using enterprise monitoring tools.
- Detect alerts, perform initial health checks, validate service availability, and initiate incident notifications as per operational procedures.
- Log, categorize, prioritize, and troubleshoot Level 1 incidents using SOPs and knowledge articles while ensuring SLA compliance.
- Restore services where possible and escalate unresolved or critical incidents to L2/L3 teams with complete documentation.
- Perform user account administration, password resets, account unlocks, and basic Windows server health verification.
- Validate Windows services and escalate server or Active Directory issues requiring advanced administration.
- Monitor Linux server availability, service status, and system logs to identify operational issues.
- Execute approved service restarts and escalate operating system issues beyond standard support procedures.
- Monitor virtual machine availability, health, and resource utilization across virtualization platforms.
- Perform basic VM recovery activities and escalate hypervisor or platform-related issues to specialized teams.
- Monitor Kubernetes clusters, nodes, pods, and services to ensure platform availability.
- Perform approved pod/service restarts and escalate cluster or orchestration-related issues.
- Monitor storage health, capacity, and backup job execution to ensure operational continuity.
- Perform authorized backup recovery actions and escalate storage or backup infrastructure failures.
- Perform basic network diagnostics, connectivity verification, and infrastructure service validation.
- Monitor network health and escalate complex connectivity or performance issues to network support teams.
- Monitor GPU server health, utilization, and hardware alerts to maintain operational readiness.
- Coordinate hardware replacement activities and vendor support for GPU-related incidents.
- Perform basic hardware health checks and assist in diagnosing infrastructure component failures.
- Support hardware replacement activities, post-maintenance validation, and asset record updates.
- Maintain accurate incident records, operational logs, and shift handover documentation.
- Follow SOPs, contribute to knowledge management, and report recurring issues for continuous service improvement.
Relevant Experience
- 1–3 years in an IT Infrastructure Operations, NOC, Service Desk, or Data Center Operations environment.
- 24×7 support operations, enterprise monitoring tools, ITSM platforms (ManageEngine, ServiceNow, BMC Remedy, Jira), and basic cloud or virtualization environments.
- 1–3 years of hands-on experience in Infrastructure Monitoring, Incident Management, Windows/Linux Administration, Active Directory, Virtualization (VMware/Hyper-V), Kubernetes, Storage & Backup Monitoring, Network Diagnostics, and ITSM tools (e.g., ServiceNow).
- 1–3 years of experience in Hardware Support, Documentation, ITIL/SOP adherence, Shift Operations, Asset & Vendor Coordination, Customer Support, and Operational Reporting
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search