Network Operations Center Engineer
Indexed description
The Associate Engineer I, Network Operations Center (NOC) role is designed as a hybrid technical position within the NOC. Candidates are not expected to possess mastery in all four focus areas (Server & Infrastructure Operations, Network & Infrastructure Monitoring, Software & Tooling, Major Incident Response). Instead, successful candidates will demonstrate:
- Strength in one or more areas of the role’s focus,
- Working knowledge across additional areas, and
- A strong desire and ability to learn, cross-train, and grow into broader responsibilities over time.
The NOC operates as a 24x7 environment where engineers collaborate closely. Depth in one specialty area combined with adaptability in others is valued over complete mastery in all domains.
ESSENTIAL FUNCTIONS:
The following duties are a representative summary of the primary duties and responsibilities. Incumbent(s) may not be required to perform all duties listed and may be required to perform additional, position-specific duties.
Server & Infrastructure Operations
- Monitor, triage, and resolve issues in production Windows and Unix/Linux servers
- Perform routine maintenance, including patching, reboots, and upgrades
- Administer and manage VMware ESXi environments: VM provisioning, health checks, and host diagnostics
- Support Windows Server infrastructure; integrate with Azure cloud-based resources
- Manage Linux systems (RedHat preferred), including user management, services, disk I/O, and performance tuning
- Use remote tooling to access and support data center and cloud-hosted servers
Network & Infrastructure Monitoring
- Monitor network performance, including bandwidth utilization and device uptime
- Respond to network-related alerts and outages using tools like SolarWinds
- Troubleshoot connectivity issues, VPN tunnels, link degradation, or port errors
- Escalate complex issues with detailed diagnostics to network engineers
- Assist with routing/switching troubleshooting, including basic VLAN and trunk diagnostics
Software, Scripting & Tooling
- Develop and maintain scripts to automate common tasks (e.g., health checks, remediation actions, alert validation)
- Use scripting to enhance troubleshooting, generate reports, and streamline operational workflows
- Configure and tune monitoring platforms to ensure meaningful alerts and actionable metrics
- Contribute to internal documentation (e.g., Standard Operating Procedures (SOPs), KBs, and other technical references)
Major Incident Response & Coordination
- Participate in and lead Major Incident bridges using Blackrock3 IMS roles (e.g., Scribe, Liaison, IC)
- Drive rapid triage and service restoration under high-pressure, high-stakes conditions
- Coordinate cross-functional response teams (infra, apps, vendors)
- Provide real-time communications to leadership, stakeholders, and internal IT teams
- Conduct post-incident reviews to capture root causes and preventive actions
- Contribute to timeline notes, flag critical events for RCA and follow-up
- Leverage ITIL-aligned practices during incident, change, and problem workflows
- Participate in a 24x7x365 on-site operations environment, including nights/weekends/holidays
- Prioritize service uptime, customer impact, and post-incident learning
- Operate as a bridge between engineering teams and business stakeholders
Required Skills:
- Intermediate knowledge of Windows and Linux server administration
- VMware hypervisor experience (ESXi)
- Understanding of patch cycles, upgrade paths, and service restart processes
- Experience with remote session tools (RDP, SSH, IPMI, etc.)
- Working knowledge of TCP/IP, DNS, DHCP, and common routing protocols
- Familiarity with SNMP-based network monitoring tools (e.g., SolarWinds, Intermapper)
- Ability to interpret switch/router logs and CLI output
MINIMUM QUALIFICATIONS:
Education and Experience:
- Bachelor’s degree in computer science, Information Technology, or related field or equivalent experience, or a combination of both education and experience.
Preferred Licenses or Certifications:
- LPI / Linux+ (strongly preferred)
- Azure Fundamentals (a plus)
- CCNA or JNCIA (highly desired)
- CCISP or equivalent (a plus)
- Automation-related coursework or practical portfolio
- DevOps bootcamp or scripting certificate (optional but beneficial)
Preferred Knowledge and Skills:
- Understanding of network protocols (TCP/IP, BGP, OSPF, MPLS, VPN, DNS, etc.).
- Familiarity with data center, cloud Azure and hybrid infrastructure environments.
- Experience with NOC tools such as SolarWinds, Splunk, ServiceNow, Nagios, LogicMonitor, or equivalent.
- Knowledge of ITIL, Blackrock 3 IMS, automation (Python, Ansible, PowerShell), and monitoring frameworks (SNMP, NetFlow, etc.).
- Exposure to Blackrock3 or FEMA ICS frameworks
- ITIL Foundations (preferred but not required)
- Previous participation in Sev1/Sev2 coordination or on-call rotation
Physical Demands / Work Environment:
- Work is performed in a standard office environment.
- Participate in 24x7x365 operational coverage, including shift rotations and on-call support
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search