Site Reliability Engineer
Indexed description
Location: Norfolk, VA
Job Type: Full-Time (on-call support as needed)
Target Salary Range*: $145,000 - 165,000
- This represents the potential salary range for this position depending on education level, years of experience and/or certifications in addition to other position specific requirements which may impact salary
The Site Reliability Engineer supports, migrates, automates, and optimizes software development and deployment processes, Infrastructure as Code, and identity and access management capabilities. This role contributes to the maturity of the Site Reliability Engineering program and supports the creation, maintenance, modernization, and refresh of capabilities and components for the Navy Enterprise Network.
Key ResponsibilitiesSRE Testing and Reliability Engineering
- Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
- Create, script, and run performance tests to measure system behavior under varying levels of load and traffic.
- Identify bottlenecks, performance degradation, and areas for optimization.
- Design, implement, and maintain automated test suites for infrastructure and application components.
- Ensure testing is integrated into the CI/CD pipeline to validate system reliability with every release.
- Build automated systems for continuous performance testing, stress testing, and load testing.
- Work closely with Site Reliability Engineers, developers, and operations teams to define reliability goals and develop testing strategies to validate those goals.
- Ensure new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.
- Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.
- Ensure Service Level Indicators and Service Level Objectives are accurately measured and tracked through automated testing frameworks.
- Test, maintain, patch, STIG, upgrade, troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager, Active Directory Federation Services, DHCP, DNS, WINS, Group Policy Objects, and PKI.
- Support the creation, maintenance, update, modernization, and refresh of capabilities and components of the Navy Enterprise Network.
- Provide technical leadership and related task knowledge.
- Coordinate with project managers, customers, stakeholders, and engineers to support ongoing activities and new projects that maintain, transform, and modernize the Navy Enterprise Network.
- Write code to automate software releases, monitor systems, and detect and resolve problems before users are impacted.
- Develop features using AI coding tools and script repositories to automate, scale, test, and secure cloud infrastructure and pipelines.
- Develop and code high-quality pipeline automation workflows to support environments inside and outside the cloud platform.
- Ensure automation workflows align with business and technology strategies.
- Contribute to ongoing SRE maturity by recommending improvements to engineering build, maintenance, automation, and reliability across the platform using SRE/DevOps tools and Infrastructure as Code.
- Work with development and operations teams to support fast and reliable software deployments.
- Monitor systems and improve overall platform reliability.
- Discover, document, and resolve system bugs.
- Enhance performance monitoring of systems using Splunk or other dashboard reporting tools.
- Identify performance bottlenecks and optimize cloud infrastructure performance.
- Maintain complex computer systems through automation, monitoring, and proactive issue resolution.
- Resolve most conflicts between timeline, budget, and scope independently.
- Escalate sophisticated or consequential issues to senior management.
- Work nights, weekends, and provide on-call support as needed.
- B.S. degree and 2–4 years of prior relevant experience; or
- Master’s degree with less than 2 years of relevant experience.
- Experience designing, configuring, and managing Active Directory, Group Policy Objects, DNS, DHCP, and WINS.
- Experience with Windows Server 2016, 2019, 2022, or 2025.
- Experience with SQL Server 2019, 2022, or 2025.
- Experience with CI/CD toolsets, such as Jenkins or GitLab.
- Experience in application administration, configuration, and integration.
- Experience creating Jira and/or Azure DevOps workflows, projects, and custom configurations.
- Experience administering and maintaining an SRE platform using Ansible playbooks, such as upgrading Jenkins.
- Experience automating tasks with scripting languages such as PowerShell or Python.
- Experience integrating and maintaining third-party CI/CD tools such as Jenkins and GitLab.
- Experience with PaaS using Red Hat OpenShift, Kubernetes, and Docker containers.
- Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
- Experience with automated provisioning and configuration tools such as Terraform, CloudFormation, Chef, Puppet, Ansible, or similar technologies.
- Working knowledge of the Risk Management Framework and DISA STIGs.
- Familiarity with Agile development methodologies.
- Good command of Linux/Unix and command-line concepts.
- Automated script design, coding, debugging, and maintenance skills using Bash, Python, or similar languages.
- Knowledge of Agile, DevSecOps, and SRE concepts and best practices, with a desire to grow that knowledge.
- Hands-on experience with Atlassian products, including Jira, Confluence, and Bitbucket.
- Ability to work with a distributed team.
- Ability to work in a highly collaborative, forward-thinking, and innovation-driven environment.
- DoD 8570.01 IAT Level II certification required prior to onboarding and must be maintained while supporting the SMIT Contract. [Required]
- Active DoD Secret security clearance, with ability to maintain the clearance. [Required]
- Must be able to support program execution in classified environments and access SIPRNet from an NMCI location. [Required]
- 100% onsite work is required. [Required]
- Must be willing to work nights, weekends, and provide on-call support as needed. [Required]
- Previous work experience providing support to the NGEN-NMCI program.
- Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation for automating test environments.
- ITILv4, Scrum Master, or Agile SAFe certification, or applicable experience.
- Familiarity with designing, configuring, and managing FIM 2010 R2/MIM 2016 synchronization.
- Familiarity with VB.NET.
- Familiarity with designing, configuring, and managing Delinea.
- Familiarity with Microsoft Active Directory Federation Services structure.
- Familiarity with cloud engineering.
- Familiarity with Public Key Infrastructure certificates.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search