Infrastructure & DevOps Engineer
Indexed description
Unlike cloud-native roles, this position sits directly at the intersection of on-premise datacenter operations and modern cloud/DevOps practices. The ideal candidate brings a solid foundation in physical infrastructure (networking, storage, virtualization, Dell servers) and has layered cloud and automation capabilities on top. You will have end-to-end ownership to move fluidly between troubleshooting firewall HA failovers, provisioning AWS infrastructure via code, and maintaining CI/CD pipelines.
About First Factory
We are a software development company with over two decades of experience, boasting a dynamic team of 175+ professionals actively engaged in diverse projects across various industries. We invite you to join us on this journey as we thrive and embrace fresh challenges.
Key Responsibilities
- Manage physical Dell server hardware deployment, hardware lifecycles, firmware updates, vendor/lease relationships, and physical datacenter operations (rack-and-stack, power, cooling, and migration/decommissioning projects).
- Configure, maintain, and troubleshoot Cisco switches/routers (Catalyst series), FortiGate firewalls (HA clusters, SD-WAN, IPsec VPN tunnels), and Mellanox datacenter switches. Diagnose ISP circuit-level issues and manage carrier fault tickets.
- Administer Pure Storage arrays (provisioning, performance monitoring, troubleshooting). Execute backup and disaster recovery workflows using Nakivo, Wasabi (S3-compatible cloud storage), and Rclone.
- Administer VMware vCenter and ESXi environments, utilizing PowerCLI for management and operational automation.
- Provision, maintain, and manage AWS cloud infrastructure (core services, IAM, VPC), manage working Azure environments, and maintain secure, seamless connectivity between on-premise data centers and cloud platforms.
- Write, review, and maintain Infrastructure as Code (IaC) using Terraform. Develop and maintain PowerShell and Bash scripts to automate routine operations.
- Own, maintain, and enhance CI/CD pipelines (GitHub Actions or equivalent) to automate build, test, and deployment processes.
- Implement and manage full monitoring coverage, alerting, and dashboards for legacy and cloud infrastructure using LogicMonitor and Datadog.
- Serve as a technical escalation point for infrastructure incidents, driving full-stack troubleshooting across network, storage, virtualization, cloud, and application layers.
- Proficiency with Cisco switching/routing (Catalyst series), FortiGate firewalls (HA clusters, SD-WAN, IPsec VPN), Mellanox datacenter switches, VLAN configuration, and routing protocols.
- Practical expertise in Pure Storage SAN/NAS administration, alongside DR and backup tooling (Nakivo, Wasabi S3-compatible storage, Rclone).
- Solid background managing VMware environments (vCenter, ESXi, PowerCLI) and Dell server lifecycles (rack-and-stack, firmware updates, hardware diagnostics).
- Strong AWS administration skills (core services, IAM, VPC), complemented by a working knowledge of Azure.
- Demonstrated use of Terraform for infrastructure provisioning, paired with pipeline management in GitHub Actions (or equivalent platforms).
- Fluency in PowerShell and/or Bash to build operational automation.
- Ability to configure end-to-end monitoring, alerting, and custom dashboards using LogicMonitor and/or Datadog.
- A track record of operating in combined environments, bridging physical datacenter operations with modern DevOps practices.
- Highly self-directed with a proven ability to take full ownership of complex technical problems through resolution.
- Sharp troubleshooting instincts across network, virtualization, storage, cloud, and application layers.
- Clear written and verbal English communication skills for technical documentation, status updates, and vendor management.
- Experience with additional network/firewall platforms (Palo Alto, Cisco ASA).
- Basic Postgres database administration (backup/restore, basic troubleshooting).
- Production on-call and incident management experience.
- Familiarity with Jira and Confluence for ticket tracking and technical documentation.
- Exposure to containerized and ECS-based deployments operating alongside traditional VM setups.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search