Senior IT & Infrastructure Engineer
Indexed description
StratusGrid is a cloud-native company building Stratusphere, a multi-agent platform that turns cloud complexity into measurable outcomes. This role is for a senior, hands-on engineer who will own and evolve the internal technology foundation that enables our team to work securely, efficiently, and at increasing scale.
You will operate across the full internal IT and infrastructure lifecycle: supporting a highly technical, all-macOS workforce; building automated onboarding and offboarding; managing identity, endpoints, SaaS systems, security tooling, and compliance controls; and improving AWS infrastructure through Terraform. This is an individual-contributor role for someone who sees recurring manual work as a problem to eliminate, uses AI and agents daily, and wants the opportunity to build or rebuild systems the right way from the ground up.
Responsibilities:
- Internal IT Ownership: Own the systems, tooling, processes, and documentation that support StratusGrid employees. Provide responsive, low-volume helpdesk support to technical and power users in an all-macOS environment while continuously reducing support needs through better systems and automation.
- Agent Tooling & Skills Development: Build and maintain the tools, skills, prompts, integrations, runbooks, and context that enable AI agents to safely perform IT and security work. Turn repeatable tasks into reusable agent capabilities for support, onboarding, access management, security triage, compliance, reporting, and infrastructure operations.
- Onboarding & Offboarding: Build and operate an exceptional employee lifecycle experience. Ensure accounts, permissions, devices, applications, and security controls are provisioned and removed accurately and on time, with automation handling the standard path and strong execution covering exceptions.
- Identity & Access Management: Design and maintain Okta as the central identity and automation layer. Implement scalable group, role, application, SSO, MFA, lifecycle-management, and least-privilege access patterns across the company’s systems.
- Endpoint & Device Management: Own Mosyle configuration and the macOS device lifecycle, including zero-touch deployment, baseline configuration, software distribution, patching, inventory, security policies, and device retirement.
- Security Operations: Configure and operate internal security tooling, investigate and respond to alerts, manage endpoint and infrastructure vulnerabilities, coordinate remediation, and improve response workflows through automation. Partner with Engineering on decisions that affect shared technical systems, with architecture ultimately accountable to the CEO.
- Cloud Infrastructure & Infrastructure as Code: Manage and improve internal AWS infrastructure using Terraform and disciplined Git workflows. Apply secure defaults, least privilege, patching, observability, change control, validation, and rollback practices to infrastructure changes.
- SaaS & Vendor Management: Own the operational administration of company SaaS systems and evaluate vendors for security, privacy, compliance, reliability, integration, and lifecycle-management requirements. Maintain accurate ownership, access, renewal, and risk records.
- Compliance & Control Operations: Operate and improve technical and administrative controls supporting SOC 2 Type II and GDPR. Keep controls effective throughout the year, maintain evidence and documentation, resolve gaps, and use Vanta and automation to reduce manual compliance work.
- Systems Reporting & Operational Visibility: Build reporting that makes system health, access, device posture, vulnerabilities, incidents, control status, and operational work visible and actionable. Establish measurable service levels and use data to prioritize improvements.
- Execution with Reliability & Urgency: Own work end-to-end. Capture commitments, communicate status and risks proactively, escalate early, document decisions, and close the loop visibly—especially for access, security, and employee-impacting issues.
Requirements:
- Broad IT & Infrastructure Experience: Senior-level experience operating across corporate IT, identity, endpoint management, SaaS administration, security operations, cloud infrastructure, and automation. You are comfortable moving between user support and deep systems work without treating either as beneath or outside the role.
- Automation & Software Skills: Strong ability to automate operational work using Python, TypeScript, shell scripting, APIs, webhooks, or similar tools. You build maintainable systems rather than collections of one-off scripts, with appropriate testing, logging, failure handling, and documentation.
- AI-Native Operator: You already use AI agents extensively to investigate problems, write and review code, automate workflows, create documentation, analyze alerts, and improve operational systems. You understand that agent output requires context, permissions, testing, monitoring, and you exercising good judgment.
- AWS & Terraform: Hands-on experience managing AWS infrastructure and delivering infrastructure changes with Terraform via pull requests.
- Identity & Access Management: Strong knowledge of SSO, MFA, SCIM, role and group design, access reviews, least privilege, joiner-mover-leaver workflows, and identity-centered automation. Direct Okta experience is strongly preferred.
- macOS & Endpoint Management: Experience managing a macOS fleet through an MDM platform, including automated enrollment, configuration profiles, software deployment, OS and application patching, device compliance, and troubleshooting.
- Security Operations & Vulnerability Management: Experience configuring security platforms, triaging alerts, investigating suspicious activity, remediating endpoint and cloud vulnerabilities, and developing practical incident-response and escalation workflows.
- Compliance & Privacy: Experience operating controls and producing evidence for SOC 2 Type II or a comparable framework. Working knowledge of GDPR requirements as they affect access, vendors, systems, data handling, and operational processes.
- Systems Thinking & Problem Solving: Proven ability to work through ambiguity, find root causes, design scalable solutions, and improve the program around a system—not just fix the immediate ticket.
- Excellent Communication: Clear and thoughtful written and verbal communication. You can explain technical risk and tradeoffs to leadership, collaborate effectively with engineers, support employees with empathy, and create documentation that others can reliably follow.
- Strong Ownership & Discretion: High standards of reliability, judgment, confidentiality, and follow-through. You will hold privileged access across the company and must handle it with care.
- Emergency Availability: Ability to respond outside normal working hours when an urgent incident materially affects company operations. This role does not own customer-facing production infrastructure or a routine customer on-call rotation.
- Remote-Work-Ready: Equipped to work effectively in a distributed team environment, including a reliable high-speed internet connection, a professional and distraction-limited workspace, and the ability to consistently communicate, collaborate, and execute independently.
- Authorized to work: Applicants must be legally authorized to work in the United States at the time of hire. This position does not offer visa sponsorship.
Nice-to-have / Differentiators:
- Hands-on experience with Okta, Google Workspace, Mosyle, Vanta, Aikido, SentinelOne, and GitHub Actions.
- Experience building identity-driven workflows using Okta Workflows, SCIM, application APIs, or event-based automation.
- Experience applying AI agents to IT or security operations, including alert enrichment, guided remediation, evidence collection, reporting, or employee support.
- Experience owning SaaS vendor reviews, security questionnaires, data-processing considerations, renewals, and application rationalization.
- Experience building dashboards or operational reporting from APIs and system data.
- Located in Chattanooga, TN and able to come into the Chattanooga HQ office on a regular basis and as needed.
Primary Systems:
- Okta for identity, access, lifecycle management, and automation
- Google Workspace for communication, documents, and collaboration
- Mosyle for macOS device and endpoint management
- Vanta for compliance monitoring and evidence management
- Aikido for code and cloud infrastructure security
- SentinelOne for endpoint detection and response
- GitHub and GitHub Actions for source control, infrastructure code, and automation
- AWS and Terraform for cloud infrastructure management
About StratusGrid
StratusGrid is building Stratusphere, a multi-agent platform that turns cloud complexity into measurable outcomes. Stratusphere coordinates specialized infrastructure agents to observe AWS and Azure environments, simulate policy-safe plans, and execute approved changes with auditability and rollback-ready safety, so savings compound over time, security improves, delivery accelerates, and teams spend less time on toil and fire drills.
We're a team of builders and operators who care about trustworthy automation and real business impact. Our work sits at the intersection of cloud engineering, product, and customer outcomes, helping customers make infrastructure changes they can measure, explain, and stand behind.
At StratusGrid, we recognize the importance of inclusion and the value of diverse perspectives. We are committed to equal employment opportunity regardless of race, color, national or ethnic origin, age, religion, disability, sexual orientation, gender, or any other characteristic.
We'd love to meet you!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search