Software Engineer, Infrastructure
Indexed description
We're a team of approximately 50 people, well-funded, growing quickly. Our work directly influences how the world deploys AI agents and systems at scale..
About
Learn more about how we work.
The Role
Gray Swan is looking for an Infrastructure Engineer to build and scale the systems that power our AI security platform. You'll design the backend services, distributed infrastructure, and cloud architecture that enable Gray Swan to build AI systems and for customers to safely deploy frontier AI models at scale.
This role is ideal for an engineer who enjoys solving infrastructure challenges across reliability, scalability, observability, and performance. You'll work closely with machine learning engineers, product engineers, and security researchers to ensure our platform remains fast, resilient, and secure as we grow.
You'll have significant ownership over foundational systems and the opportunity to influence technical direction in a rapidly evolving AI startup.
What You’ll Do:
- Design, build, and maintain highly available backend services and distributed systems that power Gray Swan's AI security platform.
- Own cloud infrastructure across Kubernetes, AWS, networking, storage, and compute to ensure reliable production environments.
- Build scalable APIs, internal platform services, and infrastructure tooling that improve developer productivity and system reliability.
- Improve system observability through logging, metrics, tracing, dashboards, and automated alerting.
- Optimize performance, latency, and infrastructure costs while maintaining reliability and security.
- Partner closely with machine learning, security, and product engineering teams to deliver production-ready infrastructure for AI workloads.
- 5+ years of experience building backend infrastructure or distributed systems in production environments.
- Strong programming skills in C/C++, Go, Python, Rust, or Java.
- Experience operating services on Kubernetes and modern cloud platforms such as AWS, GCP, or Azure.
- Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures.
- Experience designing APIs, microservices, asynchronous systems, and event-driven architectures.
- Comfortable debugging complex production issues and improving reliability through automation and operational excellence.
- Passionate about writing clean, maintainable code and building infrastructure that other engineers love using.
- Excited to work in a fast-moving startup with significant ownership and ambiguity.
- Experience supporting machine learning or LLM infrastructure.
- Familiarity with infrastructure-as-code tools.
- Experience with Kafka, Redis, PostgreSQL, ClickHouse, or similar distributed data systems.
- Experience building internal developer platforms or platform engineering tooling.
- Knowledge of cloud security, infrastructure hardening, or zero-trust architectures.
- Previous experience at a high-growth startup or building products from zero to one.
- Interest in AI safety, cybersecurity, or adversarial machine learning.
You’ll Thrive Here If You:
- You thrive on ownership and solving hard problems. You're energized by ambiguity, enjoy building systems from the ground up, and take pride in delivering reliable solutions from design through production.
- You think at scale. You enjoy designing resilient infrastructure, optimizing performance, and building systems that are secure, observable, and built to grow.
- You collaborate across disciplines. You work effectively with machine learning engineers, security researchers, and product teams, knowing that the best infrastructure enables everyone else to move faster.
- You're excited by our mission and startup environment. You enjoy moving quickly, adapting to change, and helping build the foundation for the future of secure AI.
Compensation: $180,000 - $290,000 (depending on level) plus performance based bonus and meaningful equity package.
Benefits:
- 401k with up to 4% matching
- 28 days annual leave (vacation + holidays)
- Health, dental, and vision coverage
- Catered lunches (Pittsburgh office)
- Flexible work arrangements
- Visa sponsorship available for exceptional candidates
✏️ Online technical screen (15 min). Complete a simple, job-relevant exercise.
🗣 Intro call (30 min). We learn about you; you learn about us.
🧑💻 Technical interview (90 min). Live coding with some tasks requiring AIand others not.
🗣 Experience & culture interview (60 min). Conversational exploration of the skills fit.
😇 Reference checks. We’ll reach out to 3-5 references that you provide.
📃 Offer. If it’s mutual, we move fast.
How To Apply
Submit your resume, link to your portfolio, and answer the questions on the application.
Compensation Range: $180K - $290K
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search