Sr Software Engineer (Site Reliability) Austin or Dallas, TX
Indexed description
Location: Austin (preferred), open to Dallas, TX (Hybrid)
Key Responsibilities & Essential Functions
- Develops and maintains tooling used for environment monitoring and task automation
- Engages in and improves whole lifecycle of services, including inception and design, deployment, operation, and refinement
- Analyzes and establishes efficient configurations for software and servers, DB connections / indexes, drivers, etc.
- Collaborates with development teams to design service architectures, software platforms and frameworks, capacity planning, release plans and launch reviews
- Monitors internal and vendor service level objectives (SLOs) and agreements (SLAs); identifies / resolves SLO / SLA gaps
- Serves as technical subject matter expert (SME) for cross-functional engineering Teams; assists with / troubleshoots systems-related issues and maintenance
- 5+ years experience designing, analyzing, developing, or troubleshooting distributed systems
- 3+ years of SRE experience managing Google Kubernetes Engine (preferred), K8s, or AWS environments
- 2+ years of Java (Spring) programming experience preferred
- 3+ years of using Terraform to maintain cloud infrastructure
- 3+ years of CI Pipeline experience with either Gitlab Pipelines, or GitHub Actions
- Experience with tools such as Gitlab, JIRA, Slack, Confluence and Intellij is preferred
- Experience with microservices architecture patterns
- Experience working with PostgreSQL, Kubernetes, Docker, Linux, GCP, Terraform, and APIs using REST and GraphQL
- Experience working with monitoring and visualization tools such as Datadog, Grafana, or New Relic
- Strong proficiency with scripting languages such as Python, Ruby, Groovy, Bash
- Proven track record of researching, understanding, and effectively applying Scalability and High Availability principles
- Advanced knowledge in system and data architecture, data modeling, and design and capable of architecting and designing at the application or service level using well-accepted design patterns -
- Able to review platform designs for strength of engineering solutions, namely performance, sustainability, and iterative development potential. -
- Comprehensive knowledge of Computer Science fundamentals: data structures, algorithms, design patterns, system architecture and design patterns -
- Advanced understanding of development methodologies and processes -
- High degree of personal accountability to self and team for continued growth -
- Adjust - Leverages Agile metrics to improve team performance and deliverables. Evaluates and adjusts resources, self, and team as necessary. -
- Collaborate - Ability to work on tasks which span multiple domains, requiring cross-team collaboration, which have a high impact on your project. -
- Agility - Embraces risk, change, and helps team manage ambiguity within the team's scope of work. -
- Able to drive progress without having a complete picture and can articulate potential tradeoffs and prioritize when faced with ambiguity. -
- Connect - Delivers clear, concise, effective messages across different levels; can tailor communication based on intended audience. -
- Growth Mindset - Fosters a culture of mentoring and coaching across multiple technical teams and other stakeholders. -
- Relate - Fosters a culture within their team where people are encouraged to share their opinions and contribute to discussions in a respectful manner, approach disagreement non-defensively with inquisitiveness, and use contradictory opinions as a basis for constructive, productive conversations. -
- A Computer Science degree or comparable formal training, certification, or work experience -involving software / systems engineering
- Travel by car or plane with overnight stays
- Work extended hours; sit for extended periods
- Work rotating and on-call schedules, as needed
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search